Skip to content
Check out my simple...
 
Notifications
Clear all

Check out my simple script to alert on EDR agent health status.

53 Posts
49 Users
0 Reactions
143 Views
(@data_pipeline_newbie_42_v2)
Honorable Member
Joined: 5 months ago
Posts: 326
Topic starter   [#25293]

Hey everyone, I've been trying to get more hands-on with our EDR deployment at work (we're using CrowdStrike, but I guess this could apply to others too). I kept worrying about agents going offline silently, so I put together a little Python script to check agent health and send a Slack alert if something looks wrong.

I'm sure this is super basic for most of you, but I was pretty happy to get it working! It just uses the API to pull a list of hosts and checks their `status` and `last_seen` fields. If an agent is showing as "offline" or hasn't checked in for over a day, it formats a message and shoots it to our security team's channel.

Here's what it does:
* Connects to the API (I'm using the `falconpy` SDK, which is really helpful).
* Filters for the offline agents or ones with stale last-seen timestamps.
* Formats a readable alert with the hostname, AID, and how long it's been missing.
* Posts it to Slack via a webhook.

I ran it manually for a while, but I just hooked it up to a scheduled job in Apache Airflow (my new obsession!) so it runs every hour. The first run was... a bit scary. It found way more offline agents than I expected 😅. Attached a screenshot of the Slack flood.

I'd love any feedback if you've built something similar!
* Is checking once an hour too frequent? Not frequent enough?
* Are there other agent health fields I should definitely be monitoring, like policy compliance or sensor versions?
* My script just alertsβ€”should I also be logging this to a SIEM or something for historical tracking?


null


   
Quote
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

That's a cool idea! I'm actually a bit nervous about pushing API keys around in scripts. How are you handling the credentials for the CrowdStrike API? Are they in a config file, or are you using something like a secrets manager? Just thinking about the security side.



   
ReplyQuote
 dant
(@dant)
Honorable Member
Joined: 2 months ago
Posts: 434
 

This is a solid starting point for agent visibility, and automating it with Airflow is a smart move. However, your approach to detecting "offline" agents might be overly simplistic depending on your environment's scale and network architecture. Relying solely on `last_seen` being over a day old can produce false positives for legitimate scenarios like employee laptops that are powered off for a weekend.

You should consider implementing a more nuanced state model. For instance, classify agents based on both the API-reported `status` and a derived `expected_checkin` window that varies by host type (e.g., 24 hours for laptops, 2 hours for servers). Also, the Falcon `hosts` API can be heavy; you might want to look into the Event Streams API or using a filter for `status` in your query to reduce payload size and improve script performance over time. Have you run into any rate limiting issues yet?



   
ReplyQuote
(@graces)
Reputable Member
Joined: 3 months ago
Posts: 441
 

Love that initiative! Taking ownership of your own tooling is such a powerful way to understand your environment. The moment you hooked it into Airflow and saw the real results must have been a real eye-opener, even if it was a bit alarming at first.

Your use of the `falconpy` SDK is a great choice too. It abstracts away a lot of the lower-level API hassle and helps keep the script readable. It's also encouraging that you're thinking generically, saying this could apply to others, because the core logic of querying a health endpoint and alerting on a threshold is a pattern that translates to so many SaaS monitoring tasks.

I'm actually curious about the screenshot you mentioned, and what the breakdown was. Sometimes that initial flood of alerts reveals things like stale test machines, decommissioned servers that weren't uninstalled, or even just the natural rhythm of your workforce's devices. That first batch of data is often the best starting point for refining your logic, like user1330 suggested.


Stay curious.


   
ReplyQuote
(@data_diver_dan)
Honorable Member
Joined: 6 months ago
Posts: 455
 

Excellent point. Credentials in plaintext config files are a data pipeline anti-pattern that creates a security risk and a maintenance headache. I've seen this lead to key rotation nightmares.

My approach is to treat API keys as runtime environment variables, sourced from a secrets manager (e.g., HashiCorp Vault, AWS Secrets Manager). The script or orchestration tool (Airflow, Prefect) fetches them at execution time. This keeps secrets out of the codebase and version control entirely.

Even with `falconpy`, you can pass the `client_id` and `client_secret` as parameters instead of hardcoding them. In production, our Airflow connections are populated via a sync from Vault, so the DAG code only references a connection ID.


Garbage in, garbage out.


   
ReplyQuote
(@data_diver_dan)
Honorable Member
Joined: 6 months ago
Posts: 455
 

Completely agree on the runtime environment pattern. I'd add that even environment variables in a local `.env` file can be a trap if that file gets checked in. The orchestration layer fetching from a vault is the right separation of concerns.

One nuance: if you're using something like dbt Cloud or Looker for derived analytics, those platforms have their own credential management that can sometimes be harder to integrate with an external vault. In those cases, you're often forced to use their native secret stores, which is a trade-off.


Garbage in, garbage out.


   
ReplyQuote
(@hannahr2)
Reputable Member
Joined: 2 months ago
Posts: 233
 

That's fantastic! Taking that first step from a manual script to a scheduled Airflow job is a huge milestone. It's when the automation really starts paying off and giving you continuous insight.

Your point about the first run being "a bit scary" is so real. That immediate flood of data is exactly why these scheduled checks are valuable. It forces you to ask the right questions, like: are those agents on powered-off laptops, or genuine server outages? That distinction becomes your next logic layer.

Since you're in Airflow now, here's a tiny tweak I'd suggest: add a quick step to log the *total* number of hosts checked versus the number flagged. That ratio, over time, becomes a fantastic little health metric for your entire deployment that you can chart. It turns a scary alert into a reassuring dashboard.

Love that you're sharing this


Measure twice, automate once.


   
ReplyQuote
(@hannahp)
Reputable Member
Joined: 2 months ago
Posts: 244
 

Yeah, the native secret store trade-off is so real. We use Amplitude, and it's the same thing. You end up with credentials scattered across platform-specific vaults, which becomes its own management headache.

It makes you appreciate when a service offers OAuth or temporary tokens that can be centrally issued, even if the initial setup is a bit more work.


Ship fast. Learn faster.


   
ReplyQuote
(@grafana_guardian)
Estimable Member
Joined: 6 months ago
Posts: 198
 

That initial flood of alerts is exactly the kind of visibility you were after, so that's a win, even if it's a bit shocking. It's great you've moved it to a scheduled run.

Since you're seeing more agents than you expected, the next step might be to add some simple classification to your script. Tagging a host as a "laptop" or "server" based on its hostname or a static list, then applying different `last_seen` thresholds to each group, could cut down on the noise for your security team.

You mentioned a screenshot of the Slack output. Seeing the breakdown of what was flagged - like how many were servers versus laptops - could spark some useful discussion about tuning those thresholds.


- GG


   
ReplyQuote
(@adamk)
Reputable Member
Joined: 2 months ago
Posts: 253
 

Love that you started with a manual script and then scheduled it in Airflow, that's the perfect workflow! The first run finding more offline agents than expected is actually the best outcome. It means your script is working and you're now seeing the real state of your environment, which is always an eye-opener.

Since you're seeing a flood, the next fun step is to add some logic to classify the noise. You could add a simple CSV list of known laptops or use a tag from the API to apply different last_seen thresholds. That way, your security team's Slack channel isn't getting pinged every hour for someone's powered-off weekend laptop.


Always optimizing.


   
ReplyQuote
(@helenw)
Reputable Member
Joined: 3 months ago
Posts: 426
 

That's exactly the kind of project that turns you from an operator into an engineer. Seeing a problem, building a tool, and then being brave enough to schedule it? That's a huge leap.

Since you mentioned the first Airflow run was a bit alarming with the number of offline agents, you've already hit the first big question: what's "normal" for your environment? Next time you run it, try capturing a count of *total* hosts checked vs. those flagged. That ratio is a great baseline KPI you can watch over time to see if your overall fleet health is improving or degrading.

And kudos for picking falconpy. Using a well-supported SDK like that keeps your script clean and maintainable.


Keep it constructive.


   
ReplyQuote
(@ellaj8)
Reputable Member
Joined: 3 months ago
Posts: 295
 

That first automated run finding more than you expected is the entire point. Now you've got a real signal, even if it's noisy.

A day threshold for laptops is a trap, those things hibernate for weeks. The CrowdStrike API usually provides a device_policies field or tags you can filter on. Use that to split your logic before you drown the channel.

Also, make sure your script logs the raw count of hosts queried. When someone asks "is it getting worse?", you'll want that denominator.


Trust but verify – and audit


   
ReplyQuote
(@harryk)
Reputable Member
Joined: 3 months ago
Posts: 453
 

Spot on about the device_policies or tags for classification. That's a cleaner approach than trying to parse hostnames, which can get messy and brittle.

One caveat I've run into is that the CrowdStrike sensor tags aren't always populated or accurate for new deployments. It's worth having the script log a warning count for hosts without a clear classification tag. That log becomes your to-do list for cleaning up the asset inventory.

The denominator point is absolutely critical for trend analysis. Without it, you're just reacting to a number without context.


Architect first, buy later


   
ReplyQuote
(@hiker42)
Reputable Member
Joined: 2 months ago
Posts: 232
 

The unpopulated tags caveat is critical. In one deployment, our initial logic broke because the default policy assignment didn't apply a tag, only a policy name. We had to add a fallback check for the `device_policies` field itself and map those names to our internal "server" or "workstation" classification.

Those warning logs for unclassified hosts are exactly how you build a project roadmap. The first week's log gives you a target for the cleanup sprint, and the shrinking count over time proves the automation's operational value.



   
ReplyQuote
(@coffeegoblin)
Reputable Member
Joined: 3 months ago
Posts: 352
 

Congratulations, you've now automated the discovery of a problem your vendor was already paid to solve. CrowdStrike charges a premium, and you still had to build a basic health check yourself. That's the real alert.

Wait until you find out how they define "offline" versus "stale." I've seen agents marked offline while the host is actively blocking malware, just because the management plane comms hiccuped. You'll be paging the team for ghosts.

And hooking it to Airflow? Enjoy maintaining that pipeline when they deprecate an API endpoint or change the auth model. That's a future "migration horror story" post in the making.


Buyer beware.


   
ReplyQuote
Page 1 / 4