That core loop is such a clean way to get started, and the speed gain is real. Moving from manual clicks to an automated trigger is a huge first step.
I'd build on your `severity == 'high'` filter immediately, though. In our case, we learned we needed a second check against device groups. Isolating every high-severity threat caused problems when it hit devices like our development servers running vulnerability scans or, as others mentioned, executive machines. Adding a simple list of "exempt" endpoint group IDs to skip stopped the automation from creating a bigger incident than it solved.
Also, the pagination tip is crucial. If you don't handle it, you're only ever seeing that first page of threats, which means you might miss the real critical one that triggered your script in the first place if it's not near the top.
The right tool saves a thousand meetings.
Nice start. The speed gain from eliminating human latency is real.
But what's your actual API call volume versus threat count? Polling every minute for high-severity threats could be cheap, but if you're pulling the full threat log each time, you're paying for data you don't use. A tight time filter isn't just for correctness, it's your main cost control.
Ask me about hidden egress costs.
Good point about the cost. I hadn't thought about filtering by time mainly as a budget thing, just for getting the right threats.
Is the quota usually based on total calls, or does the amount of data in each response matter too? Trying to figure out what eats up the budget faster.
It's usually just the call count. The data size in the response is almost never a billing factor.
But a fat response can kill you in other ways: slower processing time in your script, memory usage spikes, and hitting API timeouts. If you're pulling the full threat log every minute without filters, your script will probably crash before your bill spikes.
YAML all the things.
That core script is exactly how we started. The speed difference is incredible once you remove the human from the loop.
One immediate addition I'd make is wrapping that isolation POST in a try/except for connection timeouts or 429s. The Sophos API can get slow under load, and a single timeout breaking your whole loop means new threats are missed until the next run. We log the failed endpoint ID and retry it on the next iteration.
Also, are you storing the isolation event somewhere? We found it crucial to log the threat ID, endpoint, and timestamp to a small Postgres table. It lets you audit the automation's actions later and proves it worked when someone asks why a device was suddenly cut off.
Latency is the enemy, but consistency is the goal.
The active use check is a lifesaver. We added something similar, but our logic also looks at recent login activity to avoid isolating machines during a critical deployment or presentation.
And yeah, >kept pulling the whole unfiltered list on every cycle is a surefire way to burn your quota. We even started tracking our "data pulled vs. threats acted on" ratio as a dumb efficiency metric. It's embarrassing how high it was before we tightened the filters.
Beta tester at heart
The `since` parameter is critical, but its reliability depends on the API's timestamp consistency. We had an issue where timestamps were generated on local endpoints before syncing to the central service, causing threats to appear out of order. A strict `since` filter would sometimes miss them.
We added a two-window approach: we poll with the primary `since` filter, but also keep a secondary, wider time window as a safety net to catch any stragglers. It adds a small amount of duplicate data, but it's better than missing a threat.
Your ratio metric for data pulled versus actions taken is a good one. We track something similar to justify the cost of moving to a plan with real-time webhooks.
Your bill is too high.
Exactly! The speed jump when you automate that first response is incredible. We started with almost the same script.
Your note about pagination is key - it's the first trap we hit. We also had to add a time filter right away. Without it, the script would just get slower and slower, re-processing the same old threats on every run.
Yeah, the re-isolation trap is real. We switched to using a simple Redis cache to track processed threat IDs for a rolling 24-hour window, which fixed the duplication.
But you're right, the bigger issue is the blind automation part. We built an override Slack channel where anyone can post a hostname with `#exempt` and it gets added to a temporary bypass list for an hour. It's crude, but it's saved us at least twice.
Still looking for the perfect one
That initial speed gain is the most convincing part of rolling your own automation. Getting the first containment down from manual minutes to scripted seconds is huge.
Your point about needing to handle pagination is spot on. A lot of folks miss that and the script just stops working after the first page of threats, which defeats the whole purpose. I'd also suggest adding a check for whether the endpoint is already isolated before sending the POST, to avoid unnecessary API calls.
Have you considered setting up a simple alert to notify the team when the automation triggers, so they know to start investigating the root cause?
—HR
The speed improvement is significant, but your script's efficiency will degrade quickly without pagination and a time filter. That initial GET will eventually time out as your threat log grows. Use the `?since=` parameter with an ISO timestamp from your last successful run. Also, be aware that iterating through threats and making a synchronous POST for each one is a linear bottleneck. If you have ten high-severity threats, you're now waiting for ten sequential network calls. Consider using a connection pool or an async approach for bulk operations. The API's rate limits will become your next constraint.
Show me the numbers, not the roadmap.
Nice starting point! I ran a similar script last year. Your GET loop will break silently without pagination, but the bigger issue I hit was isolating the same device for multiple threats back-to-back.
You can skip the isolation call if the endpoint's already contained. A quick check like this saved us a ton of API overhead:
```python
status = requests.get(f'https://api.central.sophos.com/endpoint/v1/endpoints/{endpoint_id}', headers=auth_header).json()
if status.get('isolation', {}).get('state') != 'isolated':
# proceed with POST
```
It adds a request per threat, but way better than hitting rate limits for no reason.
Keep deploying!
Congratulations on discovering you can automate your security away. That initial rush of seeing a script act in seconds is great, right up until you find out it's been cheerfully isolating the same developer's laptop three times a day because their antivirus definitions are weird.
You've built a perfect loop: fetch threats, isolate. What happens when the threat is on the CFO's machine during the quarterly earnings call? Your script doesn't care, and now you're in a meeting explaining why finance is offline.
And "saves our team a ton of manual work" is true, until it creates a different kind of manual work: untangling automated actions. You'll spend that saved time building override systems, caches, and apology emails.
The immediate speed improvement is compelling, but you'll likely see diminishing returns as your log volume grows without that pagination and time filter. Your linear, sequential POST loop also creates a scalability bottleneck - isolating ten endpoints takes ten times as long as isolating one. For bulk operations, an async approach or connection pooling is necessary to maintain that sub-minute response time under load.
The core logic also needs to account for existing isolation state, as another poster mentioned, and a method to avoid re-isolating the same threat ID. More critically, you're trading manual containment work for the operational overhead of managing false positives. Without an override mechanism for critical systems, you risk automating a major disruption.
The async bottleneck is a critical detail once you move from proof-of-concept to production. We hit it when a widespread signature update triggered hundreds of alerts. The script took over 15 minutes to iterate, which defeated the whole purpose of rapid containment.
Our solution was to push the threat IDs into an SQS queue and have a pool of Lambda functions consume and process them in parallel, respecting the API's rate limits with a token bucket. This shifted the bottleneck from our script's loop to the vendor's API ceiling, which is the real constraint.
The override mechanism you mentioned became our next layer. We ended up tagging critical assets in our CMDB and having the Lambda check a DynamoDB table of exempted resource tags before any isolation call. It adds latency, but it's a necessary trade-off for safety.