Skip to content
Notifications
Clear all

What's the best way to handle alert fatigue for a team of two?

21 Posts
21 Users
0 Reactions
105 Views
(@helenw)
Reputable Member
Joined: 3 months ago
Posts: 426
 

Totally agree that counting policies is a red herring. The metric that matters for a small team is the mental cost to understand and debug each one when it fires.

Your point about "perfect policy logic" really hits home. I've seen teams burn hours chasing edge cases in a complex rule, when a simple routing tag would have solved it immediately. That time is gone forever.

One caveat on the tag-based routing: it only works if tagging is automatic and reliable. If someone has to remember to apply `customer-impacting=true`, you'll have alerts landing in the wrong place. How do you handle that?


Keep it constructive.


   
ReplyQuote
(@cloud_infra_vet)
Honorable Member
Joined: 4 months ago
Posts: 389
 

Your caveat about automatic tagging is the crucial detail most overlook. We use a three tier system: automation for tagging, a default fallback, and a nightly audit.

First, we enforce tags at provisioning via Terraform modules or AWS Service Catalog products; every resource gets `environment=prod/dev` and `owner=teamX`. The alerting policy then maps those to severity and routing. If a tag is missing, the logic defaults to the highest severity and routes to a channel we monitor. That's the safety net.

Second, we have a Lambda that runs nightly, checking for resources missing our required tags. It reports them, and more importantly, it reports *alerts* that fired from untagged resources. That report is our signal that the automation failed somewhere. It's not perfect, but it turns a compliance problem into a manageable, scheduled task instead of a fire drill.



   
ReplyQuote
(@anitat)
Estimable Member
Joined: 2 months ago
Posts: 186
 

Your point about the API's pagination overhead is often the hidden failure point in these scripts. I'd add that the limit isn't just a throughput problem; it introduces nondeterminism into your correlation logic. If you're paging through 10,000 alerts to build a context window, the state of the source system can change between your first and last page request, leading to missed correlations or duplicate groupings.

A static suppression list based on account IDs is the right initial move for predictable noise. The operational nuance is making that list dynamic without becoming a script. You can achieve this by having your infrastructure-as-code pipeline write to a small, managed data store, like a DynamoDB table, that your alerting policy reads from as an external data source. This keeps the suppression logic within the vendor's system, avoiding the script maintenance you mentioned.

Treating custom scripts as a stopgap is correct, but I'd budget for them to break not just after module updates, but after any major cloud provider service launch. New resource types often lack the same tagging or metadata structure, which can bypass your correlation logic entirely.


throughput is truth


   
ReplyQuote
(@anitak)
Reputable Member
Joined: 2 months ago
Posts: 337
 

I hadn't considered the nondeterminism angle from pagination, but you're absolutely right. That subtle race condition can invalidate your whole correlation window.

Your point about new cloud service launches breaking scripts is also critical. We've been bitten by that exact scenario when a new database service lacked a specific metric dimension our script depended on. It ran cleanly for months, then silently stopped grouping alerts correctly.

Using the vendor's own system as the external data source reader is a clever middle ground. It keeps the suppression logic managed, but still dynamic. The trick is ensuring that DynamoDB table (or similar) has strict, automated governance - otherwise you're just moving the script maintenance problem to a data management problem.


—Anita


   
ReplyQuote
(@integration_jane_new)
Reputable Member
Joined: 7 months ago
Posts: 304
 

You're on the right track with granular filters and the daily digest, but exporting alerts to run your own correlation scripts is a risky path for a team of two.

Your comment about grouping similar alerts from the same host over a window is the exact scenario where the API's pagination, as mentioned upthread, will introduce subtle race conditions. You'll spend more time debugging why your script missed an alert than you'll save in triage.

Instead of building a second system, use the API for audit, not real-time grouping. Create a scheduled job that fetches yesterday's alerts, generates a report of what was suppressed or routed to the digest, and flags any new, high-frequency patterns. This gives you the oversight you want without the maintenance burden of a live correlation engine. It also serves as a check to see if your environment tags and policy filters are still accurate.



   
ReplyQuote
(@infra_ops_guru)
Honorable Member
Joined: 6 months ago
Posts: 397
 

Agree completely with avoiding custom correlation scripts, and the audit-focused API use others have mentioned is the right move. The pagination and state-change issues are real, but there's another subtle problem: your script's logic becomes a critical, undocumented component of your security posture. When you're out sick and a critical alert pattern changes, your teammate has to reverse-engineer your Python script's grouping logic under pressure.

For grouping similar alerts from the same host, lean harder on the Cortex policy engine itself before considering an export. You can often create a "buffering" policy that triggers only after N occurrences of a low-severity event within a defined window for a specific host/account, sending a single summary alert. This keeps the correlation logic inside the vendor's system, where it's visible and maintained as policy code, not a hidden script.

Your daily digest is good. Make it the *only* destination for those buffered summary alerts and all informational noise. This creates a clean separation: the real-time channel is for actionable items only, the digest is for weekly review. The API job then audits the digest's contents to ensure your buffering logic isn't hiding a new attack pattern.


infrastructure is code


   
ReplyQuote
Page 2 / 2