Skip to content
Notifications
Clear all

How do you handle false positives without disabling the alert entirely?

34 Posts
31 Users
0 Reactions
173 Views
(@danielh)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Great point about the blast radius. That "assume off unless on" fail-safe is crucial.

We actually implemented something similar using a heartbeating pattern. The scheduled job emits the "maintenance on" metric every minute while it runs. Our alert logic then requires two consecutive missing datapoints from that stream before it assumes maintenance is over and re-enables alerts. It adds a small buffer but prevents a single metric hiccup from waking people up.

The central campaign flag analogy is spot on. Once you have that pattern, you start using it for all sorts of contextual gating, like deployment freezes or holiday schedules.


Keep deploying!


   
ReplyQuote
(@ava23)
Honorable Member
Joined: 3 months ago
Posts: 435
 

Exactly. This is why the "just route it" suggestions later in the thread feel like kicking the can. If you can define the logic for a separate suppression calendar or a composite alarm, you've already done the work to understand the context. Encoding it directly is cleaner.

Otherwise, you're just building a parallel maintenance system that will inevitably drift. Now you're debugging why an alert fired - is the CPU actually high, or did the calendar job fail? You've traded one alert for two systems to monitor.


Trust but verify.


   
ReplyQuote
(@devops_barbarian_v3)
Honorable Member
Joined: 6 months ago
Posts: 403
 

You can't add a time condition directly in CloudWatch. The composite alarm or notification routing hacks work, but they're just treating the symptom.

The real problem is alerting on a raw metric without context. CPU is just a symptom. Why not alert on the *consequence* of high CPU during that batch job? Like latency on user-facing endpoints dipping below SLA? That way you don't care about the nightly spike, only if it starts hurting customers.



   
ReplyQuote
(@chris)
Honorable Member
Joined: 3 months ago
Posts: 407
 

Completely agree with shifting from symptom to consequence. That's exactly the principle behind monitoring SLIs derived from golden signals rather than the raw signals themselves.

However, your suggestion to alert on user-facing latency hinges on the batch job and user traffic being on the same underlying resource. In a modern, decoupled architecture, that's often not the case. The batch job might run on a dedicated, isolated set of instances or a separate environment entirely. Its high CPU has zero bearing on the user-facing service's SLO.

In those scenarios, the consequence is purely internal: the batch job's own completion time or success rate. So you'd still need an alert on *that* SLO, but it moves the needle. Now you're tuning an alert threshold based on the job's own historical performance envelope, which is far more stable than raw CPU, and the "known event" context is inherently baked in.


—chris


   
ReplyQuote
Page 3 / 3