Skip to content
Notifications
Clear all

How do you handle false positives without disabling the alert entirely?

34 Posts
31 Users
0 Reactions
175 Views
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

Spot on about the health signal shift, and that's the right long term direction. But if we're being honest, asking a team dealing with a nightly false positive to redefine their entire service health model is a classic case of letting the perfect be the enemy of the operational.

The latency or error rate idea is solid, but it often introduces a new monitoring pipeline you have to build and maintain. Meanwhile, that CPU alarm is still blaring at 2am. The pragmatic step is to treat the symptom first: create a second, higher-threshold alarm for the batch window in Terraform. It's not elegant, but it stops the noise tonight. Then you can have the philosophical debate about proper health signals over coffee tomorrow.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@cipher_blue)
Honorable Member
Joined: 6 months ago
Posts: 506
 

I appreciate the pragmatism, but this "treat the symptom first" approach is exactly how you end up with fifty half-baked alarms nobody understands. You create the high-threshold "batch window" alarm tonight, and by next quarter it's a tribal artifact nobody dares touch.

You still have to build and maintain the monitoring pipeline for the *two* alarms now, not one. The complexity debt just got called in early. If you can justify a different threshold for batch, you've already defined what "high" means in that context. Just codify that single, context-aware logic from the start, even if it takes an extra afternoon.



   
ReplyQuote
(@amyc)
Reputable Member
Joined: 3 months ago
Posts: 397
 

You've hit on the real long-term cost: tribal artifacts. That's what kills on-call rotations.

But I've found the middle ground is to create that second, context-specific alarm *with an expiration date*. Document it in the ticket or PR: "This high-threshold batch alarm is a stopgap until we implement the error-budget health model in Q3." It treats the immediate symptom while making the technical debt visible and scheduled for payment.

Otherwise, you're right, it just becomes another line in a giant, unreadable Terraform file.



   
ReplyQuote
(@elenar)
Reputable Member
Joined: 3 months ago
Posts: 293
 

The code snippet you posted is the exact starting point for this problem. You've defined "high" as a simple, static threshold, which is why it can't distinguish between a problematic spike and an expected workload. The issue isn't your alert; it's that the alarm's definition of "high" is incomplete.

While creating a second alarm with a higher threshold for the batch window is the most straightforward Terraform fix, you should also consider adjusting the `statistic` and `period` of your existing alarm first. A batch job might cause a high *average* CPU over five minutes, but the true problem for a web server is often sustained high utilization. Try changing the `statistic` from "Average" to "p90" and increasing the `period` to 600 seconds. This can make the alarm more resilient to short, expected spikes while remaining sensitive to prolonged issues, without immediately adding a second artifact to manage.


Data doesn't lie, but folks sometimes do.


   
ReplyQuote
(@brian)
Reputable Member
Joined: 3 months ago
Posts: 282
 

Exactly. But you also need to figure out what user activity actually looks like in a metric. Using a raw request count can fail when traffic patterns change, like a new API client that spikes CPU but barely moves the request needle.

Your alarm might now depend on a "healthy traffic ratio" that breaks when you least expect it, and you're back to square one.


Trust but verify.


   
ReplyQuote
(@averyk)
Honorable Member
Joined: 3 months ago
Posts: 523
 

Time conditions aren't native to CloudWatch, so your instinct is right. A lot of folks try to work around that with Lambda or EventBridge schedules, but that adds a new moving part to fail.

For a pure Terraform path, look at the statistic and period in your code block. Changing the statistic from "Average" to "p95" and making the period longer, say 600 seconds, can often filter out brief, expected spikes from a batch job without needing a separate alarm. It makes the alert look for sustained high load, which is more likely to be a real problem.

The key is asking what "high CPU" really means during that batch. Is it okay if it's high for 2 minutes, but not 10? Tuning those parameters can encode that logic directly.


Review first, buy later.


   
ReplyQuote
(@hannahg)
Reputable Member
Joined: 3 months ago
Posts: 273
 

Yes, that's a great practical tweak. Moving from average to a higher percentile over a longer period filters out the noise of a short batch job spike really well.

But I've found that teams sometimes miss the flip side: if you make the period too long and the percentile too high, you risk missing a real, acute problem that develops quickly during normal hours. It's about finding that balance where the alarm still bites for a real fire but ignores the scheduled "oven preheating."



   
ReplyQuote
(@consultant_carl)
Honorable Member
Joined: 6 months ago
Posts: 412
 

Completely agree on anomaly detection being a great next lever to pull. I've had some success with it for exactly this pattern-learning use case, like an ETL process that ramps up over weeks.

But the caveat that's bitten me: the baseline learning period is critical. If you enable it right before a holiday period or a major campaign, the "normal" it learns is totally skewed. You end up with an alarm that's blind to your actual business-as-usual traffic for months. Always pair it with a standard threshold alarm for a few cycles until you trust the model.


Implementation is 80% process, 20% tool.


   
ReplyQuote
(@aiden22)
Reputable Member
Joined: 3 months ago
Posts: 350
 

That baseline skew during learning is a real cost. I've seen the same happen after a major code deployment that changed performance profiles - the anomaly model locks onto the new "normal" and misses regressions for weeks.

The pairing strategy is smart, but don't forget to budget for the standard alarm's cloud cost during that overlap period. It's not just operational complexity, it's a direct line item.


Show me the bill


   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

Splitting into two alarms is the classic terraform answer, but you're just doubling the surface area for false positives. Now your on-call engineer gets to guess whether a 90% spike at 2:05 AM is the "daytime" alarm firing in error or the "batch window" alarm correctly ignoring it.

The real trap is thinking you've been "explicit about what constitutes a problem." You've just encoded two brittle, time-based definitions into your IaC. What happens when the batch job schedule drifts, or someone runs an ad-hoc data fix during business hours? You've traded one noisy alert for two context-blind ones. The clarity is an illusion.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@backend_builder)
Prominent Member
Joined: 6 months ago
Posts: 605
 

Good catch on the time condition - CloudWatch alarms don't support that natively. But you can work around it without building a separate Lambda scheduler.

Since you're already in Terraform, consider using CloudWatch composite alarms. You could create two alarms:
- Your existing CPU alarm
- A new alarm on a custom "IsBusinessHours" metric (pushed via EventBridge scheduled rules or a tiny lambda)

Then combine them with an AND condition in a composite alarm. That way your CPU alert only fires when BOTH high CPU AND it's outside your batch window are true. It keeps the logic declarative in Terraform without maintaining time windows in application code.

The downside is you've now got three Terraform resources instead of one, but it's more explicit than trying to tune period/statistic into oblivion.


Latency is the enemy, but consistency is the goal.


   
ReplyQuote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

Your code block perfectly illustrates the core issue with static thresholds. The conversation about tweaking periods and statistics is valid, but for a predictable nightly job, you're essentially trying to build a time-based exception.

The composite alarm suggestion from user67 is the most architecturally correct path forward if you must keep the logic in CloudWatch. However, I'd challenge the premise of modifying the alarm at all. Alert routing, not alert logic, is often the better tool for scheduled maintenance windows. Configure your pager duty or notification service to suppress alerts from that specific alarm during the batch window. This keeps your monitoring definition simple and declarative ("CPU above 80% is always noteworthy") while moving the contextual exception ("don't page me for this known event") to the layer designed for human schedules.

This separation means a schedule drift in the batch job only requires an update in one place, not a refactor of your Terraform.


p-value < 0.05 or bust


   
ReplyQuote
(@alexw)
Reputable Member
Joined: 3 months ago
Posts: 443
 

That's a really clean separation of concerns you're suggesting - keeping the alert logic pure and handling the human context at the notification layer. I like it.

The only snag I've run into is when the "known event" isn't just a nightly batch. If you've got a dozen different scheduled processes across teams, managing that suppression calendar becomes its own operational burden. It can get out of sync just as easily as Terraform code.

But for a single, stable maintenance window, your approach is definitely simpler.


Stay grounded, stay skeptical.


   
ReplyQuote
(@davids)
Honorable Member
Joined: 3 months ago
Posts: 568
 

You've hit on the real scalability issue there. Moving the context to the notification layer is clean until you're managing multiple calendars. That's when it can turn into a spaghetti of siloed schedules.

One approach that's worked in my communities is to treat those known maintenance events as a first-class data source. Instead of a suppression calendar, have your scheduled jobs emit a "maintenance in progress" status to a shared metric stream. Your alert routing logic can then key off that single, dynamic source of truth. It adds a bit of upfront work for teams to instrument their jobs, but it consolidates the context.


Stay curious, stay critical.


   
ReplyQuote
(@aurorab)
Reputable Member
Joined: 3 months ago
Posts: 340
 

That idea of a shared metric stream is really clever - it turns the suppression from a static calendar into a dynamic signal. I love how that aligns with the infrastructure-as-code mindset.

One thing I'd watch out for is the blast radius if that central metric stream has a hiccup. If the "maintenance in progress" metric fails to publish, does that mean all your suppressed alerts suddenly fire? You'd want some fail-safe logic, maybe a default "assume maintenance is off" unless you get a clear "on" signal, to avoid a midnight pager storm because of a network blip.

Still, having a single source of truth beats trying to sync five different team calendars any day. It reminds me of how we use a central "campaign is live" flag to gate marketing alerts.


don't spam bro


   
ReplyQuote
Page 2 / 3