Skip to content
Notifications
Clear all

I think we over-engineer alert severity levels. Three is plenty.

4 Posts
4 Users
0 Reactions
23 Views
(@moderator_jane_doe)
Eminent Member
Joined: 6 months ago
Posts: 20
Topic starter   [#2469]

I've been reviewing a lot of tooling setups lately, both from vendor demos and in our internal threads, and a pattern is starting to bother me. It seems like every new alerting or on-call platform wants us to define five, seven, or even ten levels of alert severity. I'm starting to think this complexity is actively harmful to a good incident response posture.

The argument always centers on "granularity" and "precision." But in practice, what does a "Severity 3" vs. a "Severity 4" really tell an on-call engineer at 3 AM? The cognitive load of deciphering a nuanced matrix often delays the actual response. More levels also lead to endless internal debates about classification, which is a waste of cycle time.

In my view, and from what I've seen work in well-oiled teams, three levels are sufficient:
1. **Critical:** Impacts customers/production right now. Page immediately.
2. **Warning:** Needs attention but isn't currently customer-impacting. Maybe a non-paging alert, or a ticket.
3. **Info:** Log it for later review. No action required.

This forces a binary, actionable decision at alert creation: is this user-facing right now, or not? It eliminates the "grading on a curve" mentality that creeps in with more levels. Your on-call engineers have a clearer mental model, and your reporting gets simplerβ€”you can track "pages" vs. "non-paging alerts" vs. "noise."

I'm curious if others have pushed back on vendor-prescribed severity taxonomies. Have you found success with a simpler model, or does your environment genuinely require more nuance? Let's keep the discussion to practical experiences and avoid any tool-specific shilling.


Remember the rules


   
Quote
(@llm_eval_curious_42)
Estimable Member
Joined: 6 months ago
Posts: 57
 

I've been benchmarking LLM responses to structured alerts lately, and your point about cognitive load at 3 AM is key. The more granular the scale, the more inconsistent the classification becomes, even with detailed runbooks. I've seen teams waste more time reconciling severity levels post-incident than actually fixing things.

Your three-tier model maps neatly to a simple, automatable decision tree. It forces a useful binary: "wake someone up" or "don't." Where I've seen this break down is in large, siloed orgs where "customer-impacting" means different things to different teams. A database latency spike might be a "Warning" for the infra team but immediately "Critical" for payments. The debate often shifts from "what severity?" to "whose customer?"

Have you found teams need a separate, orthogonal dimension for "scope" or "affected service" to make those three levels truly universal?


Prompt engineering is engineering


   
ReplyQuote
(@new_evaluator_emma)
Eminent Member
Joined: 5 months ago
Posts: 26
 

You know, this is so refreshing to hear. As someone trying to set up a basic alerting system for our small team, I get completely lost in those vendor demos. They make it sound like you need a PhD in severity taxonomy just to get started.

I love your three-level breakdown. It feels so much more human. My one worry is the middle one - "Warning". I can already imagine us letting things pile up there because it's not "Critical". Do you have a rule for how long something can sit as a warning before it escalates, or is it purely about the type of issue?



   
ReplyQuote
(@martech_trial_taker)
Trusted Member
Joined: 5 months ago
Posts: 32
 

> This forces a binary, actionable decision

I think this is the key. We overcomplicate it. As a marketer, I see parallels with lead scoring - too many stages just means nothing moves. We ended up with "hot," "warm," and "cold." Same principle.

But for alerts, how do you handle false positives? I worry a simple system might make them seem more critical than they are.



   
ReplyQuote