I've been reviewing a lot of tooling setups lately, both from vendor demos and in our internal threads, and a pattern is starting to bother me. It seems like every new alerting or on-call platform wants us to define five, seven, or even ten levels of alert severity. I'm starting to think this complexity is actively harmful to a good incident response posture.
The argument always centers on "granularity" and "precision." But in practice, what does a "Severity 3" vs. a "Severity 4" really tell an on-call engineer at 3 AM? The cognitive load of deciphering a nuanced matrix often delays the actual response. More levels also lead to endless internal debates about classification, which is a waste of cycle time.
In my view, and from what I've seen work in well-oiled teams, three levels are sufficient:
1. **Critical:** Impacts customers/production right now. Page immediately.
2. **Warning:** Needs attention but isn't currently customer-impacting. Maybe a non-paging alert, or a ticket.
3. **Info:** Log it for later review. No action required.
This forces a binary, actionable decision at alert creation: is this user-facing right now, or not? It eliminates the "grading on a curve" mentality that creeps in with more levels. Your on-call engineers have a clearer mental model, and your reporting gets simplerβyou can track "pages" vs. "non-paging alerts" vs. "noise."
I'm curious if others have pushed back on vendor-prescribed severity taxonomies. Have you found success with a simpler model, or does your environment genuinely require more nuance? Let's keep the discussion to practical experiences and avoid any tool-specific shilling.
Remember the rules
I've been benchmarking LLM responses to structured alerts lately, and your point about cognitive load at 3 AM is key. The more granular the scale, the more inconsistent the classification becomes, even with detailed runbooks. I've seen teams waste more time reconciling severity levels post-incident than actually fixing things.
Your three-tier model maps neatly to a simple, automatable decision tree. It forces a useful binary: "wake someone up" or "don't." Where I've seen this break down is in large, siloed orgs where "customer-impacting" means different things to different teams. A database latency spike might be a "Warning" for the infra team but immediately "Critical" for payments. The debate often shifts from "what severity?" to "whose customer?"
Have you found teams need a separate, orthogonal dimension for "scope" or "affected service" to make those three levels truly universal?
Prompt engineering is engineering
You know, this is so refreshing to hear. As someone trying to set up a basic alerting system for our small team, I get completely lost in those vendor demos. They make it sound like you need a PhD in severity taxonomy just to get started.
I love your three-level breakdown. It feels so much more human. My one worry is the middle one - "Warning". I can already imagine us letting things pile up there because it's not "Critical". Do you have a rule for how long something can sit as a warning before it escalates, or is it purely about the type of issue?
> This forces a binary, actionable decision
I think this is the key. We overcomplicate it. As a marketer, I see parallels with lead scoring - too many stages just means nothing moves. We ended up with "hot," "warm," and "cold." Same principle.
But for alerts, how do you handle false positives? I worry a simple system might make them seem more critical than they are.