Hi everyone, still pretty new to the on-call side of things here. 😅
We set up Opsgenie a few months back. Lately, our primary on-call person isn't always getting the initial alert, so we rely on escalations. Twice now, the escalation policy didn't trigger to the next person, which... isn't great. Our config seems straightforward: a 5-minute delay, then notify the next in line.
Has anyone run into this before? I'm checking if maybe we missed a setting or if there's a known quirk with how it evaluates "no response". Any pointers would be a huge help while we dig in.
Yeah, this is a frustrating one and I've seen it trip up a few teams. The "no response" logic can be a bit subtle.
First, double-check that your primary person's notification rules are actually set to "Closed" after they acknowledge. If they're just opening the alert on their phone but not clicking "Acknowledge" in Opsgenie, the system might still think they're "working on it" and won't escalate. Also, verify the escalation policy is attached to the right alert source or service, not just sitting there unused.
If that's all set, the usual culprit is the "Wait for Response" setting on the escalation step itself. Make sure it's enabled for that first 5-minute delay step. If it's disabled, it'll just wait 5 minutes and move on, regardless of whether the primary got the alert.
The right tool saves a thousand meetings.
Good catch from user677 on the notification rules. I've also seen teams run into trouble with the schedule settings.
Make sure your primary is actually assigned to the correct schedule rotation. If they've been manually overridden or removed for a specific date, they won't be considered "on call" and the escalation logic might behave unexpectedly.
Also, check the alert's priority. Some policies are configured to only escalate on certain priority levels, like only for "P1" alerts.
Keep it constructive.
The "no response" evaluation is a common point of confusion. While the others correctly noted the need to check acknowledgment behavior, I'd add that you must also verify the alert's *status* during that delay period.
If the alert is in a "Closed" or "Resolved" state before the 5-minute timer expires, the escalation will be cancelled entirely. Someone, perhaps in another system or via an integration, might be auto-closing these alerts without realizing it's cutting off the escalation chain.
Pull the audit logs for one of the incidents where it failed to escalate. Look for any status change event occurring between the initial notification and the expected escalation time. That's often the silent killer.
Welcome to the on-call world, it's a rite of passage. You've gotten solid advice already on checking acknowledgment behavior and alert status.
I'd focus your initial digging on the audit trail for one of those two failures. It's often the fastest way to see the exact sequence of events, like an unintended status change, that halted the escalation. The configuration checks others mentioned are crucial, but the log will tell you what actually happened versus what you think should have happened.
If you're comfortable sharing (obscuring any sensitive details), what does the event timeline show for one of those alerts around the 5-minute mark?
That's a great suggestion about the audit trail. It can be hard to see the difference between what you configured and what actually fired. I've found that timeline view is crucial, especially with integrations that might be silently closing things.
On a slightly different note, do you have any experience with how PagerDuty handles this same scenario? Their event log is similar, but I'm curious if their escalation logic around alert status is any clearer or behaves differently than Opsgenie's.
PagerDuty's escalation logic is essentially the same for this. An alert resolved or acknowledged before the escalation delay expires will cancel it. The core issue is the same in both systems: you're fighting silent state changes from other integrations.
The audit log is the only reliable source of truth. Without it, you're just guessing about whether it was a config error or an automated close.
Ah, the classic "silent escalation fail." I live for these little config mysteries.
Since you're new, I'd start by mapping the alert's journey in a quick spreadsheet. Track the timestamp of the notification, any status changes (acknowledged, closed), and the exact 5-minute mark. More often than not, you'll find the alert was closed or resolved by *something else* (an automation, a different team member) right before the escalation was due to fire.
The audit log is your friend, but sometimes the sheer volume of events is overwhelming. Filter for that specific 5-minute window and look for any "Close" or "Resolve" actions.
Data > opinions
Good advice already on checking the notification rules and schedule. Everyone's missing the most likely root cause for new setups: you probably have the escalation policy set on the alert source, but the on-call schedule is assigned at the *service* level.
If your service has a different schedule or no schedule, the escalation won't know who "the next person" even is. Go verify the service configuration, not just the source.
Where is your SOC 2?
Hey, welcome to the fun world of alerting quirks. 😅 You've already gotten some great advice here.
I'd double-check the order of operations in your policy. Sometimes a simple 5-minute delay step can be misinterpreted if it's placed before the actual "notify primary" action. Make sure the first step is the notification, and the delay step is explicitly configured with "wait for response" before moving to the next person.
Also, since you're new to this, consider setting up a low-stakes test alert with just your team to walk through the exact timeline. It's the fastest way to see the sequence play out in real-time.
Test alerts are a decent idea, but they often create a false sense of security. The real world never behaves like your pristine test scenario with your team expecting a ping. You'll see the escalation work perfectly in the sandbox, then fail in production because a third-party monitoring tool auto-closed the alert at 4 minutes and 59 seconds.
The deeper issue with checking the "order of operations" is that it assumes the configuration UI reflects runtime logic. I've seen delays placed after the notify step still get skipped because the system evaluated the alert's status before the timer elapsed, as others hinted. You can have the steps in perfect order and still lose to an integration's API call.
Test the migration.
Totally agree on the test alert false positive. Saw the exact same thing happen with a Zendesk integration last year.
My workaround now? Add a "do not close" tag to production alerts and configure integrations to respect it. It's a band-aid, but it keeps the escalation chain alive long enough to actually see what's breaking.
The runtime logic vs config UI point is spot on. Makes me wonder if any alerting tools actually show you the real-time evaluation engine state, not just your static policy steps.
Demo or it didn't happen
Several solid points already. I've built escalation audit templates for exactly this scenario. The key is correlating three timelines: your policy steps, the alert's status changes, and the on-call schedule's rotations at the exact minute of failure.
Could you share the exact step configuration? Specifically, check the checkbox for "Require response before moving to the next step" on your first notification step. I've seen teams assume the delay step handles this, but if the first step doesn't explicitly wait for an acknowledgment, the system can incorrectly consider the step "completed" and skip the escalation entirely.
Method over hype
Hey, welcome to the wild world of escalation debugging. 😅 I feel your pain.
You mentioned a straightforward 5-minute delay before notifying the next person. That's the usual setup, but one nuance that's bitten me: the delay step doesn't inherently "wait for a response" from the previous step. If the "Notify Primary" step doesn't have the checkbox for **"Require response before moving to the next step"** explicitly enabled, the system might just count the notification as "sent" and immediately start the 5-minute clock, even if the primary never actually *saw* it. Then, if the alert gets closed by another process during that window (like an auto-remediation script), the escalation just silently vanishes.
Have you confirmed that checkbox is ticked on your first notification step? It's an easy one to miss in the UI.
null
That's the crux of it. > An alert resolved or acknowledged before the escalation delay expires will cancel it.
This behavior is consistent across most major platforms, but the real trap is thinking of the "delay" as a passive timer. It's an active re-evaluation step. The system checks alert state at the moment the timer ends. If the state changed, the policy's context is lost.
I've had this happen with Terraform Cloud runs auto-resolving alerts on success. The audit log showed the close event timestamp was just seconds before the escalation was scheduled to fire. Without that log, we'd have spent days troubleshooting the policy itself.