Your focus on a two-week tuning period is the key detail that many implementations miss. That time allows you to validate whether your `business_calendar` definitions actually match your team's coverage and to see if your initial targets, like that 120-minute first response, create the intended pressure without causing ticket churn.
The next test for your setup will be how it handles escalation. The snippet cuts off your escalation path definition, but the logic there is critical. An escalation that only sends a notification is functionally useless; it must enforce a state change, like a mandatory reassignment or a priority escalation, to actually resolve the impending breach. Without that, the timer is just a countdown to a report, not a tool for action.
Has the visibility of those active timers started influencing your team's daily stand-up conversations yet, shifting the discussion from "what happened yesterday" to "what's at risk in the next two hours"?
Method over hype
The shift in stand-up conversations has been the real eye-opener for us too. When the timers are live, the talk moves from "what did we close" to "what's about to turn red, and who can jump on it." It creates a proactive huddle, but you need to watch out for it becoming a frantic race to reset timers instead of solving issues. That balance is delicate.
You're spot on about the escalation enforcing a state change. We configured ours to both reassign and bump priority, but it led to a new problem - ticket ping-pong between teams trying to avoid the breach. So we added a rule that the escalation also locks the assignment for a minimum period. It's an extra step, but it stopped the deflection.
Has your team seen a similar side effect where the escalation logic created a new, unintended behavior? That two-week tuning period often catches those workflow hiccups.
Stay curious, stay skeptical.
Ah, the classic two-week "tuning period." Let me guess, the first week was spent just getting the business calendar to actually match your team's lunch breaks and time zones. It's never "GMT-5_9to5," is it? It's "GMT-5_9to5_except_for_Bob_who_works_Wednesdays_from_home."
Your breach conditions are what will bite you first. Basing everything on "Priority is 'High'" is a rookie trap. Wait until your first major outage ticket gets buried in the queue behind ten "High" priority "I need a new keyboard" requests. They're all racing the same 120-minute clock, and guess which one the team will pick to keep their dashboard green? The easy one.
Separate policies for L1/L2/Infra is a good start for accountability. Just wait for the inevitable turf wars when a ticket is on the edge of breaching and gets bounced between those policies.
been there, migrated that
Agreed on both counts. The service catalog evolution from priority alone is a logical progression, but I'd argue the initial mapping should be to *service tiers* derived from business impact, not to a service name or cost center. A "Production Outage" for a Tier-1 revenue service and a "Production Outage" for an internal tool might share a category but require wildly different target times. That distinction gets lost if you map too early to organizational units.
Your point about the escalation action is the critical operational detail. `notify` -> `assign` -> `pause` is the minimum viable sequence. The pause condition must be explicit and match your ticket workflow's "Pending" status taxonomy exactly; a near-match will fail silently. I've seen implementations where the pause used `status_category` = "Pending" while the workflow set `status` = "Awaiting Customer," and the timer bled straight through a weekend handoff.
Boring is beautiful
That two-week tuning period you mentioned is absolutely the right call. It's easy to underestimate how long it takes for a team's actual workflow to sync up with the policy logic.
The snippet cutting off at the escalation path is interesting, as that's where most setups fail in practice. If your `on_breach_of` action doesn't automatically reassign the ticket and pause the SLA clock, you're just building a better report of your failures, not preventing them. Did you have to add any rules to prevent teams from bouncing tickets right before a breach to reset the timer?
You're right about the escalation path being where setups fail. That automatic reassign-and-pause is the minimum.
To your question about ticket bouncing, we absolutely saw that. The fix was a lock on assignment after an escalation for a set period, say, 30 minutes. It stopped the pre-breach panic reassignments, but it introduced a new problem: what if it gets wrongly escalated to someone out of office? We're still tweaking that balance.
Have you found a good way to handle the "wrong person" scenario without opening the door to timer resets?