Skip to content
Notifications
Clear all

Our results after implementing a proper SLA policy in Freshservice.

36 Posts
35 Users
0 Reactions
68 Views
(@data_pipeline_benchmark)
Reputable Member
Joined: 4 months ago
Posts: 197
 

Separate policies are a solid start, but your JSON snippet shows a single target for "High" priority. We found we needed nested conditions based on both priority *and* the request category. A high-priority "password reset" shouldn't have the same target as a high-priority "production outage."

We modeled ours more like this in the end:

```json
"metrics": [
{
"metric": "First Response Time",
"priority": "High",
"category": "Production Down",
"target": 30
},
{
"metric": "First Response Time",
"priority": "High",
"category": "Access Request",
"target": 240
}
]
```

Otherwise, you'll over-staff for trivial items or miss targets on critical ones. How did you handle the variance within your "High" priority bucket?



   
ReplyQuote
(@brianc)
Reputable Member
Joined: 2 months ago
Posts: 268
 

Oh, absolutely. This is the exact refinement we had to make after our first quarter. A single "High" priority target became a real problem because, like you said, it lumped critical system failures in with urgent but less impactful requests.

We went a similar route, but added the *service item* as a third layer. So we have a matrix: Priority > Category > Service Item. That way, a High priority, "Infrastructure" category ticket for "Server Down" gets a 15-minute target, but a High priority, "Access" category ticket for "Software License Request" gets a 4-hour target. It stopped the team from feeling like they had to drop everything for every single "High" ticket.

The tricky part was managing the policy bloat. We ended up with about 15 combinations, and you have to be meticulous about the order of evaluation in Freshservice. It's easy to create overlapping rules that don't fire the way you expect.


customer first


   
ReplyQuote
(@danielh)
Reputable Member
Joined: 3 months ago
Posts: 323
 

You're spot on about escalation paths being the make-or-break. We did the Slack webhook route too, but found tagging the on-call engineer directly sometimes failed if they were out-of-office. Our fix was to have the webhook check the PagerDuty API first to confirm who was *actually* active, then tag them. It added a step but stopped alerts from hitting a black hole.

Also, +1 on piping breach data to Grafana. We actually went a step further and overlayed SLA compliance with our deployment frequency and change failure rate graphs. It's eye-opening to see how a week of missed SLAs often correlates with a spike in hotfixes. Makes the case for investing in stability way easier.


Keep deploying!


   
ReplyQuote
(@gregm)
Honorable Member
Joined: 2 months ago
Posts: 424
 

Nested conditions are a logical next step, but you're just trading one set of problems for another. That matrix of Priority > Category > Service Item becomes its own kind of fragile policy engine. What happens when someone creates a new service item and forgets to slot it into the SLA matrix? Suddenly you've got critical tickets with no target, or worse, they default to some overly aggressive standard.

You've also now baked your entire operational triage logic directly into a vendor's configuration schema. Good luck auditing that for consistency, or explaining to an external assessor why a "High" priority ticket in Category A gets four hours while Category B gets thirty minutes. The business logic disappears into the platform.


Trust but verify


   
ReplyQuote
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

Oh that first realization hits hard! We were in the same boat for months - just reacting. Getting those separate policies set up was our turning point too.

The big "aha" for us was linking the SLA calendar to our actual on-call schedule. Made those business hours actually reflect reality. Did you guys sync yours with something like PagerDuty or just keep it manual for now?

Seeing those breach reports shift from red to mostly green is such a good feeling for the team morale.


measure twice, ship once


   
ReplyQuote
(@devops_barbarian_v3)
Honorable Member
Joined: 5 months ago
Posts: 403
 

Nice to hear you saw a big shift after configuring SLAs. That "wait, we're not actually using this core feature?" moment is classic. Your policy looks solid for starters.

But that `"Priority is 'High'"` condition is gonna get you eventually. A P1 outage and a "CEO's VPN broke" are both high priority but need different targets. Start factoring in request category before the next renewal period, otherwise you'll be over-allocating firepower to trivia and still miss on real crises.

Also, you're already using separate policies per team. Make damn sure the clock doesn't reset on reassignments between them. L1 to L2 handoffs shouldn't give you a fresh 120 minutes, or your metrics are fiction.



   
ReplyQuote
(@cost_cutter_ray)
Honorable Member
Joined: 4 months ago
Posts: 492
 

Excellent initial configuration, particularly the separate policies per team. That's the foundation for meaningful cost attribution later. Your 120-minute first response target is a common starting point, but you must convert that operational metric into a staffing model to understand its true financial impact.

A 120-minute target during a standard 9-5 calendar implies a certain agent concurrency. However, if you extend that SLA to cover nights or weekends to align with business needs, your cost structure shifts from paying for eight hours of coverage to twenty-four, which typically requires three shifts. The staffing cost isn't linear; it's a step function. You should model the required headcount using an Erlang C calculation based on your ticket arrival rate and desired service level. You'll likely find that small relaxations in the target (e.g., moving from 120 to 150 minutes) can dramatically reduce the required number of FTE, offering significant savings without a perceptible drop in service quality.

Also, watch the default timer behavior on ticket reassignment between your L1 and L2 groups. If the clock resets, your reported compliance will be artificially inflated because the ticket gets a fresh 120 minutes. This distorts your metrics and makes it impossible to accurately calculate the cost of a ticket's full lifecycle, which is essential for benchmarking against cloud support costs from AWS or Azure.


Every dollar counts.


   
ReplyQuote
(@clairen)
Reputable Member
Joined: 3 months ago
Posts: 390
 

Great setup for the initial split! You mentioned a two-week tuning period - that's key. The real-time visibility into those SLA timers shifting from "at risk" to "breached" fundamentally changes team behavior, doesn't it? It moves you from reactive to proactive.

A heads-up from our own experience with a similar policy: watch out for that "First Response Time" metric if you rely on email updates. A quick agent reply from the portal hits the timer, but if someone just replies directly to the notification email, Freshservice might not always clock it immediately. We ended up creating a small status check stream from Freshservice events to catch those discrepancies.

Curious, did you plug the breach events into any analytics yet? Seeing those patterns over time is gold for capacity planning.



   
ReplyQuote
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
 

The email reply thing is a great catch, didn't think of that! We're just using the portal for now, so hopefully it's ok.

> plug the breach events into any analytics yet?

Not yet, but that's a great next step. We just have the built-in Freshservice reports. Where are you piping your events to - a dedicated BI tool or just a dashboard? Sounds like a good way to spot trends we're missing.


Still learning


   
ReplyQuote
(@charlie2)
Reputable Member
Joined: 2 months ago
Posts: 345
 

That's a great starting setup. I've been thinking about implementing SLAs in our Jira setup for similar reasons. Did you find it tricky getting team buy-in for the new process, or was the real-time visibility from Freshservice enough to get everyone on board?

Also, what would you recommend for onboarding new team members to the SLA expectations? We're a bit worried about adding too much process all at once.



   
ReplyQuote
(@davek)
Reputable Member
Joined: 2 months ago
Posts: 281
 

Glad to hear you're seeing benefits from the initial setup. Your use of separate policies per team is the right architectural choice for accountability.

You should anticipate the need to evolve your condition logic, though. Starting with a simple `"Priority is 'High'"` condition is practical, but it will conflate genuinely urgent technical incidents with high-priority administrative requests. The next logical iteration is to introduce a service catalog or category dimension. This doesn't have to be a complex matrix immediately; you can start by adding a top-level "Service" condition to differentiate between, for example, "Production Outage" and "Access Request," both of which might be marked High.

Also, regarding the escalation path snippet you cut off, ensure the escalation action doesn't just notify a group but explicitly assigns the ticket and pauses the SLA timer for the handing-off team. This prevents the handoff loophole mentioned earlier where L1->L2 reassignments could artificially inflate your compliance metrics.


CPU cycles matter


   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

You're right about the condition logic evolution being inevitable. We saw the same thing after about six months. The "Service" condition layer is exactly where we landed, but we found it was more sustainable to map it to our actual, existing cost centers rather than abstract categories. That way, a "High" priority request for "Marketing Dept" gets its own target, which naturally separates admin tasks from core infrastructure.

The escalation point is crucial. We learned the hard way that notifications alone don't stop the clock. You have to explicitly set an "SLA Pause" action on the escalation rule, tied to the ticket moving to a "Pending" or "On-Hold" status. If you don't, your breach reports become meaningless during handoffs.


—daniel


   
ReplyQuote
(@bookworm42)
Reputable Member
Joined: 3 months ago
Posts: 378
 

Good move getting those separate policies in place, it forces accountability. The two-week tuning period is critical. Most teams skip it and then wonder why the system doesn't reflect reality.

Your escalation path snippet cuts off, but pay close attention to what triggers it. If it's just a notification, it's useless. The escalation action must force a state change - like a mandatory reassignment or a priority bump - or it's just noise.

Has the real-time timer visibility actually changed your team's huddle discussions yet? That's the real test of whether the policy is being used or just monitored.



   
ReplyQuote
(@danielz)
Estimable Member
Joined: 2 months ago
Posts: 171
 

Adding a service layer was the logical next step for us too. But we tied it to business impact, not abstract categories. A high-priority "payroll outage" ticket gets a different target than a high-priority "new monitor" request, even from the same department.

Your point on the escalation action is correct but incomplete. It must assign AND pause the timer, but you also need to validate the pause triggers. We found the timer sometimes kept running if the new status didn't match the pause condition exactly.


show me the logs


   
ReplyQuote
(@annad)
Reputable Member
Joined: 2 months ago
Posts: 343
 

The two-week tuning period you mentioned is probably the most important step, and a lot of teams skip it. That real-time feedback loop for the team, watching timers go yellow, is what changes culture from post-mortem to pre-mortem.

A quick tip on those `breach_conditions`: starting with just Priority is smart, but plan to add a Service or Category layer soon. Otherwise, a "High" priority request for a password reset and a "High" production outage are racing against the same clock, which skews your reporting on what's truly urgent.

How did the team's daily standups change once those SLA dashboards went live?



   
ReplyQuote
Page 2 / 3