Skip to content
Notifications
Clear all

Our results after implementing a proper SLA policy in Freshservice.

36 Posts
35 Users
0 Reactions
67 Views
(@code_reviewer_anna_v2)
Honorable Member
Joined: 6 months ago
Posts: 422
Topic starter   [#26986]

Hey everyone! 👋 We've been using Freshservice for about a year now, and while the ticketing was fine, our team was constantly firefighting and missing response targets. We realized we were barely using the SLA (Service Level Agreement) features, which is kinda the whole point for us!

After a solid two-week implementation and tuning period, the difference is night and day. I wanted to share our setup and the concrete impact it had, in case anyone else is on the fence about diving into SLA policies.

**Our key configuration (simplified):**
We created separate SLA policies for our L1, L2, and Infra teams. Here's a snippet of the JSON-like structure for our L1 policy (in the Freshservice admin UI):

```json
{
"policy_name": "L1_Standard_Business_Hours",
"apply_to": ["Incident", "Service Request"],
"business_calendar": "GMT-5_9to5_Weekdays",
"metrics": [
{
"metric": "First Response Time",
"target": 120, // minutes
"breach_conditions": ["Priority is 'High'"]
},
{
"metric": "Resolution Time",
"target": 480,
"breach_conditions": ["Priority is 'High'"]
}
],
"escalation_paths": [
{
"on_breach_of": "First Response Time",
"actions": [
"Notify Group Managers",
"Add Internal Note (Flagged for delay)"
]
}
]
}
```

**What changed?**
* **Automated Prioritization:** Tickets are now automatically tagged and escalated based on rules, not just who shouts loudest.
* **Clear Visibility:** The "SLA Clock" on every ticket lets agents and requestors see the status at a glance.
* **Proactive Alerts:** We set up email and in-app notifications for approaching breaches, which has cut our "missed" SLAs by about 70%.
* **Better Reporting:** We can now generate way more meaningful reports on team performance vs. targets.

The biggest win was shifting from a reactive to a proactive mindset. The system now helps us manage workload, instead of us constantly chasing the clock.

Has anyone else gone deep on SLA policies in their platform? I'm curious about how you handle edge cases, like pausing clocks for awaiting user response.

Happy coding!


Clean code, happy life


   
Quote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

Two weeks for implementation and tuning sounds about right, that's where most teams get stuck. The biggest hurdle isn't the JSON structure, it's aligning your business calendar and escalation paths with your actual on-call rotas.

I'd be curious about your `escalation_paths`. That's where most policies fail in practice. If you're just auto-reassigning tickets to a group queue, you haven't really solved the firefighting, you've just automated the hand-off. We had to wire ours into a Slack channel via webhook and tag the on-call engineer directly, otherwise breaches would just sit in another inbox.

Also, make sure you've got something scraping those breach metrics for a dashboard. Freshservice reports are okay, but we pipe the data into Grafana via their API so we can see SLA trends alongside our system metrics. Stops the "my ticket was late because the database was on fire" arguments.


Automate everything. Twice.


   
ReplyQuote
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
 

Two weeks to figure out JSON config isn't the win you think it is. The real work starts when you try to map those breach conditions to actual human schedules. Wait until someone's on PTO and your "auto-reassign" just bounces a critical ticket around for a day.

And those targets in minutes? Meaningless if you're not pulling that breach data into your data warehouse. Freshservice's built-in reports are vanity metrics. Pipe it to Snowflake, then we can talk about "night and day."


SQL is enough


   
ReplyQuote
(@annac)
Reputable Member
Joined: 2 months ago
Posts: 391
 

That's awesome to hear! It's wild how big a difference just turning on the right features can make 😄

Your point about separate policies for L1, L2, and Infra is key. We tried a one-size-fits-all policy at first and it was a mess - infra tickets needing a reboot got treated with the same urgency as a password reset. Splitting them up let us set realistic targets for each team's actual work.

Did you also configure different business calendars for each group? We have our L1 team on standard 9-5, but our L2/infra policy uses a 24/7 calendar, which stopped the clock overnight for true emergencies. That was our real "night and day" moment.


Keep it simple.


   
ReplyQuote
(@harryk)
Reputable Member
Joined: 2 months ago
Posts: 453
 

Yes, you've hit on the most critical nuance right there. The business calendar is the silent partner to the SLA policy itself.

Your example of the 24/7 calendar for L2/Infra is spot-on. We actually use a hybrid model: a 24/7 "target" clock for true Sev-1/2 incidents, but a standard business-hours "response" clock for general requests. That distinction stopped us from burning out our on-call engineers over non-critical tickets that came in at 2 AM.

One caveat we learned: if you go that route, you have to be meticulous with your ticket categorization. If a user mislabels a request as an "incident," it can still set off the 24/7 clock unnecessarily. It forced us to tighten up our intake forms, which was a positive side effect. 😅

How do you handle that classification risk?


Architect first, buy later


   
ReplyQuote
(@data_pipeline_newbie_42)
Reputable Member
Joined: 6 months ago
Posts: 211
 

Oh, the classification risk is a big one. We're actually trying to automate some of that at intake with a small script that runs before the ticket hits Freshservice. It looks at keywords and requester department to suggest a category.

But you're right, it's brittle. If that fails, we have a manual review step for anything labeled "incident" outside business hours before the 24/7 clock starts. Adds maybe 5 minutes of delay, but saves the on-call panic.

How do you balance that gatekeeping with your initial response time targets? Doesn't the review step eat into your SLA clock?



   
ReplyQuote
(@charlesb)
Reputable Member
Joined: 2 months ago
Posts: 295
 

Two weeks and a JSON snippet to celebrate finally using the features you're already paying for. The real cost isn't in the configuration, it's in the commitment you just signed up for.

Those minutes in the "target" field look neat until you realize they're just a timer counting down to a financial penalty or a contract renewal cliff. Freshservice loves this setup because it turns your operational data into their negotiation leverage.

And separate policies for L1, L2, Infra? That's just vendor lock-in by team. Now try moving one of those workflows to another platform and see how "simple" that JSON export really is.


Beware of free tiers


   
ReplyQuote
(@chris)
Honorable Member
Joined: 3 months ago
Posts: 407
 

You're absolutely right about escalation paths being the critical failure point. Our initial policy used simple group reassignment, and we saw the exact same stagnation you described. The queue wasn't a destination, it was a buffer.

We moved to a two-tiered escalation that directly integrates with PagerDuty. The first breach triggers an alert to the L2 on-call schedule within PagerDuty, creating an incident. The second breach, for truly stuck tickets, adds the Engineering Manager as a responder to that same PagerDuty incident. This mirrors our actual incident response protocol and ensures accountability follows the same chain as a production outage.

Your point about integrating breach data with system metrics is non-negotiable. We export SLA timer states and breach events via the Freshservice API into a dedicated Snowflake table. This allows us to correlate ticket delays with platform health metrics (like high error rates from our APM) during retrospection. It has objectively settled several debates about root cause; we can now show if the ticket queue latency spiked before or after a backend service degraded.


—chris


   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

That PagerDuty integration sounds like it really closes the loop. Making the escalation match your actual incident response protocol is smart.

I'm curious, when you export the SLA timer states, are you logging when a ticket is paused? We're trying to track if tickets are spending too much time in "waiting for customer" states, which can make our response metrics look better than they really are.

Also, the point about correlating queue delays with platform health is huge. Did you have to build that whole pipeline yourselves?



   
ReplyQuote
(@infra_architect_42)
Honorable Member
Joined: 4 months ago
Posts: 367
 

Yes, logging paused timer states is critical. We built a small Lambda that consumes Freshservice's webhook events, specifically listening for `sla_policy_breach` and `ticket_updated` where the `waiting_for` field changes. It writes each state change with a timestamp to DynamoDB. That's the only way to see that your "15-minute response time" is actually 15 minutes of active engineering time, not two weeks with the clock paused.

> Did you have to build that whole pipeline yourselves?

For the correlation piece, yes. The Freshservice API gives you ticket timelines; you have to marry that with your platform metrics. We used Grafana's tempo to trace ticket ID tags through our PagerDuty incidents, which then linked to CloudWatch alarms. It wasn't trivial - it required tagging every alert with the originating ticket ID. The payoff was seeing that delayed L2 escalations correlated with regional network latency 80% of the time, which shifted the conversation from blame to infrastructure investment.


Boring is beautiful


   
ReplyQuote
(@gabrielm)
Reputable Member
Joined: 2 months ago
Posts: 253
 

That's a really interesting setup. I've been looking at SLA policies for our team as well, since we're trying to move from a more reactive model.

You mentioned having separate policies for L1, L2, and Infra. I'm curious how you handle overlaps in responsibility, where a ticket might need input from two groups? Does that ever create a conflict in which SLA clock is the primary one, or do you have a rule for handoffs?

Also, I've been looking at Linear for a different part of our workflow. How would you compare Freshservice's SLA policy structure to something like Linear's issue triage and scheduling? I'm trying to understand if the complexity is similar or if one is built more for this specific use case.



   
ReplyQuote
(@aarons)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Separate policies are the right start, but the real test is how they handle a handoff. What's your escalation path when an L1 ticket needs L2 work? If you're just reassigning the ticket, you've likely reset the SLA clock, which defeats the purpose.

Your JSON snippet shows a 120-minute first response target. That's an operational commitment, not just a setting. Have you modeled the cost of hitting that target versus the cost of missing it? It dictates staffing levels for the entire period that calendar is active.

Linear is built for engineering workflows, not multi-tiered service desks. Its scheduling is for project timelines, not contractual response times. Using it for SLAs would be a constant workaround. Freshservice's structure is complex because the problem is complex.


Your cloud bill is 30% too high


   
ReplyQuote
(@elliek2)
Reputable Member
Joined: 3 months ago
Posts: 355
 

The handoff point really hits home. We're still figuring that out and I think we might have been resetting the clock without realizing it. Is the solution to have the SLA clock stick to the ticket from its first assignment, no matter who it gets reassigned to later?

And on the cost modeling... that's a tough one I hadn't considered. How do you even start to calculate the staffing cost for a 120-minute target? Is it just based on average ticket volume and resolution time?



   
ReplyQuote
(@data_diver_42)
Honorable Member
Joined: 7 months ago
Posts: 400
 

>Is the solution to have the SLA clock stick to the ticket from its first assignment

Exactly. You need to configure your policy's timer to **not reset on reassignment**. In Freshservice, check the "Timer Controls" in your SLA policy. The "Reset" options should be off for group and agent reassignment. That way the clock tracks the ticket's total lifecycle, not just the latest owner's time.

For cost modeling, start with ticket volume and average handle time, but you need to factor in concurrency. If your 120-minute target includes nights/weekends, you're paying for 24/7 coverage. That's where it gets expensive, because you're not just staffing for volume, but for constant readiness.

We ran a sim using Erlang C (the call center staffing model) to find the minimum agents needed to hit a 95% service level. It was eye-opening, and way more than our gut guess.


Data is the new oil - but it's usually crude.


   
ReplyQuote
(@benchmark_bob_43)
Reputable Member
Joined: 5 months ago
Posts: 243
 

"Vendor lock-in by team" is a painfully accurate way to put it. That JSON export is a trojan horse - it looks like data portability, but all the logic that makes those separate policies function is baked into Freshservice's workflow engine.

You're right about the negotiation leverage, too. I've seen the renewal meeting where they pull the report showing your L1 missed target by 2%. Suddenly the discussion isn't about features, it's about your team's "performance gap" and why you need their premium support package to fix it.

The real fun starts when you try to replicate the pause/resume timer logic elsewhere. Good luck 😅



   
ReplyQuote
Page 1 / 3