After five years with PagerDuty, our team just completed a full migration to Grafana OnCall. It wasn't a decision we made lightly, but the cost-benefit for our specific needs became impossible to ignore.
For context, we're a mid-sized SaaS company running about 30 microservices. Our main pain points with PagerDuty were:
* **Cost scaling:** It got expensive fast as we added more teams and services. The per-user pricing felt punitive for a 24/7 engineering team.
* **Alert fatigue:** We lacked fine-grained control over alert grouping and snoozing logic. The same flapping alert could wake you up multiple times.
* **Runbook integration:** Everything felt siloed. Our runbooks lived in Confluence, and switching contexts during an incident added friction.
Grafana OnCall addressed these directly. The open-source core meant we could self-host for a fixed cost, which was huge. But more than that, its deep integration with our existing Grafana stack was a game-changer. Alert rules from Grafana literally *become* the on-call alerts with no extra configuration. Their "grouping" and "auto-resolve" logic feels more intelligent out of the box.
The trade-offs? Absolutely. The UX isn't as polished. Mobile experience is functional but PagerDuty's app is definitely smoother. We also had to build a few custom integrations they don't natively support yet, which required some internal dev time.
For us, the calculus came down to control and ecosystem synergy. Being able to define on-call schedules, escalation chains, and notification rules as code (yaml) alongside our other infra-as-code has been fantastic for consistency.
Has anyone else made a similar switch? I'm particularly curious how others handle the post-incident review workflow compared to PagerDuty's timeline feature.
Cheers, Carla
Benchmarking my way to better decisions