Hey everyone! I've been deep in the weeds evaluating on-call tools for our team's 2025 stack consolidation, and Grafana OnCall has been on our shortlist for a while—especially with its tight integration into the Grafana ecosystem we already love. But this new pricing model they rolled out has me doing a double-take, and I'd really love to compare notes with anyone else navigating this! 😅
For context, we're a mid-sized SaaS product team running about ~15 services. We currently use a patchwork of PagerDuty for alerts, a wiki for runbooks, and Slack threads for incident coordination. The dream is to reduce context-switching and finally get some solid post-incident workflow metrics. Grafana OnCall seemed promising because it could live right next to our metrics and dashboards.
Here’s what’s giving me pause about the new pricing:
* **The shift from per-user to "per-contact point"** is the big one. We have engineers who need to be reachable via phone, SMS, *and* Slack during an incident. Does that now count as three "contact points" for one person? Our initial math shows our costs potentially doubling if that's the case, which is a tough sell internally.
* **The included "Essential" features vs. "Advanced" tiers.** I'm all for transparent pricing, but I'm struggling to map this to our actual incident response maturity. For example, "Advanced Analytics" is a paid add-on—but without good analytics, how do we even measure if the tool is reducing our pager fatigue or improving MTTR? It feels a bit like buying a car and paying extra for the speedometer.
* **Where do runbook integrations land?** A huge part of our desired value is automating the first steps of a response and making post-mortems less painful. Is that considered an essential workflow or an advanced one?
**I'd love to hear from teams already using it under the new model:**
* Has the pricing structure changed how you configure on-call schedules or notification rules? Are you now incentivized to limit how many ways a person can be contacted?
* For those on the "Advanced" tier, is the analytics suite robust enough for real post-incident review and workflow quality tracking?
* More broadly, does this model feel aligned with the value you get, or does it feel like it's gatekeeping core incident response best practices?
We're trying to build a more resilient, less stressful on-call culture here, and the tooling cost/benefit is a huge part of that conversation. Sharing your experiences—good or bad—would be incredibly helpful!
keep building,
katiec
keep building