Hey everyone, new to the forum and really learning a lot from these rebuild stories. 😅 Thought I’d share our recent experience since it was my first time being part of a big tool switch.
We’re a small but growing dev team, and we were running both PagerDuty and OpsGenie in a kind of messy, overlapping way after a merger. The forcing function was pretty simple: costs were getting out of hand for our scale, and the alert noise was overwhelming. We decided to consolidate on a newer platform called Claw Alerting.
The sequencing was the hardest part to figure out. We didn’t want a risky big-bang cutover. Here’s what we did over roughly 12 weeks:
- **Weeks 1-3:** Set up Claw for all our non-critical, informational alerts only. This let us test the routing and get comfortable with the UI without waking anyone up at 3 AM.
- **Weeks 4-8:** Started moving over our medium-priority alerts team by team. We kept PagerDuty/OpsGenie active as a backup for those alerts for two full weeks before turning the old ones off. This built a lot of confidence.
- **Weeks 9-12:** Tackled the critical, on-call alerts. This is where things slipped a bit. We underestimated how many unique escalation policies and schedules we had tucked away in the old systems. We had to do a lot of manual reconciliation, which added about a week of delay.
The main takeaway for me was that even with a phased plan, you’ll discover hidden “tribal knowledge” in the old setup. For us, it was those one-off schedules for specific services. Overall, the phased approach saved us from any major outages, and the cost savings are already noticeable. Really curious if others have tackled similar alerting tool consolidations and how you handled the policy mapping.
We did the same thing moving off OpsGenie a while back. The escalation policy mapping is always the killer.
Did you store your policy logic in code, or was it all still in the old UIs? We found using Terraform to define the policies first made that last phase way smoother. You can validate everything before flipping the switch.
—cp
Terraform for escalation policies is the only sane way. We did the same, but found you need to pair it with automated validation against your on-call schedule API. Otherwise you miss mismatches that only show up when someone's on vacation.
We still had to manually review the rendered policy digraph in Claw's UI. The config-as-code output doesn't always match human intuition for rotations.
Exactly. That manual review step is unavoidable when you're dealing with complex rotations because the policy's logical structure can be valid while still being semantically wrong. The digraph visualization often reveals circular dependencies or unintended single points of failure that the raw config misses.
Our validation script pulled from the schedule API *and* the HR system's holiday calendar, then ran a Monte Carlo simulation across the planned migration period to flag any gaps in coverage. Even then, we caught a critical flaw in the rendered digraph: a tertiary escalation path was visually depicted as a direct line, but in practice it required a permission the target role didn't have.
The real lesson is that config-as-code gives you reproducibility, but not necessarily correctness. You still need a human to look at the runtime representation and ask, "Would I trust this to wake me up at 3 a.m.?"
Single source of truth is a myth.