A common challenge in modern incident response is ensuring that alerts are not just fired, but intelligently delivered to the right responder at the right time. Static on-call rotations often fail to account for time zones, working hours, and complex team schedules, leading to alert fatigue and slower response times. The core of the problem is moving from a simple "person A is primary" model to a dynamic routing engine that considers time-based rules and schedule overrides.
The "best" way typically involves a layered approach, combining a dedicated incident management platform with well-defined policies. I've found the most robust solutions separate the *routing logic* from the *alert generation*. Here's a conceptual architecture:
1. **Alert Source Integration:** Tools like Prometheus Alertmanager, Datadog, or AWS CloudWatch detect issues and forward them to a central routing layer (e.g., PagerDuty, Opsgenie, VictorOps). The alert should carry rich metadata (service, severity, environment).
2. **Dynamic Routing Layer:** This is the core. The incident management tool should evaluate the alert against:
* **Scheduled On-Call Rotations:** The primary, static schedule.
* **Time-Based Override Rules:** "Between 22:00 and 06:00 local time, escalate directly to the secondary on-call engineer."
* **Team Calendars:** Integration with Google Calendar or Outlook to detect out-of-office events, public holidays, or focused work blocks, and automatically find the next available responder.
* **Escalation Policies:** A cascade of actions if the primary doesn't acknowledge (e.g., page secondary -> page entire team -> page engineering manager).
For example, a Terraform configuration for a PagerDuty escalation policy might look like this, though the time-of-day logic is often handled via the UI or API when defining rules:
```hcl
resource "pagerduty_escalation_policy" "platform_critical" {
name = "Platform Critical - Follow-the-Sun"
num_loops = 2
teams = [pagerduty_team.platform.id]
rule {
escalation_delay_in_minutes = 10
target {
type = "user_reference"
id = var.primary_oncall_user_id
}
}
rule {
escalation_delay_in_minutes = 5
target {
type = "schedule_reference"
id = pagerduty_schedule.follow_the_sun_apac.id
}
}
}
```
Key considerations for implementation include the handling of follow-the-sun rotations across global teams, the definition of "business hours" per service, and the maintenance burden of schedules. It's also critical to build in redundancy; your routing logic must have a failsafe (e.g., a final escalation to a high-volume channel like a dedicated Slack channel or SMS to a manager) to ensure no alert is ever dropped.
Finally, measure the effectiveness of your routing. Track metrics like time-to-acknowledge segmented by alert route, the percentage of alerts handled during vs. outside scheduled hours, and the frequency of manual overrides by team members. This data will show you where your routing logic is creating friction or missing its objectives.
I'm a platform engineer at a mid-sized fintech (~200 engineers) where we run a polyglot microservices stack on Kubernetes. We route about 300 daily alerts from Prometheus, Datadog, and custom services through PagerDuty in production.
Core comparison of dedicated incident platforms:
- **Pricing and Scale:** PagerDuty runs $25-50/user/month for their core schedules and routing features at our scale. For a team under 10 users, Opsgenie's free tier is shockingly capable. The real cost is in enterprise SSO and advanced analytics, which can double the quoted price.
- **Dynamic Routing Complexity:** PagerDuty's "rulesets" engine can handle multi-layer time-based and on-call overrides cleanly in code (they have a Terraform provider). Opsgenie uses "escalations" and "routing rules" which are GUI-first; changes feel riskier to version control. Both support timezone-aware schedules, but PagerDuty's handling of "follow-the-sun" across geo-distributed teams is more deterministic.
- **Integration Effort:** Connecting Alertmanager webhooks to either takes under an hour. The real effort is in tagging your alerts consistently - services without a `team` or `severity` label become unroutable. We spent a week refining alert metadata before routing worked as intended.
- **Failure Mode:** Both have occasional mobile app notification delays (2-3 minutes) during carrier issues. PagerDuty fails open to SMS, which we've had to use twice in three years. The bigger limitation is handling "within working hours only" for a subset of alerts; you'll need to model this as a separate schedule, not a simple filter, in both tools.
My pick is PagerDuty for any team over 25 users or with complex rotations across timezones. If you're a single small team or strictly in a 9-5 window, Opsgenie's free tier is the obvious start. To decide cleanly, tell us your monthly alert volume and whether your on-call rotations require legal compliance logging.
sub-100ms or bust
The layered architecture you're describing is sound in theory, but you're glossing over the integration tax. Prometheus Alertmanager -> PagerDuty is a common example, but you'll spend weeks wrestling with label mismatches and payload transforms before your "rich metadata" actually populates in the routing engine correctly.
Also, calling those schedules "static" is generous. Most teams I've seen define them in YAML, dump them into the vendor's API, and then immediately have to manage overrides in a separate Google Calendar because someone's kid got sick. The real dynamic routing often happens in the Slack channel after the "correct" person gets paged at 3 AM.
>immediately have to manage overrides in a separate Google Calendar
This is the bit that always gets me. The vendors sell you on this single pane of glass, but the moment you need a human adjustment, you're back in a shared calendar that their system can't parse. So now you're maintaining authoritative state in two places, which is a compliance headache waiting to happen.
Your point about the Slack channel being the real router is painfully accurate. I've seen teams where the official on-call gets paged, immediately posts "not me," and starts an @here chain. The integration tax isn't just setup, it's the perpetual overhead of the system being too rigid for reality.
- Nina
You've absolutely nailed the core architectural separation, and that's the exact pattern I coach my clients toward. The layered approach prevents vendor lock-in at the alert source level, which is huge for long-term flexibility.
The one thing I'd add to your "rich metadata" point is a procurement red flag. When evaluating the routing layer, you must test its ability to parse custom payload fields for routing decisions *before* you sign. I've seen two major deployments fail because the vendor's marketing said "fully customizable routing" but their product required a professional services engagement to map anything beyond their standard 5-6 fields.
Always get a proof-of-concept where you send a test alert with a custom JSON field like `"priority_tier": "tier-3"` and build a routing rule that uses it. If the sales engineer hesitates, that tells you everything about the real-world flexibility you're buying.
null