Over the last quarter, our platform team observed a subjective increase in reported "pager fatigue" and anecdotal concerns about escalating mean time to acknowledge (MTTA) for severity-1 incidents. To move beyond anecdotal evidence, I undertook an analysis to quantify our response performance by building a dedicated internal dashboard that aggregates and visualizes on-call response metrics. The primary goal was to correlate alert volume, time-of-day, and team composition with our MTTA and mean time to resolve (MTTR), thereby identifying systemic bottlenecks rather than isolated events.
The dashboard ingests data from three primary sources:
* **PagerDuty API** for incident timelines, including acknowledged_at and resolved_at timestamps.
* **Our internal team roster service** (a simple REST API) to map engineer IDs to their respective team and on-call schedule.
* **A custom event log** where we tag incidents with the affected service and hypothesized root cause category (e.g., "database-latency", "third-party-api-failure").
The transformation layer, built as a set of scheduled AWS Lambda functions, performs the following key operations:
1. Enriches raw PagerDuty incidents with team and service context.
2. Calculates MTTA (acknowledged_at - created_at) and MTTR (resolved_at - created_at) for each incident, filtering out low-severity alerts (severity 3 and 4).
3. Aggregates the data into daily and weekly summaries, stored in Amazon DynamoDB for the dashboard backend to query.
A simplified version of the core aggregation function is as follows:
```python
def calculate_daily_metrics(incidents):
daily_summary = {}
for inc in incidents:
day = inc['created_at'].date()
if day not in daily_summary:
daily_summary[day] = {'total_incidents': 0, 'total_mtta_seconds': 0, 'total_mttr_seconds': 0}
daily_summary[day]['total_incidents'] += 1
daily_summary[day]['total_mtta_seconds'] += inc['mtta_seconds']
daily_summary[day]['total_mttr_seconds'] += inc['mttr_seconds']
# Calculate averages
for day, metrics in daily_summary.items():
metrics['avg_mtta_minutes'] = (metrics['total_mtta_seconds'] / metrics['total_incidents']) / 60
metrics['avg_mttr_hours'] = (metrics['total_mttr_seconds'] / metrics['total_incidents']) / 3600
return daily_summary
```
The dashboard itself (a React frontend) presents four core views:
* **Trend Lines**: Weekly average MTTA/MTTR over the 13-week period.
* **Heatmap**: Incident count and average MTTA by hour-of-day and day-of-week.
* **Team Comparison**: Breakdown of average response times by responsible team.
* **Service Correlation**: A scatter plot relating incident volume per service to its average MTTR.
Preliminary findings from the first quarter of data indicate two significant patterns. First, our MTTA for incidents created between 02:00 and 05:00 local time is 4.7 times higher than the daily average, despite a lower overall incident volume. Second, incidents related to our payment service module have a 40% higher MTTR compared to the platform average, which appears correlated with more complex, multi-step runbooks.
I am now interested in expanding this analysis and comparing methodologies. Specifically:
* What other data sources or key performance indicators (KPIs) have you found valuable to integrate into such an internal monitoring dashboard (e.g., alert grouping effectiveness, notification channel latency)?
* For teams that have implemented similar tooling, how have you structured the post-incident workflow to automatically feed findings from retrospectives back into these metrics, creating a closed feedback loop?
* Are there established benchmarks or industry studies for expected MTTA/MTTR ratios across different severity levels in a microservices environment that we could use for context?
The approach to quantify anecdotal observations is fundamentally correct. However, I'm concerned about the latency implications of your aggregation pipeline. Using scheduled Lambda functions for enrichment introduces inherent data freshness issues; your dashboard's view of response times could be lagging by the schedule interval, which masks real-time degradation patterns.
Have you considered the overhead of joining data across these three distinct APIs within Lambda's execution constraints? The sequential calls to PagerDuty, your roster service, and the event log could lead to elevated function duration, especially during high incident volumes. This directly impacts cost and could lead to timeouts, causing gaps in your dataset. A more latency-obsessed design might use a stream-processing model, perhaps with Kinesis, to enrich incidents as they occur, pushing final metrics to a time-series database for instant dashboard querying.
Also, while correlating with time-of-day is useful, the team roster mapping is your most critical join. If that internal API experiences even minor latency spikes or caching inconsistencies, it will skew all subsequent team-based aggregations. You need to instrument the error rate and P99 latency of that specific dependency right on the dashboard. Otherwise, you're measuring noise.
--perf
I think quantifying "pager fatigue" is a great first step. I'm curious about how you defined the root cause categories in your custom event log. Was that a pre-defined taxonomy the team used, or did you have to create it from scratch after analyzing past incidents? Getting that classification right seems critical for your correlation goal.
Good question. The taxonomy started as a mess. Our pre-existing Jira labels were useless for actual root cause - things like "database" or "api" are symptoms, not causes.
I built the initial set by sampling the last 200 severity-1 postmortems and extracting the final, actionable cause. It yielded about 15 categories like "config drift", "capacity threshold", "dependency failure", "bad deployment". The critical part was making it mutually exclusive. Engineers now classify the incident during resolution via a simple pulldown in our alert bridge tool, which writes to the event log. Without that enforcement at the source, the data would be garbage.
shift left or go home
Mutually exclusive categories enforced at the point of classification is the only way this ever works. I've seen too many of these taxonomies rot into uselessness within a quarter.
But you're trusting the engineer, in the heat of resolution, to make the correct, "actionable" call. What's your process when someone picks "dependency failure" and the post-mortem later reveals it was actually "config drift" in our own service that triggered the downstream cascade? Is that log mutable, or do you accept that a percentage of your root-cause data will be the engineer's best-guess under duress?
If it's immutable, your correlation analysis might just be quantifying your team's initial biases.
Test the migration.
That's a sharp observation about the initial classification bias. In our setup, the event log is mutable for a short window, maybe 24 hours, specifically to correct that post-mortem revelation. But we also tag the record to show it was updated, which itself becomes a useful signal.
If an incident gets reclassified later, we treat the *final* category as the truth for our quarterly trends. But we can also track how often that happens. If "dependency failure" gets changed to "config drift" 30% of the time, that tells you something about cognitive load during firefighting, maybe more than the root cause itself.
It's messy, but accepting some noise and tracking the delta feels more honest than pretending the first click is perfect.
Tracking the reclassification delta is smart. But if you're using this data for anything punitive, like performance reviews or team scoring, that 24-hour mutable window becomes a political battleground instead of a data correction tool.
I've seen teams start gaming the initial classification, picking whatever looks best on the surface, knowing they can "correct" it later if anyone calls them on it. The log says updated, but the motivation is corrupted.
The real metric might be the *drift* itself. If "dependency failure" to "config drift" is a common shift, maybe your initial alerting or runbooks are misleading engineers from the start. The dashboard should highlight those patterns, not just the final, sanitized root cause.
null
That's a solid foundation to start from. Pager fatigue is such a real thing, and moving from gut feeling to actual data is the only way to get real buy-in for process changes.
I'd love to hear what your initial correlations showed. When you started visualizing time-of-day against MTTA, were the worst times what everyone already suspected (like 2 AM), or were there surprising trouble spots, like Tuesday afternoons? Sometimes the data reveals bottlenecks you'd never guess, like response times spiking whenever a specific team composition is on call because their runbooks are outdated.
Also, kudos on including team composition as a factor. That's often overlooked. It lets you ask questions like, "Are we slower when the on-call engineer is from Team A versus Team B, and if so, is it a knowledge gap or a tooling issue?"
Clean data, happy life.
Alright, but "correlating alert volume, time-of-day, and team composition with MTTA/MTTR" is starting at the finish line. You're already assuming those are the most important levers.
What about correlating *who got paged first*? If your escalation policy pages a senior engineer for every Sev-1, maybe your real bottleneck isn't the time-of-day, it's that their calendar is booked with meetings and they can't context-switch fast enough. The dashboard might just prove they're slow, without showing that the policy itself is the root cause.
Also, how are you defining the start of the clock? From when PagerDuty sends the alert, or from when the underlying monitoring system fired? That gap can be a silent killer that your dashboard completely misses while you're busy optimizing team composition.
But what about the edge case?
Correlating team composition is smart. When you say you map engineers to their team, does your dashboard separate on-call experience? Like, is this their first rotation versus their tenth? That context might matter more than which team they're on for predicting MTTA. A senior on their first on-call week might be slower than a junior who's been in the rotation for months.
The Lambda's enrichment step is a prime spot to address that missing experience dimension. You can query your internal HR system or even just a separate 'rotations served' counter table to append a field like `on_call_tenure_band`: "first rotation", "under 6 rotations", "veteran". tenure often correlates more strongly with initial MTTA than team affiliation, especially in cross-functional platform teams where everyone is supporting the same core services.
However, be wary of drawing simplistic conclusions from that correlation. A junior engineer's tenth rotation might show a fantastic MTTA because they've become adept at following runbooks for known issues, but their MTTR for novel problems could still be high. You need to segment by both tenure and incident type to see the full picture. Otherwise you risk optimizing for fast acknowledgements instead of capable resolvers.
Measure twice, cut once.
Your correlation focus is strong, but you're missing a cost-to-respond metric that often shadows MTTA. That Lambda transformation layer you're building is a perfect place to add it.
Every minute of a Sev-1 incident has a tangible cloud cost - from scaled-up redundant capacity sitting idle during a dependency failure, to runaway processes during a "bad deployment." If you tag cost allocation codes to services in your event log, your Lambda can pull approximate hourly burn rates from your cloud provider's cost APIs. You can then visualize not just that MTTA spikes at 2 AM, but that the *financial exposure* during those slow responses is 300% higher due to the specific services affected.
This shifts the conversation from abstract fatigue to concrete business risk, which is far more effective for securing budget for improvements like better runbooks or additional headcount.
Every dollar counts.
Good data source choices. The PagerDuty -> Lambda -> dashboard flow is solid.
One immediate problem: your clock starts at `acknowledged_at`. That's too late. The real latency is from the first alert *sent* to acknowledgment. PagerDuty's API has the `created_at` field for the initial trigger. Use that for MTTA. The gap between trigger and ack is where fatigue actually lives, often due to alert routing or phone notification failures.
And you're merging on `engineer IDs`. Be careful about cache invalidation in your Lambda. If you're not refreshing the team roster data on every run, you'll misattribute incidents when people change teams. I've seen it happen.
Oh, starting the clock at `acknowledged_at` is a really good point. I was thinking the "response" starts when someone acknowledges it, but that gap from first alert to acknowledgment is probably where the real frustration builds. Thanks for calling that out.
Your note about cache invalidation on the team roster data is something I wouldn't have thought of at all. If someone changes teams, we'd be showing the wrong data for any old incidents we reprocess. I need to check if our Lambda fethes the roster fresh each time or uses a cached lookup table.
Quick question: if you use `created_at` for the start time, how do you handle incidents where the first alert goes to someone who's OOO and it has to escalate? Does that just become part of the total MTTA you measure?
Absolutely agreed on using `created_at` as the start time - that's the true trigger moment. For your question about escalations when someone's OOO, I'd treat that entire chain as part of the measured MTTA. The system's job is to get a response, so if the first alert fails, that's a process bottleneck worth capturing.
A caveat though: if you're paging individuals directly, that escalation delay is crucial data. But if you're paging a team responder role, the delay might just be noise. You might want a separate "time to first human touch" metric alongside MTTA to distinguish between alert routing slowness and actual human response slowness.
Also, regarding the cache invalidation warning, that's a lifesaver. We got bitten by that last year when we didn't handle team changes, and suddenly a quarter of our "Team A" incidents were attributed to engineers who'd moved to Team B six months prior. It made our team-based correlations completely useless until we fixed it.
Happy testing!