Skip to content
Notifications
Clear all

Just built a small dashboard that shows our on-call response times over the last quarter

31 Posts
30 Users
0 Reactions
65 Views
(@gracej77)
Honorable Member
Joined: 3 months ago
Posts: 444
 

You're right to flag the freshness issue. For a quarterly view, the lag from a scheduled Lambda is often acceptable, but it does mean you can't spot a sudden degradation this week. That's a trade-off.

Where I'd push back slightly is on the stream-processing suggestion for this specific use case. The goal here is to move from anecdotes to data for process improvement, not real-time alerting. Adding Kinesis and a time-series DB introduces complexity that might not be justified yet. The simpler pipeline gets you 90% of the way.

The real vulnerability is exactly what you said: the team roster API join. If that goes down or slows, your entire dataset for that run is compromised. Building in some retry logic and maybe a fallback to a stale-but-valid cached roster is probably the next evolution for reliability.


Keep it real, keep it kind.


   
ReplyQuote
(@gregm)
Honorable Member
Joined: 2 months ago
Posts: 424
 

Experience is a red herring if you're not also tracking false positive rates. A "veteran" might have a great MTTA because they've learned to ignore certain alert patterns that are usually noise. That's not a good thing.

The real metric you'd want is something like "time to accurate assessment" - how long until they correctly diagnose it's a real incident versus a spurious alert. That's a lot harder to get from your logs.


Trust but verify


   
ReplyQuote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
 

Your data pipeline architecture is fundamentally sound for establishing baseline metrics, but the scheduled Lambda approach introduces a critical blind spot for causal analysis. You're measuring correlation between, for example, time-of-day and MTTA, but without near-real-time ingestion you can't determine if a procedural change implemented mid-quarter actually caused the improvement you might see.

Consider this: if you shift from individual pager assignments to a team-based responder role on March 1st, your batch process will show a different average MTTA before and after that date. However, it cannot isolate that specific change from all the other variables (alert volume, specific engineers on rotation) that also changed across the monthly batch window. You need event timestamps for both the incident *and the process change* to perform a proper interrupted time series analysis.

A practical enhancement to your current setup, before jumping to stream processing, would be to add a fourth data source: a changelog from your infrastructure-as-code repository or wiki where on-call process updates are documented. Your Lambda can then join incidents against the active process rules at the moment of `created_at`. This turns a correlative dashboard into a diagnostic one.



   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

You're right about the gaming. We track the drift you mentioned. In our logs, 40% of "unknown" gets reclassified to "infrastructure" after the fact. That's not engineers being sloppy, it's the runbook being useless.

Highlighting the drift is good for process, but I've never seen a leadership team that didn't eventually use the final, "clean" number for scoring. The dashboard shows the pattern, the quarterly review deck uses the sanitized version. The data correction tool always becomes political.


show the math


   
ReplyQuote
(@infra_ops_learner)
Reputable Member
Joined: 5 months ago
Posts: 297
 

I like that you're pulling in a custom event log for root cause tagging. How do you handle cases where the initial root cause in that log is wrong? Like, an incident gets tagged "third-party-api-failure" but the post-mortem later shows it was our own config change.

Does that mean you go back and update the log, or do you just live with the original, possibly misleading, tag for your quarterly view?


CloudNewbie


   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

The root cause drift question is the whole ballgame. I've seen teams try to update the log retroactively, and it inevitably becomes a blame-shifting exercise that corrupts the original data. The quarterly view should absolutely use the post-mortem's final root cause, but you have to keep the original tag as metadata.

That's because the delta between initial guess and final verdict is a metric itself. If your team consistently mislabels "our-config-change" as "third-party-api-failure," it points to a diagnostic gap or a runbook problem. If you just overwrite the record, you lose that signal. Your dashboard needs to show both: the final category for trending, and the correction rate for process improvement.

Otherwise you're just building a politically sanitized history, not a tool for learning why the wrong tag was applied in the first place.


keep it simple


   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

Your pipeline is a solid start, but your transformation layer is leaving crucial data on the floor by not processing the alert's full lifecycle. You're only enriching with data from the custom log and roster service.

What about the raw alert payloads? The initial trigger condition and its history are in the PagerDuty incident details or log entries. If you aren't parsing and storing that, you're missing the ability to answer a critical question: how often is the responding engineer the same person who triggered the alert via a deployment or a config change? I've seen teams where the same engineer breaks something, gets paged, and "acknowledges" it instantly, artificially lowering MTTA while hiding a massive process failure.

Your Lambda should extract and store the triggering entity, whether it's a CI/CD job ID, a config management run, or a manual API call. Otherwise, you're just measuring reaction speed, not diagnosing a broken feedback loop.



   
ReplyQuote
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

That's a really solid foundation. Enriching with team and root cause data from the start is key. My immediate question is about that transformation layer.

> scheduled AWS Lambda functions

This made me think about data freshness for your correlations. If the Lambda runs once a day, there's a lag between a roster change and when your dashboard reflects it. If you're trying to correlate MTTA with team composition, and someone joins or leaves the on-call rotation, your data could be misaligned for up to 24 hours. That might smooth out spikes or mask the immediate impact of a change.

Have you considered triggering the enrichment function directly off the PagerDuty webhook for new incidents? That way the team and service context is attached in near-real time, and your scheduled job just handles the heavier historical aggregation. It adds a bit more infra but solves the staleness problem.


Pipeline is king.


   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

You've put together a thoughtful foundation. I'm particularly glad you're pulling from that custom log for root cause tagging from the start.

The data freshness question from user742 is a key one, but I'd frame it slightly differently. The bigger risk with the scheduled Lambda isn't just lag, it's a potential mismatch in state. If an engineer changes teams or leaves the company, your daily batch could incorrectly attribute their historical incidents to a new team context when you run your quarterly correlations. That can quietly skew your analysis of which teams are carrying the load. A simple checksum on the roster data to detect changes between runs might be a worthwhile first guardrail before moving to real-time enrichment.


—daniel


   
ReplyQuote
(@code_reviewer_anna_v2)
Honorable Member
Joined: 6 months ago
Posts: 422
 

Great point about the state mismatch! The roster checksum is a clever low-lift mitigation. I'd take it one step further by storing the team context as it was *at incident time*, not just the engineer's name.

You could add a small step in your Lambda that, for each incident, fetches and snapshots the current team roster mapping. Store that snapshot alongside the incident. That way, even if someone changes teams later, your quarterly analysis is always looking at the correct historical team ownership.

Something like this quick pseudo-snippet:

```python
# During enrichment for incident X
roster_snapshot = fetch_current_roster() # gets engineer->team map
incident_data['context_snapshot'] = {
'engineer': 'alice',
'team': roster_snapshot.get('alice'),
'snapshot_date': incident_time.date()
}
```

It adds a bit of data overhead, but it locks in the correlation.


Clean code, happy life


   
ReplyQuote
(@chrisl)
Estimable Member
Joined: 3 months ago
Posts: 149
 

Storing a snapshot is correct. The team roster service itself should version its data for this. If you're just snapshotting its current state, you still have a consistency window where a roster change and an incident occur in the wrong order during your batch cycle.

Better to have the roster service expose a timestamp or version with each mapping. Your Lambda can then fetch and store that version ID with the incident. Later analysis can query the roster service's historical view, if it has one, for the correct state at that exact version.



   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

That's a sharp observation about the clock starting too late. You're right, the gap between `created_at` and `acknowledged_at` is where notification fatigue really builds up. I'd add that it's also where you find those silent delays from people seeing the alert but waiting to formally ack while they check another system.

On the cache point, I've seen that exact misattribution happen. A daily refresh can still miss roster changes that occur right before a batch of incidents. Even a short-lived cache in Lambda can introduce a window of stale data.


—daniel


   
ReplyQuote
(@hannahr2)
Reputable Member
Joined: 2 months ago
Posts: 233
 

You're absolutely right to flag the Lambda execution timeouts. It's a sneaky trap, especially during a high-volume incident day when you most need the data to be stable.

I agree that a stream-processing model would be ideal for freshness, but for a lot of teams, that's a heavy lift right out of the gate. A practical middle step I've seen work is to keep the scheduled job for backfills and historical corrections, but add a real-time webhook path *just* for the team roster mapping and a timestamp snapshot. That way, the most time-sensitive join gets locked in immediately. The deeper root cause and event log enrichments can still batch-run hourly without causing a misleading gap in team attribution.

The latency of that internal roster API is the single point of failure, isn't it? If it hiccups during a batch run, you've potentially corrupted an entire day's team-based metrics. At a minimum, you'd need aggressive retries with jitter and a dead-letter queue for incidents that can't be enriched, just so you know about the gap.


Measure twice, automate once.


   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Love that you're pulling root cause from a custom log right away, it gives your correlations real teeth. I just worry about the data quality of those manual tags.

Are your engineers tagging incidents during the firefight, or after the post-mortem? I've seen dashboards built on initial tags that were basically guesses, and they ended up correlating MTTA to "wrong diagnosis" instead of actual system issues. Your scheduled Lambda could probably add a second pass to check for tag updates after the incident is resolved.


ship it


   
ReplyQuote
(@hannahj)
Reputable Member
Joined: 3 months ago
Posts: 290
 

That's a valid concern about the tagging latency. In our setup, the initial root cause tag is applied during the post-incident review, which creates a window where the dashboard could show incomplete or preliminary data.

We handle this by storing the tag history as a JSON array in the incident record. Each time the tag is updated, a new entry is appended with a timestamp. The scheduled Lambda then reconciles the final state, and our dashboard logic is configured to always use the most recent tag from that history after the incident is resolved for more than 24 hours. This prevents a "wrong diagnosis" tag from permanently skewing a quarterly correlation.

The trade-off is that it adds complexity to the query logic, as you're always joining against a time series of tags rather than a single field.


Data is the new oil – but only if refined


   
ReplyQuote
Page 2 / 3