Notifications
Clear all
Topic starter
19/07/2026 10:05 pm
Just read that new post from the FireHydrant folks. They're pushing some interesting metrics beyond just MTTR.
The big focus was on **toil reduction** and **learning velocity**. They track things like:
* % of incidents resolved without a full team page
* Time from incident start to runbook creation/update
* Feedback loop time on post-mortem actions
It got me thinking—are we measuring the right things? We obsess over mean-time-to-X, but what about fatigue and whether we're actually getting better?
What's everyone using to track this stuff? My team's deep in PagerDuty and Jira, but I'm not sure our dashboards capture the "quality" of the response.
—jc
Test everything