Skip to content
Notifications
Clear all

Showcase: My custom dashboard for tracking MTTD and MTTR using ES audit data.

2 Posts
2 Users
0 Reactions
17 Views
(@henry)
Reputable Member
Joined: 3 months ago
Posts: 274
Topic starter   [#25902]

Hey everyone! As someone who lives at the intersection of security ops and analytics, I'm always looking for ways to measure what matters. One of the biggest challenges I've seen is getting a clear, actionable view of detection and response effectiveness.

While Splunk ES has great out-of-the-box dashboards, I wanted something more focused on the *process metrics* my team owns: Mean Time to Detect (MTTD) and Mean Time to Respond (MTTR). I built a custom dashboard that sources directly from the `audit` index, particularly the `notable_events` and `response_tasks` data.

The core idea is to track the timeline from event generation to closure. Here's a simplified view of the key SPL components I used:

- **For MTTD:** I calculate the delta between the `_time` of the original data and the creation time of the notable event. This gives a raw "time to escalate" figure.
- **For MTTR:** I track the time between the notable event creation and when its associated response task is marked closed or resolved.

The dashboard panels show:
* Weekly/Monthly rolling averages for MTTD & MTTR
* Breakdown by urgency level (critical/high/medium/low)
* Top contributing rule names – this is gold for pinpointing which detections need tuning
* A trend chart to see if our process improvements are actually moving the needle

The biggest pitfall I had to work around was ensuring the audit data was being populated correctly (check those indexer forwarding settings!). Also, remember that these metrics are directional – they can be gamed if teams rush to close tickets. The real value is in the trend lines and using the data to ask better questions.

Has anyone else built similar operational health dashboards? I'd be curious to hear how you're correlating these metrics with things like false positive rates or analyst workload.

Cheers,
Henry


Cheers, Henry


   
Quote
(@benwhite)
Reputable Member
Joined: 2 months ago
Posts: 209
 

Interesting approach, but using the audit index for this feels fragile. What happens when that data model changes or an admin purges old audit logs to save license? Your "time to escalate" figure is only as good as the audit retention policy.

Also, you're measuring based on Splunk's internal timestamps. How does this account for the clock sync drift between the original host and your Splunk instance? You could be measuring your infrastructure's time problems, not your team's speed.

What's the yearly cost of keeping that audit data hot for your dashboard versus just using a dedicated, simple time-series database?


read the fine print


   
ReplyQuote