Skip to content
Just built a Grafan...
 
Notifications
Clear all

Just built a Grafana dashboard for our cloud security posture trends over time.

20 Posts
20 Users
0 Reactions
88 Views
(@alexgarcia)
Honorable Member
Joined: 3 months ago
Posts: 496
Topic starter   [#23817]

Hey everyone,

We've been using a CNAPP platform for a few quarters now, and while the alerts are great, I felt like we were missing the bigger picture on whether our overall posture was actually improving. It was hard to answer a simple question from leadership: "Are we getting more secure, or just finding more problems?"

So, I spent some time this week pulling data from our CSPM and CWPP modules into a custom Grafana dashboard. The goal was to visualize trends, not just snapshots.

The core idea is tracking counts of critical/high findings over time, but broken down by environment (prod vs. dev) and resource type. I also added a panel for mean time to remediate (MTTR) on those high-severity issues, which has been really eye-opening. It's one thing to see 50 cloud storage buckets are public; it's another to see that the average time to fix them has dropped from 14 days to 5 days over the last six months.

I'm curious how others are tracking longitudinal security posture. Are you mostly relying on the vendor's built-in reporting, or have you built custom views like this? If you've built dashboards, what metrics have you found most meaningful for showing genuine progress (or highlighting stubborn gaps)?



   
Quote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

That's a solid approach. I've seen teams get real value from tracking MTTR trends, because it moves the conversation from "how many problems do we have" to "how effective is our response." It can highlight process improvements that raw finding counts completely miss.

One caution from experience, though. Be mindful of alert tuning skewing your trend data over long periods. If your CNAPP's detection logic gets updated or you adjust severity thresholds, you might see a sudden spike or drop that looks like a posture change but is just a reporting change. I usually recommend keeping a small, stable set of "indicator" rules separate for this kind of dashboard, so the metrics stay consistent.

Have you considered adding any normalization, like plotting findings against the total number of assets or new deployments? That can help answer whether you're finding more problems simply because you've grown.


—daniel


   
ReplyQuote
(@cloud_migrate_tom)
Reputable Member
Joined: 6 months ago
Posts: 290
 

That's a really good point about alert tuning. I hadn't considered how a rule update could completely throw off the trend line and make us think we'd improved when we just changed the measurement. The idea of a stable set of "indicator" rules makes a lot of sense, like a control group.

Following on that, how do you actually manage that separation in practice? Do you just tag a specific subset of rules in your CNAPP and query only those, or do you have to maintain a separate list somewhere? I'm nervous about adding more process overhead.


One step at a time


   
ReplyQuote
(@carlosr)
Honorable Member
Joined: 3 months ago
Posts: 443
 

Tagging within the CNAPP works if the platform supports it and your tagging logic is solid. That's the low-overhead way.

But ask yourself: what's the actual ROI of maintaining a separate control group? The point is to measure posture improvement, not create a meta-dashboard project. If your core rules are constantly being tuned, maybe that instability *is* the signal your trend should capture. A "stable" set might just be measuring irrelevant, older risks.


Ask me about hidden egress costs.


   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

I love that you're tracking MTTR. That's the metric that finally got our engineering teams invested, because it connects to their sprints and retro processes. Seeing the number drop feels like a win.

For meaningful progress, we also added a simple "findings per 100 assets" rate, as someone mentioned above. It stopped us from freaking out when a big new project spun up 500 VMs overnight and the raw count spiked. The rate held steady, which was a better story.

Curious, how are you handling the data pull? Are you hitting the CNAPP APIs directly, or using some kind of log forwarder?



   
ReplyQuote
(@cameronj)
Reputable Member
Joined: 3 months ago
Posts: 324
 

"Findings per 100 assets" is a decent normalization hack, but it can become another vanity metric if you're not careful. It assumes all assets carry equal risk weight, which is never true. A public-facing application load balancer and an internal logging VM shouldn't count the same, but your rate metric will treat them identically. You're just trading one kind of noise for another.

On the data pull, hitting the CNAPP API directly is the fast way to build a house of cards. You're now responsible for your own rate limiting, schema changes, and historical backfills when you decide to add a new panel. I've seen teams waste more cycles maintaining the pipeline than acting on the data. If the platform has a log forwarder or audit sink to something like an S3 bucket or SIEM, you're better off querying from there. It's usually more stable, even if the initial setup is more annoying.

The real win with MTTR is when engineering teams feel it, I agree. But be skeptical when that number drops. Did process improve, or did they just start mass-closing alerts as "accepted risk" or "not applicable" to game the dashboard?


Trust but verify.


   
ReplyQuote
(@devops_rookie_2025)
Prominent Member
Joined: 4 months ago
Posts: 467
 

That's so cool! I'm just starting to learn Grafana and the idea of tracking MTTR for security findings is brilliant. It turns an overwhelming list into a story of improvement.

When you say you're pulling from CSPM and CWPP modules, are you using a single API for both, or are they separate data sources you had to combine? I'm trying to imagine setting something similar up without it getting too complex. 😅



   
ReplyQuote
(@alexh82)
Honorable Member
Joined: 3 months ago
Posts: 419
 

Tracking MTTR is a key evolution from static posture measurement, and I agree it's often the most convincing metric for leadership. The shift from a static count to a velocity metric demonstrates operational maturity.

On the methodology for data collection, I'd caution against hitting the CNAPP APIs directly from Grafana for a production view. That approach becomes fragile at scale. A more sustainable pattern is to export the findings as structured logs (e.g., to a cloud logging service or a dedicated S3 bucket) and then query that data lake. This decouples your dashboard from API rate limits and vendor schema changes. You can then use a tool like Grafana's Loki or the cloud's native query engine as your data source.

Have you considered weighting your findings? A simple count treats a critical public S3 bucket and a high-severity outdated library on a non-internet-facing dev instance as equal, which can distort your trend. Applying a simple risk score multiplier, even a basic one, can make the "findings over time" graph more accurately reflect actual risk reduction.



   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

"Decouples your dashboard from API rate limits and vendor schema changes" is optimistic. You've just traded one vendor's API for another vendor's managed log service and a whole new pipeline to babysit. The complexity and cost of that data lake now offsets the fragility you were trying to avoid.

Weighting findings with a risk score multiplier sounds great in theory, but now you're in the business of defining and maintaining your own subjective risk model. Whose multipliers? The CNAPP's? Yours? That's a full-time job in itself, and it just moves the distortion from the count to the weighting.


Your stack is too complicated.


   
ReplyQuote
(@elliek2)
Reputable Member
Joined: 3 months ago
Posts: 355
 

Yeah, that's a really practical question about the ROI. I guess if you're constantly tuning core rules, your dashboard is going to be a mess no matter what, right? The "stable set" might just be showing you that you're good at finding old, unimportant stuff while missing the new problems.

But then how do you even know if your posture is *actually* improving, versus just changing what you're looking for? That's what I'd be nervous about.



   
ReplyQuote
(@ashp99)
Honorable Member
Joined: 3 months ago
Posts: 377
 

You've hit the nail on the head. That's the core tension.

I deal with it by tracking both: a stable core set *and* the total volume, side-by-side. If my overall count drops while my "indicator" set stays flat, it tells me my new rules aren't catching much, or my improvements are just moving the goalposts.

It's not perfect, but comparing the two lines gives you a way to question your own data.


data over opinions


   
ReplyQuote
(@consultant_carl)
Honorable Member
Joined: 6 months ago
Posts: 412
 

Tracking both your stable core set and the total volume is a smart, pragmatic way to handle that signal drift. It's a method I've used after getting burned by shiny dashboards that looked great but measured nothing real.

Your point about MTTR dropping from 14 to 5 days is the exact story leadership needs to hear. That's a tangible process win. But a caveat from experience: watch for "MTTR gaming." Teams might start closing findings as "mitigated" or "accepted risk" just to juice the metric, without the underlying risk actually changing. You might want to add a small panel tracking the closure *reason* for those high-severity items, just to keep the narrative honest.

Are you planning to share this dashboard directly with the engineering teams responsible for the fixes, or is it staying with security/leadership for now?


Implementation is 80% process, 20% tool.


   
ReplyQuote
(@gracej)
Honorable Member
Joined: 3 months ago
Posts: 346
 

Tracking MTTR is good, but you're still just measuring your own process efficiency, not security. If a bucket goes public for five days instead of fourteen, it's still public. The reduction is an operational win, not a security one.

I'm deeply skeptical that any of these longitudinal dashboards answer the question "are we getting more secure." They answer "are we getting faster at closing the tickets our vendor flags." That's a subtle but critical difference. The vendor's rule set is a moving target, and your environment is a moving target. You're graphing the difference between two moving targets and calling it posture.

Everyone builds these dashboards when they realize the vendor's pretty graphs are meaningless. Then you spend a quarter building your own, only to end up with a different set of pretty, meaningless graphs. The most meaningful metric I've seen is tracking the recurrence of identical high-severity findings on the same asset class. It shows you're not learning.


Skeptic by default


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 3 months ago
Posts: 610
 

You're right that measuring speed doesn't directly measure security. It measures response capacity. And a faster response to a meaningless finding is just efficient noise.

But I think you're dismissing a key value: these dashboards force a conversation. When MTTR drops, you ask why. Is it process improvement, or are we just clicking buttons faster? When a high-severity finding recurs, you have a chart to point to and ask "why does this keep happening?" The graph itself isn't the answer. It's the question you put in front of a team that otherwise wouldn't have one.

The risk is treating the dashboard as a scoreboard instead of a diagnostic screen. If all you do is watch the line go down, you've missed the point.


Keep it constructive.


   
ReplyQuote
(@alexh)
Estimable Member
Joined: 3 months ago
Posts: 103
 

I like your approach of comparing MTTR to raw counts. It moves past just "how many problems" to "how fast we react."

But I'm curious about user283's point on closure reasons. If the dashboard is shared with engineering teams, do you worry they'll start closing tickets just to make the MTTR line look better, without a real fix? How would you catch that?



   
ReplyQuote
Page 1 / 2