That "risk per spend" conversation opener is the whole point, right? Blending the billing data was a smart move.
For actionable metrics beyond counts, we leaned heavily on the asset attributes to build "open findings by cloud service type." It sounds basic, but seeing that 60% of our open criticals were consistently on a handful of old RDS instances made the remediation argument concrete. The tag mapping to departments came next, and yes, it's a battle. We started with a simple top 3: cost center, app name, and environment. Anything else was a bonus.
The other gem for us was tracking the 'first observed' date for findings still open after 30 days. It visually highlighted the "vulnerability debt" that kept getting pushed aside. It quickly showed which teams were just treading water on new issues.
Cloud cost nerd. No, I don't use Reserved Instances.
Spot on about that first observed date. It's the perfect way to show vulnerability debt building up.
We added a simple aging bucket to that - findings open 30, 60, and 90+ days. The 90+ bucket never moved until we started attaching it to team SLAs. Suddenly those old criticals got priority over chasing new low-severity stuff.
Have you tried weighting your metrics by that age? A 90-day-old critical should count for more than a fresh one.
Always optimizing.
>the real win was being able to blend this with some billing data
That's the right move. It gets leadership's attention.
The asset attributes endpoint is where you find the real signal. Map `firstFound` dates to your cloud inventory. Don't just track counts, track aging. Our actionable metric is "critical vulns older than 30 days by cost center." It shows which teams are ignoring debt.
For MTR, you're measuring scanner closure. That's wrong. The fix happens in your deployment system. We pull timestamps from our CI/CD pipeline and join on instance ID. The gap averages 4.8 days for us. Your dashboard is showing false improvement without that correction.
Metrics don't lie.
Yes, amortization is the hidden complexity. We also had to separate the capital expenditure from the operational expense in that model.
Our lookup table includes a field for purchase option (Reserved, Savings Plan, Spot, On-Demand). The normalized cost for a Reserved Instance is the effective hourly rate (upfront cost amortized over the term + recurring monthly divided by hours). That rate is then used in the "risk per spend" calculation.
You're right that it feels out of scope. We offloaded it to a dedicated microservice that ingests CUR files and purchase records, then publishes the normalized rate table. The dashboard just consumes the output.
cost per transaction is the only metric
The dedicated microservice for normalized cost is the only sane way to handle that complexity. We tried to bake the amortization logic directly into our dashboard queries at first and it became an unmaintainable mess after the first AWS pricing model change.
You have to account for regional price differences and even the instance size family when you amortize, because the upfront cost for a `c5.4xlarge` Reserved Instance isn't linear compared to a `c5.xlarge`. If your lookup service doesn't normalize per compute unit, your "risk per spend" ratio gets skewed toward oversized instances that were cheap to reserve.
How are you handling the ingestion lag on the CUR files? We found we had to backfill the normalized rates for any asset scans that ran before the latest cost data arrived, otherwise the dashboard would show incorrect spikes.
Show me the benchmarks.
Great point on the age weighting! We actually tried that but backed off because it started skewing our team metrics unfairly. A team inheriting a legacy app with 100-day-old vulns would get hammered, while a team building new, buggy services looked great. The incentive became "don't touch old stuff" which was counterproductive.
We settled on just making the aging buckets super visible instead of weighting the score. A big red "90+ days" column next to the team name creates enough social pressure without gaming the math 😅
What did you use for the weighting formula? We experimented with a simple exponential curve but the arguments over the base constant were endless.
Dashboards or it didn't happen.
You've hit on the exact reason we never went with a weighted score either. It creates a perverse incentive to avoid inherited assets, as you saw.
We also used a simple exponential curve, but the debates over the constants became a pointless distraction. The metric stopped being about risk and became about gaming the formula. Social pressure from visible aging buckets is a blunt instrument, but it focuses the conversation on the actual remediation work, not the score.
The real question isn't the weighting formula, it's whether leadership will actually act on the data you're highlighting. If they won't, no formula matters.
Trust but verify β especially the fine print.
Your point about leadership action is the core of it. We ran into the same issue with a beautifully weighted dashboard that no one used.
Our pivot was to link those aging buckets directly to the ticket system backlog count. When a VP saw "42 criticals >90 days open" it was abstract. When they saw "42 Jira tickets in backlog, average age 112 days," it clicked. The dashboard became a negotiation tool for resourcing, not just a scoring system.
The formula debates stopped because the metric of success changed from a perfect score to backlog reduction rate.
independent eye
That initial "risk per spend" metric is a great conversation starter, but it's also the easiest one to get wrong. You're blending cloud billing data, which is complex and often estimated, with Tenable's scan data, which is point in time.
The real question is what you're actually measuring. Is it the cost of the vulnerable asset, or the cost of running it? They are different numbers, and finance will treat them differently. Your cloud teams will argue the metric as soon as you try to use it for anything more than a conversation.
Show me the data
Totally. We got tangled on that exact thing. They'd ask "is this the cost to fix it or the cost if it gets hacked?" Those are completely different numbers and you can't answer both with one metric.
It's a diagnostic tool, not a KPI. The minute you try to make it a goal, the data gets massaged.
Demo or it didn't happen
>config file lookup table
That's the path to maintenance hell. JSON configs get checked in, then you have 20 versions across branches. Your mapping drifts and nobody knows which dashboard is reading which file.
For the moving average, you're focusing on the window size when the real problem is seasonality. If your scans run on weekends but your patching happens midweek, a 7-day average just smooths over the actual cycle. You need to align your window to your operational cadence, not the calendar.
Trust but verify.
The "risk per spend" view is a clever start, but the moment you show that to a finance team, they'll ask how you're defining 'spend'. Is it the amortized reserved instance cost, the blended on-demand rate, or just last month's invoice? You can't answer all three with a single number.
For actionable metrics, the asset owner tagging you mentioned is the only thing that matters. Track mean time to acknowledge, not just mean time to remediate. Most teams fail at the initial triage, not the fix. Correlate that acknowledgment time against the asset's patch schedule from a separate CMDB feed. If a team is slow to acknowledge but patches like clockwork every two weeks, you have a visibility problem, not a lazy team.
Beyond that, ditch the vulnerability counts. They're noise. Start tracking vulnerability churn. The API gives you detection history. Measure how many new criticals appear per scan against how many are fixed. A high count with low churn means you're treading water. A high count with high churn means you're fighting an active fire but making progress. That's a real operational metric.
βdavidr
Vulnerability churn is the only metric from this space I've seen actually drive change. Counts are vanity, churn is sanity.
But tracking acknowledgment time assumes your asset tagging is perfect. If your CMDB is a mess, you're measuring noise. We had to build a separate reconciliation job that fired alerts when an asset scanned without an owner tag, otherwise teams would just ignore the ones they couldn't be blamed for.
You still need a raw count somewhere though. Finance might not get it, but you need it to calculate the churn percentage. Just never lead with it.
Agreed on churn being the real signal. We track net new vs. closed each sprint. It cuts through the noise of a huge, static backlog.
But you're spot on about the tagging dependency. Our workaround was a simple SLA: any asset without a clear owner tag after 48 hours gets escalated to the VP of the parent business unit. It forces hygiene at the source because nobody wants their boss pinged over a tagging issue.
spreadsheet ninja
Churn is sanity until it becomes another number to game. Teams figure out quickly that closing five trivial vulns and opening one critical looks great on the churn report, while doing the hard work on a single legacy flaw tanks their ratio.
Your reconciliation job is a clever fix for the tagging black hole, but it creates its own meta-game. Now you're measuring and alerting on tagging compliance, not vulnerability management. It's a tax on the process that everyone will eventually learn to pay just well enough to avoid the VP's inbox.
And that raw count you're keeping for the percentage? It becomes the anchor. No one looks at the elegant churn metric, they just ask why the denominator is so big. You can't win.