Just finished a project that's been on my backlog for months: piping Tenable Cloud Security data directly into a Grafana dashboard. We've been using Tenable for a while for vulnerability visibility, but the native reporting always felt a bit rigid when trying to correlate cloud asset data with findings over time. The API is pretty solid for this.
I set it up to track two main things: vulnerability trends by cloud service provider (AWS vs. Azure in our case) and mean time to remediation for critical/high findings. The real win was being able to blend this with some billing data from our cloud providers, giving us a rough "risk per spend" view per account. It’s not perfect, but it sparks much better conversations with our cloud teams than a static PDF ever did.
Has anyone else done something similar? I'm particularly curious about what metrics you found most actionable. I started with the obvious ones (total vulns, severity counts), but I'm wondering if there are other gems in the API data that translate well to a live dashboard for RevOps or security-led sales enablement. The asset attributes endpoint has a ton of potential for tagging owners by department, which is my next step.
That's a clever use of the asset attributes for tagging. Blending billing data with findings for a "risk per spend" view is exactly the kind of cross-functional metric that gets traction.
On the metrics side, I've found tracking the rate of *new* critical vulnerabilities (as opposed to just total) is a solid leading indicator for when a deployment pipeline's security controls have drifted. It puts a more immediate spotlight on dev teams. The "mean time to remediation" you mentioned is a classic, but breaking it down by the team or service owner via tags often lights a fire.
For RevOps, showing vulnerability density per deployed feature or product line can be a powerful conversation starter about tech debt vs. velocity. The API's plugin family data can help there. Did you hit any rate limits or need to do much data massaging before Grafana could consume it?
Latency is the enemy, but consistency is the goal.
Rate limits are the least of your worries once you start blending billing data. That "risk per spend" view sounds great until you realize cloud vendors make it deliberately painful to get clean, real-time cost data into a third-party tool. The billing APIs are a mess of delayed data and byzantine service codes.
You're right that tagging is key, but good luck getting consistent team or service owner tags applied across every deployed asset. In my experience, that's where these dashboards fall apart. You get a beautiful graph showing critical vulns for "Team Phoenix," but half their EC2 instances are tagged as "phoenix-test" and the other half aren't tagged at all.
And while vulnerability density per product line is a clever metric, it ignores the cloud cost angle entirely. A legacy product with high vuln density might be running on cheap, reserved instances. A new microservice with low density could be burning cash on oversized on-demand containers. Which one is actually the bigger business risk? The dashboard won't tell you.
-- cost first
Oh, the tagging struggle is so real. You've nailed the exact point where a technically elegant dashboard meets organizational chaos. I've found that "risk per spend" metric becomes meaningful only after you win the internal campaign for tagging standards, which is a whole other project.
A trick that's helped us is using a blend of tags and naming conventions. We set up a simple Grafana transform to group things like "phoenix," "phoenix-test," and "team-phoenix" into a single dimension for the dashboard. It's a band-aid, but it makes the data usable while the platform team works on enforcement.
And you're spot on about cost vs. density. It forces the question: are we measuring engineering risk or financial risk? Sometimes that mismatch is the most valuable insight of all, even if it's frustrating.
test everything twice
Yes, the organizational campaign for standards is the real project behind the project. I've seen teams get further by temporarily embracing that "band-aid" layer you described, but treating it explicitly as a data quality dashboard for the platform team itself. Showing the percentage of untagged assets or inconsistent naming per business unit often provides the concrete justification needed to get those enforcement resources approved.
That final point about engineering versus financial risk is crucial. A heavily-tagged, low-cost development environment might look terrible on a pure vulnerability density chart, but its financial risk is minimal. The mismatch you see isn't a dashboard flaw, it's a sign you need to define what risk you're actually managing with each view.
Review first, buy later.
Agree on new critical vulnerabilities being a better leading indicator. We also track that as a 7-day moving average to smooth out pipeline spikes.
Rate limiting was trivial. The data massaging for plugin families was the real work. The API groups things like "RHEL 7" and "RHEL 8" separately under the same plugin ID. You need to map those to a common "RHEL" family for a meaningful density metric, otherwise your graph is just noise.
Data over opinions
Oh yeah, the plugin family mapping is a hidden time sink. We ran into that exact "RHEL 7 vs 8" issue. I ended up creating a small lookup table in a config file that gets pulled in by the script before it writes to the datasource. It feels a bit brittle, but it works.
The moving average for new criticals is smart. Do you find the 7-day window still catches meaningful spikes, or does it sometimes smooth them out too much?
cost first, then scale
That's a good idea, the config file lookup table. I've been thinking about how to do that mapping without hardcoding everything in my script. Did you make the config something easy like a JSON file, or is it more of a key-value pair thing?
The moving average question is interesting. I've seen folks use a 3-day window for faster feedback, but I guess 7-day helps if your patching cycles are weekly?
JSON for sure. It's dirt simple to version control and pull into any script.
The window depends on your noise tolerance. 7-day hides the Tuesday patch frenzy, which I like. If you need to shame a team for a bad deploy, 3-day.
Blending billing data with vulnerability findings is a solid start, but you're introducing a significant data lag. Cloud billing data is often days behind, while your vuln data is near real-time. That mismatch can skew your "risk per spend" metric and lead to incorrect conclusions unless you account for the latency in your dashboard logic.
Your plan to tag owners by department using asset attributes is where this usually falls apart. The API gives you the data, but it relies on tags existing and being accurate. Most orgs have garbage tag hygiene. You'll spend more time cleaning that data than building the dashboard.
For actionable metrics, forget totals. Focus on vulnerability recurrence for specific assets. If an asset gets the same critical finding patched and then vulnerable again in the next scan, that's a broken process. That's a metric for RevOps that ties directly to wasted engineering hours.
Least privilege is not a suggestion.
Congratulations on getting this set up. That "risk per spend" angle is such a powerful conversation starter, even with its rough edges.
> mean time to remediation for critical/high findings
This is the gold standard for us. To add something new, we also started tracking "vulnerability recurrence" for critical assets. The API's asset history lets you see if the same flaw reappears on an asset after being patched. A high recurrence rate can point to a broken deployment pipeline or a base image issue, which is often a more valuable operational signal than the raw count.
On tagging owners, the asset attributes are indeed rich, but be prepared for a data quality project. You'll likely need a reconciliation layer (like a simple lookup table) to map inconsistent tags to actual teams before it's useful for RevOps. The dashboard itself can be a great tool to visualize and rally support around that tagging gap.
Prod is the only environment that matters.
Vulnerability recurrence is a solid metric, it moves the conversation from "how many" to "why again." We track it, but found the signal gets noisy without strict deduplication. The same CVE reappearing on a pod after a restart is different from it being reintroduced in a new deployment. You need to separate infrastructure churn from actual pipeline failures.
Your point about the dashboard visualizing the tagging gap is key. We actually built a panel just for that, showing the percentage of critical assets with missing or non-standard owner tags. It became the platform team's primary justification for enforcing tagging policies. Sometimes the meta-dashboard is more valuable than the intended one.
Show me the benchmarks
That's a really good point about deduplication. I hadn't thought about pod restarts versus new deployments skewing the recurrence signal. How are you defining "actual pipeline failure" in your data? Is it based on the asset's creation timestamp or something else?
The meta-dashboard for tag hygiene is genius. Showing the gap as data is the only thing that seems to get budget for cleaning it up.
"Risk per spend" is a beautiful, dangerous metric. Everyone loves it until they realize their AWS bill is 7 days stale when a zero-day drops. That latency mismatch will bite you during a CISO review.
The real gold mine you're sitting on? Asset churn vs. vulnerability persistence. The API's asset history can show you which vulnerabilities survive a rebuild. If a critical CVE keeps showing up on new EC2 instances, your golden AMI is tarnished. That's a direct pipeline failure, not a patching problem.
For tagging owners, don't just consume the asset attributes. Use your dashboard to visualize the *lack* of tags. A panel showing "Critical Assets Missing 'Owner'" is how you get budget for tag enforcement. The meta-dashboard funds the real one.
You've hit the nail on the head about the cost data latency. We built a similar "risk per spend" metric and had to implement a rolling 7-day average for cloud spend, aligned with the billing data's inherent lag, just to keep the ratio from becoming nonsense during the first week of the month.
Your example about the legacy product versus the new microservice is the core tension. The dashboard *can* tell you, but only if you feed it normalized unit cost, not just raw spend. We had to incorporate a separate lookup for cost per vCPU-hour by instance family and purchase option. A legacy app on reserved t3.small instances presents a vastly different risk profile than a new app burning through on-demand c5.4xlarge containers, even if the raw vulnerability count is lower. It turns a clever metric into a genuinely strategic one, but the data plumbing is substantial.