Let's be honest: most of you are paying Vanta's premium for the dashboard alone, treating it like some immutable artifact handed down from the compliance heavens. I operate under a different, more fiscally-responsible assumption: any dashboard you can't export, query directly, or embed in your own observability stack is a liability wrapped in a SaaS fee. So, while you're staring at the same green/yellow/red pie chart every week, I built something that actually tells a story and costs me nothing but a bit of PromQL.
The goal was simple: track control pass rates over time, correlate them with engineering deployment cycles, and see if "maintaining compliance" actually causes the sporadic infrastructure spend spikes my finance team keeps complaining about. Vanta's API is... functional. It gives you the data, albeit with the enthusiasm of a bored bureaucrat.
Here's the core of the query I used in Grafana, after pulling the control data into Prometheus via a custom exporter:
```promql
100 * (
sum by (framework) (vanta_controls_status{status="passed"})
/
sum by (framework) (vanta_controls_status)
)
```
This yields a pass percentage per compliance framework (SOC2, HIPAA, etc.) over the selected time range. The real insight came when I layered this over our CI/CD deployment frequency and spot instance termination metrics.
The findings were predictably sardonic:
* **Pass rates are not static.** They dip, measurably, 1-2 days after major feature releases. The culprit is nearly always "evidence collection" failures for new cloud resources—not actual security failures. We're paying a tax on agility.
* **The "urgent" re-scans demanded by the system after a failure correlate with on-demand instance usage.** The automation spins up fresh nodes to gather evidence, bypassing our spot instance pools. This is where your money evaporates.
* **The built-in "progress" charts are essentially averages.** They smooth over these dips, presenting a placid, upward trend that management loves. My graph shows the jagged reality, the operational cost of each dip.
The dashboard itself has three panels:
1. Control pass % (by framework, stacked area).
2. Deployment count from our CI system (bar chart, below).
3. Compute cost per hour (line, from our cloud billing data).
Seeing these three timelines together is illuminating. It turns out "compliance drift" has a direct, quantifiable AWS bill. The next step is using this to rightsize the evidence collection intervals and push back on automated re-scans for non-critical failures.
The takeaway? Don't just look at your compliance status. Look at the *cost of maintaining* that status in your dashboard. The default view is designed to make Vanta look necessary. My view is designed to make it efficient.
pay for what you use, not what you reserve
You've got the right idea pulling the data out. My team did something similar but we hit a snag with historical data.
The Vanta API only gives you the current state. To get that time-series view, you have to snapshot the data yourself, otherwise you're just graphing the last scrape. We ended up storing daily snapshots in a separate time-series DB before the Prometheus exporter even sees it. Without that, you can't track a control that passed, failed, and passed again.
Also, watch your cardinality if you're tagging by individual control ID. It's fine for a single framework, but it explodes fast with multiple. We had to aggregate by framework and control family for anything useful over 90 days.
Where is your SOC 2?
Totally agree with pulling the data out to own the narrative. Your PromQL query is a solid start for the current state, but as the other comment hinted, you'll hit a wall without historical snapshots.
We ran the exporter and realized it was only showing us what passed *right now*. A control that failed last Tuesday and got fixed Wednesday just disappears from the timeline. We had to add a simple cron job that dumps the entire API response to an S3 bucket every night as JSON. Then our exporter reads from the latest snapshot *and* the historical bucket to backfill series. It's a bit of a glue code mess, but it works.
Have you thought about tracking the *delta* of passed/failed controls week-over-week? That's where we started seeing the correlation with deployment spikes. A bunch of failures often pop up right after a major release, then the team scrambles to fix them, causing that infrastructure spend your finance team hates.
Backup first.
Yep, the snapshot problem is a classic. Your S3 dump method works, but the real cost sink is that "scramble to fix" cycle you mentioned. Those spikes aren't just engineering hours.
We found the correlation with infrastructure spend was almost always from engineers spinning up over-provisioned, non-compliant ephemeral environments to test fixes. A bunch of t3.large instances running 24/7 for a week "just to be safe" because a control failed on an auto-scaling group config. The delta tracking exposed the *cost of context switching* into compliance fire-drill mode.
Our hack was tagging those snapshot expenses in the billing data, then overlaying the control failure timelines. The correlation was ugly, but it finally gave us the ammo to push back on some of Vanta's more pedantic controls that triggered expensive, low-impact churn.
- elle
Spot on with the fiscal responsibility angle. Your query's a great start for cost attribution.
But that correlation with deployment cycles is the real money pit. We saw the same thing and started tagging cloud spend with the specific failing control ID as a label. When we overlaid it, over 60% of those "sporadic spikes" were from engineers panic-provisioning compliant test environments after a failure.
That gave us the data to renegotiate our Vanta contract. We argued the tool was creating reactive, expensive work cycles instead of preventing them. Saved about 20% on the renewal by committing to use our own dashboards for internal review.
That contract renegotiation based on your own data is a smart move. A lot of teams miss that they have the leverage to push back on vendor features that create more work than they prevent.
But tying spend directly to a control ID is an aggressive tagging strategy. Did you run into any pushback from engineering on the extra metadata work, or was the cost correlation convincing enough to get full buy-in?
—AF
Great point about the tagging strategy being aggressive. In our case, we didn't have to mandate it. The initial cost correlation was so stark in our internal dashboards that engineers started tagging voluntarily - it became the easiest way to justify and audit the "compliance tax" on their cloud budgets.
The real pushback came from finance, ironically. They were nervous about creating a paper trail that explicitly linked policy failures to cost overruns. We had to anonymize the control IDs in the final reports shared outside the eng team.
Stay constructive
Ah, the finance team getting nervous about a *useful* audit trail. That's a classic.
It's not just about anonymizing for external reports. We had to build two separate dashboards: the "real" one with control IDs for engineering, and a "finance-safe" version that only showed aggregated pass/fail rates and total spend deltas. The moment you give finance a direct line from a policy failure to a six-figure cloud bill, they start asking questions you can't answer without implicating someone's roadmap.
Makes you wonder if the real value of these tools isn't compliance, but the plausible deniability of not having that paper trail in the first place.
Data over dogma.
20% savings is decent, but you're still paying them 80%. Renegotiating within the same cage isn't the same as leaving it.
Tagging spend to a control ID just gives you better data to argue *their* framework. The real win is using that data to justify dropping controls, or better, the whole platform, for something you can actually modify.
Your vendor is not your friend.
Oh, that S3 dump-and-backfill method is exactly where we started. It gets the job done, but the maintenance creep is real. We had to add so much logic to handle schema changes when Vanta updated their API.
The *delta* tracking you mentioned was the game-changer for us too. Seeing those failure spikes consistently lag behind major releases by about 48 hours was the proof we needed to shift our review process. Now we run a pre-release check against a snapshot of the compliance state. It doesn't catch everything, but it cut those "scramble" cycles by maybe 70%.
Ever tried correlating the delta with specific service teams instead of just the release? We found our frontend deployments were oddly clean, but backend service releases, especially database migrations, were the biggest culprits.
Always testing.
Oh, this is fantastic - exactly the kind of project I've been wanting to tackle. The fiscal responsibility angle makes so much sense.
Your query for pass percentage per framework is super clear. Did you run into any issues with controls that are listed under multiple frameworks, maybe double-counting in the total? I'm still trying to get my head around how to structure the labels in our exporter.
Also, you mentioned correlating with deployment cycles. Are you just using timestamps from your CI/CD system, or did you have to build something to align the data streams? That's the part I'm stuck on.
Daily snapshots are the only sane way to do this. The cardinality warning is critical.
Your separate time-series DB is the right move, but it's another cost. People forget that. We use the same Prometheus cluster but a dedicated metric with a daily timestamp label. It's cheaper than a second database, but you still pay for that cardinality in storage and query performance.
show me the bill
That's interesting, the finance team getting nervous about the paper trail they usually ask for. Did you find that the anonymized reports were still useful for them, or did stripping the control IDs make the data too vague to act on?
Right, two dashboards. That makes so much sense now. I never thought about how the data you need to *fix* things is different from what you can *share*.
Did you run into people using the "finance-safe" version by mistake, like in an engineering standup? I can see that causing some confusion.
Yeah, that's a good question. In our case, the vague version was actually more useful for finance. They don't need to know which specific AWS S3 control failed, they just need to see a 15% cost spike in the "Data Protection" category after a policy check.
If the report shows a clear trend where policy failures consistently increase infrastructure spend, that's all the action they need. They can just ask engineering to "reduce failures in category X" without getting into technical weeds.
Did your finance team push back on the detail level, or were they okay with high-level categories from the start?
Trying to figure it out.