"super clear" means you haven't looked at it in six months. Try updating the framework mapping after a vendor update. It's never clean.
The double counting problem? Don't solve it. Pick a primary framework per control in your scraper and ignore the others. The data is fuzzy anyway. You're looking for trends, not an exact count.
Correlating with deployments is a timezone matching nightmare. We just dump a daily timestamp label from our pipeline into a metric. It's coarse, but it's close enough to see if the roof is on fire. Building a perfect alignment layer is a waste of cycles.
If it ain't broke, don't 'upgrade' it.
>Pick a primary framework per control in your scraper and ignore the others.
This is the way. We did the same - we just assign everything to CIS initially, then manually recategorize the few critical ones that matter. Trying to perfectly model their mapping will drive you insane.
And you're dead right about the timezone alignment layer. The ROI is never there. If your trend line moves before your deployment marker, you still know what happened. You just see the smoke a few hours later.
I've seen teams burn a whole sprint trying to sync timestamps down to the millisecond, when a daily label would have shown them the same 40% failure spike.
Cheers, Henry
Exactly! It's funny how the incentives work. Once engineers saw the cost of compliance failures in their own team's budget - not just some vague company policy - they became the biggest advocates for tagging.
We had a similar thing with finance. They worried about audit risk in the raw data. Our workaround was to aggregate the failures into broader categories like "data security" or "access management" before the reports left the engineering Slack channel. That way, the cost correlation is still clear, but you lose the direct link to specific controls that might spook auditors.
Dashboards or it didn't happen.
That initial PromQL you've got is a perfect starting point. The moment you start graphing that percentage over time, the story emerges. One thing I'd add, though, is to filter out frameworks that have a very low number of total controls in your time window; they'll show massive, misleading percentage swings from just one control changing state. Something like:
```promql
100 * (
sum by (framework) (vanta_controls_status{status="passed"})
/
sum by (framework) (vanta_controls_status)
) and on(framework) (count by (framework)(vanta_controls_status) > 10)
```
It keeps the noisy, sparse frameworks from distracting from the real trends in your major ones.
That initial PromQL you've got is a perfect starting point. The moment you start graphing that percentage over time, the story emerges. One thing I'd add, though, is to filter out frameworks that have a very low number of total controls in your time window; they'll show massive, misleading percentage swings from just one control changing state.
I'm curious about the "correlate them with engineering deployment cycles" part. How did you actually pull in the deployment markers? Are you using something like a pipeline event to annotate the graph, or just eyeballing it with a separate panel?
You're right that the real leverage is in justifying platform changes, not just negotiating within it. But I've seen teams get stuck on that final jump because they can't quantify the operational drag of the old system versus a new one.
Your data from tagging spend to control IDs is actually perfect for that. It's not just about arguing their framework, it's about building a business case showing how much time and money you spend *servicing* their framework compared to a more flexible alternative. That's what gets budgets approved.
—daniel
Yeah, that's exactly the trade-off we're still figuring out. At first, stripping the control IDs felt like it was making the data too vague to be useful. But our finance team actually prefers the broader categories, like "data protection" or "access management," because it lets them spot cost trends without getting lost in technical details they don't need to understand.
For example, seeing that "access management" failures went up 20% in a quarter, and that category's cloud spend also spiked, is a clear signal for them to ask questions. They don't need to know it was specifically about IAM role permissions.
Do you think there's a risk in *over*-aggregating, though? Like, if a spike is caused by one expensive but critical control, could blending it into a category hide something important?