Skip to content
Notifications
Clear all

Showcase: My custom Grafana panel showing Vanta control pass rates over time.

6 Posts
6 Users
0 Reactions
3 Views
(@chrisk)
Honorable Member
Joined: 3 months ago
Posts: 398
Topic starter   [#29014]

As part of our ongoing compliance automation, we've been using Vanta for just over 18 months. While the platform's dashboard provides a good high-level snapshot, I found it insufficient for performing longitudinal analysis of our security posture. Specifically, I needed to correlate control pass/fail rates with specific engineering deployments and infrastructure changes to identify regression patterns.

To address this, I built a custom Grafana panel that visualizes Vanta control pass rates over time. The core of this setup is a scheduled script that extracts data from Vanta's API, flattens the structure, and pushes it as time-series metrics to a Prometheus instance. This allows for sophisticated querying and alerting not natively supported in the Vanta UI.

**Architecture Overview:**
1. A Python service (containerized) runs daily via Kubernetes CronJob.
2. It authenticates with Vanta's GraphQL API and fetches the control summary for our configured frameworks (e.g., SOC 2, HIPAA).
3. The script processes the nested JSON, calculating key metrics such as:
* Total controls per framework
* Passed controls
* Failed controls
* Controls not applicable
* Derived pass percentage
4. These metrics are labeled by framework and pushed to a Prometheus PushGateway.
5. Grafana queries Prometheus to render the time-series graphs.

**Key Python snippet for data extraction and metric formatting:**

```python
import requests
from prometheus_client import CollectorRegistry, Gauge, push_to_gateway

VANTA_GRAPHQL_ENDPOINT = "https://api.vanta.com/graphql"
QUERY = """
query GetControlSummary($framework: ComplianceFramework!) {
controlSummary(framework: $framework) {
totalControls
passedControls
failedControls
notApplicableControls
}
}
"""

def fetch_vanta_metrics(api_key, framework):
headers = {'Authorization': f'Bearer {api_key}'}
variables = {'framework': framework}
response = requests.post(VANTA_GRAPHQL_ENDPOINT, json={'query': QUERY, 'variables': variables}, headers=headers)
data = response.json()['data']['controlSummary']

registry = CollectorRegistry()
g_total = Gauge('vanta_controls_total', 'Total controls', ['framework'], registry=registry)
g_passed = Gauge('vanta_controls_passed', 'Passed controls', ['framework'], registry=registry)
g_failed = Gauge('vanta_controls_failed', 'Failed controls', ['framework'], registry=registry)
g_pass_percent = Gauge('vanta_pass_percentage', 'Percentage of passed controls', ['framework'], registry=registry)

g_total.labels(framework=framework).set(data['totalControls'])
g_passed.labels(framework=framework).set(data['passedControls'])
g_failed.labels(framework=framework).set(data['failedControls'])

pass_pct = (data['passedControls'] / (data['totalControls'] - data['notApplicableControls'])) * 100 if (data['totalControls'] - data['notApplicableControls']) > 0 else 0
g_pass_percent.labels(framework=framework).set(pass_pct)

push_to_gateway('prometheus-pushgateway:9091', job='vanta_metrics', registry=registry)
```

**Insights Gained:**
* We identified a 12% dip in pass rate for a specific framework two weeks ago, which correlated precisely with a major microservice deployment. The failing controls were related to updated log aggregation configurations, allowing us to remediate within 48 hours.
* The "not applicable" controls metric has steadily increased as we've refined our scoping, providing clear data for audit discussions.
* We've set up Grafana alerts to trigger if the pass percentage for any framework drops by more than 5% in a 7-day period.

This approach transforms Vanta from a point-in-time compliance checklist into a quantifiable performance metric integrated into our engineering observability stack. The next phase is to break down failures by control category (e.g., "Access Control", "Data Protection") to further pinpoint systemic weaknesses.

I'm interested if others have attempted similar integrations, particularly if you've found efficient ways to extract more granular control-level history, which remains a challenge with the current API.

-ck



   
Quote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

This is fantastic! I've been thinking about building something similar, but for a different compliance platform. The hardest part for me has always been flattening the nested API responses into something usable for time series.

>to correlate control pass/fail rates with specific engineering deployments

That's the real win. Have you been able to set up alerts that trigger when a specific deployment causes a control failure spike? I'm picturing a Grafana alert that pings the deployment Slack channel automatically.

Also, how are you handling control definitions that Vanta updates? Does your script account for new controls being added to a framework mid-period, or do you just let the pass rate percentage adjust naturally?



   
ReplyQuote
(@charlotteb)
Reputable Member
Joined: 3 months ago
Posts: 323
 

The flattening is definitely the trickiest bit, especially when you're trying to maintain a consistent metric label structure for querying later. For our setup, I ended up creating a separate "control_definition_version" label to track when Vanta updates a control's test logic, which helps avoid false regressions.

On alerts, we do have a few critical controls tied to deployment events. The alert rule itself lives in Grafana, but we use a simple webhook to post to a dedicated Slack channel that both engineering and security monitors. It's not fully automated to ping the deployment channel yet, because we found we needed a human to triage whether it was a real failure or a data lag issue first.

For new controls mid-period, we exclude them from the aggregate pass/fail percentage calculation for that entire reporting period. It felt misleading to let them immediately drag the score down when the team hasn't had any time to implement them. We only include a control after its first full period in our scope. How are you thinking of handling that?



   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

Integrating with the GraphQL API is the right move for this, much more efficient than trying to scrape the web UI. Your approach to pushing flattened metrics into Prometheus is solid for trend analysis.

I have a specific critique on the daily CronJob frequency. For correlating with deployments, a daily snapshot often lacks the necessary temporal granularity. A deployment at 9 AM and a control failure at 10 AM will be conflated into the same data point, muddying causality. I run my collector every six hours, though this does require managing the Vanta API quota more carefully.

Your third processing step mentions key metrics. I hope you're also extracting and preserving the individual control IDs as labels, not just aggregate counts. Without that, you lose the ability to drill down or alert on the failure of a specific critical control, which is where the real operational value is.



   
ReplyQuote
(@harryj)
Reputable Member
Joined: 3 months ago
Posts: 381
 

This is exactly the kind of project I love. We did something similar but pulled data directly into our ticketing system, creating an automatic ticket for any control that flips from pass to fail. It gives us an audit trail and assigns it straight to the right team.

Your method with Prometheus is much better for the historical trends and dashboards though. Have you considered tagging the metrics with the team or service name responsible for each control? That was the key for us to get the right alert to the right Slack channel without manual triage.

I'm curious about the "De..." at the end of your last bullet point. Did you mean to include controls marked as "Deferred"? That's a tricky state to handle.


Automate the boring stuff.


   
ReplyQuote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

That's a great point about the CronJob frequency. Daily snapshots can definitely obscure the sequence of events. The trade-off between granularity and API quota is real, especially if you're pulling data for multiple frameworks or tenants. We settled on an eight hour interval as a compromise, and we only fetch the detailed control statuses for the frameworks we're actively monitoring for deployment correlation, not the entire set.

You're absolutely right about preserving individual control IDs. We do store them as labels, and we found it's also crucial to include the control's category or domain. That way, if a high-level pass rate dips, you can immediately see if it's, for example, a cluster of failures in the "Data Encryption" category versus a single outlier.

Has managing the API quota with a six-hour interval forced you to be selective about which frameworks or data points you query each run?


—HR


   
ReplyQuote