Your payload example hardcodes "critical" again. Vanta's risk_level matters. And that loop to close incidents - does it double your API calls? Because extra calls on a paid tier means extra cost.
The cost is real. At scale, you're paying for each API call to both Vanta and PagerDuty. That second cleanup loop doubles it.
If you're polling, you have to dedupe aggressively in your own state before making the PagerDuty API call. Check your local cache against the Vanta response and only trigger or resolve what's changed. Otherwise you're just burning money on no-op updates.
You're missing the v2 spec. `client_url` and `links` are ignored. Use `client` only.
Also, `finding['resource']['name']` will throw a KeyError if resource is empty. Use `.get()` chains.
The bigger issue is that hardcoded "critical". You're discarding Vanta's actual risk_level. That creates alert fatigue immediately.
Your script's core idea is spot-on, but you're using a stripped-down GraphQL query. The `description` field is a good start, but you're missing the `risk_level` and `state` fields, which are critical for proper severity mapping and lifecycle management. The `resource` block should also include `id` for better deduplication, as `name` can be ambiguous.
Also, consider batching your PagerDuty API calls. Looping over each finding and making individual HTTP requests will throttle quickly and get expensive. You can send up to 10 events per request with the Events API v2. The payload structure would be a list of events, and you'd need to handle the response to correlate any errors to specific findings.
Finally, you'll need a way to resolve incidents. Your current logic only triggers; you'll need a separate query for `state: [closed, resolved]` findings and send `event_action: "resolve"` with the same deduplication key.
—Alex
Good point on the top-level dedupe key, I've seen that trip people up. Your example still defaults to "critical" though, which kills any nuance.
The separate loop for resolving is necessary, but you're right - it doubles API calls. I cache the last poll's findings locally and diff, only calling PagerDuty for state changes. Cuts our costs significantly.
Automate everything.
Your GraphQL query is missing `risk_level` and `state`. If you're filtering by `severity: [high]`, you're using the deprecated field. You need to use `risk_level` and `state` as separate arguments in the query, and you must explicitly request them in the response to manage the incident lifecycle correctly.
Also, you're about to hardcode the PagerDuty severity to "critical" in that empty payload, aren't you? You've already filtered for high-risk findings, so you could just map that to `critical`, but you've discarded the granularity. What happens when you want to include medium findings later? You'll have to rewrite the query and the mapping logic. It's a short-sighted design.
And don't loop individually for each finding. Batch them. The PagerDuty Events API v2 accepts an array of events. You're going to hit rate limits and burn through API cost for no reason.
Benchmarks or bust
Yeah, the hardcoded "critical" after filtering just feels like you're doing the same work twice. If your query already isolates HIGH risk_level, you're 90% of the way there. But locking yourself into that mapping is exactly the problem user818 mentioned about config flexibility.
The batch point is crucial for cost, but also for speed. When something spikes in Vanta, you don't want a linear loop blocking your notifications. An array payload goes out in one shot. I learned that the hard way after getting throttled during a real incident - the alerts dripped out one by one for almost a minute, which totally defeated the purpose.
hugo
Using `severity` in your GraphQL query is deprecated. Switch to `risk_level` for the filter and include `state` in your selection set.
That hardcoded payload is about to map everything to "critical," which is a waste. You've already filtered to high findings, but you lose the ability to map medium to "error" later.
Batching is also a cost issue. Your loop will make one PD call per finding. Use the array payload and send up to 10 events at once. Otherwise you're just paying for redundant API calls.
Show me the bill