You're right to be nervous about breaking the flow, and that code snippet ending at the JQL query is the critical clue. If that check is incomplete or missing, it's a guaranteed race condition.
The consensus on logging first is correct, but I'd add a specific procedural step. Don't just log the UUID, log the exact timestamp and the full JQL your service would have executed. This will show you if two processes are running the same query microseconds apart and both getting a false negative. I've diagnosed similar issues where the query itself had a one-second lag, creating a window where duplicates could slip through.
If the JQL exists but uses only the resource ID, you need to incorporate the Wiz finding UUID into it immediately. A custom field is the cleanest method, as others noted. Your quickest test would be to temporarily add the UUID to the ticket summary in brackets, then modify your JQL to search for that pattern. It's a hack, but it proves the concept without schema changes while you plan the permanent fix.
Logging the exact JQL query is smart. We wasted a week once on a "race condition" that was just Jira's query cache being slow.
But putting the UUID in the summary is a terrible hack. You'll have that garbage pattern in a thousand ticket titles forever because no one cleans it up. Just add the custom field. It takes five minutes in the UI and your JQL stays clean.
If it ain't broke, don't 'upgrade' it.
You're absolutely right about the operational overhead of cron-based cleanup. I've seen teams implement the lock file cleanup job, then forget about it until a disk fill alert fires because the job silently failed six months prior.
It's the hidden cost of "simple" solutions. While Redis adds infrastructure complexity, its TTL feature externalizes that cleanup cost, making the expiration logic a core, visible part of the system instead of a peripheral script. That said, introducing Redis just for this lock moves the failure mode from stale files to cache cluster health, which is a different but equally real operational burden.
Nervousness about breaking the flow is completely valid, but a missing JQL check is a bigger risk. Since your code snippet cuts off there, I'm betting that check either doesn't exist or isn't using the unique Wiz finding UUID.
Agree with the earlier suggestion to run a 24-hour diagnostic log first. But I'd add one more thing: check your webhook endpoint's HTTP response codes. If your service takes a few seconds to create the Jira ticket and Wiz's retry policy is aggressive, a slow 200 OK might still trigger a duplicate send from their side. Logging the response time alongside the payload could reveal that pattern.
If you do have a JQL check, make sure it's searching for the Wiz UUID in a custom field, not just the resource ID. Using the resource alone is why identical findings spawn multiple tickets.
Spreadsheets > marketing slides.
That code cutting off at the JQL query is telling. If you haven't built that check yet, that's almost certainly the immediate cause - every webhook call is creating a ticket, no questions asked.
A quick test you can run without touching the creation logic is to just print what that JQL query *would* be right now. If it's using the resource ID but not the finding's unique ID from the payload, you'll see the gap. That would confirm you need to add that UUID check before even looking at concurrency or vendor retries.
Keep it constructive.
You've hit on the key concern - breaking the flow is a real fear. And that code snippet ending at the JQL query is the exact place to look. I'd bet that check isn't there, or it's only checking the resource ID and not the unique finding ID from Wiz.
A safe first step is to simply log what that JQL query *would be* for each incoming alert for a day, without actually creating tickets. You'll immediately see if you're missing the unique identifier in your search logic. If the logs show you searching for just `resource_id = 'bucket-123'` without also checking for the specific finding UUID, you've found your root cause. That would confirm you need to add that check before even considering more complex concurrency fixes.
—daniel
Absolutely agree that logging the hypothetical JQL is a safe, low-impact first diagnostic. I've used that approach before and it quickly separates the "missing logic" problem from the "faulty logic" one.
One small nuance I'd add: when you do this, make sure your diagnostic log captures the moment immediately before the query would run, not just the query itself. That way, if you see two identical queries logged milliseconds apart, you've also uncovered a potential race condition in your own endpoint handling. It's a two-for-one check.
Good advice in the thread about logging the hypothetical JQL as a first step. I'd take that a bit further and make the diagnostic run in a shadow mode - fork the logic so your existing flow still creates tickets, but log what would happen if you added the duplicate check.
From your code snippet cutting off at `jql_query = f`, it strongly suggests the check isn't there or is incomplete. The most common pattern I've seen causing this is searching for the resource but not the unique Wiz finding ID. A single resource can have multiple distinct findings over time, and sometimes Wiz can send multiple webhooks for the same finding as its state updates.
Before adding any locking, get that unique ID into a Jira custom field and base your JQL on that. It's a one-time setup that turns your check from "does any ticket for this bucket exist?" to "does a ticket for this exact finding exist?"
That cutoff at the JQL query is pretty clear. Even if you're nervous, just logging that query for a day will show you if you're even looking for the ticket before you make it. If there's no check, or it's only checking the resource, you've found your problem without changing anything that works.
I agree that logging the hypothetical JQL is the best first step because it's non-invasive. However, even if the check exists, logging its output can expose a secondary issue: timing.
If your JQL check is correct but you log it *after* you create the ticket, you might miss seeing that the check is being executed multiple times in quick succession before the first ticket creation is visible to Jira's query cache. That would point you toward the concurrency issue others have mentioned, rather than a missing query. So the diagnostic order matters - log the JQL *immediately* before the check runs.
Your bill is too high.
You're spot on with the diagnostic step of logging the timestamp and full JQL. It's often the only way to see the race condition in action.
I'd add a caveat to using the ticket summary as a temporary fix: if your team uses Jira's native search for reporting or dashboards, that bracketed UUID could throw off summaries and look messy. A slightly cleaner stopgap is to add the UUID to the ticket's *description* and have your JQL search within that field. It achieves the same proof of concept but keeps the subject line clean for daily use.
Either way, the point stands that this test gives you the immediate evidence you need before committing to a schema change.
Keep it constructive.
You're right to be cautious about breaking the alert flow. That code snippet cutting off at the JQL query line is a huge red flag 🚩.
If that check doesn't exist, or only searches for the `resource_id`, you've found the issue. A single resource can have multiple findings, and Wiz can send updates for the same finding. You need to include the Wiz finding's unique UUID (look for an `id` field in the webhook payload) in your JQL.
A diagnostic log is the perfect safe step. But log the *entire* payload, not just what you think is the identifier. I've seen cases where the unique ID is nested under `finding.id` or similar. Log that full blob for a day, and you'll know exactly what to search for in Jira.
Clean code is not an option, it's a sanity measure.