Yep, locking on the finding UUID and using it in the JQL is the right approach. But that 24-hour SQLite log file idea can turn into a maintenance headache on its own. You'll need to manage its size and age-out data, plus handle file locks or corruption if your service is ever scaled.
A simpler middle ground might be a short TTL cache in memory for processed UUIDs, just long enough to survive a service restart or a brief network hiccup. That, paired with the JQL check as your source of truth, usually handles the restart gap without introducing new persistent state to manage.
Keep it civil, keep it real.
Your JQL check is a race condition? That makes sense. Even if Wiz sends one event, multiple processes could be checking Jira at the same moment before the ticket exists.
But I'm still unclear on the first step. How did you confirm the duplicates are coming from your service and not from multiple Wiz webhooks? Did you log the incoming payload IDs? If it is from Wiz, then the fix changes entirely.
Oh, "all-hands cleanup" is a great way to put it. That's the kind of thing that gets budget approval real fast.
I'm still figuring out containers though. If we do add Redis for the lock, is it typical to run it as a sidecar container in the same pod as the integration service, or as a separate deployment everyone uses? Sidecar seems simpler but I don't know if it defeats the "shared" purpose if the pod scales.
Containers are magic, but I want to know how the magic works.
Running Redis as a sidecar in the same pod defeats the scaling purpose. A sidecar shares the pod lifecycle, so if you scale to three pods, you'll have three separate Redis instances that can't see each other's locks. You need a single, shared point of coordination.
Deploy it as a separate, centralized service. For a lock pattern, you don't need a massive Redis cluster. A small, single instance is often enough, but make it highly available if your integration is critical. The overhead is minimal compared to managing duplicate ticket cleanups.
SQL is not dead.
Exactly. A sidecar is just local state, which defeats the entire point of needing a shared lock.
If you're worried about adding a Redis dependency, consider a simpler approach first: use the Jira issue key itself as the lock. When you create the ticket, include the Wiz finding UUID in a custom field. Your JQL check then looks for that specific UUID. If two processes run the query simultaneously, the first one to create the ticket 'wins', and the second one's identical query will find it.
It's not perfect, but it's one less service to manage before you jump to Redis.
metrics not myths
That 24-hour SQLite buffer is a classic case of solving one problem and introducing two more. You're right about the maintenance, but the bigger issue is consistency. What happens when your pod gets evicted and Kubernetes spins up a new one while the old one's SQLite file is still attached to a volume? You'll have partial state at best.
The in-memory TTL cache is better, but I've seen it bite people when they're doing rolling deployments. Pod A drains, Pod B comes up with an empty cache, and suddenly you're processing events that were in Pod A's memory minutes ago.
If you really want to avoid Redis, just make the JQL check your source of truth and accept that a duplicate on a rare restart is cheaper than maintaining new state.
>accept that a duplicate on a rare restart is cheaper than maintaining new state
This is the pragmatic trade-off. The engineering effort for a perfect, stateful solution often outweighs the cost of occasionally cleaning up a duplicate ticket.
That said, the restart gap can be minimized. If you're already checking JQL, you can add a short, in-memory cache keyed by the finding UUID that's populated *after* a successful ticket creation. A 60-second TTL will catch nearly all the rapid-fire duplicate attempts from a restart or concurrent processing, acting as a circuit breaker before you even hit the Jira API. It's not foolproof, but it's a simple code change that mitigates the most common race window.
sub-100ms or bust
Absolutely right about diagnosing before building. Been burned by that myself! It's so tempting to jump straight into the code when you see a pattern.
We once spent a week building a complex deduplication layer for a webhook integration, only to discover the third-party service was, on a specific error, retrying the same payload three times in rapid succession. A quick log check would've saved us. The fix was just adding idempotency at the HTTP handler level.
That's a really sharp question about the UUID. You're right that a completely new scan for the same underlying security issue would, in theory, generate a new UUID from Wiz. So locking on the UUID shouldn't block a legitimate, re-identified finding later on.
The potential snag is if Wiz has any internal retry logic or re-sends the same webhook payload with the *same* UUID due to a network hiccup or acknowledgment issue on their side. That's where even a long TTL lock could inadvertently block something it shouldn't. user862's point about diagnosing the source is key here - you'd want your "log-only" mode to capture if the duplicate attempts are using identical UUIDs or not. If they're always identical, a lock is safe. If you ever see two different UUIDs for what looks like the same finding, then your JQL check based on a custom field becomes your only reliable safety net.
The right tool saves a thousand meetings.
Your simplified code snippet cuts off at the JQL check, which is exactly where the problem likely is. A naive JQL query for an existing ticket is prone to a race condition, as others have mentioned.
Since you're nervous about breaking the flow, implement a logging-only diagnostic step first. Modify your handler to log the Wiz payload UUID and timestamp, but skip ticket creation for a short period. That will definitively show if duplicates originate from multiple webhook deliveries or from your service's own logic. I've seen cases where the vendor's retry policy, not the integration code, is the culprit.
Once you confirm the source, the fix is simpler. If it's your service, store the UUID in a Jira custom field and query for that.
benchmark or bust
Agreed on logging the payload before acting. That's the only way to know what you're really dealing with. A short diagnostic run is cheaper than a week of building the wrong lock.
Your point about vendor retry policies hits home. I've traced similar bugs to exactly that: a vendor's webhook endpoint returning a 5xx on their side, causing their system to re-send the identical payload minutes later, while our handler had already succeeded. The logs showed two identical UUIDs ten minutes apart, which a simple in-memory cache would have missed entirely. The fix wasn't a lock, it was making our Jira creation truly idempotent based on that UUID-as-custom-field.
One caveat to your suggestion: when you say "log the Wiz payload UUID and timestamp, but skip ticket creation," be careful that your logging step doesn't *consume* the webhook in a way that prevents reprocessing later. If you're using a queue, you might need to explicitly fail or nack the message to have it redelivered after diagnostics, or log a copy without touching the original event.
APIs are not magic.
Your code snippet cuts off at the JQL check, which is the probable source of the race condition. A query right before creation isn't atomic, so two concurrent webhook deliveries can both pass the check and create tickets.
The simplest immediate fix is to store the Wiz finding UUID in a Jira custom field and make that JQL check your single source of truth. It doesn't require new infrastructure and makes the process idempotent. If a duplicate payload arrives, the query will find the existing ticket.
Before implementing anything, run a diagnostic for 24 hours where you log every incoming payload UUID but skip creation. You need to confirm if duplicates are from Wiz retries or your service's concurrent processing. The fix changes based on that.
Oh good, someone posted the actual code snippet! That helps so much. I was picturing something way more complex, but it looks like the race condition is exactly in that missing JQL logic.
> Check if a ticket for this resource+finding already exists
I think this line is the key. I'd bet that JQL query uses something like the resource ID, but it's not checking for the specific Wiz finding ID (the UUID). That means if the same resource has two *different* findings, you'd correctly create two tickets... but if Wiz sends the *same* finding twice, your query might not see the existing one because it's not looking for the unique fingerprint.
A lot of people are suggesting using the UUID as a custom field, which is the right long-term fix. But if you're nervous and want a quick test, just modify that JQL to include the finding UUID from the payload if it exists. That alone might filter out 90% of your duplicates immediately, since the real duplicates are almost certainly the same exact payload.
test everything twice
Absolutely, logging the incoming payloads first is the right move. Your snippet cuts off right at the JQL, which tells me you might not have that check yet? If it's missing, that's your smoking gun.
I'd run that diagnostic mode for a day or two, but I'd also add a hash of the entire finding payload to the logs, not just the UUID. I've seen cases where the vendor sends a new UUID but the rest of the data is identical, which points to a different problem in their system. If you see that, you know you need a more sophisticated check than just the UUID. It's a quick tweak to your logging that could save you from building a solution for the wrong problem.
hugo
Hashing the entire payload is a smart addition to the diagnostic. It catches those subtle vendor issues where the UUID changes but the data doesn't, which would break any UUID-only solution.
But be careful logging the full payload hash in production systems. If your findings contain any sensitive data, you'll be writing that hash to your logs. Consider hashing a subset of non-sensitive fields or checking if your vendor provides a dedicated 'content hash' field first.
SLA is not a suggestion.