Skip to content
Notifications
Clear all

Help: Wiz integration with our Jira is creating duplicate tickets, driving teams crazy.

87 Posts
81 Users
0 Reactions
378 Views
(@code_weaver_max)
Reputable Member
Joined: 4 months ago
Posts: 370
 

Totally get the nervousness about breaking the alerting flow. That JQL check is so close! The snippet cuts off, but I bet the race condition is in the gap between the query and the `jira.create_issue()` call.

Since you're new to this, a safe first step is to add a short-lived memory cache, like others said. But you can key it on something from the webhook itself to avoid even hitting Jira for duplicates. Something like:

```python
from cachetools import TTLCache

seen_cache = TTLCache(maxsize=1000, ttl=120) # 2 minute memory

@app.route('/wiz-webhook', methods=['POST'])
def handle_wiz_alert():
data = request.json
dedupe_key = f"{data['id']}" # Use the Wiz finding ID if available

if dedupe_key in seen_cache:
app.logger.info(f"Skipping duplicate webhook for {dedupe_key}")
return
seen_cache[dedupe_key] = True
# ... rest of your JQL check and create logic
```

This will squash the immediate retries without touching your Jira logic at all. It's a quick band-aid to stop the bleeding while you plan the Redis lock.


Prompt engineering is the new debugging


   
ReplyQuote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

You're absolutely right about the process boundary being the critical factor. A Python dictionary's volatility isn't just about crashes; it's about any form of horizontal scaling. Even using a process manager like Gunicorn with multiple workers, which is a default for many deployments, instantly voids any in-memory deduplication.

The "trivial cost" comparison is spot on, but I'd frame it as a risk calculation. The probability of a duplicate ticket might seem low until your traffic spikes or you have a brief network partition that triggers retries. Then you're debugging a flood of duplicates while your team's trust in the automation evaporates. The operational burden of managing a Redis instance is predictable, while the cost of a duplicate ticket incident is chaotic and multiplicative.


p-value < 0.05 or bust


   
ReplyQuote
(@infra_auditor_nina)
Honorable Member
Joined: 6 months ago
Posts: 467
 

> the real one has more error handling

I'll bet my last audit report it doesn't have enough *logic* handling. That code snippet cuts off right at the JQL query. Show us the actual query. My money's on it searching for the resource ID but not the finding type, or using a created date filter that's too narrow.

Everyone's jumping to distributed locks, which treats the symptom. Fix the query first. If your check can't reliably find the ticket it created 500ms ago, you've built a system that's fundamentally unreliable, lock or no lock.


- Nina


   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Oh, exactly right. >what happens if the process gets killed while it's holding the lock? That's the rub. The OS releases the file descriptor, so the *lock* is gone, but the lock *file* is still sitting there.

Then your script's next check - if it's just looking for the file's existence - sees it and assumes the lock is still held. Instant deadlock until someone SSHes in and rm's it. Been there at 2 AM, not fun. That's why you see folks suggesting a cron cleanup job, which just trades one operational headache for another.

It feels simple until you're the one who gets paged for the frozen integration.


it worked on my machine


   
ReplyQuote
(@heidir33)
Reputable Member
Joined: 3 months ago
Posts: 270
 

That's a really good point about the JQL query itself being the source of unreliability. Even with a perfect lock, if the check can't find the ticket it just made, the problem just shifts.

Before diving into locks, I'd want to audit the webhook payload and the actual query logic. I've seen similar issues where the deduplication check uses a field that Wiz can send in slightly different formats, like a resource ARN with and without a path suffix. Or maybe the query looks for an exact title match, but the title includes a timestamp that changes slightly between rapid-fire webhooks.

Could you share a bit more about what fields you're using in the JQL to identify a duplicate? Specifically, are you using the Wiz finding `id`, the resource identifier, and maybe the finding type? Sometimes the out-of-the-box templates miss a key immutable field.



   
ReplyQuote
(@integrations_jane)
Reputable Member
Joined: 5 months ago
Posts: 319
 

That snippet cutting off mid-JQL is the entire story. You're running the classic "check then act" race condition, and a webhook can fire multiple times for the same event faster than your script can complete one cycle.

Don't even think about locks yet. First, fix the atomicity. Your JQL check and the subsequent `create_issue` need to be a single, idempotent operation. The Jira API supports this if you use the `updateOrCreate` action, or you can structure your ticket creation logic to be inherently idempotent by using the Wiz finding ID as the Jira issue key or in a custom field you can enforce uniqueness on.

I've seen this exact pattern with cloud security webhooks. The vendor often sends multiple "status update" pings for the same finding, and your naive listener treats each as a new event. Your dedupe key should be a composite of the Wiz finding `id` and the `severity` or `status`, not just the resource ID.


APIs are not magic.


   
ReplyQuote
(@gracek)
Reputable Member
Joined: 3 months ago
Posts: 200
 

Finally, someone cuts through the distributed lock cargo cult. Idempotency is the only sane foundation here.

>using the Wiz finding ID as the Jira issue key or in a custom field you can enforce uniqueness on

This is the real answer, but you're underselling the politics. Trying to use a custom field for uniqueness means begging your Jira admins to add an index, which in most orgs is a project requiring three approvals and a change advisory board. Using the issue key is clever, but good luck convincing the security team their ticket numbers will look like `CLD-12345-WZ-abcd1234`.

The `updateOrCreate` approach is the pragmatic middle ground, but I've watched teams fail because they treat it as a silver bullet. If your mapping logic from Wiz payload to Jira fields is sloppy, you'll still get duplicates, they'll just be updates to the wrong ticket. You fix the race condition but introduce a data integrity mess.

Your last point about the composite key is the hidden trap. The moment you key off `severity` or `status`, a re-evaluation in Wiz that changes either one becomes a brand new ticket in Jira. Now you've got ticket sprawl. The real question is whether your process wants a new ticket for a severity change, or just an update to the existing one. Nobody figures that out until after the integration is live.



   
ReplyQuote
(@emilyl2)
Reputable Member
Joined: 2 months ago
Posts: 219
 

Okay, so your listener is checking for an existing ticket but then still creating one anyway? That's the part that cuts off, right? I'm new to this too, but reading the other replies, the JQL check might be failing to find the ticket it just made. What are you actually putting in the JQL query? If it's looking for the resource ID, maybe Wiz sends the same resource in slightly different ways across webhook pings?



   
ReplyQuote
(@annab)
Reputable Member
Joined: 3 months ago
Posts: 349
 

I actually hadn't considered extending the TTL like that, but it makes a lot of sense for covering deploys. The idea of a second persistence layer adding complexity clicks for me - it's another thing that could break.

But doesn't that just mean the lock is held for a longer period of false failure? If a restart happens right after a lock is acquired, the process is down but the lock is still there for five minutes, blocking any new webhooks that come in for that finding. Is the trade-off of missing some real alerts for a few minutes acceptable, or is there a way to mitigate that?



   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 363
 

Yeah, that's a good call. I've been looking at our own setup and you're right about the ARN format. Sometimes it's just the resource ID, other times it's the full ARN with the region and account. If your JQL is doing an exact match, that'll miss it.

Are you checking the Wiz finding ID itself? That should be the immutable one, right?



   
ReplyQuote
(@harukik)
Honorable Member
Joined: 3 months ago
Posts: 400
 

That's a really solid point about the API call hanging longer than the TTL. I hadn't thought about that. So even with Redis, you're still vulnerable if the external service you're calling (like Jira) has performance issues.

Does that mean the only real solution is to make the call outside the lock, maybe queue it somehow? Or is that overcomplicating things?



   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Right on about using the Wiz finding ID as the lock key, that's the way to go. Your point about pairing a lock with a persistence layer for restarts is clever, but I'd be a little wary of the text file log approach. Managing state in a file introduces its own headaches around concurrent writes and cleanup.

For a lightweight setup, I've had good luck using a simple Redis SET with an EXPIRE to act as that "processed in the last 24 hours" cache. The lock (maybe a Redlock) keys on the UUID, and the SET check is a cheap idempotent guard after the lock is acquired. If the service restarts, the SET has already been written and survives for the expire window. It's one less moving part than managing log files.


ship it


   
ReplyQuote
(@danielm)
Honorable Member
Joined: 2 months ago
Posts: 453
 

Your code snippet cutting off at the JQL query is hilariously perfect. It's the ultimate race condition metaphor.

The real trap in that logic is assuming the webhook fires once per finding. It doesn't. Wiz, like most vendors, sends multiple events as a finding's state changes (detected, confirmed, maybe a status ping). Your script sees each POST as a brand new command to create a ticket.

You're right to be nervous about making changes to a critical alert flow, but ignoring this is guaranteeing alert fatigue. The teams are already ignoring the queue. The biggest risk isn't your code breaking, it's the process being broken because everyone starts manually filtering the noise.


— skeptical but fair


   
ReplyQuote
(@hannahb)
Reputable Member
Joined: 3 months ago
Posts: 261
 

Yeah, the chaos advice made me raise an eyebrow too. But isn't adding a custom field and making it unique also a politics thing? I'm asking because we ran into that - our Jira admins said adding a unique constraint needs a full schema change review, which takes weeks. Did you find a way around that, or is it just a necessary fight to have?



   
ReplyQuote
(@gracec)
Reputable Member
Joined: 3 months ago
Posts: 315
 

That cutoff in the code snippet is a perfect snapshot of the problem - the JQL check is failing, so the service just keeps creating tickets.

You mentioned being nervous about breaking the flow, but right now the flow is already broken by the duplicates. The teams ignoring the queue is the biggest risk. I'd focus on making your JQL query bulletproof first. Like others have said, use the Wiz finding ID, not the resource ID. The resource ID format can vary, but the finding ID is a stable UUID designed for this exact purpose - tracking a specific instance of a finding.

Can you check if your incoming webhook payload actually contains that `finding.id` field and if you're using it in your current logic? That small change might resolve 90% of your duplicates without a major architectural overhaul.


The right tool saves a thousand meetings.


   
ReplyQuote
Page 3 / 6