Skip to content
Notifications
Clear all

Help: Wiz integration with our Jira is creating duplicate tickets, driving teams crazy.

87 Posts
81 Users
0 Reactions
373 Views
(@data_pipeline_rookie_42)
Reputable Member
Joined: 5 months ago
Posts: 237
Topic starter   [#22319]

Hi everyone. I'm relatively new to the data engineering side of things and I've been tasked with helping to manage our cloud security alerting pipeline, which uses Wiz. We have it set up to automatically create Jira tickets for high-severity findings so our security and dev teams can track remediation.

The problem is that we're seeing a huge number of duplicate tickets. For example, a single finding about an overly permissive storage bucket might spawn 3 or 4 identical Jira issues within minutes. It's creating a lot of noise and the teams are starting to ignore the queue, which is obviously bad.

Our integration is pretty much out-of-the-box. We're using the Wiz webhook to push to a small Python service that then uses the Jira API to create the ticket. I'm nervous about making changes here because I don't want to break the flow and miss a critical alert.

Here's a simplified version of our listener logic (the real one has more error handling):

```python
@app.route('/wiz-webhook', methods=['POST'])
def handle_wiz_alert():
data = request.json
issue_key = data['cloud_metadata']['resource_id']

# Check if a ticket for this resource+finding already exists
jql_query = f'project = SEC AND labels = wiz_{issue_key}'
existing = jira_client.search_issues(jql_query)

if len(existing) == 0:
# Create new ticket
issue_dict = {
'project': {'key': 'SEC'},
'summary': data['finding']['title'],
'description': data['finding']['description'],
'labels': [f"wiz_{issue_key}"]
}
new_issue = jira_client.create_issue(fields=issue_dict)
```

From what I can tell, the webhook might be firing multiple times for the same event. Should we be deduplicating on the Wiz side before the webhook is sent, or is our service's deduplication logic flawed? Maybe we're using the wrong unique identifier (`resource_id` doesn't seem to be enough if it's a finding that can reappear).

Has anyone else run into this? What's a safe, reliable pattern to ensure one Jira ticket per unique *finding instance*, without dropping real new alerts?



   
Quote
(@benchmark_basher)
Reputable Member
Joined: 4 months ago
Posts: 312
 

That JQL query check is your first problem. You're likely checking for an existing ticket *before* creating the new one, but you have no locking mechanism on the webhook endpoint itself. If Wiz sends two webhook calls for the same finding within a second, you'll have two requests pass that check and create two tickets.

Your Python service needs to handle concurrent duplicate detection. Add a simple in-memory cache (like a dict) of recently processed `resource_id` + finding type combinations with a short TTL, and reject duplicates instantly. Or, if you must query Jira, use a mutex or a distributed lock if your service scales.


-- bb


   
ReplyQuote
(@devops_barbarian_v3)
Honorable Member
Joined: 6 months ago
Posts: 403
 

That in-memory cache is a band-aid that'll fall off the second you scale to two pods. Wiz webhooks can be weirdly bursty too, a TTL might not save you.

Better to make the ticket creation itself idempotent. Use the Wiz finding ID as part of the Jira ticket key or stash it in a custom field. Then your "check before create" JQL can actually be reliable because the finding ID is the source of truth, not a timestamp race.

Or just let it create dupes and have a downstream cleanup job. Sometimes it's easier to embrace the chaos and fix it later.



   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

Idempotency is the right call, but using the finding ID as part of the Jira key is asking for trouble with Jira's own key formatting rules. Put it in a custom field and make that field unique.

"Embrace the chaos" is terrible advice for a security alerting pipeline. Ignored tickets due to noise is how breaches happen. You don't fix that later, you fix it now.


If it's not a retention curve, I don't care.


   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 462
 

That truncated JQL check is the heart of the problem. You're checking Jira *right before* you create, but as others said, two webhooks hitting at the same time will both pass that check.

Here's a solid middle ground: keep your JQL check but make it *much* more specific. Use the Wiz finding ID, not just the resource ID. Then, add a quick in-memory lock (like a threading lock or a Redis key if you're multi-instance) for that specific finding ID right before the JQL query. This prevents the race condition.

I've done exactly this with Optimizely webhooks, and it's stable. The lock ensures only one request per finding ID can even get to the check-create stage.


✌️


   
ReplyQuote
(@averyt)
Reputable Member
Joined: 2 months ago
Posts: 274
 

I like this approach a lot. The key is moving the lock *before* the JQL check, as you said. It stops the race condition at the door.

One tiny caveat from experience with a similar setup: if your lock TTL is too short and the Jira API call hangs for any reason, a second webhook could slip in after the lock expires but before the first ticket is actually created. Setting the lock to last for, say, 60 seconds (way longer than any API call should take) gives you a safe buffer against timeouts.

Redis is perfect for this, especially if you're already in a cloud environment. It turns that "in-memory lock" into a simple distributed lock that'll work across any number of service instances.


Automate all the things


   
ReplyQuote
(@devops_contrarian_42)
Honorable Member
Joined: 6 months ago
Posts: 479
 

"our of the box" is your first clue. Vendor integrations are rarely tuned for actual production use.

You've already got the right diagnosis - it's a race condition on the JQL check. The lock idea before the check works, but you don't need Redis for it. A file lock on a single instance is fine. You're not running a global scale security platform, you're forwarding some alerts. Keep it simple.

Just make sure your lock scope is the finding ID, not just the resource ID. And handle lock acquisition failure by logging and dropping the dupe webhook, don't retry.


Keep it simple


   
ReplyQuote
(@clara12)
Estimable Member
Joined: 3 months ago
Posts: 210
 

I agree that simplicity is a virtue here, and a file lock could be a good fit for a single instance. My hesitation is around operational resilience, though, rather than scale.

If that single service instance restarts or crashes, which can happen during deployments or unexpected failures, doesn't the file lock approach lose its state entirely? A subsequent burst of webhooks after a restart would then be vulnerable to the original race condition again, at least until the cache repopulates. Using a finding ID as a lock key in a system with some persistence, even a simple one, might guard against that brief window.

Would you consider that a significant risk, or is the expectation that the service restart would be a rare enough event that the occasional duplicate post-recovery is an acceptable trade-off for the architectural simplicity?



   
ReplyQuote
(@amelia2)
Reputable Member
Joined: 3 months ago
Posts: 261
 

Missing the end of your code snippet, but I can guess the rest.

The jql check isn't enough. You're right to be nervous about breaking it, but the noise is already breaking it. Start with the lock before the check, like the others said.

If you're a single instance, use a threading.Lock per finding_id. If you might scale, a Redis lock is trivial to add. Something like this right after your route starts:

```python
lock_key = f"wiz_lock:{data['id']}"
if not redis.set(lock_key, "1", nx=True, ex=30):
return "Duplicate", 200
```

Then do your JQL check and create. Delete the lock after. It's a small change and won't drop alerts.


Ship it, but test it first


   
ReplyQuote
(@davek)
Reputable Member
Joined: 2 months ago
Posts: 281
 

Your simplified code snippet is exactly where the race condition happens. The other posters are correct about needing a lock before the JQL check, but there's a nuance they haven't mentioned: you also need to ensure your JQL query uses a truly immutable, unique identifier from Wiz. The `resource_id` alone isn't sufficient. A single resource can have multiple findings of the same type, and a specific finding can be updated, triggering a new webhook for the same underlying issue.

You must use the Wiz finding `id` field (a UUID) in your JQL, stored in a custom field like `customfield_10100`. Your check should look for `'cf[10100] = "abcd-1234..."'`. This guarantees you're tracking the exact finding instance, not just a resource.

The lock (in-memory, file, or Redis) should key on that same UUID. This gives you two layers: the lock prevents concurrent processing for the same UUID, and the JQL check is a final idempotent guard. If you go with a file lock for simplicity, pair it with a small SQLite database or even a text file log to track processed UUIDs for, say, the last 24 hours. This survives service restarts and prevents duplicates during that vulnerable restart window user885 mentioned.


CPU cycles matter


   
ReplyQuote
(@infra_architect_rebel)
Honorable Member
Joined: 5 months ago
Posts: 544
 

Agree on the UUID, that's the only field you should key on. But the SQLite/text file log is overkill for a restart window.

If you're worried about state loss, extend the lock TTL. A 5-minute Redis lock covers deploy restarts and still fails clean. Adding a second persistence layer just creates two problems.


Simplicity is the ultimate sophistication


   
ReplyQuote
(@elliek2)
Reputable Member
Joined: 3 months ago
Posts: 355
 

That makes a lot of sense. So the five-minute lock TTL is basically a safety net for the restart scenario, and Redis handles the expiry for you cleanly.

I guess my only follow-up is, doesn't extending the TTL that long mean you could potentially block a *legitimate* second webhook? Like, if someone manually closed a Jira ticket and then re-ran the scan in Wiz immediately, would the new finding be blocked for five minutes? Or is that actually the desired behavior?



   
ReplyQuote
(@elenar)
Reputable Member
Joined: 3 months ago
Posts: 293
 

You're right to be cautious about breaking a critical alert flow. However, the current state where teams ignore the queue due to noise *is* a broken state - it's just a silent, procedural failure instead of a technical one. The risk of a missed critical alert is already present.

Your nervousness suggests a good operational instinct. Mitigate it by implementing the lock and UUID check as a phased, observable change. Start by adding detailed logging for every lock acquisition, check, and creation event. This doesn't alter the core logic initially. Then, deploy the lock mechanism but have it run in a "log-only" mode for a short period, comparing what it *would* have blocked against what actually gets created. This gives you a validation window to confirm the logic catches duplicates without dropping anything new. Only then flip it to actively block.

This transforms the change from a "big bang" fix into a measurable data pipeline adjustment, which should align better with your data engineering mindset.


Data doesn't lie, but folks sometimes do.


   
ReplyQuote
(@devops_rookie_2025)
Prominent Member
Joined: 4 months ago
Posts: 467
 

Thanks for sharing your actual code snippet! I'm learning a ton from this thread.

> I'm nervous about making changes here

That's totally fair, I'd be nervous too. user667's idea about a "log-only" mode first is brilliant - you could add the lock logic but just log "would block duplicate for finding ID: X" instead of actually stopping it. Run that for a day and you'll see exactly how many dupes it catches before you flip the switch. That might help with the confidence.

A quick question for the experts here: if you're using the Wiz finding UUID as the lock key and for the JQL check, wouldn't a re-scan for the same issue generate a completely new UUID? So a long lock TTL shouldn't block a legitimate new finding, right?



   
ReplyQuote
(@emmae)
Reputable Member
Joined: 3 months ago
Posts: 255
 

>Vendor integrations are rarely tuned for actual production use.

That's honestly a relief to hear, because I was worried I was missing something obvious! 😅 I keep forgetting that the vendor demo version isn't the same as real-life chaos.

I like the file lock idea for simplicity, but what happens if the process gets killed while it's holding the lock? Would the lock file just stay there and block everything for that finding ID until someone cleans it up, or does it get released automatically?



   
ReplyQuote
Page 1 / 6