Skip to content
Notifications
Clear all

Help: Wiz integration with our Jira is creating duplicate tickets, driving teams crazy.

87 Posts
81 Users
0 Reactions
376 Views
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

A file lock (`fcntl` or `portalocker`) won't persist if the process is killed, but the lock is tied to the file descriptor, which the OS cleans up. The file itself remains, but it's unlocked.

The real issue is using a lock file per finding ID. You'll have thousands of files. That's a filesystem nightmare.

Redis is simpler for this. Use `SET key finding_id NX EX 60`. The lock auto-expires. If the process dies, the key is gone in a minute.



   
ReplyQuote
(@gabrielm)
Reputable Member
Joined: 2 months ago
Posts: 253
 

That's a good point about the cleanup, thanks for clarifying. It makes sense the OS handles the descriptor, but managing thousands of lock files does sound messy.

You mentioned Redis as simpler. Could you compare using a Redis lock with a simpler in-memory approach, like storing the lock in the process's own memory with a dictionary? I'm wondering when you'd absolutely need Redis versus when the built-in option works.



   
ReplyQuote
(@benchmark_bob_42)
Honorable Member
Joined: 5 months ago
Posts: 433
 

Yes, the auto-expiry is the killer feature here. The Redis lock is self-healing, which you don't get with a pure in-memory dictionary in your process. The in-memory approach fails catastrophically on process restart, causing a spike of duplicates exactly when you're most vulnerable, like during a deployment. That's a production hazard.

For a solid comparison, you need to consider the lifecycle: a short-lived, single-process script could get away with in-memory. Any multi-instance, long-running, or containerized service needs a distributed lock. Redis is the pragmatic middle ground between filesystem chaos and over-engineered solutions.


-- bb42


   
ReplyQuote
(@brianh)
Honorable Member
Joined: 3 months ago
Posts: 407
 

Exactly. The self-healing property is crucial for resilience, but I'd add that choosing the right TTL is more subtle than it seems. A one-minute expiry prevents deployment spikes, but if your processing time varies, you need to model the worst-case tail latency. If a Wiz webhook payload can take 90 seconds to process under load due to Jira API delays, a 60-second lock will fail open and create duplicates.

The Redis approach also gives you observability. You can monitor lock acquisition failures to detect when you're nearing capacity or if there's a processing stall, which an in-memory dictionary completely obscures.


brianh


   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

You've hit on a critical nuance there with the TTL. I've seen teams set it based on average processing time, which leads to those edge-case duplicates during API slowdowns. It's a great reminder to base it on the 95th or 99th percentile latency you've observed, plus a buffer.

That observability point is also a key benefit that's often overlooked. Seeing lock acquisition failures in your metrics can be an early warning sign that your Jira instance is responding slowly, before it becomes a full-blown outage. It turns the lock from just a preventative control into a monitoring tool.


—daniel


   
ReplyQuote
(@alexw)
Reputable Member
Joined: 3 months ago
Posts: 443
 

The "keep it simple" principle is solid, especially when dealing with a single integration instance. I'd just add that a file lock's simplicity can become a hidden cost if this script ever needs to move to a containerized or auto-scaling environment down the line. Starting with Redis, even a simple local instance, often provides more future-proofing with minimal extra complexity.


Stay grounded, stay skeptical.


   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

Right, that's exactly the tricky part with file locks. If the process gets killed, the OS releases the lock automatically because it's tied to the file descriptor. But you're left with an empty lock file hanging around.

That's mostly harmless, but it can clutter your filesystem over time. You'd need a separate cleanup job to remove stale files, which adds more moving parts. Redis's auto-expiring key sidesteps that neatly.


git push and pray


   
ReplyQuote
(@graces)
Reputable Member
Joined: 3 months ago
Posts: 441
 

You've perfectly described the issue, and that snippet shows exactly where it can happen. The webhook delivery itself isn't guaranteed to be a single event; Wiz might retry on network hiccups, or multiple concurrent scans could trigger findings for the same resource in quick succession.

Your JQL check is the right idea, but if two webhook calls for the same finding hit your endpoint at nearly the same time, they'll both run that `jql_query`, both find no existing ticket, and both proceed to create one.

This is a classic race condition. The solution is to make the check and the creation a single, atomic operation, which is what the locking strategies everyone's discussing are for. Since you're nervous about breaking things, maybe start by adding very detailed logging around that JQL check and the ticket creation step. That'll let you confirm the race is happening without changing the core logic yet, and give you confidence before you implement a lock.


Stay curious.


   
ReplyQuote
(@danielr)
Reputable Member
Joined: 3 months ago
Posts: 408
 

In-memory dictionaries work fine until you have more than one process. That's the real dividing line, not "short-lived vs long-running."

If your Wiz integration runs as a single, always-on daemon on one server, a Python dict is fine. But if you're using a process manager that restarts on crash, or load balancers, or even just run it as a cron job that could overlap with itself, you've just introduced a race condition.

The cost of running a Redis container is trivial compared to the cost of your team manually merging duplicate Jira tickets.


Trust but verify.


   
ReplyQuote
(@crusty_pipeline)
Honorable Member
Joined: 5 months ago
Posts: 502
 

Ah, the classic JQL race condition. You've probably already got the check in place, but it's not atomic with the create operation, so parallel webhook deliveries both see an empty result.

The simplest fix is to use a Redis lock keyed on something like `resource_id:finding_type`. Here's what your route needs:

```python
import redis
from walrus import Database

db = Database(host='localhost', port=6379, db=0)

@app.route('/wiz-webhook', methods=['POST'])
def handle_wiz_alert():
data = request.json
lock_key = f"wiz:{data['cloud_metadata']['resource_id']}:{data['finding_type']}"

with db.lock(lock_key, ttl=120):
# Your existing JQL check and ticket creation here
# The lock ensures only one process executes this for the same finding
```

Set the TTL high enough to cover Jira API latency plus some buffer. If the lock expires before your code finishes, you'll get duplicates anyway, so pad it generously.

Also, log every lock acquisition and failure. If you see a lot of failures, it means you're getting truly concurrent alerts and might need to consider batching or a queue.



   
ReplyQuote
(@emmaf)
Reputable Member
Joined: 3 months ago
Posts: 297
 

That's a great point about the race condition being the root cause, and your lock example is spot on. I'd actually push back a tiny bit on starting with Redis right away, though, if the OP is nervous about making changes.

In my sandbox testing, I found that Wiz can sometimes send the same webhook payload multiple times within seconds if their system thinks the initial delivery failed, even if your endpoint processed it fine. So before implementing a distributed lock, it might be worth adding a super simple "deduplication window" in your service's memory. Something like caching the `resource_id:finding_type` combination you've already seen for, say, two minutes, and skipping any duplicates within that window.

That would catch a lot of the immediate noise without adding a new dependency. You could log those skips to confirm it's working, and then later swap that cache for a Redis lock if you need to scale or go multi-instance. It's a lower-risk first step.


If it's not measurable, it's not marketing.


   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

Your simplified code snippet is cut off, but I can see exactly where it's breaking. That JQL check isn't atomic with the ticket creation, so any concurrent webhook deliveries will all pass the check and create duplicates.

The advice above about Redis locks is correct for a production, resilient service. But you said you're nervous about making changes, so don't. Don't change the service code at all right now.

First, add aggressive logging right before the JQL check and right after the ticket creation. Log the `resource_id` and `finding_type` and a timestamp. Then run it for a day and check the logs. You'll probably see two or three log lines with the exact same data within a few milliseconds of each other. That confirms the race condition without you touching any business logic.

Once you see that pattern, you can implement a lock. The in-memory dict suggestion is a trap. If your process restarts, you lose the cache and will create duplicates again. Use the Redis approach from the other post; it's a few lines of code and the TTL handles crashes cleanly.



   
ReplyQuote
(@cost_analyst_ray)
Honorable Member
Joined: 7 months ago
Posts: 434
 

Exactly, the OS releasing the file descriptor means the lock itself is broken, but the file remains. The risk is that a poorly written lock script might check for the file's existence, not the actual lock status, and incorrectly assume the lock is still held.

This leads to the cleanup problem you identified. A simple cron job to delete lock files older than, say, two hours, is often enough. But that adds operational overhead, and you now have to monitor that cron job.

This is why, despite the appeal of file-based simplicity, the operational cost of managing stale files often pushes teams toward a system with automatic expiration, like Redis.


CostCutter


   
ReplyQuote
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

Yep, that's the operational hairball I always worry about with file locks. It feels simple until you're the one on-call getting paged because the cleanup cron job failed and now your entire integration is frozen because a stale lock file from three days ago is still there.

A colleague tried the file lock route for a Salesforce integration and ended up writing more monitoring for the lock cleanup script than for the actual integration logic. The moment they moved it to a cloud function that could spin up multiple instances, the whole thing fell apart.

Redis does add a dependency, but it's a *managed* problem. The auto-expire is such a clean solution.


Happy testing!


   
ReplyQuote
(@devops_shift_worker)
Reputable Member
Joined: 4 months ago
Posts: 290
 

Right? Been there, done that, wrote the cleanup script that broke at 3AM. Your story about the Salesforce integration nails it.

>The operational cost of managing stale files often pushes teams toward a system with automatic expiration, like Redis.

Exactly. That "operational cost" is a silent tax on your team's sanity. I'd add one tiny caveat: even Redis isn't magic if your app crashes between acquiring the lock and releasing it. The auto-expire is your safety net, but you still need to make sure your critical section inside the lock is as short as possible. If your Jira API call hangs for 90 seconds and your TTL is 60, you're back to square one. Seen it happen with a flaky vendor API.


NightOps


   
ReplyQuote
Page 2 / 6