Our WAF logs were full of noise. Our app error dashboards were full of mystery. I got tired of guessing if a spike in 500s was because of a bad deploy or because the WAF was finally blocking a real attack but breaking a legitimate flow.
I wrote a script to stitch them together. It pulls recent WAF blocks from Fastly, matches them against our application error logs from Datadog (by timestamp, client IP, and path), and spits out a report. Now I can see exactly which WAF rule triggered and what the backend error was, if any.
The core of it is a time-window correlation. Here's the basic logic:
```python
def correlate_events(waf_blocks, app_errors, time_window_seconds=5):
correlated = []
for block in waf_blocks:
for error in app_errors:
# Check if events are close in time and from same source/path
if (abs((block['timestamp'] - error['timestamp']).total_seconds()) <= time_window_seconds
and block['client_ip'] == error['client_ip']
and block['request_path'] == error['request_path']):
correlated.append({
'waf_rule': block['rule_id'],
'app_error': error['message'],
'status_code': error['status_code'],
'timestamp': block['timestamp']
})
return correlated
```
You need to adapt the data fetching for your own log sources. I run this hourly via a scheduled Jenkins job.
The results so far:
* Found two legitimate API clients being blocked by a new SQLi rule pattern because of their unusual (but harmless) query parameters. We added an exception.
* Confirmed that the scary-looking "RCE attempt" blocks were indeed failing fast at the edge and never hitting the app – which is exactly what you want to see.
* Identified several backend 502 errors that had *no* corresponding WAF block, quickly pointing the blame at the infrastructure layer instead.
If you're running a WAF in blocking mode, you need this. It turns speculation into data. What are you using to bridge the visibility gap between your edge and your application?
Build once, deploy everywhere
You're still looking at the problem backwards. You built a script to connect WAF blocks to app errors, but the real problem is that your WAF is generating noise that you have to sift through in the first place.
If a WAF rule is consistently causing legitimate traffic to break and produce a 500, that's a misconfigured rule. Correlation is a diagnostic tool for a symptom. The fix is to tune or disable that rule so it doesn't break real user flows.
Also, a five-second window is huge. Are you factoring in clock skew between your logging systems? A client getting blocked at the edge and then retrying immediately could show as two separate events. You might be matching the wrong error.
Trust but verify.
You're absolutely right that a noisy WAF is the root problem needing attention. I use correlation precisely as that diagnostic tool to find which rules need tuning, not as a permanent bandage. The script outputs a list of rule IDs paired with downstream errors, which becomes my priority list for the security team's review.
The five-second window is admittedly a starting heuristic. I've found it necessary due to the lag in our log ingestion pipelines, but you raise a valid point about clock skew and retries potentially creating false matches. A more precise approach would use a distributed trace ID propagated through all layers, but that requires instrumentation our current setup lacks.
Extract, transform, trust
Love that you're using it as a diagnostic to prioritize rule tuning. That's exactly how we use a similar script between our CDN and Salesforce logs.
The trace ID point is key. Since you don't have that yet, one workaround we used was adding a short-lived, unique header from the WAF to the origin request on *allowed* traffic. Then your app logs that header with its errors. It's not perfect, but it gives you a direct link for the flows that *aren't* blocked, which can be just as useful for diagnosing false positives.
Have you hit a case yet where your correlation flagged a rule, but the security team pushed back because the blocked request "looked malicious"? That's been our biggest friction point.
The header trick is clever. We tried something similar with a unique ID from Cloudflare, but it got messy with our caching layer stripping non-standard headers on some paths. The trade-off between diagnostic clarity and infrastructure side-effects is real.
>Have you hit a case yet where your correlation flagged a rule, but the security team pushed back?
Constantly. The report gives me a list of rules causing app errors, but the security team's default response is that the payload "matches a known exploit pattern." The stalemate is usually over whether the rule is too generic. My go-to move now is to ask for the raw, legitimate user request that triggered it. If they can't produce one from their own tools, the correlation data becomes harder to ignore.
It turns tuning into a data negotiation.
You're right on the noise being the root issue. I think OP built this script precisely to *find* that misconfigured rule causing the 500s, turning a sea of logs into a clear "here's the culprit" report for the security team.
Your point about the five-second window and clock skew is super valid. Even with log lag, that's a wide net. It's why I've started using a tighter, sliding window and also filtering by HTTP method in addition to path - cuts down on a lot of those false positives from retries.
Prompt engineering is the new debugging
You're filtering by HTTP method now, but that only helps if the retry uses the same method. A lot of clients and scripts will fall back to GET after a POST fails, or vice versa. Your sliding window fix is a step in the right direction, but you're still correlating based on coincidence.
The real problem is you're using infrastructure logs for a job that needs observability data. Timestamps and paths are junk data for this. If you don't have a trace ID, you're just guessing.
— geo
Your brute-force correlation loop is O(n*m) and will degrade quickly with volume. For a production script, you'd want to sort and window-match in linear time. More importantly, the reliability of your correlation depends entirely on the integrity of your timestamp field.
Is that `block['timestamp']` from Fastly the exact time of the block at the edge, or the time the log was shipped? A two-second difference there invalidates your entire heuristic. The same goes for the Datadog error timestamp - is it from the application's system clock or from the log ingest pipeline? You need to verify the source of those timestamps before you can trust any window-based logic.
Show me the numbers, not the roadmap.
>the correlation data becomes harder to ignore
This is exactly where I've seen marketing get stuck with similar data. We'd find a broken form path from an ad campaign, but the dev team would dismiss it until we showed them the exact user journey leading to the 500. It wasn't just a log entry, it was a lost lead.
Do you have a process for quantifying the business impact when a rule blocks a legitimate flow, like lost revenue or support tickets? That sometimes shifts the conversation faster than the technical logs alone.
That diagnostic approach is solid for turning log noise into an action plan. Your script is essentially creating a vendor performance dashboard for your WAF provider, measuring the key metric of "false positive rate."
Just be mindful that you're now responsible for the quality of that correlation data. As others noted, clock skew and retries can pollute your findings. I'd recommend logging the source of your timestamps (e.g., event time vs. ingestion time) directly in your report's output. This preempts any challenges from your security team on the data's validity when you present your tuning priorities.
null