Skip to content
Notifications
Clear all

Just built a script to correlate WAF blocks with our app's error logs.

19 Posts
16 Users
0 Reactions
2 Views
(@infra_architect_rebel_2)
Honorable Member
Joined: 6 months ago
Posts: 410
Topic starter   [#29015]

Another week, another layer of abstraction to debug. We've been running Cloudflare's WAF for a while now, mostly on managed rules, and the dashboard proudly shows it's blocking thousands of "threats." My team was ready to propose scaling up our backend, citing increased "attack traffic." I remained skeptical—correlation is not causation, and a WAF block is a *prevented* event, not load.

So I built a script to cross-reference Cloudflare's WAF log data with our application's own error logs from the same time window. The goal was simple: are the things Cloudflare is blocking actually things that would have caused errors, or is it just chaff? The results were, predictably, eye-opening. A significant portion of the blocked requests fell into two categories:

* Legitimate user mistakes (badly formatted form data, stale API clients) that our app's validation would have handled with a clean 400 error.
* Automated scan noise that our app would have trivially rejected with a 404 or 405.

The truly malicious payloads that would have caused a 5xx error? A fraction of a percent. This isn't to say the WAF is useless, but it highlights a critical oversight in how we measure "load" and "threats."

I used the Cloudflare Logs API (they call it Logpush) to get WAF events and a simple grep/pattern-matching script against our structured app logs. The key was aligning on the request identifier. Here's the core of the correlation logic, simplified:

```python
# Pseudocode essence
for waf_event in cloudflare_logs:
# Find corresponding app log entry via CF-Ray ID or timestamp + client IP
app_log_entry = find_app_log(waf_event.ray_id, waf_event.client_ip)

if app_log_entry is None:
# WAF blocked something that never reached our app. Good.
log_metrics("blocked_externally")
elif app_log_entry.status_code >= 500:
# WAF missed something that caused a server error. Bad.
log_metrics("missed_threat")
else:
# WAF blocked something our app would have handled gracefully.
log_metrics("overzealous_block")
```

This exercise proved two of my long-held beliefs: first, that much of the "complexity" we defend against is phantom, and second, that without correlating your perimeter data with your internal telemetry, you're flying blind. You might be adding servers (and cost) to handle load that never existed, or worse, tuning your WAF so aggressively it starts breaking legitimate traffic that your robust monolith would have processed just fine.

Has anyone else done this kind of correlation? I'm particularly interested in whether you used the findings to dial down WAF rule sensitivity or to re-architect logging to make this analysis more continuous. I'm leaning toward the latter—building a permanent, low-cost dashboard that shows *actual* threats versus *theoretical* ones.


monoliths are not evil


   
Quote
(@gregoryt)
Reputable Member
Joined: 2 months ago
Posts: 418
 

That's a really smart way to check the actual impact. I've been thinking about the "noise" from automated scans too, but never thought to verify it like this.

Could you share what you used to pull the Cloudflare WAF logs? Are you hitting an API or streaming them somewhere? I'm guessing you used something like Grafana Loki for the app logs? Trying to learn the setup 😅



   
ReplyQuote
(@cloud_infra_rookie)
Noble Member
Joined: 4 months ago
Posts: 552
 

Yeah, the API method is what I was wondering about too. Do you need to set up a Logpush to something like S3 first, or can you query recent blocks directly from their API? Seems like there would be latency either way.

For the app logs, we're actually just using CloudWatch for now, not something fancy like Loki yet. Makes me wonder if that's a good next step for this kind of correlation work. How do you handle the time syncing between the two different log sources? That part seems tricky.



   
ReplyQuote
(@cloud_infra_rookie)
Noble Member
Joined: 4 months ago
Posts: 552
 

Yeah, the time sync thing is exactly what I was wondering about too. Our logs also go to CloudWatch right now, and the timestamps can be a bit off depending on the log agent's buffer. I've heard Loki or something with a common ingest timestamp can help, but it feels like a big move for just this one task.

> can you query recent blocks directly from their API?

I looked into this a bit because Logpush seemed heavy. From what I read, you can't get WAF-specific logs from the general API; you need Logpush to a destination like S3 or a SIEM. That latency is real - it's not for real-time checks. Does anyone know if that's still the case? Maybe there's a newer endpoint.



   
ReplyQuote
(@grafana_knight_shift)
Reputable Member
Joined: 6 months ago
Posts: 324
 

Totally agree on the load vs. prevented event distinction. I've seen the same mindset where high WAF block counts trigger capacity discussions, but the actual resource hit on the backend is often from the requests that *aren't* blocked.

The part about legitimate user mistakes being flagged is key. It makes me wonder if tweaking the sensitivity on specific managed rules, or adding some custom allowlists for known benign patterns, could reduce that noise without dropping real protection. Did your script give you any insight into which specific rules were catching those "clean 400" cases?



   
ReplyQuote
(@danm)
Honorable Member
Joined: 3 months ago
Posts: 452
 

Yeah, you need Logpush for WAF logs, there's no direct API for them. I set up Logpush to S3, and it does have that latency, maybe 1-2 minutes. Not for real-time, but fine for this daily correlation job.

For time syncing with our app logs, I just used the CF ray ID. Both log streams have it, so I match on that instead of trying to align timestamps perfectly. Works for joins. CloudWatch should have the ray ID too if you enable it in your logging format.



   
ReplyQuote
(@evanj)
Estimable Member
Joined: 3 months ago
Posts: 189
 

That's a clever way to handle the timestamp problem, matching on the ray ID instead. It makes sense, since that's the true common thread. I hadn't considered that as the primary join key.

Your point about latency is something I run into constantly when evaluating logging tools for this kind of work. A 1-2 minute lag for a daily analysis job is perfectly fine, but it does add a step in the procurement process. Now I need to factor in whether the SIEM or data warehouse we're looking at can handle that S3 logpush flow efficiently, or if it creates another cost layer. It turns a simple script idea into a minor infrastructure decision.

Do you find the Logpush to S3 reliable enough, or does it require much babysatching?



   
ReplyQuote
(@aidenh5)
Reputable Member
Joined: 3 months ago
Posts: 312
 

Exactly. Those "clean 400s" are the real noise. You're paying for compute to block something the app would've handled for free.

We did a similar check and found rule 100174 (SQLi detection) firing constantly on garbage search strings from old bookmarks. The payloads were junk, never would've parsed. Added a narrow allowlist for that path and cut the noise by 80%.

Your last point about measuring load is the kicker. Teams see the WAF count and panic. Show them the error log correlation. If the app isn't choking, the "attack traffic" isn't real load.


Ship fast, review slower


   
ReplyQuote
(@annac)
Reputable Member
Joined: 2 months ago
Posts: 391
 

Yes, that's such a classic rule to watch for! Rule 100174 and some of the XSS detectors are notorious for catching garbage strings that look like a payload but are semantically empty.

Your point about the narrow allowlist is key. We made the mistake at first of being too broad when we saw a pattern, and it accidentally let a few sketchy ones through. The key is locking it down to the exact URI and maybe even the referrer if it's from those old bookmarks.

Showing the actual app error log side-by-side with the WAF block is what finally convinced our team to stop reacting to the big scary numbers. It turns panic into a tuning exercise.


Keep it simple.


   
ReplyQuote
(@fionap)
Reputable Member
Joined: 2 months ago
Posts: 349
 

Spot on with the distinction between prevented events and actual load. That kind of data is the perfect antidote to dashboard panic. I've seen teams waste so much energy "optimizing" against noise.

One thing to watch out for, though, is that your "clean 400" category can sometimes hide a real issue. We once found that a surge in those blocked requests was actually caused by a broken client-side SDK version on a popular mobile app. The WAF was catching the malformed requests, but the root cause was our own deployment. The script helped us spot the pattern, but we had to dig a layer deeper to find the source.

It's a great tool for sanity-checking capacity alarms and tuning rules. Have you started using the results to adjust any specific managed rule sensitivities yet, or is it still in the discovery phase?


null


   
ReplyQuote
(@ethanp23)
Reputable Member
Joined: 2 months ago
Posts: 293
 

Yeah, the script pinpointed rule 100174 like others mentioned, but also a ton from 941330 and 941100. It was shocking how many were just mangled URLs from those dodgy 'free download' sites that clutter old forums. They'd get blocked, but our app would've just thrown a 404 anyway.

The real value was seeing *which* rules were noisy. That let us dial down the sensitivity just on those specific managed rule groups, instead of a broad cut. Actually made me appreciate Cloudflare's granular controls a bit more.


Beta tester at heart


   
ReplyQuote
(@fionap)
Reputable Member
Joined: 2 months ago
Posts: 349
 

Oh, that last line hits home. We had the exact same debate about load measurement last quarter. The WAF dashboard graphs going up made everyone's capacity planning instincts kick in.

Your point about the clean 400s is so important. It's not just "noise," it's *saved compute*. Your app was going to spend cycles validating and rejecting that bad input anyway. The WAF just did it at the edge, for presumably less cost.

I'm curious, now that you have the data, are you planning to adjust your team's capacity alerts to explicitly ignore WAF block rates? Or will you use it more to tune the rule sensitivity down for those specific false-positive paths?


null


   
ReplyQuote
(@annac)
Reputable Member
Joined: 2 months ago
Posts: 391
 

Yes! That split you found - clean 400s vs. real 5xx threats - is exactly the data everyone needs to see. It changes the conversation from "we're under attack" to "our edge is working."

We use this data to adjust our monitoring dashboards. We now have a gauge that subtracts WAF-prevented events from our total "request load" metric. It stopped a few pointless scaling discussions before they even started.

Have you considered tracking the trend of that "fraction of a percent" over time? For us, seeing it stay flat while overall blocks rise is the best proof the WAF is catching chaff, not real danger.


Keep it simple.


   
ReplyQuote
(@alexr23)
Reputable Member
Joined: 2 months ago
Posts: 319
 

Glad you mentioned the granular controls. It's easy to miss, but diving into the specific rule IDs like 941330 (PL4 XSS) is what makes the difference between a useful policy and a noisy one.

I'd caution about dialing down sensitivity on an entire managed group just based on one path's noise, though. The better pattern I've found is to use those findings to create a more targeted allow rule, or to adjust the sensitivity for that specific rule ID *but only* when combined with a URI condition. That keeps the protection broad while surgically cutting the noise you identified.

Have you run into any issues yet where tuning one rule for those mangled URLs inadvertently let a similar, but malicious, pattern through on a different endpoint?


—Alex


   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

That's a great question about which rules are catching the clean 400s. In our setup, the bulk of them ended up being tied to the same rules user1357 and others mentioned - 100174 and the 941 series - but the real twist was *where* they were happening. It was almost never on our core application flows.

The script showed they were overwhelmingly hitting old, deprecated API endpoints that we kept for legacy integration support. The app was already set to return a 410 Gone or a strict 400 for malformed data on those paths, so the WAF was just adding a layer of noise. Seeing that distribution let us create a very targeted policy: we lowered the sensitivity on those specific managed rules, but only for that set of legacy URI paths. It cut down the alert volume dramatically without touching protection for our current v2 endpoints.

It turned a vague sense of "too many blocks" into a precise map of where the friction was, which made the tuning decision much easier to justify. Have you noticed a pattern in the endpoints or user agents tied to your false positives? Sometimes it's less about the rule itself and more about the context it's firing in.


Let's keep it real.


   
ReplyQuote
Page 1 / 2