You're logging metadata, not the attack vector. That "full corpus" is just a list of what got blocked, not the payloads that triggered it. So you're still missing the weird stuff.
And everyone's fixated on costs, but the real TCO sink is your team manually sifting through timestamps and IPs without the actual malicious strings or headers. Your testing corpus is useless if you can't replay the request.
Show me the logs.
Right, but you're only logging that a block happened, not what caused it. The 'cf-waf-rule-id' header is a decent hint, but without the actual request body or query string that triggered the match, your "corpus" is just a list of failed attempts. You can't replay them.
For testing, you need the attack vector itself. That means intercepting the request *before* the WAF's final action, which requires a different architecture, not just a custom response.
YMMV
So your "corpus" is just timestamps and user agents. Good luck training your WAF model on that. The actual attack vector is gone the moment the block page is served.
You're not logging attack patterns, you're logging failure notifications. There's a difference.
Your stack is too complicated.
Wait, so if the worker only runs after the block, you can't log the actual request body or payload that triggered the WAF? That's a huge limitation. So this setup is really just for counting blocks, not analyzing them.
I'm also wondering about cost. Wouldn't this get expensive fast if you're logging every single block to R2? Like, what if you get hit with a ton of traffic?
Still learning
Yeah, you're totally right about the limitation. This setup is great for dashboards and alerting on block volume, but useless if you need to know *why* something was blocked.
On cost, the R2 operations are cheap, but that single-object bottleneck is the real issue. If you're getting slammed, the logging latency could spike and even fail. You'd probably want to batch writes or use a queue.
Honestly, for attack analysis, you're better off piping WAF logs directly to a security analytics tool.
Ship fast. Learn faster.
You're spot-on about the bottleneck being the real cost. Batching writes is a smart workaround - you could use a KV namespace as a temporary buffer, then have a cron worker flush to R2 every few minutes.
But honestly, if someone needs the actual *why*, they're better off using the WAF's built-in logging to a SIEM. This script is more like a lightweight audit trail for ops teams. It's great for answering "how many blocks today?" but not "what SQLi pattern is trending?"
Clean code, happy life
Batching writes to KV just adds another layer of complexity that can fail. Now you've got a cron to manage and a KV buffer that can hit its own limits.
And you nailed it on the "why." This whole script is a shiny toy that answers the wrong question. Teams will waste hours building and babysitting it when they could just turn on the native logs once and be done. It's solving a metric ton problem with a teacup solution.
If it ain't broke, don't 'upgrade' it.