The manual approval step on the canary group is exactly where we caught a rule blocking a new health check endpoint. We pair that with automated alerts on WAF event spikes, so we don't just watch logs passively.
> How do you track what actually gets blocked after you apply?
We have a Prometheus exporter that scrapes the WAF analytics API for each zone, tracking blocks-per-rule. A simple Grafana dashboard compares the rate before and after the apply. If any rule's block rate jumps by more than a configured threshold, it fails the pipeline and rolls back automatically.
The trick is setting a sensible threshold, as some rules are designed to be noisy. We baseline each rule's normal block rate during low-risk periods.
Nice module, managing that many zones manually would be a nightmare. Workspaces are a solid start for staging.
We handle it by tagging a small group of low-traffic zones as 'test' in our module config. The module can target zones by tag, so we run the apply on just that group first. After that, we monitor the actual WAF event logs in those zones for a full business day before rolling out to the rest.
It adds a day's delay, but it caught a rule that blocked a legacy API path we forgot about. Have you considered adding zone tags or metadata to your config to segment your rollout groups?
Data doesn't lie, but dashboards sometimes do.
Oh, rule order dependency is such a good catch. We ran into that with a new OWASP rule that started blocking our login page because a previous custom rule was stripping certain parameters, changing the request body it evaluated.
> do you also run synthetic traffic against the canary zone
We do, but we found it only gets you so far. Synthetic traffic is great for testing the absolute basics, like "does the health check path still work". But it's terrible at mimicking the weird edge cases from real user traffic, which is where most of our scary false positives come from. We lean much heavier on that sampled request replay method user564 mentioned.
Prometheus and automated rollbacks sound slick, but you're putting a lot of faith in that baseline and threshold. What about rules that are dormant for months until a specific attack vector emerges? Their "normal" block rate is zero, so any jump is technically infinite and trips your fail-safe. That seems like a fast track to either a desensitized threshold or a rollback that kills a rule you actually need.
And baselining during "low-risk periods" assumes you can reliably define what a low-risk period even is anymore.
Trust but verify
I've been down that exact road, and workspaces absolutely can lull you into a false sense of security. The config diff looks clean, but it's all just theoretical until it hits the weird, malformed reality of actual traffic.
A thing that's saved us, beyond the canary ideas others have mentioned, is a pre-apply validation script. It fetches the *current* live rule set for each zone in the workspace and diffs it against the planned state from `terraform plan -json`. It sounds redundant, but we've caught config drift from manual hotfixes or other automation that never got committed back.
That scripted check, plus tagging a small blast radius of zones for a staged apply, is what made me finally sleep at night. Are you versioning the rule sets themselves, or just the Terraform module?
It's just pattern matching
You're absolutely right about synthetic traffic's limitations. The rule order dependency you described is a perfect example of a failure mode it would never catch, because the interaction between the parameter-stripping rule and the OWASP rule is entirely dependent on the specific sequence of evaluation, something synthetic scripts rarely replicate.
We mitigated a similar issue by implementing a validation step that replays a corpus of recent, real blocked requests from our logs against the new rule set in a sandbox. This doesn't just test the new rule, it tests the entire rule chain's behavior. It's computationally heavier than synthetic checks, but it caught a case where a new rate limit rule interacted badly with an old geo-block rule, creating a dead zone for legitimate traffic from a specific region.
Your point about leaning on sampled request replay is key. It moves validation from "does it work in theory" to "does it break the real traffic patterns we've already seen."
Plan the exit before entry.
Replaying real blocked requests sounds like a game-changer for catching rule interactions. But where do you store that corpus of recent requests? Do you keep a rolling buffer of logs just for this purpose, or do you query your main logging system each time? Worried about performance or cost if I'm hitting logs for every plan.
Containers are magic, but I want to know how the magic works.
We keep a separate, dedicated log sink in our logging system for a rolling 14-day window of WAF-blocked requests. Querying the main production logs for every plan was indeed cost-prohibitive and slow.
The key was structuring the export to include only the fields needed for replay - the raw HTTP request and the rule ID that triggered the block. This keeps the volume manageable. We process this sink nightly to build a compressed corpus file stored in object storage, which the validation step pulls.
You're right to worry about cost. The initial approach of live queries during plans added unpredictable latency and charges. Moving to a scheduled, pre-processed corpus fixed both. The trade-off is you're replaying requests that are at most a day old, but we found that sufficient for catching interaction regressions.
Data is the new oil – but only if refined
Totally feel your pain with the manual updates. Been there with a smaller set of zones and it was still a mess. Workspaces are okay, but I found they're too easy to get out of sync with the actual zones if someone makes a one-off fix directly in the Cloudflare dashboard.
What saved me was adding a simple tagging system to my module. I tag zones as "canary" or "prod". The apply first runs on "canary" zones only (I have like 5 low-traffic ones). Only if I see no weird block spikes in those logs for 24 hours do I roll out to the rest. It's manual, but it's a cheap safety net.
Have you thought about adding a tagging or metadata layer to your config to group zones for staged rollouts?
Love that dedicated sink idea. We do something similar but leaned into a cloud data warehouse for the replay corpus, mainly because we already had the pipeline running for other security analytics.
One caveat with the 14-day window - we found some nasty rule interactions only surface with requests that are *older* than two weeks, like monthly cron jobs or weird quarterly business reporting calls. We had to extend our window to 30 days and implement some lightweight deduplication to keep the corpus size from ballooning.
How do you handle deduping similar requests in your nightly processing? Straight comparison of the raw request string, or something more clever?
Data doesn't lie, but dashboards sometimes do.
Cost was the main blocker for us too. Querying main logs on every plan was a non-starter. We ended up with a scheduled CloudWatch Logs Insights query that dumps a daily snapshot of blocked requests to an S3 bucket. The validation step just grabs that static file, so there's zero runtime log query cost.
The performance hit is minimal, but the trade-off is you're replaying yesterday's traffic, not today's. For us, that lag is acceptable since rule deployments aren't daily events.
That's the pragmatic move, cost-wise. "Yesterday's traffic, not today's" is a good trade until you're in a reactive situation patching a zero-day. Then you really want to replay the attack pattern that just showed up.
We ended up with a hybrid setup for exactly that: the scheduled daily snapshot for general validation, plus a small, fast pipeline that can pull the last hour of raw blocked requests on demand if we need to validate a hotfix against the current attack pattern. Adds some complexity, but it beats replaying a corpus that's already irrelevant.
Data over dogma.
Your hybrid model is the logical endpoint, and it's one I've seen mature teams converge on. The cost of missing a zero-day interaction usually dwarfs the complexity budget for this.
One nuance from our implementation: the "last hour" pipeline needed careful throttling and caching. During a volumetric attack, pulling the raw stream of blocked requests can itself be expensive and slow. We had to implement a sampler that captures a representative subset, not the entire firehose, for the on-demand replay. Otherwise, the validation step would timeout waiting for the data during the very incident you're trying to patch.
It adds another layer, but it prevents the safety mechanism from collapsing under load when you need it most. Have you run into similar scaling issues with your on-demand pull during an active attack?
Sampling the firehose during an attack is smart, but you're adding a heuristic to a safety mechanism. You're now trusting that your sampler captures the "right" subset to reveal a bad interaction.
We tried that. The sampler missed a weird interaction between a volumetric rule and a new SQLi pattern because the malicious payloads were spread across thousands of identical junk requests. Sampler saw one, deemed it representative, moved on. The rule block-chained fine in isolation, but the cumulative load from the pattern across the full firehose triggered a different resource limit downstream.
Now we just take the cost hit and keep a hot, dedicated buffer with aggressive field trimming. The safety net shouldn't have holes.
-- old school
You're right that sampling introduces a blind spot. We hit a similar issue where a sampler normalized header case, missing a rule that was case-sensitive.
But I don't think the choice is purely "sampler with holes" vs "full firehose cost". There's a middle ground: buffer everything but filter duplicates *after* capture, based on a fingerprint of the normalized request plus the triggering rule ID. It keeps the weird edge cases that differ by a single query parameter but collapses the thousands of identical junk requests into one record for replay. Storage cost is closer to the sampler approach, but you don't lose the signal.
Build once, deploy everywhere