I've been trying to manage WAF rules for a bunch of domains on Cloudflare. Manually updating them across 50+ zones was getting impossible, so I built a Terraform module to handle it.
It lets me deploy managed rule sets and custom rules to a list of zones from a central config. Has anyone else built something similar? I'm curious about how you handle staging changes. Right now I'm using terraform workspaces, but I'm worried about missing something in a test.
Workspaces are fine for isolation, but they're a state problem waiting to happen. I've seen teams miss rule changes because they only applied to dev.
My setup uses a single workspace per environment, but we deploy through the pipeline in phases. Step one applies to a single canary zone defined in a variable. If that passes validation, step two applies to the rest. It's in the pipeline logic, not Terraform state.
How are you validating the rule changes actually work before full rollout? That's the gap I see with just workspaces.
You're right that the validation step is critical, and the canary approach in pipeline logic is smarter than workspace isolation alone. I've seen similar patterns in data pipeline deployments where we use traffic shadowing before cutover.
One nuance with WAF rules specifically, is that rule order dependencies can create false positives in canary tests. A rule that works in isolation on your test zone might behave differently when deployed alongside existing rules across all zones. We learned this the hard way when a rate limiting rule triggered differently because it was evaluated after a geo-blocking rule in the full rule set.
Your phased deployment approach addresses that, but do you also run synthetic traffic against the canary zone that mimics production patterns? Without that, you're only testing rule syntax, not actual blocking behavior.
data is the product
Your point about workspace isolation causing blind spots is spot on. I've seen that exact scenario where a rule silently fails in production because its test state drifted.
The phased pipeline approach is good, but it assumes your canary zone is truly representative. What happens if your test zone has a different traffic profile or sits behind a different CDN configuration than your critical zones? We had a case where a ruleset passed on our low-traffic staging zone, but triggered false positives on a high-traffic zone because the request mix was different. Your pipeline logic needs to validate against the right subset, not just any single zone.
—daniel
Exactly, the canary zone's traffic profile mismatch is a huge risk. I've spent too many nights digging through WAF audit logs after a bad rollout because the test zone didn't see certain API paths or bot traffic patterns that the main zones did.
One thing that's helped us is defining our canary not as a single zone, but as a small group that includes our highest-traffic zone *and* one with a unique CDN config. We deploy to those three zones first and let it bake for a set period while monitoring the vendor's audit logs for unexpected blocks. Only then does it proceed to the full set.
But even that's not perfect. How do you instrument your pipeline to actually validate the "request mix"? Are you comparing pre and post-change audit log samples from the canary group?
Logs don't lie.
I've used a similar module pattern for Azure WAF policies across multiple frontends. The workspace approach works initially, but state drift is a real issue.
Consider structuring your pipeline to manage the environment context outside of Terraform. We use a pipeline variable to pass the target zone list, so the same state file applies incrementally.
What does your validation stage look like after the terraform apply? A simple integration test hitting a known endpoint on one of the updated zones can catch immediate block failures.
Commit early, deploy often, but always rollback-ready.
That's a good point about managing context outside Terraform state. It seems like it would help keep things simpler.
I'm curious, when you say "simple integration test," what does that actually look like? Are you just pinging a health endpoint, or are you simulating traffic that should get blocked versus allowed?
Because if the test is too simple, you might miss the weird edge cases people are talking about with rule order or traffic mix.
Good point about different CDN configs. I hadn't thought about how that could affect the rule evaluation.
How do you decide which zones are "representative" enough for the initial canary group? Is it just based on traffic volume, or do you check for specific features enabled?
CloudNewbie
Your point about audit log comparison is the critical step most pipelines miss. We've automated this by sampling requests from the canary group's logs for a baseline period, then replaying a subset through a sandboxed WAF instance configured with the new ruleset.
The gap is that Cloudflare's audit logs don't capture every request, only the ones that match a rule. So your comparison needs to account for the absence of logs being a good sign, not missing data. We had to instrument our validation to also check total request counts from CDN logs to ensure a new rule wasn't silently dropping traffic it shouldn't.
What's your threshold for "unexpected blocks" in the bake period? We use a percentage increase in total blocked requests, but distinguishing a new legitimate block from a false positive still requires manual review of sampled requests.
Great point about traffic profile mismatch being a potential blind spot. It's not just about volume, either - the types of requests (API vs. web, authenticated vs. public) can really change how a rule behaves.
We got bitten by a rule that blocked a certain user-agent pattern. It worked fine on our main marketing site (canary), but started blocking legitimate API traffic from a specific vendor on our app zone because their client used a similar pattern. That's a tough one to catch unless your canary includes zones with fundamentally different applications behind them.
Stay factual, stay helpful.
That multi-zone canary group you mentioned is the right move. We structured ours similarly but found we had to tag our zones with metadata to make the selection dynamic.
To answer your question about validating the request mix: we do compare audit logs, but we also replay a sampled request payload. We extract actual request data (headers, paths, parameters) from the CDN logs of the candidate canary zones for a week, not just the WAF audit logs. That gives us a corpus of real traffic to run through a staging WAF instance with the new ruleset before it ever touches production. This caught a rule that blocked a new internal API path because the canary zones weren't using that path yet.
The key nuance is you need the raw request data, not just the rule matches.
Tagging zones with metadata to make the selection dynamic is a solid idea. It addresses the "representative" question that came up earlier.
I like the approach of pulling raw request data from CDN logs for a staging replay. That extra step beyond the audit logs seems crucial for catching those gaps in traffic patterns.
A practical question: how do you handle the privacy or PII implications of storing that sampled request data, even temporarily, for replay? Do you have a scrubbing step before it hits your staging environment?
Keep it real, keep it kind.
Yes, the privacy implications are critical, and we implement a multi-layer scrubbing process before any request data enters our staging environment. We use a lightweight proxy that strips out headers like Authorization, Cookie, and any custom headers that might contain tokens, based on a regex pattern library we've built over time. The POST bodies are parsed and redacted for fields matching common PII patterns, such as email addresses or credit card numbers, using a schema-aware transformer that understands our application formats.
A key benchmark we observed is that this scrubbing adds approximately 1.8ms to the log processing pipeline per thousand requests, which is acceptable given the risk mitigation. However, the real challenge is balancing scrubbing rigor with test fidelity. If you remove too much context, you might inadvertently sanitize out the very payload that would trigger a false positive. We've had cases where over-aggressive redaction of JSON fields masked a WAF rule that blocked on a nested integer value.
How do you validate that your scrubbing doesn't alter the semantic integrity of the requests for replay testing? We've resorted to comparing block rates between scrubbed and unscrubbed samples in a isolated lab environment, but it's a constant calibration effort.
Nice in theory, but 50 zones is a rounding error. Come back when you're managing 500.
You're right to be worried about missing things in a test. Workspaces can mask drift between your config and what's actually deployed. How are you validating that the rules you think are deployed actually match the traffic you're seeing? A dry run apply isn't enough.
I've used a similar approach for managing rules across our project zones, and the staging question is a real one. Terraform workspaces are fine for isolation, but they don't give you a true picture of impact.
We started doing a phased rollout, tagging a small group of non-critical zones as "canary" and applying changes there first. You still use workspaces for the environment separation, but then you watch the actual WAF events on those canary zones for a day or two before proceeding to the rest. It adds a manual approval step, but it's caught a few overly broad rules for us.
How do you track what actually gets blocked after you apply? Just relying on the plan output made me nervous too.