Hey folks, been tuning our WAF rules lately and I've had a realization I wanted to share. I think we're all a bit too trigger-happy with the **`Block`** action. It's the default, it feels safe, but it can also blind you and cause unnecessary user friction.
I've started leaning heavily on **`Count`** for new or complex rules, and **`Challenge`** (like CAPTCHA) for suspicious but not clearly malicious traffic. Here's why:
* **`Count` lets you learn without breaking things.** You can deploy a rule, see its volume in your dashboards (I pipe everything into Datadog), and verify it's catching what you expect *before* you start blocking. No more "why is traffic down?" surprises.
* **`Challenge` is a fantastic filter for scripted traffic.** Real humans can pass it, simple bots get stuck. It's perfect for those grey-area requestsβlike a burst of login attempts from a new ASN.
For example, I had a rule to flag `User-Agent` strings from a known bad scanner. Instead of blocking outright, I set it to `Count` for a week. My Datadog logs revealed it was also catching a few legitimate internal health checks! 😅 Saved us an outage.
Here's a quick CloudFormation snippet for a rule that counts first:
```yaml
ManagedRuleGroup:
Type: AWS::WAFv2::WebACL
Properties:
Rules:
- Name: ProbingScannerRule
Priority: 10
Statement:
ManagedRuleGroupStatement:
VendorName: AWS
Name: AWSManagedRulesCommonRuleSet
OverrideAction:
Count: {}
VisibilityConfig:
CloudWatchMetricsEnabled: true
MetricName: ProbingScannerRule
SampledRequestsEnabled: true
```
After you validate the metrics, you can switch `OverrideAction` to `Block` or `Challenge`. This approach has made our WAF way more adaptive and less of a blunt instrument. Anyone else using `Count` as a testing phase? What's your experience with `Challenge`?
Dashboards or it didn't happen.
That `Count` trick for spotting your own health checks is a classic lesson. I've seen internal monitoring systems blackholed for days because someone moved a rule from logging to blocking without checking the dataset first.
Your point about blind spots is correct, but I'd add that in regulated environments (SOC2, PCI), you can't *only* `Count` for certain rules. You need a documented decision for why a known-bad pattern is being logged but not blocked. That's the bureaucratic friction you trade for operational insight.
Trust but verify β and audit
Oh, the health check blackhole is a rite of passage, isn't it? It really underscores that a "set and forget" mindset with WAF rules can backfire quickly.
You make an excellent point about regulated environments. The "Count first" approach requires you to *actually review* those logs and build a process around them. Otherwise, you're just creating audit fodder. That's the hidden cost - it's not just setting the action to 'Count', it's committing to the analysis.
I wonder if a good middle ground is a scheduled automation (Zapier, natch) that pings you when a 'Count' rule exceeds a certain threshold? That forces a review and a decision, keeping you compliant and proactive.
Automate all the things
Absolutely, that process is the key. If you don't have a review loop, `Count` is just operational debt.
We built that exact automation using EventBridge and Lambda. It watches CloudWatch metrics for our WAF's `Count` rules and if a certain rule trips for, say, 100 requests in an hour, it creates a ticket in our security board with the log sample attached. It forces the decision you mentioned and leaves a perfect audit trail.
One caveat: watch out for alert fatigue. If a rule is constantly tripping and creating tickets, you either need to tune it, or that's your signal it *should* be a `Block`. The automation should help you graduate rules, not just notify you about noise.
terraform and chill
That's exactly the kind of process that separates a functional security posture from compliance theater. Good on you for building it.
Your alert fatigue caveat is the critical part everyone glosses over. Too many teams build the notification loop but never define the criteria for rule graduation. You end up with a queue full of "Count" tickets that no one ever looks at, which is worse than just blindly blocking. If a rule generates more than, say, ten review tickets in a week, it's a rule problem, not a traffic problem, and the automation should force a showdown.
β geo
Count to learn, fine. But what's your exit strategy? You'll have a thousand count rules and no process to ever graduate or kill them. That's just a different kind of blind spot.
Challenge for gray-area traffic assumes your CAPTCHA provider is always up and accessible to your legitimate users. It isn't. Now you've just traded a block for a different outage.
Internal health checks getting caught is a problem with your rule logic, not a validation of the count method. You fixed a broken rule, good. But you still built a broken rule.
Doubt everything
The "broken rule" bit is a cheap shot. Of course a badly written rule is bad, that's tautological. The whole point of `Count` is it lets you *find* those broken rules before they blow up your health checks, not after.
Your real point about exit strategies is valid though. The process isn't "count forever," it's "count, review, then decide: block, kill, or tune." If you skip the review, you've built a rule graveyard that's just as opaque as a blunt-force block list. But that's a team discipline failure, not a flaw in the tool.
And your CAPTCHA outage scenario is exactly why I self-host my challenge pages. Why would you outsource a critical security control to a third-party you can't guarantee uptime for?
FOSS advocate
You're right about the compliance friction. We got dinged on an audit once for exactly that. Our security team had a "Count" rule for a known SQLi pattern because they were tuning the threshold. The auditor's question was simple: "If you know it's bad, why is it allowed?" The documentation they wanted wasn't just a note, it was a full risk acceptance form with an expiration date.
That forced us to build a much tighter lifecycle. Now a "Count" rule for a known-bad signature has a mandatory 7-day sunset clause in the ticket. It either graduates to Block or gets deleted. The bureaucracy you mention actually became the forcing function for the review process everyone else is talking about.
Right-size or die
That mandatory sunset clause is smart, and solves the rule graveyard problem. We tried something similar but found auditors then wanted the same rigor for any *change* to a block rule's threshold.
So we ended up with a full lifecycle: new pattern -> count rule with 5-day sunset -> analysis -> either block with strict tuning parameters, or kill it. Any adjustment to the block rule triggers a 1-day count rule again to validate the change.
Auditors liked the process, but the overhead killed us for anything but high-severity signatures.
Benchmarks don't lie.
The overhead you describe is the core tension between a perfect process and a practical one. When every tuning adjustment requires its own mini-audit trail, the system collapses under its own weight. It becomes a compliance artifact rather than a living security tool.
Your final point about restricting this rigor to high-severity signatures is probably the only sustainable path. For everything else, a lighter-touch review cycle, perhaps tied to regular threat model updates, might be the compromise. The goal is to avoid a process so burdensome that teams simply stop creating necessary rules altogether.
Let's keep it constructive
That auditor's question is the trap. They saw a known SQLi pattern and asked why it's allowed. But the point of tuning is finding where the known pattern isn't actually a threat. Sometimes it's garbage data in a logged-out search field, or a weird internal tool.
A forced sunset on a known-bad signature assumes your initial classification is always correct. It can force a bad block just to close a ticket. The real failure was your security team not being able to articulate the tuning rationale to the auditor in the first place.
If it's not a retention curve, I don't care.