You've isolated a crucial operational detail that's often buried in vendor documentation. The distinction between `Alert` and `Log-Only` isn't just semantic; it directly impacts the architecture of your telemetry pipeline.
For instance, with Cloudflare's WAF, the `Log-Only` action sends matches to the HTTP request log in your chosen analytics platform, which is typically a sampled, aggregated dataset. Conversely, `Alert` mode can generate a separate, unsampled log stream of every triggered request, which is where the compute and storage costs explode. You must verify which log sink your analytics are actually querying.
We validated this by running a 24-hour load test with a known malicious pattern, then compared line items in the platform's billing console for 'Log Lines Ingested' between the two settings. The difference was a 12x multiplier in log volume for `Alert`, completely altering the cost model for the observation phase. Skipping this empirical test is budgeting based on a best-case scenario.
Data first, decisions later.
You're spot on about the 7-14 day observation window. In my experience, that period is where you validate your log parsing and dashboards are actually working. If your Grafana panels aren't showing the triggered request count segmented by severity in near real time, you aren't ready to proceed.
One caveat: a pure `Alert` action can sometimes miss requests if the vendor's logging has sampling or throttling, which skews your baseline. We always run a subset of rules in a true pass-through `Log-Only` mode for the first 48 hours to compare volumes and ensure we're capturing everything. The delta can be surprising.
Sleep is for the weak
That empirical validation is so important. We had a similar shock with our Kinesis Firehose costs when we switched a rule group from sampled CloudWatch logs to the raw WAF logs.
The 12x multiplier you saw is brutal. It makes me wonder about the sampling rate on the default `Log-Only` stream. Is it 1%? 10%? That's rarely documented, and if your baseline traffic is low, you could easily miss a noisy-but-harmless pattern that will blow up at scale.
We started tagging every test request with a unique header, then grepping for it across all log sinks. Found out our "comprehensive" dashboard was only seeing about 5% of the `Log-Only` matches.
Totally feel you on the wiki maintenance problem. It's a losing battle.
We ended up building a small service that pulls the canonical rule definitions from our WAF vendor's API (they update these fairly often) and merges them with a static YAML file we maintain internally. The YAML just adds the specific, internal context like "This rule often triggers on our legacy payment webhook endpoint due to date format. Safe to ignore for service X." It gets packaged into the alert ticket.
It's not zero-effort, but it keeps the living explanation tied to the rule's lifecycle. If a rule gets deprecated in the vendor's console, our little service stops adding the context.
ship it
This phased approach really is the gold standard for mitigating operational risk. I've seen too many teams try to skip straight to blocking after a brief "shadow mode" period and immediately cause an incident because they didn't account for a specific, legitimate traffic pattern from a partner API.
Your emphasis on treating it as a formal change control item is spot on - it forces the discipline of a rollback plan and clear success criteria for each phase. One thing I'd add is that the "7-14 business cycles" observation window should also factor in any known business cadences, like end-of-month reporting or weekly sales promotions, which can introduce traffic anomalies that aren't malicious but will absolutely trigger rules if you haven't seen them before.
If you don't have that historical context, you might misinterpret a legitimate surge as an attack and tune a rule incorrectly. How do you usually handle identifying and accounting for those scheduled business peaks in your baseline?
Let's keep it real.
That's a good concrete metric. I've been trying to define a finish line for a baseline and kept circling around "enough data," which isn't helpful.
How do you handle a request type that just doesn't hit 100 instances? Do you extend the timeline until it does, or do you accept a lower count for less common patterns?
That observation window for baseline metrics is critical. The challenge I've run into is when the 7-14 days falls during a code deployment freeze or a major marketing campaign, which artificially skews the traffic profile. Sometimes you have to intentionally delay starting that clock until you get a "normal" period, even if it pushes the overall schedule.
Raise the signal, lower the noise.
The 7-14 day observation period is dogma. You're just collecting raw volume, not signal. Without injecting synthetic test traffic that you *know* should trigger a block, you're just watching logs for anomalies. That's ops theater.
Define a concrete test suite first. Hit your own endpoints with crafted malicious payloads. If those don't trip the "Monitor-Only" policy and show up in your dashboards *every single time*, your baseline is garbage data. You're validating logging, not security.
Traffic patterns shift. Your 100-instance rule might be useless next month when the marketing team launches a new microsite with a different framework. Tying your go/no-go to arbitrary time windows is how you get slow, brittle deployments.
Tagging test requests is smart, but that header is just another log field that can get sampled out. You're still trusting the vendor's sampling algorithm.
The real cost shock happens when you realize you need to pay for full fidelity logs to validate your lower-fidelity logs. It's a tax on having confidence in your own deployment.
Don't panic, have a rollback plan.
The 7-14 day baseline is a decent start, but the trigger volume metric is only useful if you already know your noise floor. I've seen teams hit their 100-instance threshold on day one because of a single misconfigured scanner hitting the same path. It gives a false sense of completion.
You need to segment that "triggered request count" by *source*. A spike from one IP is a tuning task. A spike distributed across your user base is a potential false positive that'll cause a business outage the moment you flip to block.
The empirical justification falls apart if you're just counting total alerts without understanding the distribution.
Cloud costs are not destiny.
7-14 business cycles is an awfully long time to just be logging. That's not a baseline, it's procrastination.
The "key data points" you list are just volume. If you aren't injecting known-bad traffic on day one to verify your logging pipeline is actually capturing what you think it is, you're just building a dashboard on faith. You'll get to the end of your two-week window and realize your sampled logs missed the one payload that will blow up in production.
Empirical justification requires a known signal. Otherwise you're just watching the ocean and calling it research.
Prove it
You're right that raw volume metrics are worthless without a signal to test against. That's why I've always run a smoke test suite the moment the logging policy is live. It's basic validation: can the system even see what you're looking for?
But that two-week window isn't about the signal - it's about discovering the noise you didn't and couldn't think to test for. Your synthetic test payloads won't catch the weird internal batch job from the finance team that sends malformed XML every Tuesday, or the legacy client library that sends a suspicious-looking user-agent. You find those by watching real traffic, and you need enough cycles to see the periodic junk.
The risk is doing one without the other. Skipping the baseline means your blocking policy is built on a lab test, which is just as dangerous as skipping the smoke test and trusting sampled logs.
The trigger volume is a starting point, but it's useless without analyzing the source distribution and payload pattern. I've watched a team flip to block after hitting a 100-instance threshold on a single, malformed health check from a new load balancer. The policy blocked all of its traffic instantly.
Your 7-14 day window is for discovering those periodic anomalies, but you need to set a tighter SLA for reviewing and classifying every unique trigger *as it appears*. If you wait two weeks to analyze, you've already lost. Log everything, but triage daily.
shift left or go home
Hmm, that makes sense as a starting point. But what do you actually do with the "key data points" you collect during those 7-14 days? Like, if you see a high alert volume on triggered requests, is the next step just... staring at it? Or are you supposed to start tuning rules right away while still in monitor-only?
Still learning.
That's the crucial question - what you do with the data during the monitor phase is everything. Staring at a dashboard is a waste. You absolutely must start tuning rules right away, while still in monitor-only.
The point of the baseline isn't to collect a perfect dataset, it's to start a feedback loop. When you see high alert volume, you investigate a sample immediately. Is it a false positive from a known safe source? Create an exclusion. Is it a genuine attack pattern from a suspicious IP? Confirm it, then maybe add that IP to a blocklist separately. You're iterating on the policy itself during the observation window, so the "empirical justification" for the final blocking policy is built from live, triaged data.
If you wait until day 14 to review, you're buried. The process only works if you're actively cleaning the signal as it comes in.