Alright folks, let's talk about a claim that always makes my ops-sense tingle: "zero false negatives." I saw this in some Imperva material recently, and it took me right back to a late night in 2013, sweating over a WAF that swore it was catching everything. Spoiler: it wasn't. A clever, slow-moving exfil was sliding right through, and the logs showed... nothing.
Here's the hard truth: you can't *prove* a negative, especially in security. You can't verify you're not missing something you don't know about. The claim itself is a marketing red flag. What you *can* do is test rigorously and monitor exhaustively to build confidence.
My approach has always been to treat these claims as a challenge. Build a test bed.
```yaml
# docker-compose for a simple test rig
version: '3'
services:
vulnerable_app:
image: vulnerables/web-dvwa
ports:
- "8080:80"
# Your Imperva/WAF would be proxying traffic to this
```
Then, fire everything you've got at it. Use tools like OWASP ZAP, sqlmap, or even old-school custom curl scripts that mimic the weird, malformed stuff you've seen in real logs. The goal isn't just to see if it *blocks*, but to see if it *logs*. A "zero false negative" promise means it should log every single attempt.
But here's the kicker: you'll never know if your test suite is comprehensive. Attack patterns evolve. The best you can do is:
- Continuously run these tests (CI/CD pipeline, anyone?)
- Feed it real, anonymized traffic from a honeypot
- Correlate its logs with other telemetry (host-based, network-based)
If they say "zero false negatives," ask for their testing methodology. Ask for the raw data. In my experience, what they often mean is "zero false negatives *in our curated test suite*." Big difference.
Stay skeptical, test relentlessly, and always have a second pair of eyes (or tools) watching the watcher.
-- Dad
it worked on my machine
Exactly this. That late-2013 feeling is one I know well, and it's why I treat "zero false negatives" as an invitation to operational due diligence. You've hit on the core of it: the logging is everything. A tool can miss something new, but if its logs also show nothing, you're flying blind.
One thing I'd add to your test bed approach is to also simulate "allowed" traffic that's just... weird. Think parameter values with strange encodings, massive payloads that are technically valid, or sequences that look like failed attacks. You're not just testing if it catches known bad, but if its logging granularity gives you a trail for the unknown. A product might not block a novel exfil method, but if it logs that outlier session with full request/response details, you've got a fighting chance to detect it yourself.
Honestly, the claim itself changes the procurement conversation. It moves from "do you trust them?" to "how do we validate and monitor together?" Asking a vendor how they suggest you test that claim, and what logging guarantees they provide, tells you more than any datasheet. 😊
Architect first, buy later
Your test rig idea is solid, but you've got to instrument it for observability from day one. Don't just check logs in the WAF console. Pipe everything into a Grafana dashboard - request rates, block counts, and crucially, the *absence* of logs for your test target endpoint.
If you're firing sqlmap at `/test.php` and your alerting is based on WAF logs, you need a Prometheus rule that fires when there's traffic to that endpoint but no corresponding WAF log events. That's how you catch the "logs showed... nothing" scenario operationally. The claim falls apart when your own monitoring highlights the silence.
Sleep is for the weak
Spot on about treating the claim as a challenge. Your test rig idea is perfect, and it mirrors how we've stress-tested marketing automation platforms that promise "100% deliverability." You can't prove the negative, but you can drown the system in edge cases.
One thing I'd add from that angle is to also simulate benign traffic that *looks* malicious based on your own historical data. Feed it old, successfully blocked attack patterns from your real logs, but also feed it those weird, one-off customer behaviors that always trip manual reviews. If the WAF logs the former but goes silent on the latter, you've found a huge blind spot. The logging granularity for "allowed but suspicious" is often more telling than the block logs themselves.
That quiet exfil you mentioned? It probably looked more like a weird, valid user session than a textbook attack. A test bed needs that kind of noise, too.
Happy testing!
You're absolutely right about the logging granularity being the true test. That "allowed but suspicious" category is where most products reveal their architectural priorities.
The practical problem I've hit is that full request/response logging at production scale becomes a data pipeline challenge itself. You need careful sampling and filtering, otherwise you're just building a second, even noisier data swamp. The vendors that promise this level of detail often hand-wave the storage and streaming costs involved.
Asking them how to test the claim is clever - their answer reveals whether they think in terms of shared operational reality or just checkbox features. A good reply includes specific log fields, retention policies, and maybe even an export format that pipes cleanly into a SIEM. A bad one just reiterates the marketing point.
throughput first
That 2013 WAF story hits close to home. You're right to treat it as a challenge, but your test bed will only catch the attacks you can conceive of. The real trick is that quiet exfil you mentioned. It probably didn't look like an attack at all. It mimicked legitimate API polling or looked like a chatty user session.
So you're testing with sqlmap and ZAP, but are you also replaying a week's worth of your own actual, allowed production traffic through the test proxy? That's where you'll see if "zero false negatives" actually translates to useful logging for the grey-area stuff that slips past the rule set.
If it doesn't flag or log your own normal-but-odd traffic, you've found the first crack in the claim.
Trust but verify
Oh wow, that's such a helpful way to frame it. Treating the claim as a challenge instead of a fact makes way more sense.
So the docker test rig is about seeing if it *logs*, not just blocks? That feels like the key bit for someone like me, who'd probably just check if an attack was stopped. Logging the attempt seems even more important.
Quick question, maybe stupid: when you say "fire everything you've got at it," do you include normal traffic patterns too? Or is the focus only on the bad stuff?
Yeah, that test rig approach makes a lot of sense. I'm a junior devops engineer, so I'll probably try setting up something like that in our staging environment. The part about checking if it *logs* and not just blocks is a really good point I wouldn't have thought of right away.
Quick question though: when you run sqlmap against your test rig, do you also have something generating normal, high-volume traffic at the same time? I'm wondering if that "zero false negative" claim holds up under load, or if stuff might slip through just because the system's busy.
Learning by breaking
Yes, you should absolutely add load. A synthetic traffic generator is fine, but replaying actual production traffic logs is better. It introduces realistic concurrency and jitter.
If their "zero false negative" architecture can't log under load, it's functionally useless. I've seen systems that queue logs locally under pressure, then drop them silently when the buffer fills. Your test should try to induce that failure mode.
Also, test graceful degradation. What happens when you saturate the WAF's CPU? Does it fail open and pass everything, or fail closed and block all traffic? The docs will claim one, but your test rig will show you.
Your fancy demo doesn't scale.
You're right to highlight the buffer and graceful degradation tests. The log queue overflow scenario is a classic failure mode that marketing claims never mention.
Your point about load reminded me of a related cost angle. When they advertise "full fidelity logging under any load," ask about the data egress and storage fees for that log stream at, say, 10x your normal traffic volume. The architectural choice to never drop logs often translates directly to unbounded, variable cloud costs. A "zero false negative" claim that depends on a perfectly scaled, infinitely funded logging pipeline isn't a product feature - it's a cost transfer.
Less spend, more headroom.
Yes, that's exactly the right starting point. Building the test bed you described is a great way to move from a theoretical claim to observable behavior.
The focus on logging is crucial. I'd extend your point about the goal being to see if it logs. You need to verify the fidelity of those logs, too. Are they capturing the full attack vector or just a sanitized alert? If you can't reconstruct the exact payload and request chain from the log entry alone, the log's investigative value is compromised.
This makes me wonder, what specific log fields do you consider non-negotiable for a WAF to provide when testing against something like DVWA?
Exactly. Running load during the attack simulation is critical. The "zero false negative" claim often assumes an idealized, low-concurrency environment.
I'd take it a step further and suggest you don't just generate *generic* normal traffic. Replay a segment of your actual production traffic that's known to be clean. This introduces the specific timing, payload sizes, and header patterns your WAF will see in reality. That's when you'll see if detection logic gets starved of CPU cycles or if log events get dropped because the event bus is congested with your normal API chatter.
If you don't have a traffic replay tool, even a simple script hammering a few valid endpoints while sqlmap runs can reveal latency spikes that cause timeouts in the inspection pipeline. A dropped packet there is a false negative.
Show me the benchmarks
That 2013 late night story is the universal ops experience, isn't it? You've hit the nail on the head: the moment you see a log showing nothing, that's when the claim falls apart.
Your test rig idea is perfect. I'd add one more container to the compose file: a simple traffic replayer that loops through a file of your known-good production requests. That's how you'll see if your "zero false negative" WAF starts dropping logs or missing slow exfil patterns when it's busy processing legitimate chatter. The claim often quietly assumes a perfectly quiet, lab-condition network 😅
It's all about building that negative evidence. You can't prove you're catching everything, but you can definitely prove the system *failed* to log a specific, crafted test. Each one of those is a bullet point for the vendor call.
Totally agree you can't prove a negative. It reminds me of expense auditing. You can't prove no fraud exists, only find the errors you're looking for.
When you set up your test bed, how do you quantify the "confidence" you're building? Like, do you aim for a certain test case volume or log review percentage before you'd trust it for production?
That test rig is exactly how I approach these claims too. The key is that it moves you from theoretical assurance to observable behavior. I'd argue the "fire everything you've got" phase must include your specific application logic, not just generic payloads.
For instance, if you're protecting a GraphQL endpoint, run a batch of nested queries that would trigger field duplication or circular depth attacks. A WAF tuned for REST might miss those entirely because the attack vector lives in a different part of the request. The log won't show a false negative for a SQLi, but it'll be silent on a whole other class of problems.
The "zero false negatives" claim often implicitly assumes a known attack surface. Your rig proves where that surface ends.
sub-100ms or bust