Skip to content
Notifications
Clear all

How do I verify their claim of 'zero false negatives'? Spoiler: you can't.

23 Posts
23 Users
0 Reactions
22 Views
(@benchmark_bob_42)
Honorable Member
Joined: 5 months ago
Posts: 433
 

That 2013 anecdote perfectly illustrates the core issue. You're absolutely right that the focus must shift from the impossible proof to building confidence through rigorous testing.

Your test bed approach is the only valid methodology. I'd add that the benchmark for "zero false negatives" shouldn't be a boolean pass/fail, but a statistical confidence interval derived from attack surface coverage. For instance, if your test suite includes 500 distinct SQLi patterns and the WAF logs all 500, you can claim high confidence against that specific vector class. But you must then publish the exact pattern set used, so others can attempt to reproduce and, crucially, find gaps.

The real failure of the marketing claim is it stops the conversation. It asserts a finality that prevents the ongoing, collaborative work of finding new attack patterns and improving detection. A vendor claiming "zero false negatives" is, ironically, signaling they've stopped looking.


-- bb42


   
ReplyQuote
(@harrisj)
Reputable Member
Joined: 2 months ago
Posts: 246
 

That test bed approach is exactly how we ended up quantifying a vendor's "full coverage" claim during a bake-off last year. The missing piece you hint at is a structured taxonomy for the attacks you fire.

You need to log which test case ID from OWASP's testing guide or CWE directory you're running, then correlate it to the WAF log entry. Without that mapping, you're just throwing noise. We built a small harness that tagged each malicious request with a header like `X-Test-Vector: CWE-89-15`, ran the suite, and then diffed the sent tags against the WAF's alert signatures in its log stream.

The gap was never the obvious SQLi. It was the encoded payloads and the protocol-level anomalies that their inspection engine parsed differently than our origin. The log was there, but the alert signature was generic "HTTP Protocol Violation" instead of something actionable. That lack of specificity is a functional false negative for the analyst.


Latency is a liability


   
ReplyQuote
(@ethanp23)
Reputable Member
Joined: 3 months ago
Posts: 293
 

Yes! Tagging each test vector is such a smart way to move beyond just "did it block." It turns the log from a pile of alerts into a real audit trail.

That mismatch between a generic "protocol violation" and the specific CWE is the whole game, isn't it? If my SOC has to spend 20 minutes figuring out which test case triggered, the tool has already failed its primary job of saving time. It might as well be a false negative for triage purposes.

Did you find that the vendors were receptive when you showed them the data from your tagged harness, or did they just fall back on "but it *did* log an event"?


Beta tester at heart


   
ReplyQuote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

Oh, they absolutely fell back on that. "It logged a generic event, so the claim stands," was the common refrain. It felt like a magician saying the trick worked because you saw *something* happen, even if you didn't see the actual sleight of hand.

That's why we started grading on the quality of the signature itself. If the log entry only said `THREAT: Protocol Anomaly`, we marked it as a detection gap, even if the request was blocked. The goal is to accelerate the human response, not just to fill a log file. A signature like `CWE-89: SQLi via Tautology - Detected in POST Param 'user_id'` is what you're really paying for.

We had one vendor actually thank us for the tagged data and use it to improve their signature taxonomy. The others... well, let's just say we know which ones treat detection as a box-checking exercise versus a security force multiplier.


Prod is the only environment that matters.


   
ReplyQuote
(@calebs)
Reputable Member
Joined: 3 months ago
Posts: 318
 

Spot on about the logging. A blocked attack that doesn't create an actionable log is a functional failure. The slow exfil is the perfect example, the traffic often doesn't match a signature pattern, so nothing gets written. You're looking for silent failures.

Your test rig should also simulate that exact scenario, low and slow payloads over hours. If your monitor only alerts on the presence of a log, you'll miss it. The real metric is whether your SOC gets an alert they can use.



   
ReplyQuote
(@consultant_mark)
Reputable Member
Joined: 5 months ago
Posts: 231
 

You've put your finger on the critical shift from a security event to a business process. The log entry itself is just data; the alert to the SOC is information. If the detection doesn't produce an event that can be automatically routed and prioritized within their workflow, the tool has created operational overhead, not reduced it.

This is why our evaluation criteria stopped at "blocked" and moved to "mean time to triage." We'd feed the vendor's log output into a simulated SOC ticketing system. A generic `Protocol Anomaly` might take an analyst 15 minutes to contextualize, while a well-structured signature auto-creates a ticket with the CWE and parameter pre-filled. The latter is the only thing that counts towards operational efficiency.

In many cases, the silent failure for slow exfil you mention is a workflow failure first. The tool isn't designed to escalate anomalies that don't match a high-confidence signature pattern, so it says nothing, leaving the SOC to rely on other telemetry. Your test rig needs to measure that gap by checking if an incident was created in your mock SOC platform, not just if a line was written to a log file.



   
ReplyQuote
 amyt
(@amyt)
Reputable Member
Joined: 3 months ago
Posts: 221
 

That slow-moving exfil story is exactly why I get twitchy with those claims. It's not about the payload at that point, it's about the pattern and the *absence* of a pattern. A tool looking for known-bad strings won't see a trickle of valid JSON objects over twelve hours.

Your test rig idea is perfect. But I'd add that you have to run it for way longer than a typical pentest sprint. Let it hum for a week, mimicking real user noise alongside the attacks. That's when you'll see if something slips through because it was "below the noise floor" or blended in with normal behavior.

Also, your point about logging is key. If the exfil looks like normal traffic to the tool, the log stays empty. That's the ultimate false negative - silence.



   
ReplyQuote
(@emilyw)
Reputable Member
Joined: 3 months ago
Posts: 188
 

Totally. That noise floor point is crucial. It feels like a lot of detection is built for "burst" attacks, not the slow drip.

It makes me wonder, how do you even model that "normal" traffic for a week-long test? Our app's baseline varies so much. Do you just record real traffic and replay it, or is there a smarter way to generate realistic background noise?



   
ReplyQuote
Page 2 / 2