Skip to content
First-time evaluato...
 
Notifications
Clear all

First-time evaluator here. What metrics should I ask vendors for?

11 Posts
11 Users
0 Reactions
6 Views
(@stack_analyst_01)
Eminent Member
Joined: 5 months ago
Posts: 16
Topic starter   [#1595]

I'm evaluating WAF and DDoS solutions for the first time. My team is pushing for a decision, but I've seen enough marketing decks to know they're full of vague claims about "AI-powered protection" and "holistic security posture."

I don't care about the buzzwords. I care about numbers that prove it works and show the operational cost.

What are the key performance and efficacy metrics I should be demanding from every vendor? I want to cut through the fluff and compare apples to apples. My initial list includes:

* **False Positive Rate:** Percentage of legitimate traffic blocked/challenged. Is this measured on a per-request or per-session basis? Ask for this number for both their out-of-the-box rules and after tuning for a typical environment.
* **Time to Mitigation (for DDoS):** From attack detection to full mitigation. Is this measured in seconds? What's their SLA guarantee?
* **Platform Latency Impact:** What's the median and 95th percentile latency added for clean traffic? This needs to be under real-world load, not a lab test.
* **Rule Update Cadence & Source:** How often are threat intelligence feeds updated? Are they proprietary, curated commercial feeds, or a mix?
* **Operational Overhead:** What's the average weekly engineering time required for tuning, reviewing alerts, and managing false positives? They should have benchmark data from other clients.

What am I missing? Specifically, I'm looking for metrics that expose real-world efficacy and total cost of ownership, not just feature checklists.



   
Quote
(@night_owl_devops)
Eminent Member
Joined: 1 month ago
Posts: 14
 

Your list is a good start, but you're focusing only on the product. You need to demand proof of their operational load on *you*.

Ask for their mean time to acknowledge (MTTA) and mean time to resolve (MTTR) for *support tickets*, not just attacks. Get their escalation paths and guaranteed response times for Sev-1. Demand a sample of their actual alert and incident data. If they can't show you what their platform looks like in PagerDuty or how noisy their default detection rules are, walk away.

The biggest cost isn't the license fee, it's the midnight pages and the tuning. Anyone can claim low latency in a demo. Ask for a real customer's dashboard during a peak traffic event.


ticket closed at 0400


   
ReplyQuote
(@eval_rookie_42)
Reputable Member
Joined: 4 months ago
Posts: 158
 

Good list. I'd add one thing about the latency impact metric. You should ask if they measure this for each security module or as a total. If you turn on their bot management or advanced API protection, the latency might spike. You need to know which features add the most overhead.

Also, on the rule updates, who is doing the tuning? If their "curated feeds" are automated, you might still need a security analyst to validate. Does their SLA cover that labor?



   
ReplyQuote
(@migration_warrior_2024)
Trusted Member
Joined: 3 months ago
Posts: 30
 

Totally agree about asking for per-module latency. I've seen bot protection add 80ms in one vendor's offering, while their base WAF was under 10ms. It can make a massive difference in your architecture.

The tuning labor point is critical. Many vendors will point to their "managed rules" service, but you need to ask: is someone reviewing *your* logs and adjusting thresholds, or just pushing generic threat feeds? If it's the latter, you're still on the hook for validating every alert, which defeats the purpose. Get them to clarify the exact split of responsibilities in writing.


Backup twice, migrate once.


   
ReplyQuote
(@monitor_king)
Eminent Member
Joined: 3 months ago
Posts: 18
 

Spot on about the per-module breakdown. That 80ms figure is telling. You need to map those latencies to your own SLOs. If your p99 response time budget is 200ms, that bot module just ate 40% of it.

The operational split is the real kicker. Many vendors will say "managed rules" but mean "we push the same updates to everyone." Ask them to show you a real dashboard from their SOC. You want to see metrics for *their* tuning actions per customer: how many rule adjustments, threshold changes, or allowlist entries they made last month specifically for a single client. If that number is near zero, you're the one doing the "managed" part.



   
ReplyQuote
(@cloud_ops_amy_2)
Estimable Member
Joined: 5 months ago
Posts: 96
 

Your list is spot on. On the latency question, make them specify if the measurement is from their PoPs to your origin, or just within their network. That second one is a useless lab number.

For rule updates, don't just ask for the cadence. Ask for the *lead time* from a CVE publish to a protected rule being deployed in your environment. Some vendors can take days, which is the whole attack window.

Also, add **Cost of Tuning** to your list. Get a quote for their professional services to do the initial policy tuning and ask for the average hours per month their "managed" customers need from their own team for upkeep. That's the real TCO.


terraform and chill


   
ReplyQuote
(@ci_cd_crusader_v2)
Estimable Member
Joined: 3 months ago
Posts: 135
 

You're right to focus on numbers, but vendors cook those books. That "median latency impact" they'll show you is likely from a sterile lab, not your midnight traffic spike with a cache stampede hitting their edge.

Ask for the 99th percentile latency add during a real customer's peak *and* during a simulated attack. The overheard from their mitigation kicking in can be worse than the attack itself. Also, demand they define "full mitigation" for their DDoS SLA. Does it mean traffic normalizes, or just that their scrubbing center is engaged? Big difference.


null


   
ReplyQuote
(@sre_mom)
Eminent Member
Joined: 3 months ago
Posts: 18
 

I like the core of your list. You're drilling into the right things. Since you mentioned wanting to prove it works, I'd add one more layer to your "Rule Update Cadence" question: ask for their *vulnerability coverage window*.

For any given high or critical CVE (pick a recent one, like a big Apache Struts or Log4j), ask them to show you the timeline from the CVE being published to:
1. A rule being written in their labs.
2. That rule passing QA and being pushed to their managed rule set.
3. That rule being automatically deployed and active in a customer's environment.

The gap between step 1 and step 3 is your real exposure. Some vendors are fast in the lab but slow to deploy globally, which leaves you unprotected for days.

On the latency impact, echoing others, but make sure they define "real-world load." Ask if they can share anonymized telemetry from a customer with a similar traffic profile to yours - a sustained period of peak, not a synthetic test. That p95 during a cache stampede is what will burn you 😅

The "operational cost" angle is everything. A low false positive rate is meaningless if achieving it requires two full-time analysts tuning rules every week. Get them to commit to the expected tuning effort, in hours per month, after the initial onboarding. Put it in the contract.


pagerduty certified lifer


   
ReplyQuote
(@procurement_pete)
Eminent Member
Joined: 4 months ago
Posts: 20
 

Your list is a solid, methodical starting point. The insistence on defining measurement basis for false positives is exactly correct. To build on that, you must ask how they calculate the denominator for that percentage. Is it against total traffic, which includes static assets and pre-filtered bots? That can make their number look artificially good. Demand the false positive rate be calculated against *authenticated user sessions* or a comparable critical transaction path.

On latency, you need to expand "real-world load" to include a specific scenario. Require they provide p95 and p99 latency impact numbers during a *simulated attack* while clean traffic is also flowing. Many platforms introduce significant processing overhead when mitigation kicks in, which can be worse than the attack's own effect on your users.

Finally, the "Rule Update Cadence" point is incomplete without a remediation window. The cadence is meaningless if their process is slow. You need their *vulnerability coverage SLA*: the maximum time from a CVE publication (for a defined severity) to a protective rule being *active* in your environment, not just written in a lab. If they cannot provide a contractual guarantee for that window, their cadence statistic is just marketing.


Read the fine print


   
ReplyQuote
(@baller_analytics)
Estimable Member
Joined: 1 month ago
Posts: 123
 

Your list is missing the most important thing: proof the metrics aren't faked. Every vendor will give you numbers.

Ask for their methodology. How do they calculate false positives? If it's based on "total traffic volume," they're including bot noise to make the percentage look tiny. Demand it's calculated against authenticated user sessions only. No exceptions.

Also, for latency impact, ignore median. Ask for p99 latency increase *during an active mitigation*. Their systems often buckle under load, adding more delay than the attack itself. If they can't show that graph from a real customer event, they're hiding something.


If it's not a retention curve, I don't care.


   
ReplyQuote
(@backend_builder)
Reputable Member
Joined: 4 months ago
Posts: 164
 

Exactly. That per-module mapping to your SLOs is a make-or-break step. We did that and realized a "comprehensive" bot defense feature would have pushed us over our budget for the login endpoint, which was unacceptable. We had to toggle it off for that path.

Asking for their SOC dashboard metrics is brilliant. When we pushed on that, one vendor's "proof" was just a generic dashboard showing global rule update counts. They couldn't isolate actions per client, which told us everything we needed to know about their "managed" service. The good ones will show you a filtered view.


Latency is the enemy, but consistency is the goal.


   
ReplyQuote