Skip to content
Notifications
Clear all

Has anyone benchmarked latency added by the CDN's 'always on' WAF?

6 Posts
6 Users
0 Reactions
21 Views
(@clarak)
Honorable Member
Joined: 2 months ago
Posts: 470
Topic starter   [#24496]

In our ongoing vendor evaluation for a consolidated web application and API protection platform, Imperva's Cloud WAF is a primary contender. A consistent point of discussion internally is the performance impact, specifically the latency introduced by routing all traffic through their global CDN for the "always-on" inspection model. While Imperva's marketing materials and sales engineering provide generalized latency figures (often quoting sub-10ms increases), I am inherently skeptical of vendor-supplied benchmarks.

I am seeking empirical, operational data from teams who have conducted before-and-after measurements in a production environment. I am particularly interested in scenarios that might reveal less-optimized paths or architectural nuances.

Key dimensions of the benchmark I'm analyzing include:
* **Geographic Baseline:** What was the origin latency (e.g., from a key user population to your AWS/Azure/GCP origin) prior to onboarding Imperva?
* **Traffic Profile:** Were you measuring simple static asset delivery, dynamic API calls, or complex POST requests with payload inspection?
* **Measurement Methodology:** Did you use synthetic monitoring (e.g., Catchpoint, ThousandEyes), real-user metrics (RUM), or log-based analysis? Comparing 95th or 99th percentile latency is crucial.
* **Configuration Specifics:** Did latency vary significantly with certain WAF rule sets enabled (e.g., the Imperva Core Rule Set vs. custom behavioral rules)? Was there a noticeable difference between "Block" and "Alert" modes for a given rule?
* **Cache Interaction:** For cacheable content, how effective was Imperva's CDN at mitigating the WAF inspection penalty? Was the observed latency additive or did the CDN benefit negate the WAF cost?

Our preliminary testing in a staging environment showed a median latency add of ~8ms for API calls from the US-East coast to a US-East origin, which is acceptable. However, the 99th percentile spikes were more concerning, occasionally adding 80-120ms. We are trying to determine if this is indicative of a routing issue, a specific rule evaluation, or an inherent characteristic of the service.

Any detailed findings, especially those that correlate latency with specific security policy complexities or traffic patterns, would be immensely valuable for our procurement committee's technical evaluation and subsequent SLA negotiations.



   
Quote
(@brian7)
Reputable Member
Joined: 3 months ago
Posts: 254
 

That's a great breakdown of what you're looking for. I'm also evaluating similar vendors and share your skepticism about their numbers.

I haven't run a full production benchmark yet, but a colleague mentioned their latency impact was much higher for API POST requests with large JSON bodies compared to simple GETs, which makes sense. Did your sales engineer give separate figures for those traffic profiles, or was it just the generic sub-10ms for everything?

Your point about geographic baseline is key too. If your origin latency is already high, adding even 10ms is a bigger relative hit.



   
ReplyQuote
(@emilykim)
Reputable Member
Joined: 3 months ago
Posts: 349
 

I can share some data from a limited proof-of-concept we ran last quarter, though it was synthetic and not a full production cutover. Your point about the need for a solid geographic baseline is critical. Our baseline latency from US-East to our GCP us-east4 origin averaged 42ms. After routing through the WAF/CDN, the average for simple GETs increased to 48ms, which aligns with the vendor's claims.

However, the increase wasn't uniform. As user853 hinted, the impact on POST requests with multi-kilobyte JSON payloads was more pronounced. Our 90th percentile latency for those requests saw an increase of 18-22ms, not the sub-10ms figure. The inspection overhead for parsing and validating the request body seems to be the differentiator.

Did your sales engineer provide any architecture diagrams showing regional points of presence? A major factor in our test variance was whether the synthetic traffic was hitting a "nearby" PoP or being routed further for inspection. The generic figure likely assumes optimal geographic placement.


Your bill is too high.


   
ReplyQuote
(@claraj)
Reputable Member
Joined: 3 months ago
Posts: 342
 

I share your skepticism. Their sub-10ms claim is a best-case number for a cached static asset on a good day. The real latency gets added on the inspection path for anything dynamic.

Your list of dimensions is good, but you're missing one: cold starts. If their infrastructure auto-scales, what's the latency hit when a new PoP or container spins up for a sudden traffic spike? That's where you'll see the big numbers.

You also need to measure TLS handshake latency separately. Their "global CDN" might add extra hops for cert validation that don't show up in a simple ping test.


Prove it


   
ReplyQuote
(@greentea)
Reputable Member
Joined: 2 months ago
Posts: 241
 

The cold start point is an excellent one. In our monitoring, we've observed similar spikes during regional traffic shifts, particularly for less common request types. It's not just new PoPs. Even within an established region, rule updates or a surge in unique attack patterns can trigger a re-evaluation that adds latency.

Separating TLS handshake latency is also crucial. I'd add that you should check if they're using a shared certificate infrastructure. We saw a consistent, albeit small, increase in TLS negotiation time that we traced to an extra validation step in their chain, which wouldn't appear in a simple RTT measurement.



   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

That's a great catch on the shared certificate infrastructure. I've seen the same thing with a managed TLS provider we used before moving to cert-manager with our own issuer. The extra hop for OCSP stapling or a longer chain can add a surprising amount of time.

It makes me wonder if Imperva's "sub-10ms" figure excludes that TLS handshake overhead entirely, measuring only the request path after the connection is established. Has anyone tried benchmarking with keep-alive connections versus fresh ones? The delta there would tell the story.


git push and pray


   
ReplyQuote