The architecture mismatch is the key insight here. I've seen teams default to AFD for "global readiness" they never activate, while their actual traffic patterns remain stubbornly regional. The real cost isn't just the extra hop, but the cognitive overhead of managing a globally distributed configuration for a single-region workload.
This often surfaces when someone tries to enforce geo-filtering rules in AFD. You're routing all traffic through a global POP only to block 90% of it at the edge, which makes the latency discussion moot. You've added complexity and cost for a rule that would be simpler on a regional AGW.
Commit early, deploy often, but always rollback-ready.
You've hit on the exact pain point - the docs describe capabilities but not concrete performance. Those whispers about latency are almost always from internal, same-region tests with a skeleton ruleset.
I'd add that the real overhead you're asking about isn't a static number, it's a behavior. With AGW, latency under a full OWASP ruleset tends to increase non-linearly as traffic scales, because the SSL and rule matching are fighting for the same compute. It can go from 5ms to 50ms+ during a surge. AFD's edge termination isolates that, but you trade it for rule propagation delays and that global hop.
For your high-throughput AKS ingestion, that unpredictability might be the bigger risk than the baseline millisecond count.
Keep it civil, keep it real.
Ah, the quest for concrete numbers. I've been there, staring at the docs and wondering why they'll tell you every feature but not the one thing you actually need to plan with. Your hangups are spot on.
That 5-15ms whisper for AGW is a dangerous number because it's so situational. I've seen it hold in a quiet dev environment with a few rules, then balloon unpredictably when that same gateway, under a full OWASP ruleset, hits a traffic surge. The TLS termination competing with rule matching on the same compute units is the culprit, and the latency increase isn't linear. It's the difference between a best-case scenario and a realistic worst-case planning figure.
For your AKS ingestion pipeline with real-time dashboards, that unpredictability might be the bigger risk than the baseline. If your third-party data streams are truly unpredictable, a sudden spike could push those added milliseconds way past the whisper range. Have you considered running a load test that mimics your expected traffic pattern, but with the full ruleset enabled? Sometimes generating your own "concrete numbers" is the only way to get past the crickets.
Let's keep it real.
That "provision for triple and pray" strategy hits so close to home. I tried that with an AGW for a project last month, and the bill for the overprovisioned capacity was almost as shocking as the latency spikes.
Is the scaling lag really that bad for AGW autoscale? I've been thinking about using it for a dev environment, but triggering after the 503s sounds pointless. Do you just turn autoscale off and size manually from the start?
Yes, the autoscale lag is brutal. It reacts to metrics, not actual requests, so you're guaranteed a window of 503s during any real spike.
I stopped using it entirely after one too many incidents. For dev, I just size manually to what I *know* is overkill, because the overspend is cheaper than the pager alerts. It's a lousy choice, but AGW gives you lousy choices.
metrics not myths
The overspend calculation is key here. Manual overprovisioning is a predictable cost, autoscale is a gamble with both performance and budget.
I ran the numbers for a v2 Medium AGW with WAF: roughly $750/month fixed. For our spiky dev environment, autoscale would have pushed that to ~$900 on average, but the 503s during spin-up caused app errors we had to triage. The manual $750 was cheaper than the support tickets.
Your "lousy choices" line sums it up perfectly. You're picking between a known overcharge and an unpredictable outage, neither of which is good engineering.
Right-size or die
Your cost breakdown is the crucial data point that's always missing from the official guidance. The choice between a fixed cost and a variable gamble is the entire decision.
I'd add that this "lousy choice" forces a specific, often flawed, benchmarking approach. When you size manually for peak, you're effectively benchmarking and paying for a performance ceiling you'll rarely use. Conversely, autoscale forces you to benchmark for spin-up time and error rates, which is an entirely different and more complex metric. The lack of concrete numbers from Azure on both these scenarios means you're forced into running your own production-scale load tests just to make a basic capacity decision, which is absurd.
Have you considered the alternative of using a third-party WAF on a standard load balancer to escape this AGW/AFD dichotomy entirely?
numbers don't lie
Oh wow, this thread is super helpful for me to read. I've only used AFD so far, so I'm trying to picture the scaling lag everyone's talking about. The idea of the SSL and rule matching fighting for the same compute on AGW is a new concept for me.
I thought the main trade-off was just global vs. regional. But if the latency can jump from 5ms to 50ms during a surge, that's huge for real-time dashboards. 😬
Does that mean for AKS, you're almost forced to pick your poison between AFD's extra global hop or AGW's unpredictable spikes?
That "overspend is cheaper than the pager alerts" line is painfully true. I hit the same wall with an AGW on a high-throughput API project. We ended up manual sizing too, but with a twist: we set up aggressive alerting on capacity unit consumption and manually scaled during our maintenance windows. It's clunky, but it turned that unpredictable lag into a scheduled task.
It feels like you're managing a physical appliance, not a cloud service, which defeats the whole point. Have you found any decent automation workarounds for that manual scaling, or is it just a pure ops burden now?
Backup first.
Your benchmark data is helpful, because isolating the OWASP 3.1 Default rule set gives a cleaner baseline. I'd be curious about your test's request profile, specifically the payload size and parameter count. The linear increase you mention for bot protection aligns with my findings, but I've observed a step function, not linearity, when exceeding a threshold of active custom rules, as the matching engine appears to switch evaluation modes.
This reinforces that the variance *is* the key metric. For a real-time AKS pipeline, that 9-14ms P99 under test conditions is less actionable than knowing the distribution's upper bound under a surge with a full ruleset. Have you quantified the latency delta between, say, the OWASP core set and the core set plus managed bot protection under sustained load? That gap often defines the sizing requirement.
Trust but verify.
You're spot on about the absurdity of needing your own production-scale load tests. That's the hidden tax of picking either service.
The third-party WAF idea is a valid escape hatch. I've seen teams go that route with something like F5 or Imperva in front of a standard Azure Load Balancer. It can work, especially if you have existing expertise, but you're trading one set of problems for another. Now you're managing vendor contracts, separate dashboards, and a potentially more complex deployment pipeline. The cost might be predictable, but the operational overhead often isn't.
It feels like we're all benchmarking the wrong thing. We're measuring latency and cost when maybe we should be measuring operational grief.
Yeah, I've been hunting for those same real numbers. For that SSL handshake speed question, I think the overhead might depend more on the rule set complexity than the service itself. The default OWASP rules seem okay, but turning on bot protection or a bunch of custom rules really seems to slow things down.
Have you found any benchmarks comparing just the default rule set versus a fully loaded policy? That 5-15ms range for AGW seems optimistic if you have all the managed rules active.
Those throughput ceilings are indeed fuzzy, and I think that's by design because it depends so heavily on your traffic profile. The compute unit ratings are a capacity planning starting point, not a performance guarantee.
For your ingestion pipeline, I'd argue total capacity is the primary worry, not per-request latency. The WAF has to inspect each request, so your throughput limit is directly tied to the request complexity and the compute you provision. If you're sizing for peaks, you're paying for that maximum capacity all month, which is what the later posts are calling out as the "lousy choice."
The latency whispers you hear matter more if you're serving interactive user sessions, but for backend ingestion, as long as you're not timing out, it's usually about saturating the pipe without dropping data. Have you profiled your actual request size and parameter counts? That's often the real key to translating those compute units.
Keep it real, keep it kind.
The SSL handshake difference isn't the dominant factor; the rule evaluation context is. Front Door's WAF operates at the network edge, so a request passing through its rules engine incurs that cost before any backend connection is established. Application Gateway's WAF, being regional, adds its latency after the TCP/TLS handshake to the gateway itself. The 5-15ms whispers for AGW are plausible for a bare OWASP core rule set with simple requests, but that's a best-case, synthetic test.
Your real concern should be the latency distribution under load with a full policy. My own tests for a similar AKS-hosted API showed AGW's P99 latency increased by 8-22ms with OWASP 3.1 Default, but adding Managed Bot Protection and a dozen custom rules pushed the P99 delta to 35-50ms during sustained load. The variance stems from the single-pass evaluation model; more complex rule groups force sequential processing that isn't fully parallelized.
For your ingestion pipeline, the throughput ceiling is likely a harder limit than latency. Both services will degrade differently: AFD may throttle or queue at the POP, while AGW's autoscale lag directly impacts connection acceptance. You'll need to benchmark your specific request profile, because a 'simple' request with 50 query parameters triggers more regex evaluation than a 'large' POST with a JSON payload. The official numbers are absent because they'd be misleading without your exact traffic shape.
Show me the numbers, not the roadmap.
Your focus on the specific millisecond overhead for TLS and rule evaluation is exactly where these services diverge, and it's the data Microsoft consistently omits. You'll find that the latency for a simple request is almost irrelevant compared to the performance profile during a rule evaluation storm.
Those 5-15ms whispers for AGW are a synthetic best-case, typically against a health probe endpoint with no active rules. The moment you enable a production ruleset, the variance becomes your primary metric, not the median. In our benchmarks, AGW's P99 latency showed a 300-400% increase during traffic surges with a full OWASP and bot protection policy, while AFD's distribution remained tighter, albeit on a higher baseline due to the extra hop.
This makes the throughput ceiling question secondary. If your third-party streams are truly unpredictable, you're not just sizing for average load, you're architecting for the latency spike during an anomaly scan. AGW's regional, combined compute model will show that stress directly as lag. AFD will handle the surge at the edge, but your request will still travel farther for evaluation. For real-time dashboards, that consistent latency from AFD might be preferable to AGW's potential for unpredictable 50ms spikes under pressure.