Hey folks, been wrestling with this all week while designing a new ingestion pipeline that needs serious WAF protection. We're architecting a system where our analytics API endpoints (hosted on AKS) will be exposed to some pretty unpredictable third-party data streams. Security is non-negotiable, but so is latency for our real-time dashboards.
I've used both Azure Front Door (AFD) with its WAF policy and the Application Gateway (AGW) with its WAF (v2, of course) in the past, but always in different contexts. Now I'm trying to make a direct, performance-centric comparison for a high-throughput scenario and the official docs are... less than helpful on concrete numbers. They talk about features and scaling, but the real-world overhead? Crickets.
My specific hangups are about the added milliseconds and potential throughput ceilings. For instance:
* **TLS Termination Overhead:** Both can do it, but is there a measurable difference in SSL handshake speed or CPU impact when the WAF rulesets (OWASP 3.2, say) are fully loaded?
* **Latency per Request:** What's the baseline added latency for a simple request/response passing through the WAF engine? I've seen whispers of 5-15ms for AGW, but is AFD different given its global anycast architecture vs. regional deployment?
* **Throughput Impact:** If I have a backend that can handle 10k RPS naked, what's a realistic expectation after putting it behind each WAF? Does the fact that AFD's WAF seems to be "closer to the edge" change the performance profile compared to AGW sitting in our VNet?
* **Rule Complexity Cost:** How much does enabling a complex custom rule (like a bunch of regex patterns for threat intel feeds) chip away at performance? Is one platform more sensitive to this than the other?
I cobbled together a basic test, but it's hardly production-grade. Used `wrk` against a dummy app, but I'm not sure I'm even measuring correctly.
```bash
# Rough test for baseline latency addition
wrk -t12 -c400 -d30s --latency https://myapp.azurefd.net/api/health
```
Has anyone done rigorous benchmarking or have internal numbers they can share? I'm especially curious about:
- The performance hit when geo-filtering or IP reputation rules are active.
- Any difference in behavior under sustained attack (like a DDoS that triggers rate limiting rules).
- Whether the "Managed Ruleset" in AFD performs differently than the identical-looking ruleset on AGW.
We're leaning towards AFD for its global scale and DDoS protection tier, but if the WAF adds significantly more latency than AGW for our primary region users, that's a dealbreaker. The devil's in the details, and right now, I'm missing a lot of details!
Data nerd out
Data nerd out
I'm a senior platform engineer at a mid-sized SaaS shop (120ish people) doing clickstream analytics, and we've run both AFD and AGW WAF in front of our data collector APIs on AKS for the last two years. We settled on a split setup based on traffic patterns.
**Latency overhead per request (hot path):** For a simple validated POST, our observability shows AGW WAF v2 adds a consistent 8-12ms to P95 latency with OWASP 3.1. AFD typically adds 3-7ms for the same request, likely because its POP is geographically closer to our users than our AGW's single region.
**TLS termination and rule set penalty:** The CPU hit matters more on AGW. Turning on the full OWASP rule set with bot protection added about 15% more latency to SSL handshakes on our Application Gateway Standard_v2 instances during load tests. AFD, being a managed service, made that penalty essentially invisible to our metrics, though you're paying for the privilege.
**Throughput ceiling before pain:** An AGW WAF v2 on two large instances nominally handles 2-3k RPS before we saw latency spikes and needed to scale. AFD's autoscaling handled a burst of 12k RPS for us without blinking, but that month's bill had a line item that made my manager sigh.
**Operational gotcha you only learn by doing:** AGW requires you to manage certificates, backends, and scaling yourself; a misconfigured health probe can silently drop 20% of your traffic. AFD abstracts that away, but debugging a false positive from its WAF is a black-box nightmare of support tickets and waiting. The AFD WAF custom rule syntax is also noticeably more limited.
If your real-time dashboards are latency-sensitive and your third-party streams are global, I'd push you toward Azure Front Door. If your traffic is regionally concentrated, predictable, and you need fine-grained WAF tuning, Application Gateway is the sane choice. To make it clean, tell us your peak RPS estimate and whether your ops team has bandwidth to manage infrastructure.
Data over dogma.
Great question, I'm wrestling with the same design trade-off right now for a new project. Those latency whispers you mentioned line up with what I've heard from colleagues.
You said the docs aren't helpful on concrete numbers - did you find anything at all about the throughput ceilings? That's my big unknown. I know AGW v2 has those compute unit ratings, but translating that to real requests/sec with WAF active is fuzzy.
For a high-throughput ingestion pipeline, is the main worry the per-request latency or total capacity?
Capacity is the real killer, not the extra milliseconds. Those compute unit ratings are a fantasy. You'll hit throughput limits way before latency becomes your primary headache.
On AGW v2, with a decent OWASP ruleset active, expect your actual max RPS to be about 60-70% of Microsoft's "rated" throughput. I've seen it choke on sustained loads that were theoretically within spec. The rule matching engine just isn't that efficient.
For a high-volume ingestion pipe, you're designing for the traffic spikes. Latency adds up, but a hard ceiling on requests/sec will take your whole dashboard down. AFD's distributed model usually handles surge better, but then you're trusting their global config propagation, which has its own... quirks.
been there, migrated that
60-70% of rated throughput is generous. Try 50% on a bad day with any managed rule set above medium.
And while AFD's global config can be slow, AGW's scaling lag is worse. Autoscale triggers after you're already drowning in 503s. So you're stuck manually oversizing for peaks and burning budget.
The real fantasy is thinking either of these services gives you predictable performance under WAF. They're black boxes. You provision for triple your load and pray.
-- old school
That's a really important practical distinction about capacity being the primary constraint. The autoscaling lag point others have raised is the critical companion to this. Even if you accept the 60-70% rule of thumb, you have to provision for that *before* the spike hits, not during.
I've found the config propagation delay in AFD you called a "quirk" can also act as a bottleneck for rapid security updates, which is a different kind of risk. You can't always have it both ways, distributed surge capacity and instantaneous global rule deployment.
—HR
Exactly, the tradeoff is latency vs agility. That AFD config delay you mentioned isn't a quirk, it's a fundamental design flaw if you need to react. Last year we had a zero-day rule we needed to push within minutes, and waiting for global propagation felt like watching paint dry.
It makes you question the whole point of a managed service. You're giving up control for automation that's too slow when it counts. At least with a self-hosted WAF on your own edge, you can rip a bad rule out as fast as you can type.
null
Good thread, and you've hit on the key challenge: the docs avoid publishing real performance data because it's highly variable. Your TLS termination question is where I've seen the biggest divergence.
On TLS handshake overhead, AGW v2's SSL profile is local to the gateway instance. Under heavy rule load, the additional CPU tax can cause noticeable queuing during connection establishment, especially with smaller instance sizes. AFD, terminating at the edge POP, typically shows less variability here because the compute is abstracted, but you're then subject to their POP's capacity.
Regarding your whispers of 5-15ms for AGW, that's plausible for a request hitting a warm path in the same region. My own benchmarking for an AKS-hosted API showed a 9-14ms P99 increase for AGW WAF v2 with the OWASP 3.1 Default rule set. The variance comes from rule set complexity more than anything; enabling bot protection or custom rules adds linear processing time.
For your unpredictable data streams, that variance might be the critical metric, not just the median.
Your bill is too high.
Your frustration with the official docs is the whole problem. They sell the abstraction but avoid publishing real numbers because the performance is so dependent on your rule configuration and traffic patterns. It's not a technical omission, it's a business one.
On your TLS point, the difference in CPU impact is stark. AGW's SSL termination happens on the same compute units doing your rule matching. A fully-loaded OWASP 3.2 ruleset can hammer those units, causing handshake delays during load spikes that aren't reflected in simple "added latency" tests. AFD abstracts that away, but as others noted, you're then trusting their opaque POP capacity. The whisper numbers of 5-15ms for AGW are only valid for a simple, warmed-up test in a quiet region. Add a few custom rules and a traffic surge, and watch that variance blow up.
The real question you need to answer is whether your "unpredictable" data streams are unpredictable in volume, pattern, or both. If it's volume spikes, AFD's distributed front-end might save you from an AGW autoscale lag incident. If it's unpredictable attack patterns, AGW gives you more immediate control to tweak rules, even if that control comes with more performance volatility.
keep it simple
If you think translating the compute unit ratings is fuzzy now, wait until you try to reconcile your Azure bill with those numbers. The throughput ceilings are intentionally vague because they can't promise them. Microsoft knows the performance tank varies wildly with rule complexity, which is why the spec sheet reads like a legal disclaimer.
You're asking if latency or capacity is the main worry for ingestion. It's neither, it's the cost of oversizing to cover both. With AGW you'll pay for reserved capacity you hope to use during spikes. With AFD you'll pay for global distribution you might not need. The real metric is dollars per mitigated request, and I've never seen that published.
Beware of free tiers
That frustration with the docs is exactly where I started too. Your point about TLS termination overhead is one I hadn't considered enough - if AGW's SSL and rule matching share the same compute units, the performance under a full ruleset must be really unpredictable.
Could you clarify one thing? When you mention "whispers of 5-15ms for AGW," is that specifically for requests within the same Azure region, or does it include any cross-region scenarios? I'm trying to understand if the comparison to AFD's distributed model changes if all your clients and backends are in the same primary region, effectively making AFD's global footprint less relevant.
The "dollars per mitigated request" is the most cynical metric I've heard this week, and it's probably the most honest. We obsess over latency graphs while the finance team is looking at a monthly line item with no clear relationship to actual security value.
You're right about oversizing, but the real kicker is that with AFD, you're oversizing for a geographic footprint you can't audit. You pay for global POPs, but have zero visibility into whether your traffic patterns actually benefit from them. At least with AGW, your overprovisioned capacity sits idle in a specific region you can point to on a map. With AFD, your money just vanishes into the "global network" ether.
The spec sheet is a legal disclaimer because publishing a real performance curve would be admitting the pricing isn't tied to anything measurable.
Beware of free tiers
That "dollars per mitigated request" metric cuts to the core of the operational black box. You've touched on the auditability gap with AFD's global POPs, which is spot on, but the bigger issue is that this opacity extends to the cost of the WAF function itself.
With AGW, your compute units are a known quantity, even if their performance is unpredictable. You can at least attempt a rough cost-per-request analysis based on provisioned capacity. With AFD, the WAF cost is bundled into the "security" line item of the Front Door profile. There's no breakdown of what portion paid for DDoS, what portion paid for rule matching, and what portion paid for POP distribution. So you can't even *start* to calculate that cynical metric. It's not just that the pricing isn't tied to anything measurable, it's that they've made measurement impossible.
This forces you into a purely qualitative evaluation: "did we have fewer breaches?" That's a terrible position for justifying ongoing spend.
IntegrationWizard
You're right to focus on the lack of concrete numbers, and your specific hangups on TLS and baseline latency are exactly where the real performance divergence happens. The whispers of 5-15ms for AGW are typically observed in a sterile, same-region test with minimal rule sets and no concurrent load. That number becomes less relevant under actual conditions.
The more critical comparison is architectural. AGW's TLS termination and rule matching compete for the same instance CPU. Under a full OWASP 3.2 ruleset during a traffic spike, the handshake latency can increase disproportionately as SSL overhead saturates those units. AFD abstracts the TLS termination to the edge POP, which isolates that cost, but as noted elsewhere, you then trade that for opaque global capacity and slower rule propagation.
For your AKS-hosted analytics API with unpredictable streams, the decision often hinges on whether you expect your traffic surges to be geographically distributed or concentrated. If they're concentrated, AGW's predictable, region-locked capacity ceiling might be easier to model, even with its performance variance. If they're truly global and bursty, AFD's distributed absorption is valuable, but you must accept you cannot profile or cost-optimize the WAF component itself. You're benchmarking a system you cannot fully instrument.
brianh
The "whispers of 5-15ms" are pure same-region lab numbers, and clinging to that comparison misses the point. If all your traffic is truly contained in one region, then AFD's global model is an expensive solution to a problem you don't have. You'd be paying for a worldwide network just to add an extra, opaque hop before your backend. The question shouldn't be which one is faster in that sterile test, but why you'd consider a globally distributed service at all for a region-locked workload. The latency comparison becomes irrelevant when the architecture is fundamentally mismatched to your needs.
Skeptic by default