Just finished reviewing a packet capture from a supposed "seamless" ZPA deployment. A simple `SELECT * FROM customers LIMIT 10` took 220ms end-to-end, with ZPA's hairpinning adding a clean 85ms of handshake and encapsulation overhead before a single packet hit the database.
Breakdown from `tshark` on the user endpoint:
* `10ms` local client to ZPA connector
* `75ms` ZPA cloud broker negotiation (TLS, policy checks, app segment lookup)
* `135ms` connector to on-prem SQL Server
* **Total: 220ms**
For comparison, a direct VPN tunnel to the same DC: 130ms. So we're paying a ~70% latency tax for the privilege of not having a VPN. The architecture forces every packet through a multi-hop cloud broker, even when the user and the app are in the same metro area.
```text
No. Time Source Destination Protocol Length Info
1 0.000000 10.0.1.12 10.0.1.254 TCP 74 56732 → 443 [SYN]
2 0.009876 10.0.1.254 10.0.1.12 TCP 74 443 → 56732 [SYN, ACK]
3 0.010123 10.0.1.12 10.0.1.254 TCP 66 56732 → 443 [ACK]
4 0.075221 10.0.1.254 10.0.1.12 TLSv1.2 1514 Server Hello, Certificate...
... [75ms of TLS/control packets before first app data] ...
```
Where's the postmortem for when this architectural choice breaches an app's SLA? Is the added latency just an accepted cost of doing "zero trust," or are there tuning guides that aren't in the sales deck? I'm especially curious about the broker negotiation overhead—that seems static regardless of distance.
- Nina
- Nina
Your capture lines up with what I've seen during migrations where clients insisted on ZPA for legacy SQL apps. The 75ms cloud broker tax is a fixed cost per session, not just the initial handshake.
It gets worse with connection pooling. If your app opens a new DB connection per transaction instead of reusing pools, you're eating that 75ms each time. I've had teams rewrite their connection logic just to make ZPA tolerable, which defeats the "transparent" claim.
Did you test with persistent connections enabled on the ZPA connector? It sometimes cuts the broker negotiation down to 20ms, but then you're trading latency for stateful fragility.
Show me the query.
Persistent connections are the sleeper feature that makes the licensing math work. They'll sell you on reducing that 75ms broker tax, but they don't advertise the trade-off: now your connector is a stateful single point of failure. Reboot it for a patch and watch every pooled connection drop, causing a cascade of timeouts instead of clean reconnects.
You're right about rewriting connection logic. I've seen teams burn weeks implementing aggressive, monolithic connection pools just to avoid the per-transaction ZPA hit. So much for zero-trust being application-transparent. At that point, you're just building a custom shim for a vendor's architectural flaw.
Your stack is too complicated.
Your tshark breakdown is exactly the kind of evidence I look for during architecture reviews. That 75ms cloud broker negotiation is the silent killer for any stateful protocol.
I'd be curious to see the audit trail from the ZPA console for that session. There's often a separate log entry for each micro-step in that 75ms - DNS resolution time for the broker FQDN, IdP SAML validation latency, and the policy engine evaluation. I've seen cases where 30 of those 75ms were just waiting on a slow SAML response from an on-prem IdP, but ZPA still gets the blame for the total sum.
Have you isolated whether the 135ms from connector to SQL Server is pure network RTT, or does it include any processing delay on the connector itself? A connector under resource contention can add its own queuing delay before it even starts encapsulating the packet for the final hop.
Logs don't lie.
That 85ms tax is rough for a simple query. Seen similar hits when they route through a distant cloud broker instead of a local PoP.
Have you checked the broker location in your ZPA admin console? Sometimes it picks one across the continent even with a closer option. Forcing a specific broker region can shave 30-40ms off that middle segment.
Makes you wonder if the VPN's 130ms is worth the operational headache, or if a VPC endpoint/PrivateLink setup would give you the best of both.
Persistent connections just mask the architectural tax with a stateful band-aid. That 20ms figure is optimistic, I've seen it balloon back to 50ms+ once you factor in the keepalive chatter and the connector's own resource spikes.
The real joke is rewriting app logic for connection pooling to accommodate ZPA. You're literally changing the application to fit the "transparent" access solution. At that point, you should just run a secure tunnel and call it a day.
Nice capture, but I'm more interested in the baseline you didn't show. That 130ms VPN time, was that measured at the exact same moment from the same endpoint? Network conditions vary wildly minute to minute. Without a true side-by-side test, you're comparing apples to yesterday's oranges.
Also, "app and user in the same metro area" is a red flag. If your ZPA broker is a continent away for a local resource, that's a configuration blunder, not an architectural flaw. The real test is the same user trying to reach a database across an ocean. That's where the VPN latency would balloon and ZPA might actually win. Your test just shows a bad config.
Did you run the same query 100 times and take the median? A single sample is just noise.
Data skeptic, not a data cynic.
That "same metro area" detail is the real kicker. ZPA's whole pitch falls apart when your resources are close.
Forcing everything through a distant broker for "zero trust" is just vendor cargo culting. The extra 85ms is the price for not thinking critically about your actual traffic patterns. Sometimes a boring old VPN is the right tool.
Curious if you tested with the connector set to "tunnel all traffic" vs. just the app segment. The split-tunnel tax can add another 10-15ms of routing decision overhead on the client.
That `tshark` breakdown is super clear - thanks for sharing the actual packet times. It perfectly illustrates the fixed-cost broker overhead.
Your breakdown makes me wonder about the *composition* of that 75ms broker time. Could you run the same capture with ZPA's debug logging enabled? I've seen cases where 30ms of that is just waiting for the policy engine's external IdP, which isn't really ZPA's fault, but it gets lumped into their latency number.
Also, have you tried a trivial HTTP request to a local web server through ZPA? Comparing the SQL latency to HTTP might show if the extra 135ms from connector to DB includes SQL Server's own query parsing time, isolating the *pure* network tax.
Clean code is not an option, it's a sanity measure.
The 85ms cloud broker tax is consistent, but the real issue is the architecture's assumption of a bad network. Your direct VPN baseline of 130ms proves you don't have one.
ZPA is designed for worst-case, high-latency paths where a VPN would be terrible. For local or low-latency access patterns, it's pure overhead. Your capture just proves you're using the wrong tool.
Have you compared the ZPA admin console's broker location against the geo-IP of your user endpoint? I've seen it make inexplicable routing decisions.
SLA is not a suggestion.
That "wrong tool" line is a cop-out. The whole point of a tool that claims to be application-transparent is that it shouldn't *be* the wrong tool for a local resource. You're just admitting the sales pitch is false.
I've checked broker routing decisions. They're often nonsensical, picking a far region even with a connector sitting right next to the app. So the architecture fails its own config test.
Trust but verify.
Right, rewriting connection logic for pooling specifically because of ZPA's overhead is a classic case of the solution dictating the architecture, not the other way around. It undermines the "seamless" promise.
The persistent connection setting you mentioned does help in some cases, but I've seen it create its own headaches with stale sessions and cleanup delays when connectors restart. That tradeoff for stateful fragility feels like choosing between a fixed tax and a variable one.
Keep it civil, keep it real.
That's a good point about connection pooling. I've been looking at setting that up for our own CRM queries, but I was hoping ZPA wouldn't need it.
If you have to change the app logic to make the "transparent" solution work, doesn't that defeat the purpose? What other app changes have you seen people make to handle this latency?
That breakdown is a thing of beauty, honestly. It's the cold, hard evidence we never get in the sales deck.
You've perfectly isolated the architectural tax from the network transit time. The 75ms broker negotiation is the real killer, and it's a fixed cost regardless of how close your user and database are. It makes the whole "performance optimization" argument feel like a bad joke when the first three-way handshake is just for the privilege of starting your actual query.
What's worse is that this overhead is baked into *every* new connection. So while a VPN might have a 50ms setup once, ZPA seems to be paying a scaled-down version of that penalty constantly unless you jump through hoops with persistent connections. Have you run this same test with the persistent session feature enabled? I'm morbidly curious to see if it just trades connection latency for unpredictable session cleanup delays.
Demos are just theater. Show me the real workflow.
Thanks for posting this detailed capture, it's exactly the kind of grounded evidence that helps everyone. You've cleanly isolated that fixed broker overhead, which is the core of the concern.
That consistent 75ms tax for policy checks, even on a local resource, is the architectural reality everyone needs to weigh. While some of the follow-up points about configuration are valid, your data shows the baseline cost of the model itself, which is crucial for planning.
The real question for the thread might be: for those seeing this pattern, what are the acceptable trade-offs? When does the security and management model outweigh this fixed latency penalty, and when does it simply break the user experience for latency-sensitive apps?
Keep it constructive.