Alright, let's talk about the numbers that actually matter. Everyone's sales deck has those glossy throughput charts, usually with a 1ms latency footnote in size-2 font, claiming 5 Gbps for their SASE/SSE data plane. Then you deploy it and your users start complaining about "the VPN" being slow again.
We just finished a six-month bake-off between two major vendors (names withheld, but you can probably guess). Their spec sheets claimed near-identical performance for a mid-range gateway instance: 2 Gbps with "advanced security" features on. Our synthetic testing in their lab? 1.8 Gbps. Not bad. Reality, once we deployed in our actual AWS us-east-1 region, with our actual traffic mix (a messy blend of tiny API calls, large file transfers, and video streams)? 650 Mbps sustained before latency went to hell. The CPUs on their cloud nodes were pinned at 90%+.
The culprit wasn't the raw packet forwarding. It was the "advanced security" stack—specifically, the TLS decryption/inspection and the IDS engine. The vendors assume a perfect distribution of large, sequential flows. We have thousands of concurrent, short-lived connections. Their data plane falls apart.
Here's a snippet from the load test we built to mimic our traffic pattern, which their own test suites didn't account for:
```python
# Simulates our app's chatter - not just big, clean flows
def generate_mixed_workload():
# 70% small request/response < 16KB
# 20% medium 1-10MB
# 10% large 50MB+
# All over TLS 1.3, different cipher suites
# With concurrent connections from same source IP
```
The moral of the story? You must test with *your* traffic profile, not theirs. The published throughput numbers are essentially a best-case scenario under lab conditions with large, single-threaded flows and often with inspection features turned *off*. The moment you enable the very things you're paying for—threat prevention, data loss, etc.—the performance collapses in a real-world scenario.
So, what did you all actually measure when you went live? And did anyone manage to get a vendor to commit to a throughput SLA that includes *your* specific security stack and traffic mix, not just their idealised "IMIX" packet blobs?
Senior platform engineer at a ~500 person fintech. We've run Zscaler and Palo Alto Prisma in production for 3 years each.
* Real-world throughput tax: The 80% drop OP saw is normal. The advertised "maximum inspection" throughput assumes 8KB+ objects and 10% SSL inspection. Our traffic is similar. Expect a 60-70% reduction once TLS 1.3 decryption and full IDS/IPS are on. A "2 Gbps" node will do 600-800 Mbps sustained. If your vendor disagrees, ask for the object size and inspection rate in their spec.
* Cost model landmine: Both major players quote $7-12/user/month. That's the access license. The actual bill is 2x once you add the data plane compute. For Zscaler, that's your ZIA Private Service Edge VMs in AWS/Azure. For Palo, it's their compute gateways. Unmonitored, these auto-scale and will explode your cloud bill during a DDoS or a data migration. Budget for the compute separately.
* The east-west blind spot: These are built for north-south. They break down inside the cloud. Any traffic between two private subnets that needs inspection either hairpins out to their nearest PoP (adding 80ms+) or you deploy their heavy VMs as NAT gateways per-VPC. Neither works. We gave up and use native AWS security groups and Network Firewall for internal traffic.
* Cold-start latency: For serverless workloads (Lambda, Fargate) that spin up thousands of new connections, the first packet pay-to-play is real. Adding the SSE provider adds 150-250ms of handshake and proxy negotiation before your app even talks. It murders auto-scaling response times. You have to maintain warm pools just for the proxy, which defeats the purpose.
I'd pick Zscaler if your workforce is mostly remote/branch and you need the fastest route to the nearest POP. I'd pick Palo if you're already a Palo shop and need consistent policy across firewall and SSE. For mostly cloud-to-cloud or server-heavy traffic, tell us your cloud provider and if you own the IP space.
You're hitting the exact failure mode I've documented in postmortems. The spec sheet numbers assume a steady state of optimal packet sizes.
Your bottleneck is likely connection setup rate and context switching, not just inspection overhead. When you have thousands of short-lived connections, the per-connection TLS handshake and session teardown dominate CPU. The vendors' "maximum throughput" tests use persistent connections.
Demand they provide their CPS (connections per second) rating for the same instance size. That's the real constraint for API-heavy traffic, and they often obscure it. If their CPS is low, the node falls over under churn regardless of the gigabit rating.
Five nines? Prove it.