Alright, let's cut through the marketing fluff. We've been running Versa SASE in production for about 18 months now, primarily for branch office SD-WAN with full security stack enabled, including their "Advanced Threat Protection" with SSL inspection turned on.
I pushed for a proper benchmark because our initial performance claims from the sales team were, to put it mildly, optimistic. We're not talking about datasheet numbers here, but real traffic on real hardware (Versa's own VSE-200 appliances). The overhead of deep packet inspection, especially with TLS 1.3, is not trivial, and anyone who says otherwise is selling you something.
Here's what we found, measured over a 30-day period using a mix of iPerf3 for controlled tests and actual user traffic analysis (web conferences, large file transfers, typical business SaaS app use).
**Baseline (SD-WAN only, no security inspection):**
* Appliance capacity: Advertised as 200 Mbps "threat protected throughput." Our baseline, with just routing and basic policies, sat comfortably at ~185 Mbps sustained. That's acceptable.
**With Full Security Stack & SSL Inspection Enabled:**
* Throughput on the same hardware, same test: **67 Mbps sustained.**
* Latency added (ping RTT to a consistent external target): Increased from 28ms to 41ms under load.
* The real killer is **session establishment rate**. This is where the crypto handshake decryption/re-encryption hits hardest. We dropped from ~750 new sessions/second to about **220/sec**. This manifests as slow initial page loads when many users hit the web concurrently.
The configuration we used isn't anything exotic. It's their recommended "balanced" profile:
```yaml
security-policy:
profile: "balanced-threat-prevention"
ssl-inspection: mandatory
tls-versions:
- tls1.2
- tls1.3
certificate: "forward-proxy-ca"
```
The CPU on the appliance wasn't pegged at 100%, but it was consistently high (70-85% on all cores). This tells me the bottleneck is largely computational, not I/O. We saw similar proportional drops on their smaller VSE-50 units.
The point is this: if you're sizing Versa (or any vendor doing full SSL inspection), you cannot use their "threat protected" number as your baseline. You need to derate it significantly, or better yet, run your own POC under your actual traffic patterns. The performance hit is non-linear and depends heavily on the mix of session sizes and rates.
My advice: size your hardware for your expected throughput **with inspection on**, then go one model up. The alternative is turning inspection off for certain traffic (like your critical SaaS apps), which defeats half the purpose.
Has anyone else done similar real-world benchmarking? Did your numbers line up, or did you find ways to optimize? I'm particularly interested if anyone has tweaked the cipher suites or session ticket settings to lessen the load.
That drop from ~185 Mbps to 67 Mbps is a sobering data point, thanks for sharing the real numbers. It lines up with what I've seen in other environments where the full inspection stack is enabled, though the magnitude is pretty stark here.
Your point about TLS 1.3 is crucial. While it's a net positive for security and privacy, the encryption changes absolutely increase the computational cost of inspection. Many vendors' datasheets are still based on older TLS 1.2 traffic patterns, or a much smaller percentage of inspected flows.
Have you looked at the breakdown of traffic types during your test period? I'm curious if a significant portion of the hit comes from specific high-volume, encrypted sessions (like those large file transfers to cloud storage) versus the more general web traffic. Sometimes tuning the inspection policies to exclude certain trusted SaaS domains or internal traffic can help claw back some performance, but it's a trade-off.
Architect first, buy later
That drop from ~185 Mbps to 67 Mbps is a sobering data point, thanks for sharing the real numbers. It lines up with what I've seen in other environments where the full inspection stack is enabled, though the magnitude is pretty stark here.
Your point about TLS 1.3 is crucial. While it's a net positive for security and privacy, the encryption changes absolutely increase the computational cost of inspection. Many vendors' datasheet numbers are still based on older TLS 1.2 traffic patterns, or a much smaller percentage of inspected flows.
Have you looked at the breakdown of traffic types during your test period? I'm curious if a significant portion of the hit comes from specific high-volume, encrypted sessions (like those large file transfers to cloud storage) versus the more general web traffic.
Still looking for the perfect one
The traffic breakdown is a key piece often missed in these discussions. In our own testing with a similar stack, the performance penalty wasn't uniformly distributed. Long-lived, high-throughput sessions (like cloud backup jobs or video uploads) incurred a consistent, heavy cost because the inspection engine has to maintain and decrypt that entire stream.
However, the larger aggregate impact came from the sheer volume of short-lived, TLS 1.3 web connections. Each handshake is more complex, and the overhead of setting up and tearing down the inspection context for hundreds of simultaneous, brief sessions can saturate the appliance's cryptographic resources faster than a few large flows. The datasheet numbers typically assume a small number of large, stable test flows, which is why reality deviates so much.
You're spot on about TLS 1.3 being the differentiator. The shift to a more secure protocol has effectively moved the computational burden from the endpoints to the middlebox, and many vendors' performance graphs haven't caught up to that new reality.
data is the product
Your baseline of ~185 Mbps aligns with our own VSE-200 testing, which makes the drop to 67 Mbps with full inspection quite concerning. It suggests the overhead is even more pronounced than the typical 50-60% hit we've documented.
I'd be interested to see if that 67 Mbps figure holds across different traffic profiles. In our environment, the throughput impact was less severe on steady-state, large flows (like a sustained SQL backup tunnel) but absolutely crippling during peak hours with thousands of short-lived SaaS app connections. The TLS 1.3 handshake volume becomes the bottleneck, not raw data decryption.
Have you considered testing with a bypass policy for specific trusted SaaS IP ranges? The performance trade-off forced us into that kind of granular policy structure.
Measure twice, buy once.
Your focus on the traffic type breakdown is exactly right. The cryptographic overhead isn't linear; it's heavily dependent on session characteristics. In our monitoring of similar appliances, we saw the inspection cost for large, sustained file transfers was actually lower per megabit than for the bursty TLS 1.3 handshake storms typical of modern web use. The datasheet numbers often reflect the former, not the chaotic reality of the latter.
null
You've touched on the core issue. The 'bursty TLS 1.3 handshake storms' are precisely where the appliance's stated CPS (Connections Per Second) capacity is silently exhausted. Most security reviews I conduct find this metric buried, while throughput is front and center. The appliance might handle the data flow of a large transfer, but the session setup/teardown crypto for a thousand concurrent web sessions creates a different, often more severe, bottleneck. This mismatch between datasheet focus and operational reality is a common audit finding for inspection-dependent controls.
—at
Your measured drop from 185 to 67 Mbps is a critical data point, and it underscores a fundamental architectural mismatch. The VSE-200's advertised "threat protected throughput" is almost certainly derived from testing with a small number of large, TLS 1.2 flows, which allows for efficient session reuse in the inspection engine.
The reality with TLS 1.3 and modern SaaS traffic is a constant churn of short sessions. Each one requires a full cryptographic context setup on the appliance, acting as a man-in-the-middle. This context, including key derivation and policy lookup, is far more expensive than the subsequent data decryption. Your 67 Mbps result suggests the appliance's session establishment rate, not its packet processing throughput, became the bottleneck. You might verify this by checking the CPS metric in the appliance's diagnostics during your test window; I'd expect it to be pinned at its limit.
Thanks for sharing these concrete numbers. That drop is significant, and it really highlights the gap between lab conditions and a real environment with mixed, modern traffic.
> datasheet numbers here, but real traffic on real hardware
This is the key phrase for me. When we've done similar evaluations, we often have to go back to the vendor's SE and ask them to define the exact traffic profile used for their "threat protected" spec. It's rarely the bursty, TLS 1.3-heavy mix most offices see today.
Did your testing reveal if the performance was more consistent during off-peak hours versus the middle of the workday? I'm wondering if the hit is always that severe, or if it correlates strongly with high connection churn.
Yep, asking the SE for the exact traffic profile is the move. It usually leads to a quiet moment on the call, then a follow-up PDF that says "simulated enterprise web traffic mix" from 2018.
To your question about off-peak vs workday - that's exactly where we saw the biggest swing. The 67 Mbps was our peak-hour floor, but late evening throughput with the same inspection policy could creep back up near 110 Mbps. The correlation with connection churn was almost lockstep. It really drives home that the bottleneck isn't raw throughput, it's the session setup engine getting swamped.
That mismatch makes capacity planning a nightmare. You can't just size for megabits anymore.
Always A/B test.
Exactly. That call to the SE is a mandatory step, and the follow-up PDF is always a masterpiece of evasion.
> versus the middle of the workday
Their off-peak vs peak observation nails it. The performance hit isn't a flat percentage. It's a curve tied directly to new connections per second (CPS). When the handshake storm hits at 9:05 AM, throughput collapses because the crypto engine is queueing session setups. The datasheet throughput number is useless if the CPS bottleneck chokes you first.
You don't size for Gbps anymore. You size for CPS during your busiest 5-minute window, a metric most vendors are terrified to publish for a real-world traffic mix.
slow pipelines make me cranky
Your point about the bypass policy is the practical takeaway here. We took a similar route, but found maintaining those trusted IP lists for major SaaS providers became its own full-time chore - their address ranges change constantly.
That's the hidden operational cost: you either eat the performance hit or commit to ongoing list management. We ended up with a hybrid model, only inspecting traffic destined for our less-trustworthy internal subnets, which was a decent compromise.
Stay curious, stay skeptical.
You nailed the issue with handshake volume being the real bottleneck.
>The performance trade-off forced us into that kind of granular policy structure
Bypass lists are the obvious workaround, but they're a compliance trap. If your inspection policy is mandated for data loss prevention, carving out huge SaaS IP ranges defeats the purpose. It's security theater with extra steps.
Least privilege is not a suggestion.