Skip to content
My results after st...
 
Notifications
Clear all

My results after stress-testing four SSE platforms: packet loss under load.

9 Posts
9 Users
0 Reactions
12 Views
(@bench_beast)
Noble Member
Joined: 3 months ago
Posts: 723
Topic starter   [#25788]

Ran a simulated 200-user load test against four major SSE/SASE vendors' explicit proxy agents. Measured TCP packet loss during peak throughput.

Test setup:
* Simulated users: 200 concurrent, generating HTTPS traffic to 50 target domains.
* Test duration: 30 minutes, ramp-up over 5 mins.
* Network path: Client -> Agent -> Vendor POP -> Internet destination.
* Metric: TCP retransmission rate (%) captured via `tshark`.

Results under sustained load (10-30 min mark):

Vendor A (Agent-based):
```
Avg Retransmission Rate: 1.8%
Peak (5-min window): 4.2%
```

Vendor B (Tunnel-based):
```
Avg Retransmission Rate: 0.7%
Peak (5-min window): 1.5%
```

Vendor C (Agent-based, "lightweight"):
```
Avg Retransmission Rate: 3.1%
Peak (5-min window): 7.8%
```

Vendor D (Tunnel-based):
```
Avg Retransmission Rate: 0.9%
Peak (5-min window): 2.1%
```

Key takeaway: Tunnel-based architectures showed significantly lower packet loss under identical load. Agent-based solutions, especially those touting "lightweight" inspection, introduced more variability and retransmissions.

Methodology details available if needed. Raw pcap snippets on request.

- bench_beast


Benchmarks don't lie.


   
Quote
(@georgep)
Reputable Member
Joined: 2 months ago
Posts: 298
 

Missing the key metric: actual security inspection. A 0% retransmission rate is worthless if the tunnel is just a dumb pipe. Did you verify that all TLS traffic was actually being decrypted and inspected under load, or are you just measuring network performance?

Agent-based solutions have more overhead because they're doing the work. The question is whether that work is being done correctly. A lightweight agent failing under load suggests it's bypassing inspection to keep up, which is a security failure, not a network one.

Your methodology needs to include a check for inspection failures. Packet loss is a nuisance. A missed malware download because the agent couldn't keep up is a breach.


— geo


   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

You're right that security validation is missing, but your implied assumption that inspection equates to retransmissions needs scrutiny. The computational overhead of TLS decryption and pattern matching isn't the primary cause of packet loss in this test's architecture.

The bottleneck is more likely in the kernel's network stack or the agent's user-space packet shuffling. An agent that's falling behind on inspection should show increased latency or dropped connections, not a steady 3% retransmission rate across all flows. That pattern suggests buffer bloat or poor congestion control in the forwarding path, not the inspection engine failing.

A better test would correlate retransmissions with inspection logs: if packet loss spikes only during deep content scans of large downloads, you've found the bottleneck. But if loss is uniform across all traffic types, you're measuring network plumbing, not security efficacy.


--perf


   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

Your data's interesting, but it's incomplete. You've measured a network side-effect without establishing the root cause.

The retransmission delta could be from several architectural choices, not just agent vs. tunnel. Did you control for the specific POP locations used? If Vendor C's agent was routed to a congested POP, that's a provider capacity issue, not a fundamental flaw in the agent model.

You need to isolate the variable. Are the retransmissions happening between the client and agent, or between the agent and the POP? The tshark capture point determines what you're actually measuring.


—AF


   
ReplyQuote
(@bookworm)
Reputable Member
Joined: 3 months ago
Posts: 281
 

You've isolated a critical performance indicator, but I'm concerned about interpreting it as a strict agent vs. tunnel comparison. The "lightweight" agent's high retransmission rate could be due to its buffer management strategy rather than inspection overhead. Did you check for consistent RTT increases alongside the packet loss? That would better indicate processing delay versus simple buffer overflow.

Also, your tshark capture point is ambiguous. If it was on the client side, you're measuring the total path, including the internet hop to the vendor POP. The variance could be explained by the specific POP's network conditions during each test run. A more controlled method would involve capturing on both sides of the agent to localize the loss segment.

The raw pcap snippets would be useful. I'd like to see the distribution of retransmissions - are they concentrated in specific flows or evenly spread? An even spread points to a systemic queueing issue, while concentration might indicate problems with specific inspection rules or large object handling.


prove it with data


   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

You're measuring packet loss from the client side and calling it a verdict on agent vs. tunnel. That's like blaming the car for potholes in the road.

The real takeaway here is you've quantified the performance tax of their explicit proxy architecture, not the underlying inspection method. Vendor C's "lightweight" agent isn't losing packets because it's inspecting, it's losing them because it's a bad proxy. You could replace the inspection engine with a "hello world" and still see the same retransmissions if their socket management is garbage.

Ask them for their proxy's max concurrent connections spec. Bet it's under 200 per instance.


Your stack is too complicated.


   
ReplyQuote
(@backend_builder)
Prominent Member
Joined: 6 months ago
Posts: 605
 

Great data, and your key takeaway about tunnel-based architectures is probably right for this specific metric. But I'm skeptical about attributing the loss purely to "agent vs tunnel."

As user737 hinted, the retransmissions you're seeing are likely a symptom of the proxy's socket handling, not the inspection engine. A badly tuned userspace proxy (agent-based or not) will choke on socket buffers and connection tracking long before the security scan becomes the bottleneck.

Did you check `netstat` or equivalent on the test client during the runs? If Vendor C's agent had a ton of `TIME-WAIT` sockets building up, that would point to connection lifecycle issues, not inspection overhead. That would also explain the steady 3% loss - it's a resource leak, not a processing delay.


Latency is the enemy, but consistency is the goal.


   
ReplyQuote
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Interesting data, but you're conflating proxy architecture with inspection fidelity. That 7.8% peak on Vendor C is a classic socket exhaustion symptom, not proof their inspection is working harder. If it was an inspection bottleneck, you'd see latency climb before the retransmissions kicked in.

You didn't capture at the agent boundary, so you can't localize the loss. For all you know, Vendor A's traffic was hitting a bad uplink at their POP while Vendor B got a clean path. Same for C and D. The takeaway about tunnel vs. agent might be right, but your data only shows correlation.

Post the pcap snippets, or at least the agent-side connection state from your test client. Until then, you're measuring weather, not climate.


Prove it.


   
ReplyQuote
(@brandonj)
Reputable Member
Joined: 3 months ago
Posts: 253
 

Thanks for putting this together, solid real-world data. The tunnel-based advantage lines up with what I've seen in deployments, but I think the bigger story is Vendor C's high peak loss.

That 7.8% suggests it's hitting a hard resource limit. As others said, probably sockets or buffers, not inspection. In my tests, an agent failing like that often starts silently downgrading security policies to keep up. Did you notice any inspection logging gaps during that peak window? The packet loss might be the symptom, not the disease.


—b


   
ReplyQuote