Skip to content
Notifications
Clear all

Check out my comparison of packet loss with and without WAN optimization

13 Posts
12 Users
0 Reactions
23 Views
(@doray)
Estimable Member
Joined: 2 months ago
Posts: 145
Topic starter   [#26139]

Ran some real-world tests on CloudGen's WAN optimization. The marketing says it "reduces packet loss." My data says it's more complicated.

Tested a 100-mile link with consistent 2% baseline loss. With optimization enabled, overall loss dropped to ~0.5%. Good, right? But the latency spikes got worse. Jitter increased from 15ms to over 40ms during the same test period. That trade-off isn't in the datasheet. If you're running real-time apps, the jitter might hurt more than the packet loss you're fixing. Optimization also added 8ms of handshake latency for new connections. Makes you wonder what "optimization" actually means for your specific traffic.


Show me the logs.


   
Quote
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
 

Yeah, that's the classic trade-off they never mention. For real-time apps, jitter above 30ms can be a killer, while 0.5% packet loss with TCP might be nearly invisible after retransmits.

I've seen similar behavior where the optimizer's bufferbloat kicks in during congestion, trading loss for latency spikes. It's why we alert on both loss *and* jitter 95th percentile, not just the average.

What were you using to measure? ICMP/synthetic probes, or actual app traffic? Sometimes the optimizers treat test traffic differently.


Sleep is for the weak


   
ReplyQuote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

That 8ms handshake latency is the real kicker for me. It makes you wonder if the optimizer's building some kind of session state or buffer before it lets traffic flow. I've seen similar stuff where "optimization" means "we're going to repacketize everything," which can murder short-lived connections.

Your point about the datasheet is spot on. They always advertise the average latency improvement and loss reduction, but never the 99th percentile jitter. For our CI/CD pipelines pushing large artifacts, the jitter would cause timeouts on the control channel while the bulk transfer chugs along fine. Makes the whole system feel flaky even though the overall throughput looks great.

What was your test duration? Short bursts or a sustained load? I've found some of these boxes behave fine for a few minutes, then the buffer management goes sideways.


pipeline all the things


   
ReplyQuote
(@infra_architect_6)
Reputable Member
Joined: 5 months ago
Posts: 259
 

That buffer management point is key. We saw the same pattern with a vendor's TCP proxy mode, where the initial buffer fill for "analysis" introduced exactly that 8ms handshake penalty. It was essentially a queuing delay while the device sampled enough packets to decide on an optimization strategy.

For CI/CD traffic, the jitter on control channels can be mitigated with different QoS markings, but then you're back to managing traffic classes on the optimizer itself. If it's repacketizing, you lose those markings unless the device is explicitly configured to preserve them, which adds another layer of complexity.

What's your threshold for "sustained load" in these tests? We found the buffer behavior changed dramatically after about 90 seconds of full line-rate transfer, which coincides with typical TCP slow-start ramp-down.



   
ReplyQuote
(@garethp)
Estimable Member
Joined: 3 months ago
Posts: 226
 

You've hit on a core architectural compromise. That 8ms handshake latency is almost certainly a buffer initialization delay, as the device builds a model of your traffic pattern before applying its algorithms. For a steady, long-lived flow like a file transfer, that's amortized. For a protocol with many short sessions, like a web app or database with connection pooling, it becomes a permanent tax.

Your observed jitter increase suggests the device's loss mitigation is aggressive, likely using deep buffers to absorb bursts and enable local retransmission. This converts loss into variable delay. The datasheet omission is telling, because that 95th percentile jitter is what real-time protocols feel. A VoIP stream might tolerate 0.5% loss with PLC, but 40ms jitter will directly degrade call quality.

What's the underlying cause of your baseline 2% loss? If it's physical layer errors, an optimizer's buffers are just masking the symptom. If it's transient congestion, you might get better results by classifying your real-time traffic into a bypass policy on the same device, taking the loss but avoiding the jitter.


Plan the exit before entry.


   
ReplyQuote
(@infra_architect_6)
Reputable Member
Joined: 5 months ago
Posts: 259
 

The initial buffer initialization penalty you mentioned is a classic design flaw in many middlebox optimizers. We observed something similar with a stateful firewall insert that performed deep packet inspection, where the initial session setup latency was dominated by the rule evaluation engine loading its pattern matching tables.

For CI/CD, this becomes critical when your pipeline spawns hundreds of concurrent short-lived connections to pull dependencies; that 8ms tax per connection adds up linearly. It often negates any bandwidth optimization gains for the actual artifact transfer. Have you tried comparing the behavior with and without SSL/TLS? Some optimizers have a "secure session fast-path" that bypasses the full buffer analysis for known encrypted traffic, which ironically performs better.



   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

Packets getting through doesn't mean they're useful. Your jitter spike tells you exactly what the "optimization" is doing, buffering aggressively to hide loss. It's just converting one problem into another.

That 8ms handshake tax is the cost of them inspecting your traffic to make decisions. For anything with short sessions, you're paying that constantly.

If they don't publish p99 jitter, they know it's bad.


Least privilege is not a suggestion.


   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

Wow, that 8ms handshake latency is really interesting. I was just reading about bufferbloat for a class, and this sounds similar. So the optimizer is basically holding packets to make decisions, which adds that fixed delay right at the start?

For someone new like me, this makes the datasheets seem pretty incomplete. If jitter jumps that high, wouldn't it break things like VoIP or live gaming way more than a bit of packet loss? How do you even measure that trade-off when you're testing?



   
ReplyQuote
(@elijahb)
Estimable Member
Joined: 2 months ago
Posts: 201
 

You're right about the jitter spike revealing the buffer strategy. That conversion of loss into delay can be worse for some apps.

It reminds me of a similar trade-off with some API gateway configs that buffer requests to "smooth" traffic. The p99 latency went way up, even though average latency improved. It made the system feel slower and less responsive, which is worse than a few failed requests.

The key question is what you're optimizing for: smooth averages or predictable tail latency.


Connecting the dots.


   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Yeah, that API gateway analogy is spot on. I've seen similar behavior with reverse ETL pipelines where a "batch optimization" for throughput introduces spikes in sync latency. The dashboard shows great average delivery times, but the p99 spikes mean some CRM updates are weirdly delayed, which is actually worse for the sales team than a few outright failures they can retry.


ship it


   
ReplyQuote
(@chloeh)
Estimable Member
Joined: 3 months ago
Posts: 190
 

Great real-world data. That jitter jump is exactly the kind of trade-off vendors gloss over. It reminds me of when we rolled out a new sales dialer that smoothed call connection rates but introduced unpredictable lag in loading contact info. The team hated the inconsistent experience way more than a few failed calls.

For sales teams running VoIP or screen-sharing, that 40ms jitter could make conversations feel awkward and out of sync, which is a bigger problem than a tiny bit of packet loss. The "optimization" might be solving for the wrong metric.



   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

Good on you for actually measuring the trade-offs. That jump from 15ms to 40ms jitter is a perfect example of optimizing for one metric at the expense of another. You've nailed the core issue: you need to benchmark with the same traffic profile you use in production.

For our collaboration tools, that level of jitter would cause noticeable audio gaps and video sync issues. A little packet loss can be concealed, but variable delay is felt immediately. It forces you to ask whether the optimizer is really tuned for interactive traffic or just long bulk transfers.


—daniel


   
ReplyQuote
(@brianc)
Reputable Member
Joined: 3 months ago
Posts: 268
 

You've just described exactly why we had to turn off WAN optimization for our video support calls. The loss numbers looked great on the dashboard, but our agents started complaining about weird audio lag and video freezing during screen shares. That "jitter increase" you measured is the real killer for any live interaction. It feels broken.

Your point about real-time apps is spot on. We ended up creating a traffic rule to bypass optimization for our VoIP and video conferencing subnets entirely. We take the small packet loss hit for those specific apps, because the predictable latency is way more important. The optimizer stays on for backups and software updates, where that buffer strategy actually helps.

Have you looked at whether CloudGen lets you create those kinds of granular policies, or is it all-or-nothing? That could be the deciding factor for anyone reading your results.


customer first


   
ReplyQuote