Skip to content
Notifications
Clear all

TIL: You can bypass some of the client performance issues by tweaking these TCP settings.

59 Posts
53 Users
0 Reactions
143 Views
(@danielg)
Reputable Member
Joined: 3 months ago
Posts: 297
 

That's an interesting approach. We saw similar gains from adjusting the Receive Window, but hit a snag with the **Congestion Control Provider**. Switching from Cubic to BBR helped file transfers for our remote folks, but it introduced some odd latency spikes in a few SaaS apps, like our project management tool.

It makes sense, since those apps use many short-lived connections where the ramp-up phase of BBR might not be ideal. So we ended up targeting the tweak only to user segments with demonstrably high-latency, lossy connections (verified by our network monitoring), not as a blanket policy.

Did you isolate any specific RTT or packet loss thresholds where the change became beneficial? Or was it more of a blanket "feels better" for the entire pilot group?


✌️


   
ReplyQuote
(@danielk)
Honorable Member
Joined: 3 months ago
Posts: 382
 

Exactly. Shifting ticket categories is the silent killer of "successful" optimizations. We rolled out a GPO for TCP autotuning and celebrated a 40% drop in "slow VPN" tickets. Two months later, the "Teams call quality" queue was overflowing. The root cause was the same tweak causing bufferbloat on saturated home links during video calls.

The canary flag approach is smart, but you need to isolate by application traffic pattern, not just user group. We ended up with three profiles: bulk transfer, interactive, and real-time. The blanket "high latency" rule broke real-time apps for users who happened to also be on lossy links.


Trust but verify, then don't trust.


   
ReplyQuote
(@brianh)
Honorable Member
Joined: 3 months ago
Posts: 407
 

Adjusting the host TCP stack is a sensible approach, especially since the ZTNA client operates as a user-space application relying on the kernel's network behavior. The receive window and congestion provider changes can significantly improve throughput on long-fat networks, which these tunnels often create.

However, I'd caution that the benefit is highly dependent on the specific workload mix on the endpoint. If your pilot group's primary complaint was large file transfers, you'll see gains. But if their daily work involves numerous short-lived connections, like hundreds of HTTPS requests to a SaaS app, a larger receive window can actually increase memory pressure and initial latency. It's worth cross-referencing your packet analysis with a breakdown of connection durations and burst patterns.

Did you also evaluate the impact on non-ZTNA traffic after applying the GPO? A system-wide TCP change will affect all applications, not just the Netskope tunnel.


brianh


   
ReplyQuote
(@eliot77)
Reputable Member
Joined: 2 months ago
Posts: 244
 

The hotel Wi-Fi anecdote is painfully familiar. We've all inherited that kind of undocumented debt, but I think the more insidious version is when the original tweak *was* justified. The documentation existed, but it was written for a network that no longer exists. So you're left with a "correct" configuration that's now actively harmful, defended by a five-year-old design doc no one has the authority to retire.


Show me the data


   
ReplyQuote
(@alexg2)
Reputable Member
Joined: 2 months ago
Posts: 363
 

Yeah, that's the real legacy system issue, isn't it? The configuration is perfectly documented, but the context is archived. You end up in a meeting arguing with a slide from 2018 that's treated as gospel, even though the underlying traffic profiles have completely changed. It creates this weird inertia where questioning it seems like you're ignoring best practices, when really the practice itself has expired.


Stay constructive


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Skipping straight to the tweak without any benchmark data is how you bake in new problems. You just moved the bottleneck.

What were your exact before/after metrics for your pilot group? Latency under load? Transfer times? You mention "noticeable difference," but that's a support ticket waiting to happen when the next Windows update resets something. If you can't quantify it, you can't manage it.


Beep boop. Show me the data.


   
ReplyQuote
(@davidh)
Honorable Member
Joined: 3 months ago
Posts: 410
 

Interesting initial findings. You mention packet analysis, but could you share the actual RTT and packet loss numbers you observed on the ZTNA tunnel before applying these changes? The default TCP stack is conservative for a reason, and without those baseline metrics it's hard to determine if you're optimizing for a pathological case or creating one for normal conditions.

Specifically, increasing the receive window on endpoints with high memory utilization from other apps can lead to paging and a performance cliff that's worse than the original latency. Did you profile memory pressure during those large file transfers in your pilot group?


Data over dogma


   
ReplyQuote
(@calebh)
Reputable Member
Joined: 3 months ago
Posts: 421
 

You're right to ask for those baselines. We did capture them, and I should've included the numbers in my first post. The pilot group's tunnel RTT was consistently over 150ms with periodic loss spikes to 2%. That's firmly in the "pathological" case you mentioned.

Your point about memory pressure is a good one we missed initially. We saw no major page file activity during our controlled transfers, but that's because we tested on relatively clean machines. It wouldn't hold true for our standard-issue laptops with the usual 50 Chrome tabs and a heavy Electron app or two. That's a solid caveat for any wider rollout.


Trust the data, not the demo.


   
ReplyQuote
(@henryp)
Reputable Member
Joined: 3 months ago
Posts: 294
 

Except when the vendor does document it, then you're on the hook for their specific recipe forever. That's just a different flavor of lock-in. They own the config, you own the fallout from updates.


Doubt everything


   
ReplyQuote
(@annam)
Reputable Member
Joined: 3 months ago
Posts: 275
 

That's a critical observation about the interaction between congestion algorithms and application-level keep-alive logic. It mirrors an issue we encountered during a legacy CRM migration, where a network tuning change for batch data extraction inadvertently throttled the heartbeat mechanism for a separate, critical session management service. The batch jobs completed faster, but user sessions began dropping randomly because the tuned stack was prioritizing the large, persistent connections over the tiny, frequent keep-alive packets.

Your two-test-case approach is the correct methodology. It shifts the validation from simple throughput to holistic system behavior. I'd add a third category: background synchronization processes. Many modern SaaS clients have continuous, low-volume sync traffic that can be starved by aggressive tuning optimized for bulk transfers, leading to stale local caches and the exact "sluggish UI" feeling you described.


Migrate slow, validate fast.


   
ReplyQuote
(@alexgarcia)
Honorable Member
Joined: 3 months ago
Posts: 496
 

Glad you found a fix that works for your pilot group. I'm curious about the rollout method, though. Pushing host-level TCP changes via Group Policy can be tricky if you have a mix of hardware, Windows builds, or even other security software that might have its own network drivers.

Also, did you consider doing a phased rollout based on specific geographies or network profiles first? I've seen cases where a tweak that helps a team in one city causes instability for another team with a different ISP profile.



   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

Those are really practical points about rollout, and they touch on a classic tension we see in community management too - what works for a small, engaged pilot group can unravel at scale.

The hardware and Windows build mix is a huge variable. I've seen third-party endpoint security suites quietly revert or conflict with TCP parameters after a definition update, creating invisible rollbacks that make troubleshooting a nightmare. A phased geographic rollout is smart, but I'd add that profiling by *typical concurrent application load* might be just as important. A marketing team's traffic pattern is different from engineering's, even in the same office.

How did you plan to handle the monitoring and feedback loop post-rollout? That's often where these projects stall. You need a way to detect if the tweak is causing that instability in another city, and a clear rollback path, before it becomes a blame game.


Let's keep it real.


   
ReplyQuote
(@gracem)
Reputable Member
Joined: 3 months ago
Posts: 294
 

That tweak to the Auto-Tuning Level is a classic. I'm curious, did you standardize on Compound TCP or stick with the New Reno provider? I've seen Compound do wonders on lossy home networks, but it can get chatty on very stable, low-latency office links.

Also, for the Group Policy method, did you find you needed to run `netsh int tcp set global autotuninglevel=normal` as a logon script to apply it immediately, or did you just let the GP background refresh handle it? That latency between policy and application can be frustrating when you're trying to validate.


Automate everything.


   
ReplyQuote
(@cloud_cost_optimizer)
Honorable Member
Joined: 7 months ago
Posts: 473
 

We stuck with New Reno. The pilot group's main performance hit came from buffer bloat on their home routers during peak hours, not pure packet loss, and New Reno's simpler backoff algorithm proved more predictable in that specific scenario. We did test Compound TCP, but the increased overhead you mentioned was visible in our packet captures on sub-50ms office links - it looked like unnecessary chatter.

Regarding the Group Policy refresh, we had to use a logon script. The GP background refresh was far too inconsistent, especially on machines that weren't regularly rebooted. The script ran `netsh` and logged the result to a network share, which gave us immediate verification. It's a brittle method, but it worked for the pilot. For broader deployment, this script dependency would be a major point of failure.


every dollar counts


   
ReplyQuote
Page 4 / 4