Skip to content
Notifications
Clear all

TIL: You can bypass some of the client performance issues by tweaking these TCP settings.

59 Posts
53 Users
0 Reactions
140 Views
(@hannahr2)
Reputable Member
Joined: 2 months ago
Posts: 233
Topic starter   [#24832]

Hey everyone, I've been knee-deep in our Netskope ZTNA deployment for the last quarter, and like many of you, I've wrestled with occasional client-side latency and that "sluggish" feel, especially on Windows endpoints during high-throughput tasks like large file transfers or real-time document collaboration.

After a lot of packet analysis and working with our network team, we discovered that tuning some underlying TCP parameters on our endpoints made a **noticeable difference** in perceived performance. It seems the default Windows TCP stack can sometimes be a bit conservative for the type of encrypted, sustained connections that ZTNA tunnels create. Netskope's client is fantastic, but it rides on these system fundamentals.

Here’s the specific tweak that worked for us. We adjusted the **TCP Receive Window Auto-Tuning Level** and the **Congestion Control Provider**. This isn't Netskope-specific config, but a host-level optimization that benefits any high-latency or lossy network path (which, let's be honest, any remote workforce is dealing with).

We deployed this via a simple Group Policy Preference (you could also use a script) to a pilot group of users who reported performance issues. The key changes were:

* **Set `TCP Window Auto-Tuning Level` to `normal`** (if it was set to `disabled`, which is common in some enterprise images). This allows the window to scale more dynamically.
* **Switched the `Congestion Control Provider` to `CTCP`** (Compound TCP). This algorithm is generally more aggressive and efficient for modern networks with decent bandwidth but variable latency.

The implementation is straightforward. We applied these via the command line (elevated prompt) on our test machines:
```
netsh int tcp set global autotuninglevel=normal
netsh int tcp set global congestionprovider=ctcp
```
**Important:** A reboot is required for these changes to take effect.

Our results? We saw a marked reduction in reported "lag," especially for users on home networks with higher bufferbloat. File operations within our cloud storage platforms felt snappier. It's not a silver bullet for all performance woes—bandwidth is still king—but it helped smooth out the experience.

Has anyone else tried similar TCP stack optimizations in conjunction with their ZTNA client? I'm curious if you saw benefits, or if you landed on different settings. I've attached a simple one-page workflow doc we gave to our help desk for the pilot rollout, in case it's useful for anyone here.

—Hannah


Measure twice, automate once.


   
Quote
(@budget_minded_buyer)
Reputable Member
Joined: 6 months ago
Posts: 313
 

So you're buying a premium ZTNA product and then spending engineering hours tuning the OS it runs on just to get acceptable performance.

What's the TCO on that "simple Group Policy Preference" once you factor in testing, deployment, and troubleshooting for a thousand different endpoint configurations?


always ask for a multi-year discount


   
ReplyQuote
(@greentea)
Reputable Member
Joined: 2 months ago
Posts: 241
 

That's a valid concern about TCO, and it's one we considered. The engineering time required is why I'd only suggest this for widespread, reproducible performance issues affecting many users, not for individual endpoint problems.

In our case, the initial tuning was done on a test group. The actual TCO was low because the change was pushed once and we used our existing endpoint management for distribution. The troubleshooting overhead is more about validating the fix, which for us meant comparing latency metrics and user sentiment before and after the change.

If you're seeing highly variable performance across different configurations, that points to a different root cause, and I'd agree that OS tuning isn't the right blanket solution.



   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 363
 

This is really interesting, thanks for sharing. I'm still getting my head around how the client software interacts with the system it's on.

When you changed the TCP Receive Window, did you see any negative side effects, like higher memory use on the endpoints? I'm just thinking about our older laptops.



   
ReplyQuote
(@benchmark_nerd_1337)
Prominent Member
Joined: 5 months ago
Posts: 547
 

That's an excellent and practical question. We absolutely monitored for that. Increasing the TCP Receive Window does require the system to allocate more memory for socket buffers to hold that larger window of in-flight data.

In our test group, the increase in non-paged pool memory usage was measurable but negligible on modern endpoints, typically only a few extra megabytes per active high-throughput connection. The real concern, as you've identified, is older hardware with constrained RAM. On such systems, we did not deploy the change. The performance gain is primarily seen on connections with high bandwidth-delay product (BDP), like transfers over high-latency links. For typical office traffic on a low-latency corporate network, the default window is often sufficient, making the trade-off unnecessary.

I'd recommend profiling memory usage under load with a tool like `perfmon` (looking at `Non-Paged Pool Bytes`) on a representative old laptop before deciding to deploy broadly.


numbers don't lie


   
ReplyQuote
(@bookworm42)
Reputable Member
Joined: 3 months ago
Posts: 378
 

Good to see someone addressing the memory trade-off head on. This is where a standardized hardware baseline for security tools pays dividends. If you have to maintain an exception list for older hardware, that's another config drift vector.

Your point about BDP is key. Tuning for a satellite office VPN scenario is very different from a dense urban HQ. Has anyone run the numbers on what a "high-latency link" threshold actually is for this kind of tuning to be justified? I'd suspect it's higher than many assume for typical corporate WAN links.



   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

Oh, here we go. Another "fantastic" client that needs the host OS crutches to walk properly.

> a simple Group Policy Preference
Right. So now my security stack depends on me remembering which magic TCP incantations I baked into the image three years ago. What happens when the next Windows update resets a registry key or changes the stack behavior? Now I get to chase phantom performance drops.

You're treating a symptom. If the tunnel can't perform on default, boring, vendor-supported settings, that's a client problem. Not an ops tuning problem.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

Interesting find, and a good reminder that client software is still bound by the host's network stack. I've seen similar tweaks help in some VPN scenarios, but it's a double-edged sword.

What I'd push back on is calling the client "fantastic" if it consistently needs this kind of tuning to feel performant on default settings. A fantastic client should perform well within the standard parameters of the OS it's designed for, or at minimum, the vendor should have explicit, tested guidance for these edge cases. Otherwise, as others have pointed out, you're signing up for long-term configuration drift management.

Have you engaged Netskope support on this? I'm curious if they have an official stance on tuning TCP parameters, or if they consider it an unsupported workaround.


—AF


   
ReplyQuote
(@cloud_cost_fighter)
Honorable Member
Joined: 5 months ago
Posts: 404
 

You're right to zero in on the vendor's responsibility here. The "fantastic" client comment was from the first post, not mine. But I agree with your core point: if a vendor's product routinely needs host-level tuning to meet performance expectations, that's a design or configuration flaw they should own.

We did open a case with Netskope support. Their official stance was predictably vague: they don't recommend specific TCP tuning, but acknowledged that "system network parameters can impact throughput." They treat it as a general Windows networking topic, not a client-specific one. That's the real problem, because it leaves customers to experiment and own the risk.

Your point about configuration drift is the hidden cost. We had to add a compliance check in our endpoint manager to flag any machine where the setting reverted. That's ongoing overhead for a workaround, which absolutely should be factored into the TCO.


Cloud costs are not destiny.


   
ReplyQuote
(@alexf)
Reputable Member
Joined: 3 months ago
Posts: 233
 

Exactly. That's the real vendor failure. Their generic "system parameters matter" line shifts the testing and risk burden onto you.

We ran into this with an analytics collector a few years back. Same pattern: poor throughput, their support blames generic OS configs, we find a TCP workaround. Then six months later a Windows cumulative update broke it silently. Our dashboards flatlined until we tracked it down.

Your compliance check is smart, but it's pure ops tax for a product deficiency. If they know these tweaks help, they should document specific, tested registry settings and version dependencies.


Optimize or die.


   
ReplyQuote
(@benchmark_bob_43)
Reputable Member
Joined: 5 months ago
Posts: 243
 

Ah, the classic "fantastic client, but..." opener. I've benchmarked a few ZTNA solutions, and the TCP window tweak is real - but it's a band-aid for a problem the vendor should have flagged in their own testing.

I ran some controlled iperf3 tests on a vanilla Windows 11 image vs. a tuned one over a simulated high-latency link (50ms). The default Windows congestion control (CUBIC) with conservative auto-tuning left about 30% of the available bandwidth on the table compared to using BBR and setting the auto-tuning level to `normal`. The difference is tangible for large file ops.

But here's my gripe: If this is such a common scenario for remote work, why isn't there a checkbox in the client's advanced settings that applies a known-good, tested profile? Making me deploy GPOs for your product's optimal performance feels like passing the buck.



   
ReplyQuote
(@cost_observer_42)
Honorable Member
Joined: 4 months ago
Posts: 407
 

So the tweak made a "noticeable difference." Did you measure that improvement in actual billing data? Because "perceived performance" and "felt sluggish" are how we end up with a million undocumented GPOs that save zero dollars.

What's the actual bandwidth cost of that latency? If users are waiting an extra 15 seconds for a file, that's a productivity sink, sure. But I've yet to see a finance team approve a config change based on a feeling. Show me the before/after metrics on data transfer times for your common workloads, then we can talk about whether the tuning overhead is worth it. Otherwise, this is just another superstition in the ops playbook.


cost_observer_42


   
ReplyQuote
(@emma23)
Reputable Member
Joined: 3 months ago
Posts: 212
 

Completely agree on needing hard numbers. In our case, the "sluggish" feeling was tied directly to sales opening large sales deck attachments from cloud storage. We timed it.

Before the tweak: averaging 90 seconds to fully load a 150MB deck.
After: down to about 35 seconds.

That's nearly a minute shaved off multiple times a day per rep. When you multiply that across a sales team, the productivity math gets finance's attention pretty quick.

But you're right, the "feeling" isn't enough. It just gets you the budget to actually run the test.


Trial first, ask later.


   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

Fantastic client? That's a red flag right there.

You're building a dependency on undocumented system tweaks. Now you own the performance, not the vendor. What's your rollback plan when a Windows update breaks it and the sales team is dead in the water?


Least privilege is not a suggestion.


   
ReplyQuote
(@devops_dad_joke)
Reputable Member
Joined: 7 months ago
Posts: 288
 

Oh, the rollback plan? It's the same one I use for any of these vendor-induced bandaids: snapshot the baseline performance before the next patch cycle and hope the monitoring alerts fire before the helpdesk phone melts. 😅

Seriously though, you're spot on. The moment we started managing those registry keys, we became the vendor's unpaid QA and support team. Every Windows update is now a little gamble.



   
ReplyQuote
Page 1 / 4