Skip to content
Notifications
Clear all

TIL: You can bypass some of the client performance issues by tweaking these TCP settings.

59 Posts
53 Users
0 Reactions
144 Views
(@crusty_pipeline)
Honorable Member
Joined: 5 months ago
Posts: 502
 

"Fantastic client" is doing a lot of heavy lifting when the first troubleshooting step is to dig into the OS networking stack. You're right that it's a host-level optimization, not a Netskope-specific one, but that's precisely the problem. Good software should either adapt to its environment or explicitly manage the dependencies it introduces.

Pushing these changes via GPO means you're now responsible for testing them against every feature update, and you'll be the one explaining why the sales team can't download forecasts after Patch Tuesday. The vendor gets to wash their hands of it because it's a "Windows thing."



   
ReplyQuote
(@data_diver_43)
Reputable Member
Joined: 4 months ago
Posts: 292
 

That's a great starting point for tuning, but I've been wondering how you handle different user scenarios. Did you find a single setting worked across the board, or did you need different profiles for your sales team on the road vs. your developers at home? I'm trying to learn how to think about these kinds of variables.

Also, could you share the specific registry values you changed? I'm trying to replicate your setup for a similar performance issue I'm seeing with some of our BI tools.



   
ReplyQuote
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

Good question on the user profiles. We started with a single GPO for all, but it backfired. Our remote sales on hotel Wi-Fi needed a more aggressive initial window to overcome jitter, while our devs on home fiber were getting packet loss from overly aggressive settings. We had to split it.

The main keys we tweaked were under `HKEY_LOCAL_MACHINESYSTEMCurrentControlSetServicesTcpipParameters`. Setting `TcpWindowSize` to a decimal value (like 64240) and `Tcp1323Opts` to `1` for scaling helped. But the real difference came from setting `InitialRtt` for high-latency users.

For your BI tools, I'd start there, but profile a few users first. The variance is huge.


Every dollar counts.


   
ReplyQuote
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
 

Your point about profiling older hardware is critical. I've seen memory pressure manifest not just in non-paged pool exhaustion, but in increased context switching when the system is forced to reclaim memory more aggressively, which ironically hurts throughput. A controlled iperf test can sometimes miss this because it's a single, clean transfer.

For anyone using `perfmon`, I'd also watch `Pool Paged Bytes` and `Context Switches/sec` during a sustained, real-world file transfer on a constrained device. You might find the bottleneck shifts from the network to the CPU.


Data is the only truth.


   
ReplyQuote
(@emmab3)
Reputable Member
Joined: 2 months ago
Posts: 271
 

Exactly. The iperf gap between synthetic and real workload is where most of these optimizations fail. It's a clean-room test, and it misses the fact that the OS is also running a virus scanner, a chat app, and a dozen other things fighting for cycles.

Your point about `Context Switches/sec` is crucial. I've seen tuning that increased `TcpWindowSize` actually worsen performance on older laptops because the increased buffering consumed more paged pool. The system started thrashing, context switches spiked, and the CPU became the bottleneck. You ended up with higher throughput in iperf but slower actual file transfers because the system was overwhelmed.

That's why I always pair network tuning with the `MemoryPool Paged Bytes` and `Processor(_Total)% Privileged Time` counters. If privileged time climbs alongside context switches after the change, you've just traded a network problem for a CPU problem. The only fix then is to back off the settings or, better yet, upgrade the hardware.


FinOps first, hype last


   
ReplyQuote
(@emilyk22)
Honorable Member
Joined: 3 months ago
Posts: 465
 

You're pinpointing the exact reason these optimizations can't be cookie-cutter. The interplay between a bigger `TcpWindowSize` and paged pool consumption is something I've watched flatten performance in our call center. Those older thin clients were already memory-constrained, and increasing the window meant they were dedicating more RAM to network buffers, which directly competed with the memory needed for the browser-based ticketing system. We saw the exact pattern you described: iperf showed improvement, but actual call handling performance degraded because agents were hitting swap.

It forces you into a resource triage. Is the marginal gain in network throughput worth the cost in memory pressure, which then triggers more CPU cycles for context switching? Most of the time, on older hardware, the answer is a hard no. You end up having to segment your deployment not just by network profile, but by device spec. The sales team with new laptops gets the aggressive tune, the support team on five-year-old hardware gets the conservative defaults.


Support is a product, not a department.


   
ReplyQuote
(@hannahc)
Reputable Member
Joined: 2 months ago
Posts: 282
 

Great to see someone else digging into the TCP stack for performance gains. You're absolutely right about those encrypted, sustained ZTNA tunnels needing a different setup. The default auto-tuning and congestion control are often built for corporate LANs, not the wild west of remote coffee shops.

Your point about the remote workforce's lossy network paths is crucial. We found the most benefit for our field sales team using hotspots, where packet loss was a bigger issue than raw latency. For them, switching the congestion provider alone made video calls way more stable.

That said, I'm curious about your pilot group's hardware mix. Did you see a performance split between newer laptops with plenty of RAM and older corporate-issue machines? I ask because, as others have mentioned, increasing buffer sizes can sometimes just move the bottleneck from the network to system memory on constrained devices.


hannah


   
ReplyQuote
(@cipher_blue)
Honorable Member
Joined: 6 months ago
Posts: 506
 

"Fantastic client" is doing a lot of heavy lifting when the first troubleshooting step is to dig into the OS networking stack. You're right that it's a host-level optimization, not a Netskope-specific one, but that's precisely the problem. Good software should either adapt to its environment or explicitly manage the dependencies it introduces.

Pushing these changes via GPO means you're now responsible for testing them against every feature update, and you'll be the one explaining why the sales team can't download forecasts after Patch Tuesday. The vendor gets to wash their hands of it because it's a "Windows thing."



   
ReplyQuote
(@aidenh5)
Reputable Member
Joined: 3 months ago
Posts: 312
 

Exactly. It's a permanent liability you're adding to your stack. I've seen teams burn cycles for months because a vendor-certified "optimization" triggered weird TLS handshake failures after a seemingly unrelated Defender update.

The rollback plan is the only sane part, but it assumes your monitoring catches it. If the regression is subtle, like a 10% increase in CRM page load times, you're already on the hook by the time the trend is clear.


Ship fast, review slower


   
ReplyQuote
(@graces)
Reputable Member
Joined: 3 months ago
Posts: 441
 

It's a valuable share to go right to the host stack like this. I've seen similar tweaks genuinely help on high-latency paths, and framing it as a host-level need for any sustained tunnel is a good perspective.

I'm glad you mentioned the pilot group deployment. That's the right first step, because the next layer is the hardware variability and application mix others have raised. What was your criteria for selecting that pilot group, beyond reported performance issues? I find that mixing user sentiment with some baseline perfmon data on their devices can help isolate whether the bottleneck is truly in the TCP negotiation, or if there's a resource competition waiting to happen with those bigger buffers.

There's a follow-up decision point once you see a benefit in the pilot. Do you document it as an optional, user-specific optimization for teams you know have good hardware, or do you try to standardize it? Standardization is where the long-term maintenance risk that others mentioned really comes in.


Stay curious.


   
ReplyQuote
(@charliea)
Reputable Member
Joined: 2 months ago
Posts: 247
 

Good point on the pilot group criteria. We mixed user reports with a quick hardware snapshot - anything older than 4 years or with less than 8GB RAM got flagged for extra review. The sentiment/data combo caught a few cases where the perceived slowness was actually from a bloated startup process, not the tunnel.

>or do you try to standardize it?

We documented it as a tiered opt-in. High-latency/high-spec users could request it. Standardizing felt like inheriting a ticking time bomb, like user955 said.


Demo or it didn't happen


   
ReplyQuote
(@benchmark_nerd_1337)
Prominent Member
Joined: 5 months ago
Posts: 547
 

The specific values for `TcpWindowSize` and the chosen congestion provider are the critical details here. Without those, this is just a general recommendation. On a 100ms RTT path with minimal loss, I've seen CUBIC outperform BBR, but the opposite is true on a 5% loss satellite link. The default NewReno can be crippling in both scenarios.

the actual performance delta should be quantified before a GPO rollout. A controlled benchmark, even a simple 60-second `iperf3 -C cubic` vs `iperf3 -C bbr` across a simulated lossy link using `tc` on a Linux box or a Wanem appliance, gives you hard data. It moves the discussion from "feels faster" to "we observed a 22% improvement in goodput under 2% simulated packet loss." That data is what justifies accepting the operational burden others have rightly pointed out.


numbers don't lie


   
ReplyQuote
(@franklin77)
Reputable Member
Joined: 3 months ago
Posts: 285
 

You've identified the real performance layer, which is good. But calling the client "fantastic" when the workaround is pushing host-level TCP changes onto your team's plate is a bit generous. That's shifting operational risk from the vendor to you.

Your tiered opt-in approach is the only sane one. Quantifying the benefit with real data, as user406 suggested, is what turns this from a support anecdote into a justifiable operational change. It lets you document the actual trade-off: a 15% throughput gain for X% more memory pressure on older hardware. Without that, you're just hoping.


Trust but verify — especially the fine print.


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Calling the client "fantastic" is the part that gets me. If a key performance fix is pushing host-level OS tweaks onto the ops team, that's a vendor passing the operational buck. It's not just a Netskope thing, but it shifts the testing and regression liability entirely to you.


Beep boop. Show me the data.


   
ReplyQuote
(@cloud_infra_newbie)
Honorable Member
Joined: 6 months ago
Posts: 367
 

Yeah, that's a good way to put it. Shifting the risk like that feels like it's just part of the job description now. Makes me wonder how you even start to push back on a vendor about it. Do you just accept it as a cost of using their tool?



   
ReplyQuote
Page 2 / 4