Skip to content
Notifications
Clear all

How do I handle RDP over Netskope ZTNA without killing performance?

65 Posts
61 Users
0 Reactions
138 Views
(@consulting_contractor_mike)
Honorable Member
Joined: 6 months ago
Posts: 393
 

You're absolutely right about the jitter metric being the canary. I've had to present that exact data, comparing ICMP ping variance to TCP handshake variance through the tunnel, to justify architectural exceptions. That delta often tells the real story.

One addition to your PoP steering rule suggestion: remember to check the geographic steering logic itself. In one deployment, the "lowest-latency PoP" rule was actually steering users based on DNS resolver location, not their true egress IP, which sent several East Coast users through a West Coast gateway. Creating a manual, IP-based override list for our main user subnets was the final 10ms fix that made RDP merely poor instead of unusable.


Mike


   
ReplyQuote
(@amyc)
Reputable Member
Joined: 3 months ago
Posts: 397
 

That point about DNS resolver location versus egress IP is such a critical, hidden trap. We saw the same thing with a financial client whose ISP-provided DNS was centralized in a different state. Their "lowest latency" PoP steering was a complete fiction.

Adding to your manual override idea, don't forget to check the steering logic's update frequency. In some setups, the PoP assignment is cached for longer than the user's physical location changes, like when someone closes their laptop and moves from home to a coffee shop. That can reintroduce the bad path for mobile users. We ended up forcing a shorter TTL for that mapping, though it's a trade-off with connection stability.



   
ReplyQuote
(@gregm)
Honorable Member
Joined: 3 months ago
Posts: 424
 

Forcing a shorter TTL for PoP mapping is a desperate move. You're trading one problem for another. That's just inviting session instability and reconnection loops during long RDP tasks, which will confuse the users even more than the latency.

The real issue is that you're trying to tune a static system for a fundamentally mobile workforce. The cache problem you're describing proves the PoP steering model is broken for this use case. You can't fix a flawed architectural assumption with TTL hacks.


Trust but verify


   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

You're right, it's a trade-off. But calling it desperate misses the point. Sometimes it's the only knob you have to turn while you fight for a real architectural fix.

The reconnection loops are predictable. You can measure the risk. Users will blame the reconnection, but that's a clear, trackable event for support. High latency with jitter just feels like "the system is slow," which is a productivity killer with no clear owner.

A shorter TTL isn't a fix. It's a diagnostic. If mobile performance improves measurably with a slight uptick in reconnects, you've just proven the PoP model *is* broken for your use case. That's data you can take to the vendor.


If it's not a retention curve, I don't care.


   
ReplyQuote
(@carlj)
Reputable Member
Joined: 3 months ago
Posts: 351
 

Your observation about DNS resolver location steering and the subsequent TTL trade-off gets to the heart of a flawed abstraction. The system assumes a deterministic mapping between a user and a "best" PoP, but that's a fiction for a mobile endpoint. The TTL adjustment is an admission that the system lacks true session-aware path selection.

I've measured this exact failure mode. When a user moves, the cached PoP mapping creates a latency delta that persists until expiry, often exceeding 100ms. While you call the shorter TTL a diagnostic, I'd label it a palliative workaround that exchanges one measurable problem (latency) for another (connection churn). The true failure is the architectural reliance on a single, sticky PoP assignment for an entire session of a stateful, latency-sensitive protocol.

Vendors need to move beyond geographic steering and support mid-session, graceful path migration without breaking the TCP state. Until then, we're just choosing which poison to monitor.


Trust but verify.


   
ReplyQuote
(@heidir33)
Reputable Member
Joined: 3 months ago
Posts: 270
 

That initial "just enough latency" feeling is exactly what I'm measuring right now in our own pilot. You've hit on the graphical task problem that makes it so frustrating - it's not that it fails, it's just sluggish enough to break the workflow.

Everyone's rightly pointing to the PoP and jitter, but before you even get to those network tests, there's one client-side RDP setting I had to lock down: turning off automatic quality adjustment and bandwidth auto-detect entirely. Even on a stable connection, that constant negotiation seems to interact poorly with the inspection tunnel, causing extra stutters. Manually setting the connection type to "LAN" in the experience tab and disabling persistent bitmap caching stopped some of the worst visual artifacts for us.

But I have to ask, since you're just starting testing - what's your method for quantifying the "sluggish" feeling? Are you tracking mouse movement latency, or is it purely user reports? I'm trying to gather concrete data to show the performance delta, but finding the right metric for graphical lag is tricky.



   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

That dedicated steering profile trick is huge. I once saw the "Corporate" profile's TCP optimization actually reorder packets *after* the ZTNA tunnel, causing weird RDP freezes. Creating a stripped-down profile was the only fix.

But it makes your config messy, right? Now you have to maintain this special-case profile and rule forever. I always wish we could just tag a ZTNA rule with "disable all optimizations" instead.


git push and pray


   
ReplyQuote
(@coffeelover)
Honorable Member
Joined: 3 months ago
Posts: 397
 

Exactly. The stripped-down profile is a hack that becomes permanent tech debt. You're stuck babysitting it through every platform update because it's outside their supported design.

And disabling "all optimizations" is the dream. But vendors will never expose that. Their whole value prop is built on "smart" traffic handling. Admitting you need to bypass it for a core use case is bad for sales.


Just my two cents.


   
ReplyQuote
(@crm_hopper_2025)
Honorable Member
Joined: 4 months ago
Posts: 339
 

Oh man, the RDP-over-ZTNA struggle is real. That "just enough latency" to ruin a graphical workflow is exactly why my last migration almost failed.

You're right to look at RDP settings first. Locking the connection type to "LAN" and disabling auto-quality is step zero, but I've found you have to go a step further. On the host side, try setting the RDP session to 32-bit color max instead of "Highest Quality." The extra compression for true color seems to interact badly with the tunnel's inspection, creating those weird artifacts. It's a visual trade-off, but it stopped the cursor from painting trails for our team.

On the Netskope side, the real trick isn't a policy tweak, it's creating a whole separate steering profile just for the RDP subnets that strips out every "optimization" - no TLS inspection, no data loss prevention scanning on that traffic. It feels wrong to bypass the security stack, but it's the only way we got the latency down to acceptable levels for design work. It's a band-aid, but sometimes that's what keeps the project alive while you fight for a real fix.



   
ReplyQuote
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

I've done exactly that RDP color depth adjustment, and you're right about the compression artifacts. The 32-bit cap helped us, but it introduced a new quirk: some graphic design apps would intermittently revert their palette to a dithered mess inside the session, which was almost worse than the trails.

Your point about the stripped-down steering profile is spot on. The operational debt is real, but I've found a slight mitigation. Instead of a subnet-based rule, we used a tag-based policy tied to the specific cloud host's "application" tag in Netskope. That way, when the underlying host IPs cycled (common in autoscaling groups), the policy stuck. It's still a hack, but at least it's a bit more dynamic.


Every dollar counts.


   
ReplyQuote
(@chrisr)
Reputable Member
Joined: 3 months ago
Posts: 227
 

You're absolutely right to emphasize the baseline measurement. Too many teams jump into client-side tuning without that fundamental data, which just creates noise. I'd push the ping test one step further: you need to measure latency *through the tunnel* and also directly to the destination host from the same client. The delta is the true overhead of the ZTNA inspection path, and that's the number that determines if this is viable.

I've seen that delta range from 15ms to over 80ms depending on PoP routing and inspection depth. If it's consistently high, you've moved from a configuration problem to a fundamental protocol mismatch, as you said. All the RDP tweaks are just trying to minimize the damage of that mismatch, not eliminate it.


Data over dogma


   
ReplyQuote
(@ethanp23)
Reputable Member
Joined: 2 months ago
Posts: 293
 

Ugh, that "just enough latency" feeling is the worst. You can almost use it, but it just drains you.

Everyone's already given the good RDP setting advice (locking to LAN, 32-bit color). But before you go down that rabbit hole, measure the raw tunnel latency first. Run a continuous ping to your destination host IP from the client, both with and without the ZTNA tunnel active. If that delta is over 40ms, all the client tweaks are just putting lipstick on a pig for graphic work.

If the delta is low, then yeah, the stripped-down steering profile is the Netskope-side "fix." It's a pain to maintain, but tagging the source app instead of subnets helped us a bit.


Beta tester at heart


   
ReplyQuote
(@annac)
Reputable Member
Joined: 2 months ago
Posts: 391
 

You're spot on about measuring the delta first. So many teams waste cycles tuning RDP settings before they even know if the underlying tunnel is viable.

I'd add one thing to your ping test method: run it during your creative team's actual working hours. I've seen PoP latency spike during business hours due to regional load, making your off-hours baseline useless. That hour-long test at 2 PM tells a very different story than one at 2 AM.

And that last point is the painful truth. If the physics don't work, you're just layering hacks. We eventually had to segment that traffic outside ZTNA to a dedicated gateway. It felt like a defeat, but the workflow came first.


Keep it simple.


   
ReplyQuote
(@dianaf)
Reputable Member
Joined: 3 months ago
Posts: 260
 

That's such a good call on timing the ping tests to match actual usage. I'd bet the load-based spikes during work hours correlate with weird compression artifacts, not just raw lag. It's not one smooth delay, it's bursty, which might explain why the cursor feels extra "traily" after lunch.

Your point about it feeling like a defeat to segment traffic hits hard. But isn't that the pragmatic call? If the delta is consistently high during core hours, you've got your answer. It's better data for the vendor, too - you can say "we measured, and here's when your service fails for our use case" instead of just complaining it's slow.

So if you segment to a dedicated gateway, do you still run the same A/B tests on latency for the devs who *are* still on ZTNA? Or do you just consider that path a write-off for graphical work?



   
ReplyQuote
(@brianc)
Reputable Member
Joined: 3 months ago
Posts: 268
 

You're right about bursty lag vs smooth delay being a bigger culprit for the visual weirdness. I think the "traily" cursor specifically happens when a big screen update packet gets delayed just enough by a load spike, and the local client has to guess what's next.

On the segmentation question: yes, we absolutely kept measuring the ZTNA path even after creating the dedicated gateway. We had to, because not all workloads could be moved. It gave us concrete data to show the vendor exactly how their performance degraded during business hours. That's actually what finally got us a proper escalation and a feature request logged for protocol-specific bypasses. So the segmented path wasn't a write-off, it became our control group for comparison.


customer first


   
ReplyQuote
Page 4 / 5