Skip to content
Notifications
Clear all

Why is Netskope so slow for some SaaS apps? Troubleshooting thread

2 Posts
2 Users
0 Reactions
23 Views
(@devops_grandad)
Reputable Member
Joined: 4 months ago
Posts: 354
Topic starter   [#6092]

Alright, let's cut through the marketing fluff. I've been deploying and managing Netskope for a few years now across several orgs, and there's a consistent pattern I see: performance for most internal apps is fine, but certain SaaS applications become absolute dogs. We're talking about latency that makes a 1998 dial-up modem look snappy, specifically with apps like Salesforce, Workday, and some custom-built platforms on AWS.

This isn't a simple "the internet is slow" issue. It's intermittent, it's user-specific, and it kills productivity. The vendor's first line is always to blame your endpoint or your network, but after you've ruled those out, you're left staring at a very expensive ZTNA gateway that's supposed to be "cloud-native and fast."

Based on getting my hands dirty with packet captures and some old-school troubleshooting, here are the usual suspects I start with. This is a checklist for anyone else banging their head against the wall.

* **The TLS Inspection Conundrum:** This is culprit number one. Netskope, by default, wants to decrypt and inspect everything for its CASB and SWG features. Some SaaS applications use non-standard TLS implementations, certificate pinning, or massive numbers of connections. The overhead of decrypting/re-encrypting every single packet for these beasts adds significant latency.
* **Action:** Check your SSL Decryption policies. You might need to create exceptions. Don't just blindly bypass—do it strategically. Use the Netskope UI to see the specific app instances causing the most handshake delay.

* **Geographic Steering (or lack thereof):** Netskope has POPs everywhere, but the steering logic isn't perfect. If your user in London is being routed through a POP in Frankfurt to hit a Salesforce instance in Dublin, you're adding hops. Then there's the hairpinning: traffic goes user -> Netskope POP -> SaaS DC -> Netskope POP -> user, for inspection.
* **Action:** Use `traceroute` and Netskope's own session diagnostics (from the client) to map the actual path. Compare it to a direct connection. You might need to engage support about POP affinity rules.

* **Client Configuration & Suboptimal Path Selection:** The Netskope client's "best path" algorithm can get it wrong, especially with newer "Direct-to-Net" or "Private App Access" modes. A misconfigured `gateway.json` or tenant settings can force all traffic through a tunnel when it doesn't need to be.
* **Action:** Review the client logs. Look for path switching. A simple test is to force a direct connection (if policy allows) and compare.

* **Application-specific quirks:** Modern SPAs (Single Page Applications) like Salesforce make hundreds of concurrent API calls. Each one is a new TLS session that Netskope has to broker. The cumulative overhead is huge. Some apps also use long-lived TCP connections or specific WebSocket protocols that don't play nice with the proxy.

**Basic troubleshooting steps I run through:**

1. **Isolate the issue:** Have the user reproduce the slowness. Immediately pull the Netskope client logs.
```bash
# On macOS, the logs are typically here
tail -f /Library/Logs/Netskope/nsagent.log

# On Windows, use Event Viewer or the log path in ProgramData
```
Look for errors, session timeouts, or frequent path changes.

2. **Bypass test:** Temporarily add the problematic application to the SSL Bypass list in your policy. **This is a test, not a solution.** If performance returns to normal, you've confirmed TLS inspection is the core problem. Now you have to work out if you can live with the risk of not inspecting that traffic.

3. **Path analysis:** From the affected endpoint, run:
```bash
# To the SaaS app domain with Netskope connected
traceroute app.problematic-saas.com

# Then, disconnect the client or use policy bypass and run it again
```
Document the difference in hops and latency at each stage.

The bottom line is that ZTNA isn't magic. It's a complex proxy chain, and complexity introduces latency points. Your job is to find which piece of the chain is the bottleneck. Start with the TLS decrypt/encrypt cycle and work your way out.



   
Quote
(@jordanf)
Trusted Member
Joined: 3 months ago
Posts: 42
 

You're absolutely right about TLS inspection being the primary bottleneck. It's not just non-standard TLS or pinning, though. The order of operations in the inspection chain matters a lot. If DLP or threat protection policies are configured to scan uploaded content *before* it's passed to the service, you'll see huge delays with Salesforce or Workday when users try to upload large reports or documents.

A follow-up question: in your packet captures, have you noticed if the latency spikes correlate with specific API endpoints for those apps? I'm wondering if the issue is more pronounced with APIs that use long-lived connections or specific authentication handshakes that the gateway might be terminating and re-establishing poorly.



   
ReplyQuote