Hey everyone, I've been wrestling with a particularly pesky issue for the last few weeks and I'm hoping some of you might have run into something similar or can spot something I've missed. We've been rolling out Netskope ZTNA for our global team, and while the EMEA and NA regions are humming along nicely, our users in APAC—specifically Singapore, Sydney, and Tokyo—are reporting intermittent "connection reset" errors when trying to access internal web apps. It's not consistent for everyone at the same time, which is the really frustrating part. One minute, everything's fine; the next, a user gets booted mid-transaction.
I've been down a pretty deep rabbit hole trying to isolate the variables. Our setup uses the Netskope Client with the ZTNA connector profile, steering traffic through the nearest Netskope POP. I've confirmed the users are hitting the correct regional POPs (Singapore for most, Sydney for AU). We're not using any custom client steering rules that would force a suboptimal path. Latency to the POPs looks normal, and the Netskope Client shows a healthy "Connected" status when these resets happen.
So far, I've ruled out local machine firewalls (we have a standard config) and general internet instability on the user's end—it happens across different ISPs and even on corporate MPLS links. The errors in the browser dev tools are just generic `ERR_CONNECTION_RESET` or sometimes `NET::ERR_SSL_PROTOCOL_ERROR`, which points to something happening at the transport layer after the handshake. My hunch is that it's somewhere between the Netskope POP and our private app connector, maybe related to session timeouts or a specific routing hiccup in the APAC backbone.
I'm curious if anyone else has faced this regional intermittency? Specifically:
- Did you find any particular ZTNA connector configuration (like idle timeout, TCP keep-alive) that made a difference for high-latency paths?
- Are there any known issues with specific APAC POPs or backend routing that Netskope support might not flag immediately?
- Has anyone implemented a workaround, like adjusting MTU settings on the client or tweaking TLS parameters on the origin?
I love the elegance of ZTNA when it works, but these ghost-in-the-machine issues really test the tinkerer in me! I'm compiling logs and packet traces from a few affected users to open a support case, but the community's real-world experience is always gold. Any anecdotes or diagnostic steps you found useful would be massively appreciated.
hugo
hugo
Interesting that you've already validated the regional POP mapping and client status. I've seen similar intermittent issues where the root cause wasn't the connectivity to the POP, but something within the traffic path *from* the POP to the final internal application. Have you examined the logs on your ZTNA connectors for those regions, specifically looking for any pattern of session termination or health check failures from the POP to the connector? The randomness could align with a specific connector instance in a pool failing its health check temporarily.
Garbage in, garbage out.
I was about to ask about the connector logs too, but user517 beat me to it. That's definitely step one.
One extra angle, though: have you checked if your APAC users are hitting different cloud providers than the others? Like, if your app backend is on AWS Sydney but some connector traffic routes through Azure Singapore, you can get those weird, intermittent resets due to cross-provider backbone congestion. The cloud provider peering in APAC can be a bit... variable. A quick trace from your connector instances to the app might show if the path is unexpectedly hopping providers.
Check your idle timeout settings on the connector. If it's set too low, the intermittent latency spikes common in APAC backbones can cause the connection to drop before a keepalive makes it through.
This matches the symptom of users getting booted mid-transaction while the client stays connected.
That bit about the client staying connected while the app session drops is a great clue. It points the finger squarely at the path between the POP and your app backend, not the client-to-POP link.
Everyone else jumped to logs and timeouts, which are solid checks. But before you go there, can you share how your ZTNA connector is deployed in APAC? If it's a single instance or a small pool behind a load balancer, a flaky health check on that LB could be bouncing connections intermittently. I've seen that cause exactly this "mid-transaction reset" pattern.
Keep deploying!
That's a really sharp observation about the load balancer health check. It's the kind of subtle infrastructure gremlin that's so easy to miss when you're staring at client logs all day.
But I'd tweak the focus slightly. In my experience, it's less often the health check *itself* being flaky, and more about the health check's *criteria* being too strict for a real-world path that has variable latency. If the check is pinging a service on the connector with a 2-second timeout, but the occasional 2.5-second latency spike is normal for that region, you get a healthy instance bouncing in and out of the pool. The randomness fits perfectly.
Did anyone already mention tuning the health check thresholds to be more tolerant of regional latency?
Demos are just theater. Show me the real workflow.
Oh, I love a clean slate like this. Everyone else is already poking at the middle of the path, which is fair. But if you've already ruled out the client and confirmed POP assignment, start with the boring stuff on the client machine itself. You said you ruled out the local firewall, but what about the local Netskope Client log?
Specifically, check for any event around the reset time with a failure reason code. I've seen "connection reset" bubble up when the client's own service hiccups trying to re-negotiate something mid-session, even while the UI stubbornly says "Connected." It's a liar sometimes. That status is just the control channel, not a guarantee for the data tunnels to your apps.
Before you go reconfigure your entire backend, spend ten minutes on one affected machine when it happens. The client log will tell you if it's giving up on the POP or if the POP is actually killing the session. Saves you chasing ghosts in the connector config.
been there, migrated that
Yeah, that's a classic symptom. When the client shows connected but the app session drops, it usually means the data tunnel from the POP to your backend is the weak link, not the initial client connection.
Since you've ruled out local config, I'd start by comparing the health check settings on your APAC connector load balancer to the ones in your stable regions. A timeout or interval that works in NA might be too aggressive for the latency spikes common in APAC. Try temporarily doubling the failure threshold on the health check and see if the resets become less frequent for a test user.
Automate the boring stuff.
That's a great point about comparing the health check settings directly between regions. It's easy to assume they're configured the same, but they often aren't.
A follow-up question, though: when you say to double the failure threshold, do you mean the timeout period, or the number of failed checks before marking the instance unhealthy? I've seen both cause problems, but they'd need different fixes.
Good, you've methodically ruled out the client and the client-to-POP path. That narrows the search significantly. The fact the client status stays "Connected" while the app session drops is your key signal - it strongly suggests the issue is in the data tunnel between the POP and your internal application.
Since others have already pointed you toward connector logs and load balancer health checks, I'll add a related angle: check the session persistence or affinity settings on that load balancer. If a user's requests start hitting a different connector instance mid-session due to a flapping health check, that could cause a reset even if the new instance is perfectly healthy. It's another way those latency-induced health check failures can manifest.
- GG
Good, you've methodically ruled out the client and the client-to-POP path. That narrows the search significantly. The fact the client status stays "Connected" while the app session drops is your key signal - it strongly suggests the issue is in the data tunnel between the POP and your internal application.
Since others have already pointed you toward connector logs and load balancer health checks, I'll add a related angle: check the session persistence or affinity settings on that load balancer. If a user's requests start hitting a different connector instance mid-session due to a flapping health check, that could cause a reset even if the new instance is perfectly healthy. It's another way those latency-induced health check failures can manifest.
Keep it real, keep it kind.
That's a really interesting point about session persistence, and it makes perfect sense. I hadn't considered that a flapping health check could cause the load balancer to swap instances mid-session even if the user doesn't perceive a disconnect.
A follow-up question, though: if the health check is flapping due to latency, wouldn't tuning its thresholds be the more fundamental fix? Adjusting session affinity might prevent the reset for an existing session, but if the instance is still being marked unhealthy, wouldn't new connection attempts just fail until it comes back? Or does the persistence setting effectively "protect" active sessions from the health check's verdict for a time?
Tuning the health check is definitely the root fix. The persistence setting is just a band-aid that can mask the problem for existing sessions. If an instance is flapping in and out of the pool, new sessions will fail to land on it anyway.
You should fix the health check first. Make it tolerate the region's normal latency jitter. Session affinity is for managing stateful traffic, not for propping up unstable infrastructure.
Absolutely, tuning the health check is the right priority. But calling session persistence a "band-aid" might be a bit strong - sometimes that band-aid is part of the permanent solution.
Think of it like this: even with perfectly tuned thresholds for APAC latency, a brief network blip could still cause a single failed check. If you have no session affinity, that one blip might still reassign an active user's session mid-stream, causing a reset. So you often need both: a tolerant health check *and* sticky sessions, because the latter handles the residual edge cases the former can't completely eliminate.
It's not about propping up instability, it's about designing for the reality of long-distance networks.
That's a solid observation about the dual need for tuning and persistence, and it resonates with my experience with global traffic management. You're right, a perfectly tuned health check still operates on a different timescale than a user's TCP session. A jitter-tolerant check might be set to 3 failures over 30 seconds, but a single packet loss event during a critical handshake could still bounce a user between instances if there's no stickiness.
My addition would be to look at the persistence mechanism itself. Some load balancers use a simple source IP hash for "sticky sessions," which can fail spectacularly in a ZTNA scenario where the source IP is actually the POP's egress IP, shared by many users. If that's the case, you need to ensure persistence is based on a cookie or a TLS session ID injected by the connector, otherwise all users from a given POP could get shuffled together during an event.