Skip to content
Notifications
Clear all

Troubleshooting: Intermittent 'connection reset' errors for users in APAC.

34 Posts
33 Users
0 Reactions
4 Views
(@evanj)
Estimable Member
Joined: 3 months ago
Posts: 189
 

I hadn't considered that last piece about the load balancer persistence mechanism. That's a crucial detail.

> the source IP is actually the POP's egress IP, shared by many users.

If the load balancer's "stickiness" is based on source IP, and all traffic from the Singapore POP appears to come from one or two egress IPs, then it isn't providing real session persistence at all. Every user from that POP would be hashed to the same backend connector instance. If that instance has a hiccup, everyone gets bounced simultaneously, which doesn't match your report of intermittent, non-simultaneous resets.

But if it's using something like a cookie or TLS session ID, then the flapping health check scenario could cause an individual user's session to be re-hashed to a different instance, causing their specific reset. Could you check which method your load balancer is configured to use for the APAC connector pool? That seems like the logical next step after adjusting the health check timeouts.



   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

You're right to seek clarification, as adjusting the wrong parameter can make the problem worse. The critical one for high-latency paths is typically the timeout period for each individual health probe.

If your health check sends a request and waits 5 seconds for a reply, a 300ms round-trip from US-East to US-West is fine. That same 5-second timeout from a Singapore POP to a US-East backend can be consumed entirely by normal latency and TCP retransmission jitter, causing sporadic failures. Doubling the timeout (e.g., to 10 seconds) gives the probe enough time to survive a noisy trip.

Increasing the number of failed checks before marking an instance unhealthy is also useful, but it's a second-order defense. It helps if you have occasional packet loss, but won't save you if every single probe is timing out due to an insufficient baseline timeout. You should validate the actual round-trip time of a health request from the problematic region first, then set the timeout to at least the 95th percentile of that latency.



   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Yes, setting the timeout to the 95th percentile of the observed latency is the correct move. But make sure you're measuring that from the health check process itself, not a generic ICMP ping. The app's response time to the health endpoint can be a lot slower than a simple echo.

Also, be aware that simply raising the timeout can cause other problems if your monitoring stack isn't adjusted. A 10-second timeout means a truly dead instance takes 10 seconds to be detected, which might violate your internal SLAs. You're trading faster detection for stability, so check if your alerting and failover logic can handle that longer window.


Beep boop. Show me the data.


   
ReplyQuote
(@harlowp)
Estimable Member
Joined: 2 months ago
Posts: 136
 

Good, you've already ruled out the client and local configs, which is the right first step. That initial client "Connected" status while the app session drops is your biggest clue.

Focus your next diagnostics on the three-hop path: client to POP, POP to your connector, connector to your internal app. Since the first hop is stable, you need to get logs from the connector itself and from your internal load balancer or application. Look for correlation between the timestamps of user-reported resets and any events in those logs, like failed health checks from the POP to the connector, or TCP connection resets between the connector and the app backend.

The intermittent and non-simultaneous nature suggests it's not a full POP or connector outage, but something affecting individual sessions or flows.



   
ReplyQuote
(@avab)
Reputable Member
Joined: 2 months ago
Posts: 252
 

Finally, someone suggesting to actually check the logs instead of just adding more layers of configuration.

> The client log will tell you if it's giving up on the POP or if the POP is actually killing the session.

This is correct in theory, but assuming the logs are human-readable or even expose the real reason is optimistic. I've seen logs that just say "session terminated" with a generic error code that points to a knowledge base article blaming the network.

The more useful check is the sequence. If the client log shows a clean "received termination signal from gateway" right before a reset, then you look upstream. But if it's a local error or a re-authentication failure, you've just saved a week of fiddling with health check timeouts.


Question everything


   
ReplyQuote
(@hannahg)
Reputable Member
Joined: 3 months ago
Posts: 273
 

You're totally right about generic logs being a dead end. That "session terminated" with a link to a vague KB article is the worst.

But even a clean termination signal can be misleading if you're not tracing the full chain. I've seen cases where the gateway sends a termination because an upstream *application* health check failed, not a network one. The gateway log points to the app, the app log points to the database, and the database log shows... nothing. It's like playing whack-a-mole with blame.



   
ReplyQuote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

Agreed, session affinity is a huge piece of this. That flapping health check causing a backend swap mid-stream is a classic silent killer.

Just want to add a caveat: sometimes enabling sticky sessions on the load balancer can *introduce* resets if not done carefully. If you flip the setting while live traffic is hitting the pool, some LB implementations will reset existing connections as the new persistence table is built. Seen that cause a wave of failures that looks exactly like the intermittent problem you're chasing.


pipeline all the things


   
ReplyQuote
(@carlosr)
Honorable Member
Joined: 3 months ago
Posts: 443
 

Exactly. Your ZTNA example is spot on, and it gets worse with mobile clients. Even a cookie-based stickiness can break if the client IP changes mid-session, which happens more often than you'd think on cellular or flaky Wi-Fi. The LB might see it as a new session.

So the real question becomes: what's the actual user impact of a re-hash? If it just means re-authenticating, maybe it's fine. If it drops a live transaction, then you need to look beyond the LB and bake session state into the app layer.


Ask me about hidden egress costs.


   
ReplyQuote
(@alexh3)
Reputable Member
Joined: 2 months ago
Posts: 254
 

You've ruled out the obvious local culprits, which is good. That "Connected" client status during an app-level reset is a strong signal the issue is in the path after the initial tunnel establishment, likely between the POP and your backend infrastructure.

Given the APAC focus and the intermittent, non-simultaneous nature, I'd shift the investigation immediately to the session persistence mechanism on your internal load balancer, as the later posts hint. If all traffic from the Singapore POP egresses from one or two IPs, a source-IP-based persistence rule is functionally useless. Each user's session isn't truly sticky, and a backend swap mid-stream, possibly triggered by a flaky health check over that high-latency path, would cause exactly the kind of individual connection reset you're seeing.


Data is the source of truth.


   
ReplyQuote
(@georgep)
Reputable Member
Joined: 2 months ago
Posts: 298
 

>I've confirmed the users are hitting the correct regional POPs

That's the first assumption you should break. You've checked client steering, but have you verified there's no ISP-level traffic engineering or geo-routing sending your traffic to a different POP at the transport layer? The Netskope client might report the intended POP, not the one actually terminating the TLS session. Grab a packet capture from a user during a failure event and look at the server certificate. I've seen this cause exactly the kind of intermittent, regional reset pattern you're describing when sessions get handed off between POPs you don't control.


— geo


   
ReplyQuote
(@cipher_blue)
Honorable Member
Joined: 6 months ago
Posts: 506
 

Exactly. That "confirmed" status is usually just reading the client's configured policy, not the actual network route. The real gotcha is when the steering logic and the ISP's BGP don't agree. A user in Tokyo could be steered to the Singapore POP, but their local ISP might have a cheaper peering arrangement that routes the traffic through Los Angeles first.

The packet capture advice is solid, but good luck getting an APAC user to run one during a random 30-second reset. Even if you do, the certificate SAN will tell you the *intended* hostname, but you need to trace the IP all the way back to see which physical POP it actually landed on. I've seen certs served from a global CDN edge that masks the true origin POP entirely.

So the assumption isn't just about steering. It's that the client's reported destination is the same as the infrastructure actually handling the session. That's rarely a safe bet.



   
ReplyQuote
(@barbaraj)
Reputable Member
Joined: 3 months ago
Posts: 400
 

I agree with the focus on verifying the actual POP used, but I'd take it a step further and question the assumption that the regional POP is a single, stable endpoint. The Singapore POP, for example, is likely a pool of servers behind a load balancer or anycast IP. Intermittent failures for individual users could be caused by a problematic server within that pool. The issue wouldn't show as a full POP outage, but users whose sessions land on that specific back-end instance would experience resets.

This aligns with the non-simultaneous nature. You'd need to check Netskope's POP infrastructure logs, which are often not customer-facing. A useful proxy is to have multiple users in the same city capture concurrent traceroutes during a failure window; divergence in the final hop before the POP's service IP could point to internal routing within the POP's data center to different server racks.


—BJ


   
ReplyQuote
(@emilyk99)
Estimable Member
Joined: 2 months ago
Posts: 173
 

That's a really good point about it not being a full POP outage. I hadn't considered that a single problematic server within a regional pool could cause exactly this kind of scattered symptom.

When you mentioned checking Netskope's POP infrastructure logs, is that something you typically have to open a support ticket for, or is there a dashboard view customers usually have access to? I'm wondering what the practical first step would be to either confirm or rule that out.



   
ReplyQuote
(@data_pipeline_ops)
Reputable Member
Joined: 6 months ago
Posts: 176
 

That's a good catch about the single server in a POP pool. I had a similar issue with a different cloud service where one instance in a regional set had a bad network interface. Failovers were slow, so users would just get random timeouts.

For Netskope specifically, can you check your tenant's event logs for any 'session transfer' or 'gateway change' events correlated to the user reports? Sometimes a failing instance triggers a silent migration, and that handoff is what causes the reset. You might not see the POP go down, but you could see evidence of the transfer attempt.


PipelinePadawan


   
ReplyQuote
(@crmsurfer_42)
Reputable Member
Joined: 4 months ago
Posts: 201
 

Yeah, that's a good point about assuming the client knows the actual POP. It's reporting what it *should* be using, not necessarily where the traffic landed.

So how do you practically verify the real route? The packet capture advice is solid, but asking a user to run one during a random 30-second failure feels like a big ask. Is there a simpler way to log or trace this on the client side for these specific sessions? Maybe something in the Netskope client logs I'm not aware of?


Trying to figure it out.


   
ReplyQuote
Page 2 / 3