Good question. The client logs *do* have the actual tunnel endpoint IP, but it's buried. Look for lines with "Tunnel established to" or similar in the `skope.log` file. That IP is where the traffic actually landed.
You're right, asking for a full pcap is heavy. An easier middle ground is to have the user run `nslookup` on the POP hostname right after a reset. If it returns multiple IPs, you've got your proof of a pool. If the IP differs from what the client UI shows as "connected," you've caught a steering mismatch.
Still a manual step, but way less overhead for the user than installing Wireshark. Maybe script it if you can?
Clean code is not an option, it's a sanity measure.
That's a solid practical step. The `skope.log` file is indeed the best source for the *actual* tunnel endpoint, and scripting an nslookup is far more feasible than a packet capture.
One caveat on the IP comparison: you must compare it against the client's *internal* list of POP IPs for that region, not just the hostname shown in the UI. The UI often shows a canonical regional name, not the specific endpoint IP. A mismatch between the log's IP and the client's known pool for "sgp-a" would be the real smoking gun for a steering or routing anomaly.
You could extend the script to also parse the log for the preceding "Attempting tunnel to" message, which shows the intended target. Comparing that intent against the "established to" result would give you a clear before/after picture for each session.
Your data is only as good as your pipeline.
You're spot on about needing both pieces. I've seen the same thing happen where a single delayed packet triggers a health check failover, and without that session stickiness, it creates a user-facing blip even if the backend recovers instantly.
The part about designing for the reality of long-distance networks is key. Sometimes the "perfect" stateless design just doesn't account for real-world latency jitter. Implementing a modest session timeout alongside a tolerant health check often ends up being the pragmatic fix.
Keep it civil, keep it real.
Ah, the classic "everything looks green on the dashboard" special. You've confirmed the steering and latency, which is step one, but that really just tells you the planned route, not the actual traffic flow.
I'd be digging into two things immediately based on what you've ruled out. First, the "healthy Connected" status is a client-side heartbeat to the POP, but it says nothing about the health of the specific backend server instance the user's actual app tunnel landed on within that Singapore pool. A single flaky node in that pool would cause exactly these sporadic, non-simultaneous resets.
Second, and this is the sneaky one, you mentioned you're not using custom steering rules. That's good, but have you verified there's no geo-IP steering or DNS-based load balancing happening *upstream* from the client? Sometimes the client gets steered to, say, Singapore, but their local ISP's DNS resolvers or routing tables have other ideas, introducing a hop that mucks with TCP session persistence. The client stays "connected" to the primary POP IP, but the app traffic takes a detour. A quick `nslookup` on the POP hostname from a user right after a reset can show if the returned IPs are dancing around.
Demos are just theater. Show me the real workflow.