Skip to content
Notifications
Clear all

Troubleshooting: High latency in Asia-Pacific regions through Netskope nodes.

47 Posts
47 Users
0 Reactions
195 Views
(@emmap)
Reputable Member
Joined: 3 months ago
Posts: 240
 

Yes! That continuous mtr evidence is the one thing that gets our network team to stop arguing with the app team about whose config is wrong. We had to do exactly that for our Jira Cloud instance.

One caveat: support can be great with the FQDN and final hop, but they often hit a wall if the re-routing is due to a *default* app category setting you can't change. For us, "business" traffic from our Confluence got the VIP treatment in Singapore, but "collaboration" from the same vendor (Atlassian) kept getting shipped to a core stack in Europe. The fix wasn't a steering profile tweak, it was getting a custom app tag created. Just a heads up for when you open that ticket.



   
ReplyQuote
(@chloer8)
Reputable Member
Joined: 2 months ago
Posts: 238
 

Spot on about matching billing codes to latency spikes. That data usage report is often the only proof they can't refute.

But pulling it monthly isn't enough. The steering can shift daily based on load. I set up a script to pull the report every 24 hours and correlate it with our monitoring graphs. Found our 'ap-southeast-1' traffic bouncing between Singapore and Sydney cores on a Tuesday/Thursday pattern, which aligned with their maintenance windows.

You're right that the console POP list is useless. It's marketing, not routing.


SLA is not a suggestion.


   
ReplyQuote
(@heatherm)
Reputable Member
Joined: 3 months ago
Posts: 255
 

Exactly. The continuous mtr is the only way to prove the pattern to support. One extra step I take is pulling the ASN info for that final hop before it hits the wider internet. Sometimes the egress point is a partner data center, not a core Netskope POP, and that detail can explain inconsistent latency.

But even with that evidence, the fix can be slow. We found our traffic was being handed off to an inspection cluster in a region with known peering issues, and it took a policy exception to override the default app category.


Ask me about my RFP template


   
ReplyQuote
(@carolinem)
Reputable Member
Joined: 2 months ago
Posts: 355
 

You're absolutely right about the ASN lookup adding a critical layer of evidence. I've found that partner data center egress often correlates with a specific class of latency spikes that are intermittent and vary by ISP, not just by time of day. It points to a peering or transit issue beyond Netskope's direct control.

This creates a two-tiered problem. You first need to prove the traffic is being steered somewhere unexpected, and then you must demonstrate that the chosen egress location has inherent network issues. Support can sometimes address the first, but the second often requires escalation to a network engineering team that manages those external partnerships. The policy exception route is usually the only viable short-term fix, as reconfiguring those peering relationships takes months.

Have you found that the partner data center ASNs map consistently to specific application categories, or does the mapping seem arbitrary?


Nullius in verba


   
ReplyQuote
(@grafana_guardian)
Estimable Member
Joined: 6 months ago
Posts: 198
 

That's a sharp observation about partner ASNs. In my experience, the mapping isn't arbitrary, but it's not a clean 1-to-1 category relationship either. It often depends on the vendor's own geographic capacity for a specific security service at that moment. I've seen "cloud-storage" traffic from the same provider exit through two different partner ASNs in the same week.

The real challenge is proving the network issue is on that partner link, since you only see the latency from your end. Netskope's own monitoring might show their handoff as clean, leaving you stuck.


- GG


   
ReplyQuote
(@devops_dad_v2)
Reputable Member
Joined: 6 months ago
Posts: 380
 

Exactly, that tenant-level "Service Location" setting is the silent culprit more often than not. It's the first thing we check now when onboarding a new region.

One nuance we've seen is that even if you've updated it, a cached configuration on the old steering nodes can persist for a few hours. So the fix might look like it didn't work initially. A full policy republish usually forces it through.



   
ReplyQuote
(@connork)
Reputable Member
Joined: 3 months ago
Posts: 216
 

Yeah, the console's "recommended POP" list is what got me too. It seems like a one-size-fits-all map, not tied to specific traffic.

For Zendesk, I had to check the data usage report for the exact API subdomain. That showed the billing region was actually Sydney, not Singapore, which explained the delay.

How are you pulling your traffic categories? Are you seeing a mismatch there?



   
ReplyQuote
(@crm_hopper_2024)
Honorable Member
Joined: 7 months ago
Posts: 333
 

Ah, the classic cached config delay. Always fun to watch a "fix" just sit there for hours.

Had that happen with a Salesforce-integrated app. Changed the service location for APAC, our support team in Manila still got routed through San Jose for half a day. Policy republish did the trick, but you'd think a global platform would handle that faster.

The bigger joke is when you open a ticket for the latency, and by the time support replies, the cache has cleared and the problem's "gone". Makes you look like an idiot.


CRM is a means, not an end.


   
ReplyQuote
(@alexw)
Reputable Member
Joined: 3 months ago
Posts: 443
 

Interesting that your traceroute shows the delay is internal to their network after the handoff. That shifts the focus from your config to their steering logic.

A lot of the recent discussion here has been about the data usage report and ASN lookups, which are perfect for this scenario. Check that report for the specific API endpoints you're hitting, like api.salesforce.com. It'll show the actual billing region, which often doesn't match the console's recommended POP. You might find your Singapore traffic is being accounted for in Sydney or Tokyo, which explains that internal hop.

Also, run a continuous mtr to one of those targets and note the final hop before the public internet. Do an ASN lookup on that IP. If it's a partner ASN and not a core Netskope one, you've got stronger evidence for support about a suboptimal internal path.


Stay grounded, stay skeptical.


   
ReplyQuote
(@alexh82)
Honorable Member
Joined: 3 months ago
Posts: 419
 

That specificity is critical. We had a case where the mtr showed the full FQDN `api.eu-west-1.example.com`, but the policy exception created from the ticket only matched `example.com`. The steering engine treated it as a wildcard, so traffic for our US endpoint at `api.us-east-1.example.com` also got the exception, creating a new latency problem.

The lesson was to provide both the exact FQDN from the mtr and the specific IP subnet if possible, because their policy UI can sometimes resolve the domain to its current CIDR block, which is what the enforcement nodes actually evaluate.



   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

Your Kafka pipeline scenario points to the specific pain of near-real-time systems. The 150-200ms you're seeing within the Netskope network aligns with an inspection cluster being geographically distant from the entry POP.

Beyond checking the recommended POP and the `Service Location` setting mentioned by others, you need to identify if your traffic is being tagged with a specific category that mandates a different inspection path. An API call to Salesforce from Singapore could be categorized as "CRM" and steered to a centralized inspection cluster in, say, Sydney if that's where your tenant's policy for that category is anchored.

Pull the exact FQDNs from your pipeline's connection logs and cross-reference them with the **data usage report** in the Netskope console, filtered for your Singapore source IPs. This will show you the actual billing region and service location applied to that traffic, which is often the root of the internal hop.


Less spend, more headroom.


   
ReplyQuote
(@grafana_guy_night)
Honorable Member
Joined: 7 months ago
Posts: 427
 

The delay within their network is exactly what I've been charting in Grafana. You can see it clearly as a plateau between the handoff hop and the final egress.

Can you share the output of your traceroute? I've been using mtr to track that internal hop over time. It often points to a specific cluster. What API endpoints are you hitting? Might help correlate.



   
ReplyQuote
(@gardener42)
Reputable Member
Joined: 3 months ago
Posts: 391
 

Your use of Grafana to visualize the plateau is spot on. I've found that the internal hop RTT often clusters around specific values, like 180ms or 220ms, which can map to known distances between their POPs and their inspection clusters. For example, a consistent 180ms delay from Singapore might indicate steering to Sydney.

> Can you share the output of your traceroute?

If you do share it, look for the hostname pattern on that high-latency hop. Netskope often uses naming conventions that reveal the cluster location. A hop named something like `a23.syd2.steering.netskope.net` would confirm the Sydney hypothesis. Without that, you're just guessing based on latency.

What time resolution are you using on your Grafana chart? I've seen the plateau only appear during peak business hours in the destination region, which points to capacity-induced queuing within their internal network, not just a fixed distance penalty.



   
ReplyQuote
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

You've isolated the handoff as fine and the delay as internal, which is the right start. Stop looking at the recommended POP list. The console's map is often wrong for specific services.

Run a continuous mtr to your actual target FQDN, like `api.salesforce.com`, not just the POP. Check the hostname pattern on the high-latency hop. If you see something like `a23.syd2.steering.netskope.net`, your Singapore traffic is being hauled to Sydney for inspection.

Then immediately pull the data usage report filtered for your Singapore source IPs and that exact FQDN. It'll show the billing region the traffic is being accounted in, which dictates the inspection cluster location. Mismatch is common.



   
ReplyQuote
(@gracel)
Reputable Member
Joined: 3 months ago
Posts: 227
 

> even if you've updated it, a cached configuration on the old steering nodes can persist for a few hours.

That's so good to know, thanks for sharing! I've been pulling my hair out waiting for config changes to take effect. I never thought to do a full policy republish to speed it up, that's a great tip.

Is there a best time to run that republish, or does it not really matter? I'm worried about kicking my team off for a minute.



   
ReplyQuote
Page 3 / 4