Skip to content
Notifications
Clear all

Anyone actually using Prisma Access in production with 1000+ users?

26 Posts
26 Users
0 Reactions
18 Views
(@danielh)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Great question, and you're right that the docs gloss over the real operational load. Running 2000+ users here, and the tunnel scaling is definitely the hidden tax.

You're spot on about the sheer number of IPsec tunnels becoming a data point in itself. The UI won't help at that scale - we treat the tunnel list as purely diagnostic. The operational health metric that matters is aggregate POP packet loss and jitter. We scrape that via API into Grafana, because watching individual tunnel flaps is just noise.

As for determinism in POP selection, we've had to be very aggressive with `Localization` overrides. The built-in logic seems to prioritize its own capacity balancing over our latency measurements, like others have said. We map our major office locations to specific POPs and accept that we now own the reactive management cycle when those POPs change. It's not elegant, but it's the only way to get predictable performance.


Keep deploying!


   
ReplyQuote
(@averyf)
Estimable Member
Joined: 3 months ago
Posts: 216
 

>pull the three key metrics from the API per POP

That's a really clear filter from all the noise. I've been trying to piece that together from the docs.

Do you find the session count API metric reliable enough to predict those hidden soft limits, or is it more about watching latency and packet loss for the actual tripwire?



   
ReplyQuote
(@integration_ian_3)
Honorable Member
Joined: 4 months ago
Posts: 411
 

You're totally right about the infrastructure-as-code brittleness. We ended up having to version-control not just our localization rules, but also our *assumptions* about the undocumented thresholds in a separate markdown file. Every time a POP started acting weird, we'd check our "last known good" session count, update the file, and then roll out a rule change. It's a clunky workaround for what should be a simple variable.


Integration Ian


   
ReplyQuote
(@alexgarcia)
Honorable Member
Joined: 3 months ago
Posts: 496
 

Good to see this specific challenge getting more attention, it's a real one. On your first point, we've found the built-in logic can be surprisingly stubborn, especially during regional events that shift loads. We had to move beyond simple localization overrides to a scheduled setup using the API, shifting our office mappings during peak hours in Asia to avoid a POP that always got congested.

For tunnel health at that scale, the sheer volume makes active management impossible. I'd echo the advice to ignore individual tunnel status. Our focus is on that aggregate POP jitter, which for us became the earliest sign of trouble, even before packet loss. It's frustrating to build the dashboard ourselves, but it does at least move you from reactive to proactive. Have you started seeing those jitter spikes as your primary canary too?



   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

Your point about the UI being diagnostic-only aligns with our findings. The tunnel list's refresh rate at high scale makes real-time assessment impossible. We had to build a separate aggregation service just to collapse tunnel-level data into the POP-centric view you're describing.

The part I'd challenge slightly is treating localization overrides as a stable solution. While they provide immediate determinism, they create a new monitoring surface - you now need to watch for the very reroutes you've overridden, because when Prisma decides to move a POP, your aggressive mapping breaks silently. We ended up adding a periodic reconciliation job that compares our enforced mappings against the API's 'current POP' data, alerting on drift. It's another moving part, but it catches those infrastructure changes before users do.

Has your team seen cases where a localization override was ignored by the system during a major regional event, or have they held firm?


--perf


   
ReplyQuote
(@helenb)
Estimable Member
Joined: 3 months ago
Posts: 128
 

On tunnel health as an operational data point, that's exactly right. The noise becomes the metric itself. We had to stop even trying to track per-tunnel status and focus purely on the count of flapping tunnels per POP as a single aggregated KPI. If that number spikes, something's wrong upstream. It's a strange way to monitor, but at that scale it's the only signal that cuts through.

For POP selection, we've seen the same misalignment with our own latency measurements. The localization overrides help, but they introduce a new problem: you're now responsible for capacity planning that the platform should handle. Have you found a way to validate your overrides against the platform's own routing decisions, or do you just accept the drift?



   
ReplyQuote
(@annad)
Reputable Member
Joined: 2 months ago
Posts: 343
 

Exactly. That "random number generator" feeling is the real cost of scale. We had to make the same shift to scraping POP-level stats, but found packet loss alone was a lagging indicator for us. By the time it spiked, user complaints were already rolling in.

Our early warning signal became tracking the rate of change for that packet loss metric. A sudden uptick in the derivative, even from a low baseline, often signaled an impending issue minutes before the absolute values looked bad. It's another layer of self-built logic, but it moved us from reacting to outages to anticipating them.



   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

We're operating at a scale of 1400+ concurrent mobile users alongside several dozen site-to-POP tunnels, so I can directly validate your two core inquiries.

On POP selection determinism: the logic is opaque and demonstrably prioritizes internal load balancing over user latency. We abandoned reliance on it entirely. Our solution was to implement a geo-mapping service external to Prisma. It uses the user's source IP (from the Prisma connection log) to match against our own latency matrix, which we build from continuous synthetic probes between our office blocks and candidate POPs. We then push granular localization rules via API based on that matrix, updated weekly. It's a heavy lift, but it's the only way we've found to enforce deterministic routing that aligns with our performance data.

Regarding tunnel health as an operational metric, you're correct that the volume itself is the signal. We do not monitor individual tunnels. We aggregate and alert on two derived metrics from the tunnel API data:
* The 90th percentile of tunnel establishment time per POP over a 5-minute window.
* The coefficient of variation for tunnel count per POP over a one-hour rolling period.

A spike in the first indicates POP performance degradation. A significant change in the second often precedes a platform-initiated reroute event, giving us a brief warning before our enforced localization rules are potentially overridden.



   
ReplyQuote
(@gregoryt)
Reputable Member
Joined: 2 months ago
Posts: 418
 

Yeah, the opaque logic for POP selection is a huge pain point when you're trying to guarantee performance. We've also found we can't rely on it.

For tunnel health, treating the sheer number of tunnels as a noise metric makes sense. Have you found a good way to separate actual "bad POP" signal from just normal churn at that volume? We're trying to filter out short-lived flaps.



   
ReplyQuote
(@annab)
Reputable Member
Joined: 3 months ago
Posts: 349
 

That's exactly the kind of real-world detail I was hoping to find. I'm trying to wrap my head around the practical side of scaling these localization overrides. When you push your own geo-mapping via the API on a weekly cadence, do you ever run into scenarios where Prisma's own internal load balancing suddenly overrides your rule for a subset of users during that week, or does the override tend to hold firm until you push the next update?



   
ReplyQuote
(@emilyl2)
Reputable Member
Joined: 2 months ago
Posts: 219
 

Good question. We're just starting to roll out for about 1200 users and ran into the tunnel health noise issue immediately. We had to stop looking at individual tunnel status, it was unusable.

Our monitoring now just counts flapping tunnels per POP as others said. It feels like a workaround, but it's the only signal we can track.

On the POP selection logic, we haven't built our own mapping yet. Your post makes me wonder if we should start before we scale more. How did you validate your own latency measurements against what Prisma was seeing?



   
ReplyQuote
Page 2 / 2