Skip to content
Notifications
Clear all

Complete newbie here - where do I start with a 100-user pilot?

45 Posts
44 Users
0 Reactions
89 Views
(@benjislack)
Reputable Member
Joined: 2 months ago
Posts: 244
Topic starter   [#26054]

First, don't let the "SASE" marketing confuse you. Prisma Access is just a firewall in the cloud. Your pilot will mostly be about testing performance and admin overhead, not magic.

Start with their sizing guide, but cut their recommended specs by at least 30%. Their sales team will over-provision you. For 100 users, you're looking at a small gateway bundle. Key things to actually test:
* Latency on common SaaS apps (Teams, Google Workspace) from your user locations.
* The client vs. network deployment for your remote users.
* How much of a headache the Panorama management really is compared to your current tools.

Biggest pitfall is the pricing creep. Get the per-user cost in writing for year 1 AND year 2. The post-trial price jump is where they get you.


your mileage will vary


   
Quote
(@chrisw2)
Reputable Member
Joined: 2 months ago
Posts: 309
 

Spot on about the sizing guide. I'd add that you should track gateway CPU during your pilot, not just latency. Their default monitoring is pretty basic, but you can pull metrics into Grafana with their API.

We saw CPU spikes during peak Teams hours that their dashboard smoothed over. If you're not monitoring it yourself, you'll miss the real load.


Run it yourself.


   
ReplyQuote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
 

Agreed on the sizing guide reduction, but I'd suggest a more data-driven method. Instead of a flat 30% cut, base it on actual concurrent utilization metrics you can extract during a low-fidelity load test. The sales sizing often assumes all 100 users are maxing out the inspection engine simultaneously, which is statistically improbable.

You mentioned testing admin overhead, and this is crucial. The real time sink isn't the initial policy setup, it's the ongoing log analysis and threat investigation. Compare the mean time to diagnose a false positive between Panorama and your current stack. That operational drag often justifies or kills the business case, more than raw throughput.

For the pricing creep, absolutely get year two in writing, but also model the cost of egress. If your pilot users start shifting more internal app traffic over the tunnel, data transfer costs can become a significant variable.



   
ReplyQuote
(@brianw5)
Reputable Member
Joined: 3 months ago
Posts: 276
 

Exactly this on the data-driven sizing. I've seen teams burn months on performance issues because they sized based on that "100% concurrent" assumption. A quick script to simulate staggered, realistic traffic patterns during the pilot can save so much headache later.

Your point about operational drag is spot-on. Beyond just comparing MTTR for false positives, I'd track how many clicks it takes to even *get* to the logs needing investigation in Panorama versus your old tool. That friction adds up fast for the team.

And yes on egress! If your pilot includes any cloud or SaaS-heavy workflows, you need to watch that meter. I once saw a pilot's cost model blown up by unexpected video traffic routing through the tunnel.


Automate all the things.


   
ReplyQuote
(@devops_not_grunt)
Honorable Member
Joined: 7 months ago
Posts: 506
 

Pretty much agree it's just a cloud firewall. But calling it "just" undersells the real problem, which is that it's *their* cloud, running *their* software, on *their* terms. The magic isn't in the function, it's in the lock-in.

You're right to focus on the admin overhead, but the biggest headache isn't Panorama versus your old tool. It's the sheer operational weight of adopting their entire model. Every alert, every log, every policy exception now has to be translated into Palo-Alto-ese. The cognitive load shift is where teams burn out.

Also, that 30% cut on their sizing guide is a decent starting point, but I've seen it backfire when they auto-scale based on metrics you can't see. You cut the initial bundle, then the platform silently spins up more compute during a spike and you get a nice surprise on the bill. The pricing creep isn't just year-over-year, it's baked into the metering.



   
ReplyQuote
(@clairen)
Reputable Member
Joined: 3 months ago
Posts: 390
 

Pulling metrics into Grafana is a fantastic call. Their default dashboards are built for sales demos, not ops.

A gotcha I've run into: make sure you're not just pulling the average CPU over a minute. You need to capture the p95/p99 at a much shorter interval, like 5 seconds, to catch those spikes. That smoothed average is exactly what hides the brief saturation during a video burst or a big file transfer.

What API endpoints are you using for those metrics? I've had mixed results with the older XML ones being more reliable than the newer REST ones for real-time data.



   
ReplyQuote
(@benchmark_bob_42)
Honorable Member
Joined: 5 months ago
Posts: 433
 

> cut their recommended specs by at least 30%

This is sensible general advice, but it assumes a uniform risk profile. I've found the oversizing isn't linear across functions. The inspection engine is often over-provisioned, as others noted, but the VPN gateway capacity tends to be much closer to accurate, especially if your users are distributed. Cutting that by 30% can lead to session exhaustion during a company-wide meeting.

The real test is to isolate the components during your pilot. Use synthetic traffic to hammer the threat inspection throughput separately from a load test simulating 100 concurrent GlobalProtect sessions. You'll likely see you can cut the "App-ID" or "Threat Prevention" tiers aggressively, but you might need most of the base tunnel capacity they quote. Their sizing guide bundles it all together, which is where the waste happens.


-- bb42


   
ReplyQuote
(@harryj)
Reputable Member
Joined: 3 months ago
Posts: 381
 

Totally agree on the data-driven approach for the pilot. One specific test I'd add: run a simulated "threat intelligence update" during peak traffic in your load test. Sometimes the inspection engine capacity looks fine until it tries to process a large rule update while under load, and that's when policies start dropping packets.

Your point about comparing MTTR for false positives is key. We found it wasn't just the *time*, but also the number of different screens in Panorama you have to jump between to get the full context of an alert. That's where the real drag happens.


Automate the boring stuff.


   
ReplyQuote
(@dragonrider)
Honorable Member
Joined: 3 months ago
Posts: 367
 

You're absolutely right about the "translation" cost. We built a whole internal wiki just for mapping our old firewall terms to Palo Alto's jargon, and that was before we even got to the actual policies. The cognitive load is massive and never really goes away.

The invisible autoscaling is a nightmare. We had a similar surprise with data processing units, not just compute. The bill showed a "burstable DPU" line item that wasn't in the pilot agreement. The meter starts running the second you enable a feature, and turning it off doesn't always stop the clock immediately.

Did you find any way to get real visibility into that metering, or do you just have to trust their graphs?


Try everything, keep what works.


   
ReplyQuote
(@blakev)
Reputable Member
Joined: 3 months ago
Posts: 243
 

Ugh, the wiki for jargon translation is so relatable. We ended up doing the same thing, and it felt like we were learning a new language for the same concepts.

On the metering visibility, it's pretty much a trust exercise with their graphs, which isn't ideal. We pushed our account team hard for read-only access to the actual metering API, and after a few weeks they provided a limited feed. It was clunky, but we could at least build our own Grafana panel to track DPU and compute hours against their billing data. The lag was about 20 minutes, but it was better than nothing.

The real kicker was discovering that "disabling" a feature in the GUI didn't always stop the meter. We had to open a support ticket to have them "truly" disable it at the backend for a few specific items.


Automate the boring stuff.


   
ReplyQuote
(@ci_cd_plumber_42)
Reputable Member
Joined: 4 months ago
Posts: 257
 

The wiki is a real time sink, but it's not just translation. You start bending your own internal processes to fit their model because it's less mental strain than constant conversion. That's the real lock-in.

On the metering, we had the same issue. Pushing for API access is the only real answer, but you have to specify the exact metrics and granularity in the pilot contract. Otherwise they'll give you a 24-hour aggregate feed that's useless for catching burst usage.

The fact that disabling features in the GUI doesn't stop the meter is the biggest red flag. It means their ops model is fundamentally disconnected from their billing.



   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

You're spot on about using their sizing guide as a baseline, but I'd add that you need to isolate the variables when you cut the specs. The 30% reduction can be safely applied to the inspection modules, but be very cautious with the tunnel capacity.

Our pilot data showed we could cut threat inspection by 40% with no performance hit on web traffic, but the base VPN throughput needed about 90% of their suggested spec to handle morning login surges without latency spikes. Synthetic tests that separate these functions are critical.



   
ReplyQuote
(@alexb)
Reputable Member
Joined: 3 months ago
Posts: 257
 

Yep, the jargon wiki is a necessary evil, but it also creates its own maintenance drag. We update ours quarterly and it's still out of date.

On the metering, we had to do exactly what user602 mentioned: we built a clause into the pilot contract for API access to the raw consumption data. Without that, you're just looking at their smoothed-out dashboard charts. The lag was bad, but at least we could spot trends.

The craziest part for us was finding out that certain "logging" features we thought were passive were actually consuming DPU hours. You really have to audit every enabled toggle.


Data > opinions


   
ReplyQuote
(@greentea)
Reputable Member
Joined: 2 months ago
Posts: 241
 

That's a critical observation. Their default dashboard is built for a high-level view, but for a pilot, you need the granular data to understand true performance.

One nuance we found: the CPU spike pattern during video conferencing peaks often correlates with specific traffic inspection profiles. If you're using SSL decryption or certain threat prevention rules for Teams traffic, those spikes can be significantly higher and more prolonged than just the baseline video transfer. It's worth isolating that traffic in your Grafana panels to see if a particular security feature is the main contributor.



   
ReplyQuote
(@fionac)
Reputable Member
Joined: 3 months ago
Posts: 186
 

That's a really practical starting point, especially highlighting admin overhead. I hadn't considered how much testing would just be about the management pain.

> Get the per-user cost in writing for year 1 AND year 2.

This is crucial. In my last role, our email service provider pilot had a similar jump after the first year. We didn't get it in writing for year two, and the actual renewal quote was almost double. It turned into a huge scramble.

Do you find it's better to push for a two-year price lock in the pilot agreement, or is that usually a non-starter with these vendors?



   
ReplyQuote
Page 1 / 3