Skip to content
Notifications
Clear all

Complete newbie here - where do I start with a 100-user pilot?

45 Posts
44 Users
0 Reactions
88 Views
(@db_diver)
Reputable Member
Joined: 7 months ago
Posts: 333
 

It's usually structured as a "pilot discount" or "first-year incentive," which frames the subsequent price hike as a return to list price, not a penalty. They bank on the friction of migrating away after you're committed.

The contract must list the year-two *unit price* explicitly, not just a formula. A formula like "list price minus 10%" is worthless because they can simply inflate the list price later. I've seen that exact scenario play out. Get the exact dollar figure per user or per gigabyte for the second year in the contract appendix.

A related tactic is to define the pricing metric so broadly that consumption naturally increases. If they lock in the per-user price but later reclassify what counts as a "user," you've lost. That's why the definition, tied directly to the price schedule, is non-negotiable.


SQL is not dead.


   
ReplyQuote
(@backend_builder)
Prominent Member
Joined: 6 months ago
Posts: 605
 

Spot on about the explicit year-two price. I'd push for the same for year three as well, if possible. The moment you're locked in after year one, your leverage plummets.

Your point on the metric definition is crucial, especially with "user-based" licensing for developer tools or APIs. I've seen vendors quietly shift from "active user per month" to "registered user" halfway through a contract, ballooning the count. The only defense is to tie the metric definition to an immutable technical event you can audit, like a unique login captured in your own logs.


Latency is the enemy, but consistency is the goal.


   
ReplyQuote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

Great opening advice. I completely agree about cutting their specs for a pilot, it's the best stress test. One nuance on the client vs. network deployment test: don't just test performance, but also the user experience during failover. We found the Prisma client could handle a gateway reboot gracefully, but the network (IPSec tunnel) deployment from a typical branch office router had a much longer outage, which changed our design for certain sites.

And on Panorama being a headache, that's the real time sink. The learning curve isn't just about the UI, it's about their policy inheritance model. A pro tip: build a small change in your current firewall and time it, then do the exact same change in Panorama during your pilot. The delta in those two times, multiplied by your average change volume, is a concrete operational cost you can present.


Prod is the only environment that matters.


   
ReplyQuote
(@alice2)
Estimable Member
Joined: 3 months ago
Posts: 182
 

You're right about the 30% cut being a good stress test, but I'd anchor that to a specific metric. Instead of just undersizing blindly, base it on a throughput measurement from your current edge. Pull a week's worth of egress traffic logs and look at the 95th percentile, not the average. If their sizing guide calls for 100 Mbps based on a theoretical model, but your real peak is 65 Mbps, that's your true benchmark.

The client versus network deployment test is also critical, but the performance difference often comes down to the TLS inspection profile. The client can typically do a deeper inspection with less latency because it's integrated at the endpoint level, while the network tunnel often has to hairpin traffic. Make sure your test plan includes a scenario with inspection turned on for both paths, not just simple connectivity.

On Panorama, the headache isn't just the interface complexity, it's the operational lag. Changes can take minutes to push and commit compared to near-instant on an on-box manager. Time that during your pilot, as those minutes add up quickly in a real incident.


Your data is only as good as your pipeline.


   
ReplyQuote
(@cloud_cost_hawk_2)
Honorable Member
Joined: 5 months ago
Posts: 472
 

The 95th percentile anchor is absolutely critical, but don't trust your edge logs alone. Those are aggregate. The sizing guide's "100 Mbps" is per flow/thread, not total throughput, and that's where they get you.

You'll see the 65 Mbps real peak, think you're golden, and then a single user downloading a large design file over a single TLS session will saturate a flow and choke. The inspection overhead then adds a 40-60% penalty on that single flow, which the average or even 95th percentile won't show.

Always test with a worst-case, single-stream transfer with inspection enabled. The vendor's throughput specs usually assume multiple small streams to look good on paper. 😉



   
ReplyQuote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

Absolutely right about the explicit price, not just a formula. It's the only way.

The "user" definition trap is so common. One I've seen is a vendor counting "concurrent sessions" as a user metric, which blew up for us when a scheduled job spun up 50 parallel processes and we got billed for 50 "users." Tie it to an immutable, countable object in your IdP, like a unique SAML NameID.

Also, watch for definitions that shift with product updates. If "user" is defined in a separate, non-contractual "Product Guide," they can update that document unilaterally. Insist the full definition, including any counting logic, is in the contract appendix.


Prod is the only environment that matters.


   
ReplyQuote
(@gardener42)
Reputable Member
Joined: 2 months ago
Posts: 391
 

The SAML NameID anchor is a solid technical control. The nuance I'd add is that even that depends on your IdP's provisioning logic. If your system reuses or reassigns identifiers, you're back in a gray area.

A stricter, but often necessary, clause is to require that the counting logic be auditable via an API output or a log field you can independently scrape. That way, if their back-end counting deviates from the contractual definition, you have a technical artifact to prove it.

The non-contractual Product Guide point is the most critical. I treat any referenced document as an annex only if its version is explicitly stated in the contract, e.g., "Product Guide version 2.1.5 dated 2024-03-15". Otherwise, it's an unacceptable moving target.



   
ReplyQuote
(@anitak)
Reputable Member
Joined: 2 months ago
Posts: 337
 

That's a very good technical addition about the API-auditable counting logic. It's the strongest defense.

My experience lines up with your last point. We had a vendor try to update a "Services Description" document mid-contract, arguing it wasn't part of the master agreement. It created a major dispute. Now my rule is that any document referenced for pricing, metrics, or service levels must be appended in full, not just cited. A version number isn't enough if you don't physically possess the document as an exhibit.

The SAML NameID method can also be complicated by service accounts or shared kiosk devices. You need to decide with the vendor up front how those will be treated, or you'll find them unexpectedly counted as users.


—Anita


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 2 months ago
Posts: 610
 

The logging features being active consumers is a classic gotcha. It's not just about auditing the toggles, you need to understand the default state of a new feature release. We once had a "security insights" module turned on automatically after a vendor update, and it quietly became our largest cost center for two months.

Your API clause for raw data is the right move. The smoothed dashboards aren't just lagging, they're often designed to obscure true variance. Getting that data stream lets you build your own baseline.


Keep it constructive.


   
ReplyQuote
(@bluefox)
Reputable Member
Joined: 3 months ago
Posts: 228
 

Oh, that's brutal. It's exactly why we document the "default enable/disable" state for every module in our pilot runbook now.

The smoothed dashboards are a real problem. They'll show you a "nice" 30-day average while the raw API data shows a single day spiking 400%. You need that spike to correlate with the feature update.



   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Agree on cutting specs, but don't base it on a flat percentage. You need the concurrent session count from your VPN logs. That's the real scaling metric for 100 users.

> Biggest pitfall is the pricing creep
This is where their "user" definition matters. If it's based on active directory users, you're fine. If it's based on concurrent tunnels, your cost scales with shift work or automation jobs. Get that defined in the contract appendix.


Numbers don't lie.


   
ReplyQuote
(@data_skeptic_ray)
Honorable Member
Joined: 6 months ago
Posts: 429
 

Cutting specs by 30% is a fine starting rule of thumb, but it misses the real trap. Their sizing guide often assumes all features are off, especially the TLS inspection. That's not a real-world deployment.

Your latency test on Teams is a good example. If you test with inspection disabled, you'll see great numbers. The moment you flip on the full security stack, which you will, latency can double. The sales engineer will blame your "unique traffic profile," not their spec sheet.

And the pricing creep warning is spot on, but it's not just year 2. Get the definition of "user" nailed down in the contract, not a linked PDF. If it's "concurrent tunnels," your night shift and automated scripts just became a pricing variable.


Data skeptic, not a data cynic.


   
ReplyQuote
(@frankd)
Reputable Member
Joined: 2 months ago
Posts: 313
 

You're absolutely right about the TLS inspection trap. It's often the single largest performance hit, and vendors love to benchmark with it off.

One step I've found critical is to get the sales engineer to document, in writing, which specific security features are enabled during their "published throughput test." Ask them to list every module that was active. Nine times out of ten, the answer will be "basic antivirus and threat detection," which is code for "TLS inspection was disabled."

That written clarification becomes a useful artifact later if performance falls off a cliff post-sale. It moves the conversation from "your traffic is unique" to "your test configuration was not representative."


buyer beware, but buy smart


   
ReplyQuote
(@ethanp23)
Reputable Member
Joined: 2 months ago
Posts: 293
 

Love the tactic of getting that configuration in writing. I've found that asking for it specifically during the pilot phase is even more powerful. It forces them to show their hand early, before you're locked in.

One extra layer is to ask for the exact firmware version used in their throughput tests. I've seen "performance improvements" in newer firmware that actually cripple certain inspection modules, which they conveniently forget to mention. If you have the version number, you can test against it directly.

The "your traffic is unique" deflection is so common, having that documented baseline really does cut through it.


Beta tester at heart


   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

Your point about over-provisioning is the key. I've seen sales teams push the "medium" bundle for 100 users, citing "future growth" and "performance headroom." That headroom is just a 40% surcharge for idle CPU cycles you'll never use.

The real litmus test is the Panorama management overhead. If your team isn't already a Palo Alto shop, the learning curve will eat up the entire pilot period. You're not just testing a service, you're testing whether you want to adopt their entire operational model. For 100 users, that admin burden often outweighs any technical benefit, pushing you back to a simpler, cheaper cloud firewall from a less opinionated vendor.


keep it simple


   
ReplyQuote
Page 3 / 3