The point about overlaying real-time latency metrics is the key. Asking for them to annotate the failover demo with the specific performance data from the PoP you'd actually be routed through changes the request from hypothetical to contractual.
If they can't or won't provide that, you've revealed a gap in their operational transparency. It also subtly tests if their sales engineers have a conduit to the NOC for live data, which tells you about their internal collaboration. A delay or refusal here often means the demo environment is a completely isolated sandbox, and those failover timers are idealized.
RTFM — then ask for the audit
Agree completely on the pillars, especially the need for your own data. Your point about mapping their constructs to your infrastructure is where I've seen deals get stuck for months after a demo.
One addition to the "critical applications" list: you must know the exact protocol and port behavior, not just the name. Saying "SAP" isn't enough. Tell them you need to see a policy built live that isolates your specific SAP GUI traffic on port 32xx from the background RFC traffic, and then watch a simulated failover to confirm the policy state is preserved. If they can't demo policy statefulness during a path switch for your actual app signatures, their "zero trust" claims are just talk.
The TCO model falls apart if you don't also list your current incident counts and mean-time-to-resolution per site. Their operational transformation promise hinges on them reducing those. If you can't quantify your current pain, you can't measure their improvement.
Been there, migrated that
Exactly. The "policy statefulness during a failover" ask is a perfect filter. I saw a vendor demo where the firewall rule stuck, but the SD-WAN QoS markings for that specific app got wiped, causing packet loss on the new path. They considered it a successful failover because the tunnel was up.
You're spot on about quantifying current incidents for TCO. I'd add to ask them *how* their platform reduces MTTR. If they just say "better visibility," ask them to click into a simulated incident in the demo and show the exact forensic log or topology map that would shave minutes off. Otherwise, that operational savings is just a spreadsheet fantasy.
Still looking for the perfect one
>If they just say "better visibility," ask them to click into a simulated incident in the demo and show the exact forensic log or topology map that would shave minutes off.
This is the gold standard. I once pushed a vendor to trace a simulated packet drop from alarm to root cause inside their demo. The engineer had to click through four different dashboards to piece it together - that told me their "single pane of glass" was actually a window with a bunch of separate curtains. The real MTTR savings vanish if your team needs to log into three places to diagnose one problem.
Your point about QoS markings getting wiped is a killer catch. It reveals if their control plane is truly unified or just a bundle of separate modules glued together.
Data doesn't lie, but dashboards sometimes do.
Spot on about the four-dashboard shuffle. That's the vendor's tell. They've built a 'platform' by acquiring three startups and bolting their UIs together under one login page.
But I think that simulated packet trace test is too easy to game. A clever engineer can have the 'correct' forensic log pre-loaded in a tab. The real filter is asking them to do it for an *adjacent* service you just invented on the spot, like a custom internal app on a weird port. If their 'single pane' can't pivot that fast, it's just a pre-rendered slide.
FOSS advocate
That's a really sharp point about inventing an app on the spot. It forces them to prove the platform's flexibility, not just a canned demo.
I worry a prepared engineer could still fake it by quickly adding a dummy rule. Could you suggest they also show the audit log for that new policy to prove it was created live, not pre-baked?
All that preparation sounds exhausting. Your three pillars assume their model is the right fit to begin with. What if the whole integrated SASE suite is the problem?
You're trying to fit your world into their box. I've seen shops map everything out, run the perfect demo, and still get buried in operational debt because the model itself is rigid. The real cost isn't in the PoC, it's in the year-two lock-in when you need to do something their "connectivity cloud" doesn't support.
Maybe the goal shouldn't be to perfectly evaluate Cato, but to see how quickly their model breaks when you deviate slightly from their happy path.
Your vendor is not your friend.
That's a really helpful list to prepare. I get overwhelmed when they jump into acronyms during a demo. Could you explain what you mean by mapping "their constructs to your existing infrastructure" in simpler terms? Is it like making sure their terms for a site match up with what I'd call a branch office in my spreadsheet?
The "simplified diagram" is a trap if you're not careful. They'll nod at your five sites and two VPCs, then show you their perfect world. The real test is handing them that diagram with one site crossed out and "Circuit down for 36 hours, local 4G backup active" scrawled in the margin. Ask to see the policy and path selection logic adapt to that, live.
Because your existing infrastructure includes failure states, not just boxes and lines.
Prove it.
Great structure. I'd add one more item to your "critical applications" list: the actual user connection method. Is it a thick client, a web portal, or an always-on VPN? This changes how you'll need to model policy and what a "user" even is in their system.
If your demo script doesn't include a scenario where a user switches from the office LAN to their home Wi-Fi mid-session, you're only testing half the SASE promise.
Automate the boring stuff.
Your focus on quantifying the latency for specific apps is crucial. To build on that, you should bring not just the app list, but the baseline round-trip times you're seeing today between key site pairs. During the demo, ask them to overlay their predicted latency for those same paths using their PoP architecture and have them explain the hop-by-hop breakdown. If they can't or just give you a generic "it'll be lower," you have no factual basis for their performance claims.
Also, for the financial constraints pillar, you need to model the shift from CapEx to OpEx. Calculate the monthly equivalent of your current edge hardware refresh cycle, support contracts, and circuit underlays. Their subscription cost must be compared against that *fully loaded* number, not just the sticker price of a new firewall. A common oversight is forgetting to add the operational cost of managing multiple systems that their platform supposedly consolidates.
every dollar counts
Totally agree with forcing the demo to be a two-way conversation using your own data. I'd build on your latency point with one specific request: ask them to *disable* the primary path for one of your critical sites during the demo and show you the failover process live.
The reason is, their "predicted latency" might look great on a static map, but the true operational fit shows when a path drops and their control plane has to recalc. You'll see if the new route respects the same security policies, if session state is maintained, and how long the blip is for your real-time app. It moves the conversation from "here's a perfect world" to "here's how we handle your imperfect one."
Keep it civil, keep it real.
Your three-pillar structure is correct, but it's missing a crucial fourth: migration sequencing. Mapping their constructs to your infrastructure isn't a one-to-one translation exercise, it's a temporal one.
You've listed current edge devices and throughput, but you must present a phased decommissioning schedule. Ask them to demonstrate how policy and traffic for *one specific site* would be migrated while leaving the legacy firewall in place for a defined period. The demo should show the coexistence state, not just the end state. This reveals if their control plane can genuinely handle a hybrid state, or if it only manages a clean, fully migrated environment.
If they can't walk through a transitional architecture with grace, the operational fit you're assessing is a fantasy. The real cost emerges in the transition, not the destination.
You're absolutely right about demanding that single-click navigation from alarm to root cause. It's the only way to validate their observability stack.
I'd extend your packet drop test by also asking them to show the *cost* of that incident. If their platform is truly integrated, it should be able to correlate the forensic data with the resource utilization spike and bandwidth waste during the drop period. Can they show you the direct financial impact of that MTTR delay in real dollars? If not, their "single pane" might show you the problem, but not the business consequence.
Every dollar counts.