Just finished a year-long Palo Alto rollout across our primary and DR sites, serving about 3000 users. Management is thrilled about the "next-gen" marketing slides and the projected "security posture." Naturally, I was tasked with the post-mortem on the HA failover test we ran last weekend.
Here's the thing everyone glosses over: the datasheet promises "sub-second" or "zero downtime" failover. Our reality? A 47-second outage for roughly 60% of users. Not exactly seamless. The vendor's response was, predictably, "that's within expected parameters for a stateful failover with your session count." Right.
So, what actually broke? The usual suspects didn't fail—the BGP peering came back, the tunnels re-established. The pain points were more subtle:
* **SSL Decryption sessions:** Anything undergoing full TLS inspection died and had to be renegotiated. The failover didn't preserve the decryption state, obviously. The resulting flood of new handshakes choked the newly active firewall for a good 30 seconds.
* **GlobalProtect "sticky" connections:** Users on VPN were told they'd be "uninterrupted." Instead, their tunnels held just long enough to become unresponsive, forcing a manual reconnect. The session state transferred, but the underlying IKE/IPsec tunnels didn't follow cleanly.
* **App-ID dependencies:** A few critical internal apps that rely on specific dynamic App-ID detection just... hung. The passive unit's app cache wasn't as warm as the marketing led us to believe, causing a delay in policy resolution post-failover.
I'm looking for concrete data from anyone who's done a large-scale HA flip under load. Not the vendor's lab scenarios, but real production with thousands of concurrent sessions. How much of the "state" is truly preserved? What metrics did you track in your billing/usage data that showed the actual impact? I have a sneaking suspicion our "resilient" architecture just moved the single point of failure from the hardware to the session synchronization layer.
Show me your logs and graphs, not your slideware.
cost_observer_42