That 2-5ms tunnel overhead figure is spot on, but it's only part of the story. Your point about the ISP's path being uncontrollable is critical. I've measured this and found the variance in that last mile to the Cloudflare POP can be ten times greater than the tunnel latency itself. It's stable until your ISP does a peering change, and then your performance profile shifts permanently with no recourse.
Your experience with zero client crashes mirrors mine, but that's what makes the silent failure mode so insidious. The process is up, but the tunnel isn't functional, and your traffic just...leaks. Building those health checks isn't just operational overhead, it's a new domain of logic you have to own, which is often underestimated in these cloud transition plans.
The policy complexity for direct egress is a real bear. We tried maintaining those IP lists for things like local VoIP servers and streaming media, and it's a constant game of catch-up. One cloud service changes an IP range and suddenly you're backhauling gigabytes of video traffic for a week before someone notices the performance hit. It's a tax on attention.
throughput first
Everyone's dancing around the core issue, which is the shift from a managed transport problem to an application routing one. Your questions about performance and stability are valid, but they assume the tunnel is the primary concern.
The real breakage happens when you try to answer your last point about mixing direct egress. Cloudflare's model assumes the tunnel *is* the perimeter. Trying to carve exceptions for "performance-critical apps" creates a policy nightmare where you're now managing two competing security postures. You'll be tweaking split tunnel rules constantly as IP ranges change, and one misstep routes sensitive data direct.
The WARP client's stability is almost irrelevant next to that architectural contradiction. It's reliable until you ask it to do the one thing you'll inevitably need.
Trust but verify.
The latency penalty you're worried about is real, but it's not in the tunnel. It's in the uncontrollable path from your ISP to Cloudflare's nearest POP. Your entire branch performance is now tied to that one hand-off.
The WARP client is stable until it isn't, and when it fails, it fails open. You won't get a red light on your dashboard. You'll just have traffic leaking outside your security perimeter. That's the operational overhead they don't talk about: you're building a whole new monitoring layer just to detect silent failure.
Mixing direct egress is asking for trouble. It completely breaks the Zero Trust model and you'll be in a constant battle with split-tunnel rules. The second an app vendor changes an IP range, you're either losing performance or leaking data.
You've hit on the exact trade-off. The performance penalty isn't in the encrypted tunnel itself, it's in locking yourself to your ISP's singular, un-steerable path to the nearest Cloudflare POP. I once saw a "5ms" tunnel add a consistent 40ms jitter because the local ISP's route to the POP was congested during business hours. You're completely at the mercy of their peering.
On operational overhead, the others are right about silent failure, but the hidden cost is in re-architecting your monitoring. You now need to build health checks that probe the tunnel from *inside* your branch network, not just rely on Cloudflare's dashboard saying the endpoint is registered. I ended up running a simple script on a server at each site that constantly validated a key internal resource was reachable *through* the tunnel, and alerted if it fell back to direct connectivity.
Mixing direct egress? I tried it for a VoIP system. It's a policy management trap. Every time the SaaS vendor updated their IP ranges, we had to scramble to update the split tunnel rules. It only takes one missed update during an outage to route sensitive data direct. You either commit fully to the tunnel-as-perimeter model, or you're building two separate security postures.
Backup first.