Let's be clear: Zscaler is sold as a silver bullet for the "internet is the new corporate network" mantra. The demos are slick, the zero-trust buzzwords are perfectly aligned, and the promise of decommissioning your MPLS and on-prem firewalls is intoxicating. Having been through a multi-year deployment and now living with the operational reality, I'm here to tell you the sales deck conveniently glosses over the profound architectural and financial commitments you're signing up for.
The most glaring omission is the assumption that your entire workforce has flawless, high-bandwidth, low-latency internet at all times. Zscaler's model makes the public internet your corporate backbone. When a user's home internet has a hiccup, their connection to the office is now broken. We're not talking about a slow VPN; we're talking about a complete loss of access to internal applications because the entire security stack is now a cloud service. This shifts the blame from your network team to your user's ISP, but the business impact is identical. You are now in the business of troubleshooting residential internet connections and explaining why the corporate security architecture depends on the quality of a user's home Wi-Fi.
Then there's the true total cost of ownership. Yes, they'll show you the capex savings from retiring hardware. They will not adequately factor in the operational complexity and hidden costs. Your help desk needs retraining for a completely new set of issues. Application owners must now be deeply involved in defining micro-segmentation policies, a political and technical nightmare for legacy apps. Every internal-to-internal traffic flow that was once free on your LAN now hairpins out to the nearest Zscaler node and back, adding latency and consuming bandwidth. You will need to purchase additional, expensive premium support to get timely help with policy anomalies. The licensing model is a rabbit hole of add-ons for basic visibility or data loss prevention features that were included in your old stack.
On the subject of lock-in, it is absolute. Once you architect your entire network around their private access and internet gateway, migrating away is a multi-year re-platforming project. Your policies, your user identities, your application definitions are all living in their proprietary ecosystem. The tooling to export and translate these configurations to another vendor simply doesn't exist in any meaningful way. You are betting the company on Zscaler's long-term pricing power and architectural direction. If you think your negotiation leverage is good at renewal time after a full deployment, I have some bad news for you.
Finally, consider the security model itself. You are consolidating all your eggs into one very attractive basket. A misconfiguration in their admin portal, a compromised tenant, or an outage in their cloud region can have global impact on your organization. The principle of defense in depth is replaced with a single, massive choke point. While they have a good track record, the architectural risk is fundamentally different from a layered, hybrid approach. You are placing an immense amount of trust in a single vendor's infrastructure and operational security, with very little practical recourse if something goes wrong.
Just my two cents
Skeptic by default
You're hitting on the single biggest headache we found. The internet-as-backbone assumption is brutal. It forced us to create a whole new internal "ISP liaison" playbook and provide stipends for dual-homing home internet. The cost savings from ditching MPLS got quietly eaten by that.
And it's not just residential. Mobile users hopping between cell towers, or hotels with those awful captive portals, become massive support tickets. The Zscaler client can be fragile in those handoff moments. You end up needing a full VPN *behind* Zscaler just for troubleshooting access, which feels like admitting defeat.
Their model also assumes every app you own can tolerate the added latency of an extra hop to their nearest POP. We had to move several latency-sensitive internal apps to public clouds just to make them usable, which was another project nobody budgeted for. The sales pitch always shows the POP right next to the user... not so much when your user is in Omaha and the traffic routes to Chicago.
Automate all the things.
The mobile handoff issue is particularly under-documented. We instrumented our client logs and found the failover time during tower switches averages 8-12 seconds, which breaks any persistent TCP session. You mentioned the troubleshooting VPN, a pattern we call the "Zscaler shadow infra." We now maintain dedicated, non-Zscaled jump hosts in each region just for break-glass access, adding operational overhead that's never in the TCO model.
The POP latency assumption also ignores the internet's middle-mile unpredictability. Even with a POP in the same city, traffic can take a suboptimal BGP path before entering their private backbone. We had to implement synthetic monitoring with ThousandEyes to prove to Zscaler support that the 45ms added latency wasn't just "internet weather" but a consistent routing issue they eventually had to fix with their upstream providers.
Moving latency-sensitive apps to public cloud is a common workaround, but that introduces another cost layer. Did you find your cloud egress costs increased because now all app traffic from those clouds also routes through Zscaler IPSec tunnels to reach internal users?
Data over dogma
Your point about the troubleshooting VPN is spot on, that "admitting defeat" feeling is so real. We set up the exact same workaround, a VPN tunnel into a tiny non-Zscaled enclave for network teams.
It gets worse though, because now you have two completely separate security postures to manage. That shadow infra needs its own set of firewall rules, MFA, and logging, which adds a sneaky compliance risk. Our auditors flagged it as an uncontrolled bypass path, and we spent months documenting the "exceptional use only" procedures.
The pop latency issue you mentioned also bit us with SaaS apps, not just internal ones. Some vendors IP-restrict by region, and if your Zscaler pop is in a different state, their security systems flag it as a login from a new location. We had a wave of account lockouts that took forever to diagnose.
ship it
This is the exact conversation I have to have now when someone can't get into our CRM. "Can you try a speed test from your personal laptop? Okay, now do it from your work machine with the client on..." Suddenly my marketing team is diagnosing home network issues.
The shift in responsibility is real. I can't just tell someone the VPN is up, I have to explain their entire route to a cloud service is broken. It makes our support feel helpless.
You nailed the core shift. When the VPN is down, it's a single, known internal circuit to fix. When Zscaler is the issue, you're suddenly tracing hops across the public internet and coordinating with an ISP you have no contract with.
We've had to build an entire new troubleshooting playbook that starts with, "Turn off your Zscaler client and try the direct connection test." It completely reframes what "corporate network uptime" means. The internal SLA we could guarantee with MPLS is gone, replaced by a best-effort internet path.
And the cost doesn't just shift, it multiplies. You're paying Zscaler's subscription, plus you're now funding home office internet stipends, and you're still managing that "shadow infra" VPN for break-glass. The TCO math gets messy fast.
Cloud cost nerd. No, I don't use Reserved Instances.
That "internet as backbone" assumption isn't just an engineering challenge, it's a direct financial transfer. You're not eliminating a cost center, you're just moving it from a corporate P&L line to your employees' home utility bills.
We calculated the true cost after factoring in the mandatory ISP stipends. Our "MPLS savings" were a mirage. The network team's budget shrinks, but HR's budget for remote work benefits balloons by the exact same amount, plus the new Zscaler subscription. It's a shell game.
The sales TCO models never include that your new single point of failure is now a thousand different residential ISPs with no SLA.
Cloud costs are not destiny.
You've put your finger on the fundamental architectural pivot. It's not just shifting blame to the ISP, it's the complete loss of network observability and control. When an MPLS circuit degrades, I have a guaranteed path to a provider NOC with shared SLAs and traceable ingress/egress points. With Zscaler, a user's traffic traverses an opaque last mile before I even get visibility. My tools can only see the hop from the Zscaler POP onward, which makes diagnosing that critical first segment a guessing game reliant on user-run speed tests. The network team becomes a help desk for consumer ISP issues we cannot fix.
--perf
You're absolutely right about the blame shift, and I'd add that this fundamentally changes your support team's skillset. Suddenly, your Level 1 helpdesk needs basic network engineering knowledge to ask about bufferbloat on a user's home router or local ISP congestion. It's a massive, unplanned reskilling cost.
The "complete loss of access" point is the real killer, though. With a traditional VPN, a flaky connection might drop packet flow but the tunnel and its security context often persist or reconnect. When the Zscaler client loses that clean path to its POP, the entire security session is invalidated. The user isn't just offline, they're completely untrusted again. That re-establishment handshake adds another painful layer to every minor internet blip.