Spot on about the undocumented dependencies. The monitor-only period you suggest is wise, but it's not a panacea. We did the same and generated a firehose of logs so voluminous they were nearly useless without another tool to parse them.
And that 70/30 split is telling. It means the platform's main value prop, securing external SaaS, is the smaller part of the initial pain. Most of the work is just building an internal map you should have had already. Zscaler becomes the world's most expensive dependency discovery tool.
So you pay for the platform, then you pay your team to document your own estate with it. Clever business model, I suppose.
cg
The spreadsheet method only works if you have buy-in from the security and networking teams to *maintain* it. In my experience, it becomes stale within weeks.
We handle change by forcing the update request through the same Jira ticket that's used to request the Zscaler policy change itself. The application owner's requirement shift *is* the ticket. If they don't provide the updated details, the ticket doesn't move. It's bureaucratic but effective.
The real problem is when the requirement is vague, like "make it faster." That's when you need that senior engineer with the tribal knowledge to interpret what "faster" actually means in Zscaler policy terms.
Build once, deploy everywhere
Coupling the policy ticket to the requirement doc is the right idea, but your Jira method still depends on humans writing accurate requirements. That's the same problem.
We automated the dependency capture. The CI pipeline that deploys the app also generates a manifest file. The Zscaler policy change ticket won't even open without that manifest attached. It's rigid, but it eliminates the interpretation step for the senior engineer.
The "make it faster" problem usually means a policy is doing SSL inspection on a latency-sensitive internal API. That's not tribal knowledge, that's a platform design choice that needs to be documented as a rule.
Build once, deploy everywhere
You're right about the policy granularity being a double-edged sword. That "months-long tuning process" you mentioned is often just the start. Once your app teams realize what's possible, you get a flood of requests for highly specific rules. What saved us time initially now demands a permanent policy review cycle.
That said, the Okta integration you called out is a game-changer for controlling the flood. We built user group syncing from Okta directly into our policy framework, so access changes automatically when someone's role changes. It cut down on one-off requests dramatically.
The reporting clunkiness is real, though. I find the API is the only sane way to pull data for anything specific, like that blocked traffic report you mentioned. The UI just isn't built for it.
The learning curve is less about time and more about exposure to specific failure modes. Our team could build basic policies after a few weeks, but true comfort came only after the first major incident, like when we learned the hard way that a "Deny" rule for a subdomain doesn't implicitly block its parent domain. You don't grasp the policy inheritance hierarchy until you've broken a core service with it.
So the answer is about six months, but that's contingent on running a parallel "shadow" policy set in monitor-only for all non-critical traffic during that period. It lets you build and break things safely. The comfort comes from seeing your own mistakes in logs, not from training modules.
Also, building from scratch is a misnomer. You're almost always forking and modifying an existing policy. The skill is in auditing the chain of existing rules to predict how your change will interact, which is a separate and deeper layer of knowledge.
infrastructure is code
Your point about justifying renewal with ROI mapping resonates deeply. Many enterprises I've consulted for hit that same wall. The initial savings from decommissioned VPN concentrators and data center egress are clear, but the ongoing justification is harder.
The hidden cost isn't just the line item; it's the operational expense of maintaining the policy framework you described. That "months-long tuning process" has a direct FTE cost. I've seen teams spend 20-30% of a senior engineer's time just on policy lifecycle management, which rarely gets factored into the TCO. The platform's power creates its own policy debt that requires constant payment.
Your ROI mapping likely focused on capital avoidance. For the next renewal, I'd recommend building a parallel model showing the cost of an equivalent security stack assembled from point solutions. Factor in the labor for integrating those tools and the performance tax of backhauling traffic to a cloud proxy you manage. Zscaler often still wins, but the justification becomes about comparative operational complexity, not just hardware savings.
Always check the data transfer costs.
>the reporting feels clunky.
You're spot on. The UI for reporting is my biggest daily gripe. It feels like it was designed for a checklist audit, not for finding out *why* something broke in the moment.
We ended up building a small Power BI dashboard pulling from their API to get the visibility we actually need. It shouldn't be necessary, but it's the only way we can get quick answers.
That policy granularity is a blessing and a curse. It saves time up front, but now everyone wants hyper-specific rules. We have a permanent backlog of policy reviews because of it.
dk
Your experience with the "months-long tuning process" mirrors our own. The critical mistake, in my view, is treating the initial policy set as a one-time build. It isn't. It's the first version of a constantly evolving schema for your entire network's communication patterns.
The cost of maintaining that schema is the real TCO that often gets omitted. You trade VPN hardware CAPEX for a continuous policy engineering OPEX. We found that operating Zscaler at scale requires a dedicated, product-like approach to policy management, including version control, staged rollouts, and automated testing of rule changes against a known set of critical traffic flows.
On reporting clunkiness, we bypassed the UI entirely. The API is stable enough to build your own abstraction layer. We wrote a simple service that translates questions like "show me blocked traffic for department X last week" into the necessary API calls, then caches the results. It shouldn't be necessary, but treating their UI as a read-only admin console and building your own operational interface is the only sane path forward.
That's exactly the shift we made. Treating policies like IaC was the only way to stay sane.
We started using Terraform for Zscaler a year ago. It forces that version control and staging you mentioned, but it also surfaces drift. Nothing like a `terraform plan` to show you which ad-hoc UI changes your networking team snuck in last week.
Did you run into issues with API rate limits when building your abstraction layer? We had to get clever with caching to handle our volume.
Demo or it didn't happen
Terraform for Zscaler is a good step, but the real headache starts when you try to apply proper CI/CD to it. You can't just run `terraform apply` on a schedule without potentially breaking production.
We ended up with a two-repo approach: one for the module definitions (treated as a library) and one for the actual policy declarations (which references the modules). The deployment pipeline has to run a series of synthetic traffic tests against a staging ZIA tenant before promoting changes. It's more ops work than the vendor likes to admit.
On the rate limits, absolutely. The official provider doesn't handle them gracefully. We built a wrapper in Go that respects the Retry-After headers and implements an exponential backoff. The bigger issue we hit was API timeouts on bulk operations when managing thousands of URL categories.
That two-repo setup is smart. I'm just starting with Terraform and hadn't thought about splitting modules from declarations like that. Does your CI pipeline block the apply if the synthetic tests fail? I can see that being a big hurdle to get right.
The API timeout issue is a bit worrying. I'm only dealing with hundreds of rules now, but planning for growth. Did you move your bulk operations to a different schedule, or did the wrapper fix it completely?
Your point about the ROI mapping focusing on hardware savings is crucial. We found a more compelling renewal case when we shifted the justification to risk reduction and user productivity, metrics that were harder to quantify initially. The cost of a single security incident avoided because of granular TLS inspection, or the aggregate time saved by thousands of users not having to connect to a VPN, started to outweigh the line-item comparison.
The policy management cost you mention is real, but it also consolidates work that was previously fragmented across firewall, proxy, and VPN teams. It's a centralization tax. We started billing that engineering time back to application teams as a shared service cost, which had the side effect of making them more thoughtful with their policy requests.
On the reporting clunkiness, we built a dedicated channel in our operations Slack that pipes in alerts for specific high-risk block categories from the API. It's not a full report, but it gives the support team immediate visibility without needing to log into the portal, which helped bridge that gap.
Support is a product, not a department.
Your point about treating the policy set as a version one, not a finished product, is the key insight most implementations miss. That months-long tuning is really just the initial discovery phase for your organization's actual traffic patterns.
We made the same mistake, and it led us to a similar conclusion: you need a dedicated owner for the policy schema. It's not just a network or security task anymore, it's a product with its own lifecycle. The operational cost is real, but framing it as a centralized service, as you mentioned, is the only sustainable model.
Have you found that formalizing the policy management process actually reduced the volume of ad-hoc requests from application teams, or did it just make the backlog more visible and manageable?
>justifying the renewal required a lot of ROI mapping around saved VPN hardware and support time.
Yeah, that's the trap. They sell you on eliminating hardware, but you're just trading a capital expense for a much larger, recurring operational one. You'll spend more time managing the policy "schema" than you ever did on firmware updates.
That reporting clunkiness is classic. It's an audit box-checker, not an ops tool. You end up building your own dashboards anyway, which adds even more to the real cost.
SQL is enough
That's a really sharp point about trading one type of cost for another. It's exactly the kind of TCO blind spot I'm trying to map out in my own evaluation process right now.
My worry is that the operational cost you mention isn't just labor hours, it's also expertise. You can't just hand this "policy schema" over to a junior admin. It seems like you're trading predictable hardware refreshes for a dependency on a few senior engineers who understand both networking and this specific platform's logic.
Did your team find that the policy management workload plateaued after a certain point, or does it just keep growing with every new application the business adopts?