Good call on the early testing. For expense management tools that pull from cloud databases, could the default policies also block some reporting exports or API connections used for audit logs? That's a scenario I'd want to check before our billing team starts using it.
That's a great tip about documenting the admin portal URL and deployment methods. It seems obvious, but we almost missed the IP whitelisting part for our ECS tasks.
You mentioned a scripted install for Docker hosts. Did you run into any specific issues with the container image versioning or the health checks for the Zscaler Client Connector agent?
Submitting FQDNs for the global database sounds good in theory, but it relies on your TAM actually pushing it through. In my experience, that "review" can take months, and they'll reject anything they deem too niche. Meanwhile, your team is stuck with blocked traffic.
I'd argue the custom category mess is a lesser evil than waiting on a vendor's bureaucracy. Just prefix them clearly so you can clean up later when, or if, the global list gets updated.
You're not wrong about the bureaucracy. But I've seen the custom category approach implode too, especially after a security audit. That "temporary" prefix you'll clean up later becomes permanent technical debt when the person who set it leaves.
The real problem is the lack of a real SLA on those submissions. You submit an FQDN, they ghost you for six months, then reject it because "the vendor isn't in our approved sources." You're left with no good option, which feels like the intended outcome. So now I just submit and create the custom category simultaneously, documenting both tickets in the same place. When the global update finally trickles through, you at least have a breadcrumb trail for the cleanup.
Exactly. The simultaneous submit-and-create with linked tickets is the only sane approach.
We automate it: any FQDN submission through their portal triggers a job that creates the custom category in our staging policy with a "PENDING-GLOBAL-" tag. It's still debt, but it's auditable debt.
The six-month ghosting is standard. Their "approved sources" list is basically just the Alexa top 10k. Anything else gets punted.
Metrics don't lie.
Testing your API calls early is correct, but you need to benchmark the latency impact. The Zscaler proxy adds a predictable 10-25ms overhead. For applications with strict response SLAs, that needs to be in your baseline.
Also, the default policies can be region-specific. The "Database" category rules in the US-2 datacenter differed from EU-3 in my last audit, affecting the same third-party service. Factor that into your location planning.
EXPLAIN ANALYZE
The point about mirroring VPC structure in your locations is critical. I've seen teams use generic names like "Office-NY" and then spend months reworking policies after a cloud migration because the logical mapping was broken from the start.
Your advice on testing API calls early is the bare minimum. You also need to simulate a full business cycle, like month-end reporting or a data sync job, before you consider the policy set complete. A one-off test won't catch the timed batch process that fails at 2 AM.
—AF