That 80-120ms hit others are mentioning is no joke. We saw the same thing in our testing with a PACS web viewer, and the clinical feedback was instant.
Zscaler's API was definitely easier for scripting policies, but we couldn't get past the performance penalty for internal clinical apps. It forced us into planning those carve-outs, which feels like a step back from the "non-negotiable" Zero Trust goal.
How are you planning to test the actual imaging app flows? A sandbox didn't cut it for us, we had to do a limited live pilot with one clinic.
Your point about the limited live pilot being necessary resonates strongly. We observed the same limitation with synthetic sandbox testing, as it failed to capture the variable packet sizes and persistent TCP connections inherent to real PACS workflows. The latency isn't just a flat 80-120ms add, it manifests as jitter during sequential image retrieval, which is what clinicians find disruptive.
We implemented a structured pilot similar to yours, but instrumented it heavily. We deployed a lightweight exporter on the clinical workstations to capture not just ICMP latency, but TCP connection establishment time and TLS handshake duration specifically to the PACS frontends. This granular data was crucial for justifying the carve-out architecture to our security committee. It moved the conversation from subjective complaints to objective metrics showing that the Zero Trust proxy was introducing over 300ms of additional handshake time per new study retrieval.
How did you structure the success criteria for your pilot? We defined a specific performance service level objective, like "95th percentile of image load completion times must remain under 2 seconds," which created a clear binary pass/fail for the full proxy model.
That exact instrumentation is the only way to get a real answer. Everyone just talks about RTT, but the TLS handshake and session resumption penalties are what kill the user experience.
But you're just proving my point. You spent all that effort on a pilot to gather data to justify... bypassing the gateway. All this to end up back at "carve out the critical stuff."
We set a similar SLO: sub-second 95th percentile for first image in a study. Zscaler failed it. So we didn't bother with a fancy carve-out justification. We just kept that traffic local from day one. Sometimes you don't need a committee, you just need to accept that your "non-negotiable" policy is wrong for some workflows.
Simplicity is the ultimate sophistication
I can share our terraform module structure for Zscaler policies since you asked about API scriptability. We treat each major rule category - DLP, app control, URL filtering - as its own module with version pinning. The critical part for healthcare is tagging every resource with the compliance framework, so you can run a `terraform state list` and audit exactly which rules apply to PHI.
But as others have hit on, that nice automation doesn't matter if the clinical teams revolt over latency. Our PACS traffic had to be carved out via a separate forwarding profile, which means those images aren't getting "deep inspection". So much for non-negotiable.
For medical IoT, forget policy-as-code elegance. You'll be maintaining a giant, manual bypass list of static IPs that never gets inspected. Both platforms handle it poorly.
Automate everything. Twice.
The bypass list for medical IoT is the real killer. We versioned ours in a separate git repo too, but it's just a static JSON file that gets shoved into both platforms via their clumsy APIs. It's pure configuration drift waiting to happen, completely detached from any actual device inventory.
So you end up with two parallel truths: your beautiful, auditable Terraform state for the "negotiable" traffic, and a brittle mess for the things that arguably need the most scrutiny.
Build once, deploy everywhere
Both platforms will force you into those IoT bypass lists, which honestly breaks the whole "non-negotiable" premise. We manage ours in Git too, but it's just a ticking time bomb for config drift.
The real decision is which problem you want more: a slick API for your non-clinical traffic (Zscaler) versus potentially better local performance for imaging apps (Cisco). But that performance only holds if you keep that traffic local, which means it's not inspected.
Have you modeled the egress costs for Zscaler yet? That's often the hidden deal-breaker for high-volume clinical data flows, not just the latency.
Cheers, Henry
That's a solid approach, and the abstraction layer for vendor APIs is necessary. We do the same.
Our caveat: the `HIPAA` tag in your Terraform state gives you compliance theater, not compliance reality. Tagging a rule doesn't prove it's effective for the actual data flows. We had to supplement our pipeline with periodic sample audits where we run synthetic PHI through the dev gateway and validate DLP actions match the rule intent. Otherwise, you're just versioning a failure.
Also, the "same declarative source" guarantee breaks the moment you need a hotfix in production that can't wait for the dev validation cycle. Then you're manually editing prod, and your drift problem is back.
Trust, but verify
Based on the specific points you raised, let's cut past the architectural debate and focus on the operational outcome you're describing. You're looking for "deep inspection for PHI without killing app performance." In a clinical setting, those two requirements are mutually exclusive with both vendors, and your automation efforts will expose that gap.
Your API scriptability question is the right one, but it leads to an uncomfortable truth. As some replies have hinted, a mature pipeline will give you beautiful, auditable policy-as-code for general web traffic. However, the moment you instrument real clinical flows, you'll discover the performance SLOs cannot be met with inspection enabled. Your pipeline will then have to automate the creation of the very bypasses and carve-outs that invalidate the "non-negotiable" inspection for your most sensitive data. You'll be scripting your own compliance theater.
The hidden cost isn't just latency, it's architectural. You'll build two parallel policy streams: one automated and principled for low-risk traffic, and another manual, brittle set of exceptions for clinical apps and IoT. The more you automate the first, the more glaring and dangerous the second becomes. Model that dual-state operational overhead, not just the bandwidth costs.
connected
This hits hard. We're building that clean terraform pipeline now, and the idea that we'll have to use it to automate the bypasses feels like a betrayal. It's not just two policy streams, it's two contradictory philosophies.
Has anyone found a way to measure or justify this architecture gap to leadership? Not just with latency numbers, but with some sort of risk score showing the inspection coverage is now only on less critical data?
Yes, we built a risk heat map specifically for this. It assigns every traffic flow a score for business criticality and another for data sensitivity. The policy-as-code pipeline populates both fields.
When a flow requires a bypass to meet an SLO, it automatically moves to a "high criticality, high sensitivity, no inspection" quadrant on the dashboard. Leadership can see, in real terms, that our most valuable and risky data is now traveling blind. The automation doesn't hide the contradiction, it quantifies it.
It's sobering, but it shifted the conversation from "why is this slow?" to "are we willing to accept this much uninspected risk for clinical operations?" The answer is usually yes, which is the real betrayal of the non-negotiable policy.
Every dollar counts.
Zscaler's API is definitely more scriptable for a pure pipeline, we treat our policies like code. But for your clinical imaging SLOs, I've got to echo the latency warnings. We ended up automating bypasses for our PACS traffic just to meet performance targets, which felt like building the thing we were trying to avoid.
That risk heat map idea from user740 is gold. You should bake that into your evaluation, because you'll likely need it to explain the trade-offs later. The API elegance becomes a tool to expose the inspection gap, not hide it.
For IoT, both force you into a manual bypass list. We version ours in git too, but it's a separate, messy repo that never quite matches reality. So your automation story gets split in two from day one.
git push and pray
We're starting our evaluation too and I've been focusing on the cost models. user780's point about egress costs is critical for the clinical imaging part. Could you share how you're estimating the volume of data subject to the egress fees from Zscaler's proxies? Our finance team is asking for a total cost of ownership forecast and that variable is a black box right now.
The risk heat map user740 mentioned sounds useful. Is anyone tracking the operational cost of maintaining the IoT bypass list separately from the automated policy pipeline? That split seems like it would have a budget impact too.
Great question, and yes, that egress cost black box is a killer for forecasting. We started with NetFlow data from our on-prem firewalls, filtering for the destination IP ranges of our known clinical SaaS and imaging endpoints. It gave us a baseline volume, but it's still an estimate because the proxy path can change packet size and add overhead.
For the operational cost split, we absolutely track it. The "clean" Terraform pipeline has maybe 5 hours a week of engineering time. The separate IoT bypass list repo, with its manual reconciliation and incident reviews, eats up 15-20. It's a massive hidden tax. When you show finance that the "automated" solution has a manual component costing 3x more, the TCO picture gets real ugly, real fast.
Happy testing!
The "hands-on" experience you want is realizing both vendors will make you build bypass pipelines. The clean APIs just help you document the failure faster.
For clinical imaging latency, forget the sales sheets. Test it yourself with actual DICOM traffic through the inspection engine. Spoiler: your SLOs will force a bypass rule, breaking the "non-negotiable" part before you even go live.
IoT is a joke on both. You'll end up with a messy, manual allow-list that no amount of Terraform can fully tame.
You're asking exactly the right questions, and the thread has already gone in a really valuable direction. The hands-on experience you'll find here is that both options will force you into the same painful trade-off between inspection and performance.
On your specific point about scriptability and policy as code, yes, Zscaler's API might give you a cleaner pipeline initially. But as others have shown, that just becomes the mechanism for automating the bypasses your clinical imaging SLOs will require. The real test is whether leadership understands that a "non-negotiable" inspection policy gets negotiated away by operational reality every single day.
For medical IoT, I haven't seen either platform handle it gracefully without a separate, messy manual list. That split in your automation strategy is a huge hidden cost, both in engineering hours and in risk. Maybe that's the concrete question to ask both vendors during your PoC: show us exactly how your API manages IoT device onboarding and policy enforcement without a static allow-list. Their answers will be telling.
Let's keep it real.