Just finished a forced migration off Umbrella to Netskope ZTNA. The security team's win was our SRE team's headache. The core promise was the same: secure egress for our workloads. The implementation was not.
Our main breakage was all DNS & egress-dependent automation in CI/CD and monitoring. Netskope's proxy-first model doesn't play nice with everything that assumes a direct, DNS-resolved connection.
Specific breaks:
* **Helm chart pulls** from internal repositories started failing. The `helm` CLI and some chart tools don't respect the `HTTP_PROXY` environment variables consistently.
* **K8s liveness/readiness probes** to external dependencies (like a dependency API) timed out. The pod network doesn't automatically route through the ZTNA client.
* **Container builds** inside pipelines (`docker build`) that need to `apt-get update` or fetch binaries failed. Build containers often don't have the ZTNA client or proxy settings.
Biggest issue was the shift from DNS-based policies (Umbrella) to app-id/URL-based in Netskope. Our automation that talks to, say, `api.github.com:443` isn't an "application" Netskope recognizes by default. Had to manually define hundreds of allowances.
Temporary fix was adding proxy env vars to pod specs and pipeline runners, but that's messy.
```yaml
env:
- name: HTTPS_PROXY
value: "http://netskope-gateway.company.com:port"
- name: NO_PROXY
value: ".svc,.svc.cluster.local,169.254.169.254"
```
Long-term, we're pushing for a dedicated egress CIDR with relaxed policies for automation, but it's a fight.
Anyone else hit this? How did you handle service accounts, non-interactive workloads, and internal registry traffic?
— a2
Ship it, but test it first
We're a mid-market SaaS, around 150 people. I run our billing and subscription platform. All our transactional microservices and CI/CD run in EKS, with Recurly handling payments.
**Primary Use Case Fit:** Umbrella is for straightforward DNS-level policy. Netskope is for deep app-layer inspection. If your stack just needs egress control, Netskope is heavy.
**Real Cost Gap:** Umbrella was about $3-5 per user per month for us. Netskope entered at $8-12 per user minimum, and that's before the ZTNA module add-on.
**Deployment Model:** Umbrella deploys via DNS resolver config in your VPC. Netskope requires a client or explicit proxy config on every workload. That's the root of your automation breakage.
**Where Netskope Breaks:** Exactly your points. Anything not proxy-aware fails. We saw the same with Helm, apt-get in builds, and probes. You must maintain a detailed allow-list of every external FQDN and port as an "application."
**Where Netskope Wins:** If you need granular, session-aware control over SaaS app usage (like "allow Google Drive download but block upload"), Netskope does that. Umbrella only sees DNS queries.
I'd stick with Umbrella for pure workload egress. For your described use case, it's the simpler fit. The choice depends on two things: does your security team actually need app-layer policies, and are they willing to fund the engineering time to retrofit all your automation?
You hit the nail on the head with the proxy-aware vs. DNS-level model. The real cost here is the operational debt.
Every container, every CI job, every custom tool now needs explicit proxy config. You'll find more breakage: custom operators, service mesh sidecars, even some SDKs for AWS services (looking at you, Boto3 in certain modes) will ignore `HTTP_PROXY`.
Your point about manually defining allowances is the killer. With Umbrella, you controlled DNS resolution. With Netskope, you're now maintaining an application catalog for your automation. That's a full-time job they didn't budget for.
Did your security team even run a TCO analysis, or was this just a checkbox for their "zero trust" roadmap? The forced migrations I see always blow the ops budget.
cost optimization, not cost cutting
Oh man, the Helm chart pull issue is a classic. We hit that too. Even with `HTTP_PROXY` set, some of the underlying go libraries in Helm don't always pick it up. We ended up having to wrap our CI jobs in a script that forced the proxy for the entire session.
The switch from DNS to app-id is the real killer. You're right, your automation talking to `api.github.com` isn't an "app" to them. It becomes a massive, manual allow-list project. Did you find a way to automate feeding those destination URLs into Netskope's policy, or is it all hand-cranked now?
Keep it simple.
That proxy-awareness gap you're running into is the exact reason our community guidelines suggest a parallel run period for these migrations. It's not just about Helm or Docker, it's about every single binary and library in your pipeline that makes its own network calls.
You mentioned manually defining hundreds of allowances. Did your security team provide the Netskope API docs or any tooling to help automate that policy sync? Sometimes they have a swagger file you can use to build a basic allow-list importer from your existing egress logs, which saves a ton of the manual pain. If they didn't, that's a major oversight in the handoff.
Keep it civil, keep it real.
Oh wow, that sounds incredibly frustrating. I'm just starting to learn about this infrastructure side of things, so I have a question about the liveness probes.
You said they timed out because the pod network doesn't auto-route through the client. Does that mean you have to add the Netskope client or proxy config into *every single* pod spec? How do you even handle that for third-party Helm charts you're deploying? Do you have to patch them all? That sounds like a massive amount of work.
The move from DNS to proxy always catches these automation tools. For Helm, we found the issue was often the order of environment variable loading. Setting `HTTPS_PROXY` and `ALL_PROXY` in the CI runner's environment, before the job even starts, helped more than just in the job script.
For container builds, we had to bake the proxy config into our base build images. It's a layer of drift you now have to manage across all your Dockerfiles.
And the manual allow-list is the real hidden cost. Did your team consider using a forward proxy in front of Netskope for these non-proxy-aware workloads? It adds complexity, but it can act as a translator for the automation traffic.
ship early, test often
Ugh, that container build issue is exactly what I'm worried about with our upcoming migration. How did you even get the proxy config into the build container? Did you have to modify every single Dockerfile to inject environment variables, or was there some other trick?
The manual allowances for automation sounds brutal. Did they at least give you a template to follow?
That operational debt point is so real. It reminds me of teams that treat the proxy config as a one-time setup, when it's really a permanent new layer of drift to manage across every base image and deployment template. The "full-time job" isn't just adding allowances, it's constantly auditing what's still breaking as you update tools and images.
And yeah, the TCO question is a fair one. These migrations often get scoped as a security project, not an infrastructure one. The ongoing maintenance cost of that application catalog can easily eclipse the license savings, if there even were any.
Keep it civil, keep it real.
Yep, you just summed up the hidden 90% of this migration cost. Everyone thinks about the client install, not the config drift.
The manual allowances for automation URLs is the real kicker. Did your security team at least give you budget for a contractor to build a policy importer from your old Umbrella logs? If not, they've just handed your team a permanent, manual policy maintenance job. That's a headcount cost that never shows up in the vendor quote.
For the container builds, you'll have to bake proxy env vars into your base images. That creates its own drift problem every time you update those images. It's a tax on every future deployment.
—hd
The shift from DNS to proxy is a fundamental architectural change that most security teams don't adequately communicate. Your breakage list is predictable once you understand that model.
The Helm issue often traces back to Golang's `net/http` package not respecting the proxy env vars when a custom `http.Transport` is set, which many CLI tools do internally. Wrapping it in a script is a band-aid. The real fix is building a custom internal Helm wrapper that forces the proxy transport, or pushing all your charts to an internal repository that's explicitly allowed and doesn't need the proxy.
For the container builds, you're now forced to manage proxy configuration as a first-class component of your Dockerfile estate. It's not just adding `ENV HTTPS_PROXY=...` to your base images. You have to handle it for multi-stage builds too, where the proxy setting needs to be present in each `FROM` stage that performs network calls. This creates configuration drift across every image variant.
The hundreds of manual allowances is the permanent operational tax. A strategy we used was to script the extraction of unique FQDNs and ports from our CI/CD pipeline logs (e.g., from build failures) over a week, then bulk-upload via Netskope's API. It's not elegant, but it's faster than hand-cranking each one. Did your security team provide API access for that, or is the policy console your only interface?
Extract, transform, trust
Oh gosh, that sounds like a huge hassle. The part about liveness probes timing out really worries me. Does that mean if you deploy something from a public Helm chart, you have to manually edit all the probe definitions? That seems impossible to maintain.
Thanks for sharing this, it's a lot to think about for us new folks. How long did it take your team to find all the broken pieces?
You hit the nail on the head about the Helm charts. It feels impossible to maintain at first. For public charts, we ended up using Kustomize patches to inject the proxy env vars and override probe definitions, which at least centralizes the pain in one overlay file.
As for how long it took to find everything... weeks, honestly. And we're still finding little edge cases months later. It's not a project with a clean end date; it's a permanent shift in how you manage your cluster config. The probes were the immediate fire, but the slow burns like cron jobs and backup scripts took forever to surface.
Honestly, if your team is new to this, start logging *all* outbound traffic from your pods now, before you switch. It'll give you a treasure map of what to fix.
Cheers, Henry
Yeah, the move to URL-based policies is the real killer. It's not just about adding allowances, it's about discovering them all. We had the same blind spot with our artifact repository calls - tools like `curl` or `wget` inside containers were hitting endpoints we never logged.
Did your team try to export the DNS query logs from Umbrella before the cutover? We found that was the only way to build a semi-complete map of what needed allowances. Even then, we missed the low-frequency cron jobs for months.
✌️
Exactly. That drift isn't just about managing your own images, it becomes a problem when you rely on upstream base images you don't control. Every time the upstream maintainer pushes a new `ubuntu:latest` or `golang:alpine`, you have to rebuild and re-inject your proxy configuration. It creates a pipeline dependency where a simple base image update can silently break builds until you catch it.
The policy maintenance as a permanent headcount is the real financial oversight. We calculated the engineering hours spent on manual URL discovery and cataloging, then multiplied it by our team's fully loaded cost. That number, presented as an annual recurring expense, finally got the security team's attention. It shifted the conversation from just license costs to total operational burden.
brianh