Another year, another platform migration, and here I am again staring at a cloud bill that makes me wince. This time it's Zscaler Private Access, specifically the cost of those Connector VMs that never seem to scale down when you actually need them to. You know the drill: you provision for peak, and then you're paying for that idle capacity 18 hours a day, like renting a banquet hall to eat a single takeout meal.
The official line is to manually adjust your Connector fleet, but who has the time? And if you miss a spike, your sales team is on Slack screaming about not being able to reach the legacy deal desk app. So, after the last billing cycle shock, I built a system to auto-scale the ZPA Connectors in Kubernetes. It's not for the faint of heart, but it cut our monthly ZPA infra costs by roughly 40%. The core idea is simple: treat the Connectors as stateless, ephemeral pods and let K8s handle the scaling based on real-time metrics.
Here’s the gist of the architecture:
* **Containerized Connectors:** You need to build a Docker image for the ZPA Connector. This involves scripting the provisioning process (using the ZPA API) so a pod can bootstrap itself and join the correct App Connector Group. The secret sauce is ensuring it cleans up after itself on termination.
* **Custom Metrics & HPA:** The vanilla CPU/Memory metrics are useless here. You need to scale based on concurrent active tunnels or pending requests. I set up a sidecar that scrapes the Connector's local stats API and exposes them to Prometheus. Then, a Prometheus Adapter lets the Kubernetes Horizontal Pod Autoscaler make decisions.
* **Tricky Parts & Pitfalls:**
* **Boot Time:** A new Connector takes 2-3 minutes to provision and be ready. Your scaling policy needs a buffer, or you'll get throttled during a fast ramp-up.
* **Session Persistence:** You must ensure existing user sessions aren't massacred when scaling down. A graceful termination period (coupled with ZPA's own failover) is critical. We use a `preStop` hook to signal the Connector to start draining.
* **Cloud Costs vs. Licensing:** Remember, ZPA licensing is separate. You're saving on the underlying compute/storage, not the Zscaler SKU. This makes the most sense on a cloud with per-second billing.
The actual scaling policy in our cluster looks something like this, targeting an average of 150 active tunnels per pod before spinning up another:
```yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: zpa-connector
minReplicas: 2
maxReplicas: 10
metrics:
- type: Pods
pods:
metric:
name: zpa_active_tunnels
target:
type: AverageValue
averageValue: 150
```
Is it worth the operational overhead? If your Connector count is static and small, probably not. But if you're supporting a global, fluctuating workforce with a dozen-plus Connectors, the savings are substantial. You're essentially trading manual VM management for a more complex, but automated, orchestration layer. Just be prepared for a few nights of debugging why a pod won't join the provisioning group before you get it right.
You say it cut costs by 40%. I'll believe it when I see the billing dashboard screenshots. The ZPA connector AMIs aren't free, and neither is the K8s control plane you're now running.
What metrics are you using to scale? Connection count? CPU? If you scale in too aggressively during a lull, you'll strand user sessions and cause those exact Slack screams you mentioned.
show me the bill
That 40% figure definitely gets your attention! I can see where user389 is coming from with the skepticism though. The K8s control plane cost is a real factor, but in our case, we were already running a sizable cluster for other workloads, so adding the connector pods was just a marginal increase. The bigger savings came from replacing those always-on, oversized VMs.
The metric question is crucial. We started with CPU but found it lagged behind actual user need. What worked for us was scaling primarily on active TCP connections per pod, pulled from the connector's own metrics endpoint, with a CPU-based safety net to catch any runaway processes. You have to set the thresholds with a buffer - never scale in on a single quiet minute at 2 PM, or you *will* get those Slack screams.
It requires some tuning to get right, but the banquet hall analogy is spot on. Why pay for the whole hall when you only need a few tables most of the time?
Clean data, happy life.
Your point about treating the Connectors as stateless pods is the critical leap. The main complication we ran into wasn't the scaling logic, but the bootstrapping time. A new pod takes 90-120 seconds to pull the image, execute the provisioning script, and fully register with the ZPA cloud. If your scaling metrics are too sensitive, you'll have a thundering herd of pending pods while active ones are being terminated.
You need to build a significant cooldown period into your HPA or KEDA scaler, and set your minimum replicas to cover the baseline load plus the time-to-bootstrap buffer. We also found pre-pulling the container image onto our worker nodes was necessary to shave off those initial 30 seconds.
—Alex
That bootstrapping time is the part I'm stuck on myself. My team wants to try this but that 90-120 second delay makes me nervous. If a lunchtime surge hits, we'd be down until the new pods are ready, right?
Do you have that provisioning script for the Docker image anywhere? I'm not sure how to get the connector to self-register from inside a container.
Containers are magic, but I want to know how the magic works.
Oh, that comparison to renting a banquet hall is painfully accurate. Been there, and the wasted capacity is the whole reason I started looking at this path.
You're absolutely right that treating them as stateless pods is the key mental shift, but the real battle in my experience wasn't the HPA config-it was convincing the security team that an ephemeral, auto-scaling connector fleet met their "always available" compliance requirements. We had to build in a 'minimum replicas' floor based on our busiest expected hour for each region, which still saved a ton versus our old 24/7 peak provisioning, but it meant the savings weren't a pure 40% across the board. It's a negotiation more than a technical fix sometimes.
The scripting for the container image is the other beast. Getting the API calls right so the pod can self-provision and join the right segment is fiddly, and if you have multiple ZPA tenant IDs, your configuration map gets complicated fast.
Implementation is 80% process, 20% tool.
That 40% savings is a strong claim. Even if it's real for your specific peak-to-trough, you need to post the actual scaling config and your metric thresholds. Otherwise this is just a story.
The real cost isn't just the connector pods, it's the engineering hours to build and maintain that custom container and the scaling policies. Most teams would burn that 40% savings just getting it to run reliably.
Beep boop. Show me the data.
> The official line is to manually adjust your Connector fleet, but who has the time?
This hits home. I'm just starting with Terraform and my first thought was to maybe write a script with the AWS CLI to schedule instances? But your K8s approach sounds way more dynamic.
When you built the Docker image, how did you handle the secrets for the ZPA API? Did you mount them from a K8s secret when the pod starts? I'm trying to figure out how to make that self-provisioning script work without hardcoding keys.
> treat the Connectors as stateless, ephemeral pods
That's the right starting point. The biggest gotcha is assuming they're truly stateless. They're not. Each connector holds active sessions.
If you scale in and kill a pod with live sessions, those users get cut off. Your scaling policy needs a graceful termination period, long enough for those TCP sessions to naturally end or migrate. We set a 5-minute terminationGracePeriodSeconds and configure the HPA stabilization window to prevent rapid scale-in during short dips.
You're right about the container image being the first major hurdle. The real friction point we encountered was the licensing nuance. That provisioning API call you need to embed in the Dockerfile requires a specific provisioning key tied to a *connector group*, not just your admin credentials. Hardcoding that in the image layer is a security audit finding waiting to happen.
Our solution was a two-stage init container in the pod: the first pulls the ephemeral provisioning key from a secrets vault using the pod's service account, and the second uses that key to execute the registration. The image itself only contains the generic connector binaries. This keeps the image reusable across environments and doesn't bake in any secrets, even transient ones.
Mike