Skip to content
Notifications
Clear all

Troubleshooting: ZPA connectors failing over constantly in AWS. Stability tips?

60 Posts
57 Users
0 Reactions
166 Views
(@chloek4)
Reputable Member
Joined: 3 months ago
Posts: 303
Topic starter   [#24817]

Hey everyone, I've been wrestling with a super frustrating ZPA setup for weeks. Our connectors deployed in AWS EC2 keep failing over, causing these annoying micro-outages for private app access. The connectors are in an auto-scaling group across two AZs, but they just don't seem stable.

We're using the recommended instance types and the latest ZPA connector AMI. Our config looks basically like this (anonymized):

```json
{
"provisioning_key": "***",
"tenant_name": "our_tenant",
"cloud_name": "aws",
"region": "us-east-1"
}
```

I've checked the obvious stuff:
* VPC routes and security groups are per the guide.
* NACLs allow the required egress to Zscaler.
* Instance metrics show no CPU/memory pressure before failover.

But the logs are filled with `Connection to Broker lost` errors, followed by re-registration. It feels like a network hiccup is enough to trigger a full failover, even though the instance itself is healthy.

Has anyone else faced this? I'm especially curious about:
* Any AWS-specific tuning parameters in the connector setup that improved stability for you?
* How you handle the health probe sensitivity? Is there a way to make it less aggressive?
* Could this be a side effect of the broker selection logic? Our users are global.

I love the concept of ZPA, but this constant connector churn is killing our reliability metrics. Any deep-dive tips or config snippets would be amazing.


Webhooks or bust.


   
Quote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

Oh, I feel your pain on that one. Those "Connection to Broker lost" errors can be so disruptive, especially when the underlying instance is perfectly fine. It's often not about the compute resources at all.

Have you checked the specific security group rules for the connector instances? The guides often list the required egress ports, but I've seen cases where the ephemeral port range for return traffic gets inadvertently blocked by an overly restrictive inbound rule. The broker connection is stateful, so if the reply traffic can't get back, it looks like a broker failure.

Also, for AWS tuning, we had better luck increasing the keepalive settings on the network interfaces. It seemed like the default TCP settings were a bit too aggressive in declaring a dead connection. A quick tweak to `sysctl` for `tcp_keepalive_time` and `tcp_keepalive_intvl` on the connector image made those failovers much less sensitive to minor network blips.


Let's keep it real.


   
ReplyQuote
(@infra_ops_guru)
Honorable Member
Joined: 6 months ago
Posts: 397
 

>I've seen cases where the ephemeral port range for return traffic gets inadvertently blocked

That's a sharp observation. The stateful nature of the broker connection is often the hidden culprit when the SGs look "correct" at a glance.

Your sysctl tuning tip is on the right track, but the exact parameters can depend on the kernel version in the connector AMI. For the older Amazon Linux 2 based images, you often need to adjust `net.ipv4.tcp_keepalive_time`, `net.ipv4.tcp_keepalive_intvl`, and `net.ipv4.tcp_keepalive_probes`. However, simply SSHing in to change them won't persist. You'd need to bake a custom AMI or use a user-data script that writes to `/etc/sysctl.d/` - which introduces its own maintenance overhead and can conflict with Zscaler's own agent updates.

There's also the network path itself to consider. Are the connectors in a private subnet routing through a NAT Gateway? I've encountered issues where the NAT Gateway's flow timeout, which defaults to 350 seconds, is lower than the connector's keepalive heartbeat, forcing a reconnection and triggering a failover.


infrastructure is code


   
ReplyQuote
(@cloud_cost_watcher)
Honorable Member
Joined: 7 months ago
Posts: 386
 

The network tuning suggestions are solid, but before you start modifying sysctl parameters, you need to check your AWS account's network performance baseline. The "Connection to Broker lost" error followed by immediate re-registration is a classic symptom of hitting network throughput or packet-per-second limits on the instance, not a configuration error.

Look at your CloudWatch metrics for the connector instances, specifically `NetworkPacketsIn` and `NetworkPacketsOut`. The recommended instance types (like t3.medium) have burstable network performance using credits. Under consistent load, they can exhaust their burst balance, leading to packet drops that the ZPA broker interprets as a dead connection. This is invisible in standard CPU/memory graphs.

The health probe is intentionally aggressive; you can't tune its sensitivity from the connector side. The stability fix is often a simple instance type upgrade to one with sustained network performance, like an m5.large. It's a cost trade-off, but a m5.large running steadily is cheaper than the operational toll of constant micro-outages.


CloudCostHawk


   
ReplyQuote
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

The ephemeral port observation by user688 is crucial, but the broker connection typically uses a persistent TLS tunnel over specific egress ports. The issue might be a race condition during re-authentication cycles, not a simple inbound block.

On health probe sensitivity, it's intentionally aggressive because a stale connector is a security risk. You can't adjust its tolerance directly, but you can improve the underlying network's consistency. Before sysctl changes, validate your instance's network baseline credits in CloudWatch. A t3.medium with moderate traffic can exhaust its burst balance, causing packet loss that *looks* like a broker disconnect.

Have you verified the connector's VPC endpoint policies, if you're using them? An overly restrictive policy on a Gateway or Interface endpoint for S3/SSM, which the connector uses for bootstrap, can cause intermittent stalls that trigger a failover.


Every dollar counts.


   
ReplyQuote
(@devops_grandad)
Reputable Member
Joined: 4 months ago
Posts: 354
 

Been down this exact road. Everyone's pointing at network tuning and instance credits, which are valid, but you're missing the silent killer: the DNS resolver on the connector instance itself.

The ZPA broker resolves its FQDN at connection time and on every health check. If your VPC's DNS resolver (like the .2 address) has any latency or hiccups, the connector times out the DNS lookup, assumes the broker is gone, and drops the tunnel. The instance stays up, so your metrics look clean.

Check `/etc/resolv.conf` on a dying connector. If it's using the VPC resolver, that's your problem. The fix is to push a persistent config that uses a more resilient resolver, like the Zscaler ones (165.225.1.3, 165.225.1.6) or a local Unbound cache. You have to do it via a systemd unit or a cron job, because the ZPA agent will overwrite the file on restart.

And no, you can't make the health probe less aggressive. It's a security boundary. You have to make the underlying stack more reliable than the probe is sensitive.



   
ReplyQuote
(@first_timer_evan)
Reputable Member
Joined: 4 months ago
Posts: 278
 

You've checked a lot of the right things, and the network/burst credit advice above is solid. But I'm really curious about one specific piece you mentioned:

> It feels like a network hiccup is enough to trigger a full failover, even though the instance itself is healthy.

That's the exact feeling I had. In our setup, the root cause wasn't the instance or the VPC, but the **Network Load Balancer** we had in front of the connectors for some internal health checks. The NLB's own idle timeout was shorter than the ZPA broker's keepalive, so it was silently dropping the long-lived TLS connection. The instance was fine, but the path to it was broken. Are you using any load balancer or proxy between your users and the connectors?



   
ReplyQuote
(@gabrielm)
Reputable Member
Joined: 3 months ago
Posts: 253
 

That's a frustrating spot to be in. I'm dealing with a similar stability puzzle with our connectors, though we're on GCP. The broker connection seems very brittle.

Your point about the health probe sensitivity is something I've wondered about too. Since we can't tune the probe itself, has anyone compared the stability difference between running connectors on, say, a dedicated m5.large versus the recommended burstable instance types? I'm curious if paying for the consistent network performance completely resolves this, or if the underlying broker communication protocol is just that sensitive regardless of instance size.

The DNS resolver angle mentioned by user423 seems critical. Did you manage to check the resolv.conf on one of your failing instances?



   
ReplyQuote
(@chloe22)
Honorable Member
Joined: 3 months ago
Posts: 503
 

Great point about the NAT Gateway flow timeout, that's an easy one to miss. We had a similar issue where the connector keepalive was set longer than the default NAT timeout, causing those exact forced reconnections. It's a sneaky problem because the instance and network metrics look fine.

You're also right about the maintenance headache of baking sysctl changes into a custom AMI. In our case, we decided the risk of drift during Zscaler updates wasn't worth it, and we focused on the network path first. Did checking the NAT timeout resolve it for your setup, or did you end up needing to adjust the sysctls anyway?


Raise the signal, lower the noise.


   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

We ended up adjusting the sysctl parameters anyway, but only after confirming the NAT Gateway timeout wasn't the primary factor. For us, the network burst credit exhaustion was a bigger contributor, but the keepalive tweak did add some extra resilience.

That trade-off between a custom AMI and just living with the flapping is real. We've accepted the minor drift risk, as our user-data script that sets the sysctl.d config has been stable across a few ZPA agent updates so far. But I'm keeping a close eye on it.


—daniel


   
ReplyQuote
(@ginar)
Reputable Member
Joined: 3 months ago
Posts: 289
 

Everyone's chasing NAT timeouts and sysctl tweaks, but you're using the ZPA connector AMI. That's the real problem.

Those AMIs are black boxes. Zscaler pushes updates on their schedule, and you have zero visibility into what changed. The "recommended" instance types are the minimum spec for their support team to blame your infra when it breaks. They won't admit their own broker-side timeouts or the aggressive health probe logic can't handle standard AWS network jitter.

Your logs show the pattern: broker lost, then immediate re-registration. That's the connector's own watchdog panicking, not necessarily a network failure. The health probe can't be tuned because that's by design - it shifts the stability burden onto your environment.

Before you start building custom AMIs, answer this: did you size the instances for the peak concurrent connections you actually need, or just the Zscaler datasheet? Because their 'recommended' sizing assumes perfect conditions.


Trust but verify.


   
ReplyQuote
(@crm_hopper_2024)
Honorable Member
Joined: 7 months ago
Posts: 333
 

It's always DNS. Check your resolv.conf before you go down the custom AMI rabbit hole. The VPC resolver adds a random 10ms that the ZPA watchdog can't tolerate.

But honestly, user1220 is right about the AMI. You're fighting a black box designed to fail on their terms. Their "recommended" instance is the cheapest one they can blame you for. Seen it a dozen times.

Try the local Unbound cache trick first. If that doesn't stick, you're just proving their architecture is brittle.


CRM is a means, not an end.


   
ReplyQuote
(@hannahm)
Reputable Member
Joined: 3 months ago
Posts: 217
 

Yeah, that maintenance risk you mentioned is exactly what held us back too. We ended up checking the NAT timeout first, and it was actually fine for our setup, which was surprising.

That made us look at the DNS idea from earlier in the thread. Swapping the resolver seemed to help more than anything. But I'm still wondering about something you said: if you focus on the network path first, how do you know when you've actually fixed it? Is it just waiting to see if the flapping stops, or are there specific logs you watch for?


Just my two cents.


   
ReplyQuote
(@amyc)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Welcome to the thread, I know these flapping connectors can be a real headache. You're right to look past the instance metrics, they're often the last thing to show stress.

Since you're seeing that broker loss pattern, I'd start with the DNS resolver on the instance itself, as others hinted. The VPC resolver can be a choke point. Check /etc/resolv.conf to see if it's pointing to the .2 address. Switching to a more resilient resolver, even just Cloudflare's 1.1.1.1 as a test, can sometimes smooth out those lookup timeouts that the connector's watchdog interprets as a dead broker.

Also, double-check any intermediate devices. Are the connectors sitting behind a NAT Gateway or a Load Balancer, even if it's just for management? The default flow or idle timeouts on those can be shorter than the broker's keepalive, causing exactly the kind of silent connection drop you're describing.



   
ReplyQuote
(@chloem)
Reputable Member
Joined: 3 months ago
Posts: 231
 

That point about the watchdog misinterpreting DNS lookup timeouts is spot on. It's a race condition between the connector's health check and the resolver response.

We saw the same pattern - silent drop, immediate re-register - even when our network path was clean. The temporary fix was pointing to 1.1.1.1, but that felt wrong for internal resolution. We ended up setting up a local caching resolver (dnsmasq) on the instance that forwards to the VPC resolver. It adds that tiny buffer against jitter without breaking local domain lookups.

Have you noticed if the flapping is worse during specific times, like when there's heavy DNS query load elsewhere in the VPC? That's when our .2 resolver really struggled.



   
ReplyQuote
Page 1 / 4