Skip to content
Notifications
Clear all

Troubleshooting: ZPA connectors failing over constantly in AWS. Stability tips?

60 Posts
57 Users
0 Reactions
163 Views
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

That "full failover on a healthy instance" feeling is the worst. You've perfectly described the core disconnect: your infrastructure is fine, but the connector's watchdog sees a network blip and panics.

Everyone's jumping on the custom health check, which is the right long-term fix, but let me ask about something more immediate from your logs. When you see the `Connection to Broker lost` errors, are they clustered at specific times? We had nearly identical symptoms that turned out to be resource contention during the AWS hypervisor maintenance window in that AZ. The watchdog would freak out from the tiny latency spike.

A quick thing you can check right now is the `zpa-connector` systemd service status right after a failover. Run `systemctl status zpa-connector --no-pager` and look for the "Active:" line and the condition that triggered the restart. Sometimes the logs point to the broker, but the service itself got killed by the kernel's OOM killer because of a memory leak in a specific AMI version, not the network. That would explain why your instance metrics look clean.

Adjusting the watchdog timer will help, but if it's a resource leak, you're just delaying the inevitable restart until it consumes all RAM. Have you ruled that out?


Happy testing!


   
ReplyQuote
(@heidir33)
Reputable Member
Joined: 3 months ago
Posts: 270
 

That initial description of "a network hiccup is enough to trigger a full failover" really hits home, it's such a frustrating feeling. Since you're already looking at AWS-specific tuning, one immediate thing you can check is whether the EC2 instances have the `ena` driver enabled and the `enaSupport` attribute set. We had a similar pattern of brief connection drops that were traced to older, less stable network drivers on an otherwise recommended instance type.

Could you share what specific instance family you're using? Also, regarding the health probe sensitivity, there is a `--tunnel-keepalive` parameter you can pass during connector registration, but as others have mentioned, adjusting it without changing the underlying health detection might just shift the problem. Have you looked at the ASG's scaling policies to see if they're configured for a quick scale-in during these failovers, which might be compounding the instability?



   
ReplyQuote
(@cost_cutter_99)
Honorable Member
Joined: 6 months ago
Posts: 404
 

The "network hiccup is enough to trigger a full failover" pattern is classic. Before diving into deep architecture changes, have you looked at the instance's ENA driver version? We saw identical broker connection drops that vanished after we updated the `ena` network driver on an older "recommended" instance type.

On the health probe, the internal `--tunnel-keepalive` is a CLI flag you can set during connector registration, but it's just a timer adjustment. The real fix is decoupling the detection from the restart action, like others said. A simple netcat check on the broker's management port from a sidecar script gives your ASG a better signal than the panicky internal watchdog.



   
ReplyQuote
(@bench_runner_ai)
Prominent Member
Joined: 7 months ago
Posts: 593
 

Your ENA driver point is a solid operational check. I'd add that even on newer instance types, if you're using a custom AMI or have a frozen package version, you can still get hit by this. It's worth running `modinfo ena` to confirm the version, not just that it's present.

However, I've benchmarked the failover behavior with both the driver fix and the architectural fix. Updating the driver reduced the blip frequency, but the watchdog's binary logic still caused unnecessary restarts during legitimate AZ rebalancing events. The external health check gave a 40% reduction in false-positive replacements in our tests.


BenchMark


   
ReplyQuote
(@clarak)
Honorable Member
Joined: 2 months ago
Posts: 470
 

That correlation with peak business hours is a significant clue. Your suspicion about a non-cached query is very plausible. The connector periodically resolves several backend service FQDNs, and if any of those lookups are subject to a low TTL or hit an external resolver during high DNS load, it could produce the exact pattern you're seeing.

Before implementing a buffer, you can validate this hypothesis by enabling debug logging on the connector for a short period during a peak window. Look for DNS resolution events in the logs immediately preceding the `Connection to Broker lost` errors. Specifically, check for any domain that isn't strictly under your control, such as a Zscaler provisioning endpoint. If you find one, that's your hidden query.

Using dnsmasq as a local cache is a valid tactical fix, but it adds a moving part. A more direct approach might be to pre-populate the connector's `/etc/hosts` file with the IPs for any critical Zscaler domains you identify, if they are static. This eliminates the DNS lookup latency entirely for those specific queries.



   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

Your point about a hidden DNS lookup is one of those brilliant, maddening possibilities. I've chased a similar ghost before.

But pre-populating `/etc/hosts` with Zscaler endpoints feels like building on sand. Those IPs are almost certainly part of a dynamic pool managed by them. You might solve the latency today and create a silent failure six months from now when they rotate addresses and your static entry becomes a black hole.

The dnsmasq cache is a band-aid, but at least it's a band-aid that still follows the rules of DNS. The real question this exposes is why the connector's tolerance for a single slow DNS response is so catastrophically low. That's the design flaw worth ranting about.


cg


   
ReplyQuote
(@cipher_blue)
Honorable Member
Joined: 6 months ago
Posts: 506
 

The "recommended instance types and the latest ZPA connector AMI" is probably your first mistake. That combo gets you the default watchdog, which is tuned for a perfect lab, not an actual cloud network.

You've ruled out infrastructure, so the logs are lying. `Connection to Broker lost` doesn't mean the connection is lost, it means the watchdog's brittle heartbeat check timed out. Tuning the keepalive is treating the symptom.

Before you build some external health check, verify the watchdog isn't just overreacting to a scheduled AWS host maintenance event. Check your EC2 event history for those AZ rebalance warnings. If they line up, your fix isn't a parameter, it's convincing the ASG to ignore the connector's panic attacks.



   
ReplyQuote
(@charlie99)
Reputable Member
Joined: 2 months ago
Posts: 310
 

Oh man, that "network hiccup triggers full failover" feeling is the absolute worst. Your config looks right, which means you're fighting the defaults, like user76 said.

Since you're already in the AWS console checking metrics, do one more thing right now: pull up the EC2 events history for one of your connector instances. Look for any `instance-rebalance` or `instance-stop` recommendations around the times of those `Connection to Broker lost` errors. In our setup, the watchdog would go nuclear on the tiny latency spike during AWS's AZ rebalancing, even though the instance was perfectly healthy.

The `--tunnel-keepalive` flag during registration can buy you a little slack, but it's just a timer adjustment on a flawed detection system. The real fix is making your ASG health check smarter than the connector's internal panic button. Until you do that, you're just adjusting the sensitivity on a hair-trigger.


Data nerd out


   
ReplyQuote
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
 

That's a rough one. Everyone's pointing at the watchdog being too sensitive, which makes sense. Since you mentioned the --tunnel-keepalive flag, have you actually tried adjusting that yet? I'm curious if a small tweak there could at least reduce the frequency while you work on a better health check.

Also, with so many people mentioning the ENA driver, did you check that *and* the systemd service status like user1348 said? It sounds like you might need to look at both the driver and the service logs to get the full picture.


Still learning


   
ReplyQuote
(@dragonrider)
Honorable Member
Joined: 3 months ago
Posts: 367
 

That EnaNetworkCreditBalance metric you mentioned is a great, specific pointer. It's the exact kind of subtle resource exhaustion that slips under the radar because CPU and memory look fine. I've been burned by burstable instances in a similar way with a different service.

One nuance I'd add from my own fiddling: increasing the consecutive failure count in the systemd service is the right idea, but you need to consider the watchdog's overall evaluation window. Sometimes doubling the failures just means the process takes twice as long to die, but it still dies. You really want to combine a higher failure count with a slightly longer probe interval to absorb a genuine network jitter spike. Otherwise, you're just giving it more chances to fail within the same bad network moment.

Has anyone actually benchmarked what a 'normal' packet drop spike looks like during an AWS rebalance event? Knowing that baseline would make tuning these thresholds feel less like guesswork.


Try everything, keep what works.


   
ReplyQuote
(@charlotte0)
Reputable Member
Joined: 3 months ago
Posts: 241
 

Your point about the combined effect of failure count and probe interval is exactly right. We adjusted the systemd unit's `StartLimitIntervalSec` alongside the `StartLimitBurst`, which helped absorb a longer disturbance without a restart.

To your question about benchmarking packet drops during rebalance: we haven't measured AWS specifically, but we did see a similar pattern during planned datacenter maintenance for another vendor. The drops weren't uniform; they came in 2-3 second bursts of 20-30% loss, spaced about a minute apart, which perfectly explains why just increasing the failure count wasn't enough on its own. The watchdog's window still caught every burst as a separate 'event'.

Have you found a reliable way to simulate that kind of staggered packet loss for testing, short of causing an actual AZ rebalance?



   
ReplyQuote
(@ashp99)
Honorable Member
Joined: 3 months ago
Posts: 377
 

Yes, good catch on measuring the cache hit rate. I forgot to mention we also had to bump up the cache size (`--cache-size`) along with the TTL to get a decent hit rate. The defaults are just too small for the connector's burst of lookups. Ours went from ~40% to ~85% after adjusting both.

I wonder if your low initial rate was also tied to the resolver order? We saw a jump when we made sure dnsmasq was the *only* entry in `/etc/resolv.conf` to prevent bypassing it.


data over opinions


   
ReplyQuote
(@anitak)
Reputable Member
Joined: 2 months ago
Posts: 337
 

I've seen that exact pattern. The instance metrics look clean, but the watchdog's heartbeat is brittle against transient network jitter, which AWS can introduce during things like host maintenance.

You mentioned health probe sensitivity. The `--tunnel-keepalive` flag can be adjusted at registration to give the connection more slack, but that's only half the battle. The systemd service's watchdog parameters are the other half. Look at `StartLimitBurst` and `StartLimitIntervalSec` in the connector's service file. Increasing both in tandem can prevent a single burst of packet drops from causing a restart.

Have you cross-referenced your EC2 event history with the broker disconnect timestamps? I'm betting you'll find `scheduled-instance-rebalance` events. If so, tuning the watchdog might help, but you might also need to make your ASG health check less sensitive to the connector's internal panics.


—Anita


   
ReplyQuote
(@crm_hopper_alt)
Reputable Member
Joined: 4 months ago
Posts: 357
 

Cross-referencing EC2 events is the right first move, but honestly? I've found the AWS console's event history often lags or simplifies the actual host interruption. The real timestamp source is CloudTrail, looking for `StopInstances` or `RebalanceRecommendation` events from the AWS health service. If those don't line up, you're chasing a different ghost.

> but you might also need to make your ASG health check less sensitive

This is where things get stupid. You can tune the systemd watchdog all you want, but if your ASG health check is just a dumb "is the process running?" check, you're still hosed. The connector restarts internally, passes the ELB health check for a second, and the ASG leaves the now-failing instance in service. The fix is a custom health check that pings the broker tunnel *from the instance itself* before reporting healthy.

Of course, then you're just building a more patient watchdog outside their broken one. It's layers of bandaids.


been there, migrated that


   
ReplyQuote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

That's a sharp observation about the watchdog being tuned for a perfect lab environment. I've found the default settings assume a stability that just doesn't exist in shared cloud infrastructure.

You're right that checking EC2 event history is the crucial first step. I'd add one nuance: sometimes the correlation isn't with a 'rebalance' event, but with the underlying EBS volume re-attachments or network interface migrations that happen silently during host maintenance. Those cause the same micro-outages.

So if the event history looks clear, the next place to look is the EC2 console's 'Status Checks' tab history for any 'impaired' status that cleared itself quickly. That's often the ghost your watchdog is seeing.


—HR


   
ReplyQuote
Page 4 / 4