That's a really clever solution, using dnsmasq as a buffer. I hadn't thought of that middle ground.
>Have you noticed if the flapping is worse during specific times
We haven't correlated it with VPC DNS load yet, but we do see more drops during our peak business hours, which lines up. It's not massive, just enough to be annoying. Makes me wonder if there's a hidden query the connector does that isn't cached well.
That broker loss/quick re-register pattern is frustratingly common. Since you've already checked the network path, let's look at the connector's own state management.
The health probe can't be tuned directly, but you can indirectly influence it by making the instance's network stack more tolerant. Besides the DNS tweaks mentioned here, check the TCP keepalive settings on the instance itself. The defaults are often too sparse for ZPA's expectations.
Run this on your connector instance:
```bash
sysctl net.ipv4.tcp_keepalive_time net.ipv4.tcp_keepalive_intvl net.ipv4.tcp_keepalive_probes
```
If `tcp_keepalive_time` is 7200 (2 hours), that's likely too long. A temporary test is to set it lower via sysctl, but you'd need to bake it into your launch template user-data for persistence. We saw improvement dropping it to 300 (5 minutes) with more frequent probes.
It's a workaround, but it buys the connection more chances to recover before the watchdog panics.
security by default
The `Connection to Broker lost` error you're seeing is almost certainly the connector's watchdog overreacting to transient network conditions. Since you've already validated the network path, the next step is to systematically rule out the specific failure points others have mentioned, starting with the most likely culprits.
Run these two checks on a live connector *before* it fails over:
1. Query the DNS resolver performance: `dig time @169.254.169.253 zscaler.com` (or your broker domain). Watch for response times consistently over 10ms, especially during your peak business hours.
2. Verify your NAT Gateway's flow idle timeout if egress traffic routes through one. The default is 350 seconds; a misconfigured NACL or routing could cause flows to be considered idle prematurely, dropping the broker connection.
You asked about tuning the health probe. There is no direct way, but you can buffer its dependencies. If the dig test shows high latency, implement a local caching resolver like dnsmasq or Unbound, as user772 suggested. If the NAT timeout is the culprit, you're looking at an architectural change, possibly moving the connectors to a public subnet with EIPs.
Can you share the output of those two checks? The pattern will tell us if you're fighting a DNS/NAT issue or if you've hit the deeper problem of the AMI's watchdog being fundamentally incompatible with AWS's network behavior.
That NLB timeout mismatch is a classic silent killer. It's not just the idle timeout you need to check, but also the TCP flow idle timeout, which is a separate, lower-level setting in an NLB. Even a correctly configured TLS keepalive can get caught by it.
We had a similar issue where the brokers were fine, but an internal Application Load Balancer for metrics scraping had a default idle timeout of 60 seconds. The connector's management API would get a connection reset, and the watchdog would interpret that as a broader failure, triggering a panic reset of the main data plane tunnels. The lesson was to audit *all* load balancer and proxy layers, even those not in the primary data path.
Data over dogma
Our test showed the same pattern during peak VPC DNS load. The dnsmasq buffer helped, but only after we tuned its cache TTL. The default was too aggressive, so it was still hitting the .2 resolver more than we wanted.
Have you measured the actual cache hit rate on your dnsmasq setup? We found ours was low initially.
Interesting point about the cache hit rate being low initially. That makes sense if the default TTL settings don't match the connector's actual query patterns. We saw a similar thing and started logging the dnsmasq queries to figure out which domains were being looked up most frequently. Turned out a couple of non-critical telemetry domains were responsible for most of the fresh lookups, and we could safely increase their cache TTL aggressively.
How did you go about measuring the hit rate? We just grepped the dnsmasq log for queries and cached responses, but I've heard there are better metrics hidden in the stats.
Keep it civil, keep it real.
Right, grepping the logs gives you the raw data but misses the real-time trend. If you start dnsmasq with the `-q` flag, you can actually pipe the queries to a metrics agent. We set up a small Telegraf plugin that counts queries vs. cached answers, but the built-in stats are easier.
If you enable the built-in DNS server stats by sending `SIGUSR1` to the dnsmasq process, it dumps a full cache and query breakdown to its log, including hit rates for each cached domain. You'll want to look for the `cache size` and `queries answered from cache` lines. That's a clearer picture than just counting log lines. The tricky part is correlating the flapping events with a drop in that cached answer percentage.
Yeah, the watchdog sensitivity is a real pain. I ran into something similar last month.
We also had that broker loss error, but our fix was different. We found the default TCP retransmission settings on the EC2 instance were too low. Increasing `net.ipv4.tcp_retries2` from the default 15 to something higher made the connection more forgiving to short packet loss, which seemed to calm the watchdog down. Have you checked your kernel network parameters?
Are you using the Amazon DNS server at .2, or your own resolver? I've heard the .2 server can be slower under load.
That specific pattern - network hiccup triggers full failover despite healthy instance - usually points to the watchdog's internal timing being misaligned with your network's true recovery time. Since you've ruled out CPU/memory, I'd focus on the kernel's TCP stack tuning within the EC2 instance itself, not just the VPC.
The sysctl parameters mentioned by user142 and user217 are correct, but they're part of a broader set. The connector's watchdog often monitors a control TCP socket, and if that socket hits a retransmission limit (`tcp_retries2`) before the network recovers, it declares the broker dead. Increasing `tcp_retries2` buys more time. However, you also need to tighten the keepalive intervals (`tcp_keepalive_time`, `tcp_keepalive_intvl`) so the socket proactively checks for liveness before hitting those retransmission limits. It's a balancing act.
For a quick test, run this on a connector and monitor for a few hours. You'll need to bake this into your launch template.
```bash
# These make the kernel detect a dead socket faster and allow more retries before giving up.
sysctl -w net.ipv4.tcp_keepalive_time=60
sysctl -w net.ipv4.tcp_keepalive_intvl=10
sysctl -w net.ipv4.tcp_keepalive_probes=6
sysctl -w net.ipv4.tcp_retries2=10
```
Did you instrument the control plane traffic to see if the packet loss correlates with the VPC's default DNS resolver at .2? That's often the hidden culprit, even with correct NACLs.
Garbage in, garbage out.
Totally agree about the ephemeral ports in security groups. That one's bitten us before too. Your point about the connection being stateful is so important, it's easy to overlook the inbound side of that rule.
Quick question on your sysctl tweaks - did you find a sweet spot for those keepalive values? We lowered tcp_keepalive_time but then got worried about adding too much overhead. Curious what worked for you.
You're right about the NAT timeout being an architectural problem, but moving to public subnets with EIPs just swaps one bill for another. You're now paying for a dedicated IP per connector, which on a large fleet makes the AWS cost guys twitch.
A more cynical, and often cheaper, fix is to run a small proxy in the private subnet that handles the long-lived connection to the broker. Let the connectors talk to the proxy over a short, stable local link. The watchdog only sees the local connection, and you only pay for one EIP on the proxy.
-- cost first
Ah, the classic `Connection to Broker lost` merry-go-round. Been there, it's a special kind of pain when the box is healthy but the watchdog freaks out.
You've ruled out the infra stuff, so you're right in the territory of tuning the connector's own tolerance. The health probe sensitivity is basically baked into the watchdog's heartbeat and timeout logic. We had some luck adjusting the tunnel keepalive intervals by modifying the service parameters - not officially documented, but you can tweak them in the connector's systemd service file to make the heartbeat more frequent but less likely to declare death on a single missed ack.
Also, for the AWS-specific angle, double-check the instance's *network credit* balance on the burstable instance types, even if CPU is fine. A short network credit exhaustion can cause just enough packet delay to trip that broker connection.
Try everything, keep what works.
Great catch on the network credit angle for burstable instances. That's an often-missed metric that can cause exactly the kind of micro-pause that upsets a sensitive watchdog.
On the service parameter tweaks, we found modifying those heartbeat intervals can backfire if you're not careful. Making them more frequent increases the chance of a packet getting lost in a normal blip, which can actually cause *more* false failovers unless you also adjust the consecutive failure threshold. It's a bit of a balancing act.
Did you have to coordinate those service changes with a specific broker-side configuration, or did the connector adapt on its own?
—daniel
That's a good point about tuning the keepalive. When we tried lowering the time value, we saw fewer broker drops at first. But later we got flagged for having too many connections open because the more frequent probes kept sockets alive longer. Is there a known safe range for these values, or does it just depend on your specific VPC setup?
All this TCP tuning is just bandaids.
The real problem is probably your ASG health check. If it's ELB-based, it's checking a port that can be open while the broker connection is dead. So the ASG never replaces the faulty node, you just keep it alive and flapping.
Switch the ASG to use the EC2 instance status checks. It's slower to fail, but it actually means something. Or add a custom health check that pings the broker directly.