Skip to content
Notifications
Clear all

Troubleshooting: High CPU on the Linux connector in Docker - any config tweaks?

11 Posts
11 Users
0 Reactions
27 Views
(@cloud_ops_amy)
Honorable Member
Joined: 7 months ago
Posts: 453
Topic starter   [#21803]

Hi everyone, I've been running Twingate's Linux connector in Docker for a few months now, and overall it's been great for our zero-trust setup. Lately though, I've noticed the container process is consistently using high CPUβ€”often sitting at 30-50% even with minimal network traffic. This is on a t3.medium EC2 instance that's otherwise idle.

Has anyone else run into this and found any configuration tweaks or Docker settings that help? I'm wondering if it's related to the health check intervals, logging verbosity, or perhaps a known issue with a specific version.

My current setup is pretty standard:
```yaml
version: '3.8'
services:
twingate-connector:
image: twingate/connector:latest
container_name: twingate-connector
restart: unless-stopped
environment:
- TWINGATE_ACCESS_TOKEN=${ACCESS_TOKEN}
- TWINGATE_REFRESH_TOKEN=${REFRESH_TOKEN}
- TWINGATE_NETWORK=${NETWORK_ID}
- TWINGATE_LOG_LEVEL=info
ports:
- "443:443"
```

I've already tried pinning to a specific tag instead of `latest` and setting `TWINGATE_LOG_LEVEL=warn`, but the CPU usage didn't change much. I'm curious if adjusting the health check frequency or the resource limits in Docker might be the next step. Any insights or experiences would be super helpful!

-- Amy


Cloud cost nerd. No, I don't use Reserved Instances.


   
Quote
(@harperk)
Honorable Member
Joined: 3 months ago
Posts: 537
 

That log level change you tried is a red herring; it mostly just throttles what gets written to disk, not the internal processing overhead. The culprit is almost always the default health check probing, which can get aggressive.

You're on the right track pinning a version, but the better move is to explicitly set the health check intervals with something like `--health-interval=30s` in your docker run or the compose healthcheck definition. The default is often much faster than needed for a stable deployment.

Also, check if you're on a kernel that's had issues with TLS handshake overhead in userspace networking modes. I've seen similar spikes vanish by forcing the connector to use the host network (`network_mode: host`), though that's a bigger config change.


Data over dogma.


   
ReplyQuote
(@garethp)
Estimable Member
Joined: 3 months ago
Posts: 226
 

The network mode switch can indeed reduce overhead, but I've seen it introduce routing complications in multi-NIC environments. If you go that route, verify your outbound traffic paths match expectations, especially if you're using custom iptables rules.

The health check interval adjustment is valid, though I'd recommend correlating it with your actual recovery time objectives. If you're running in a high-availability pair, shorter intervals might be necessary despite the CPU cost.

One additional factor I've observed is memory pressure triggering frequent garbage collection cycles in the connector's runtime. It's worth checking if your container has a memory limit set that's too restrictive, forcing aggressive cleanup.


Plan the exit before entry.


   
ReplyQuote
(@freddiem)
Reputable Member
Joined: 3 months ago
Posts: 295
 

Good call on the memory pressure angle - it's easy to overlook. I've seen the same garbage collection churn when the connector is crammed into a tiny memory limit.

If you're running in Docker, check if you have a `mem_limit` set from an old compose file. The default in newer Docker versions is usually fine, but if you've capped it below 512MB, you're asking for trouble. The connector's JVM needs a bit of breathing room for its heap.



   
ReplyQuote
(@emilyr)
Reputable Member
Joined: 3 months ago
Posts: 295
 

The log level and tag pinning attempts confirm the issue isn't superficial, but the provided compose snippet is missing the critical `healthcheck` and `cpuset` parameters, which are the primary levers here. The default health check likely runs every few seconds, causing constant probe evaluation overhead.

You should define a custom health check with a significantly longer interval, like 30s, and a longer timeout. Combine that with pinning the container to specific CPU cores using `cpuset` to prevent it from bouncing across vCPUs and incurring context switch penalties, which is particularly impactful on AWS's t-series instances.

Also, add `--cpus 0.5` or similar in your compose under `deploy.resources.limits` to directly throttle the container's CPU share. The high idle usage you're seeing suggests the process isn't being constrained by the scheduler, allowing it to consume more cycles than it strictly needs.



   
ReplyQuote
(@annas)
Honorable Member
Joined: 3 months ago
Posts: 542
 

Pinning with `cpuset` is a heavy-handed fix that treats the symptom, not the disease. It can create artificial bottlenecks if your traffic profile changes and the pinned core becomes saturated while others sit idle. The kernel scheduler is generally smarter than we are.

The real problem is what the health check is actually *doing*. Setting a longer interval just spreads out the pain. You need to check what the default probe command is; if it's something expensive like a full authentication handshake, then even a 30-second interval is too often. You should override it with a lighter check, perhaps a simple port probe, before you start messing with CPU affinity.

Also, blindly using `--cpus 0.5` as a throttle is a good way to introduce latency spikes under load. You're capping total processing time, not smoothing utilization. If the process has a burst of work, it'll just take longer, which might violate timeouts elsewhere in your stack.



   
ReplyQuote
 danf
(@danf)
Estimable Member
Joined: 2 months ago
Posts: 168
 

Everyone's jumping straight to the container runtime, but you're running on a t3.medium. Those are burstable instances. Your high idle CPU is likely the connector hitting the baseline and tapping into credits, not a health check. Have you checked your CloudWatch CPU credit balance? If you're constantly burning credits at idle, the instance is under-provisioned for the base workload. Throwing config tweaks at a resource problem is a waste of time.


Anecdotes aren't data.


   
ReplyQuote
(@harukik)
Honorable Member
Joined: 3 months ago
Posts: 400
 

Oh, that's a really good point about the burstable credits. I hadn't even considered that. I just assumed the CPU usage reading was real.

So if I'm seeing a high CPU % in `docker stats`, but it's actually just eating into credits on a t3.medium, does that mean the process is still using more baseline CPU than it should? Or is the percentage itself kinda misleading in that scenario?



   
ReplyQuote
(@alexb)
Reputable Member
Joined: 3 months ago
Posts: 257
 

Yeah, the percentage itself is misleading in that scenario. `docker stats` shows the raw CPU time the container is using, but on a burstable instance, that high percentage might be okay if you have credits to burn. The question is whether you're consistently draining your credit balance down to zero.

If your credits are stable or refilling overnight, you're fine. But if you're constantly hovering near zero credits, then yes, the connector is using more *sustained* CPU than your instance type's baseline can handle. In that case, you've got a real resource issue, not just a misleading metric.

I'd check CloudWatch for `CPUCreditBalance` over, say, a week. That graph tells you everything.


Data > opinions


   
ReplyQuote
(@benjaminc)
Reputable Member
Joined: 3 months ago
Posts: 246
 

That's a good point about the burstable credits being misleading. But even if the instance can handle the sustained load with credits, isn't the high usage still a cost signal? A t3.medium has a baseline of 20%, so a container idling at 30-50% means you're paying for the burst capacity you're using. That seems like a configuration problem worth fixing, even if the instance isn't crashing.



   
ReplyQuote
(@dianar)
Honorable Member
Joined: 3 months ago
Posts: 487
 

Your compose snippet cuts off. Post the full one, especially any `healthcheck` or `deploy.resources` sections you didn't include. Without that, all the advice here is speculative.

The burstable instance angle is valid for cost, but a container idling at 30-50% CPU on an otherwise idle host points to internal churn, not just a cloud billing quirk.

First, rule out the health check. The default might be doing something heavy. Set a custom one with `interval: 60s` and a trivial command like `CMD-SHELL echo 1` as a test. If CPU drops, you found the source.

If not, move past the generic Docker advice. You need to see what the process is actually doing inside the container. Use `perf` or `strace` to sample it. It's likely constant TLS renegotiation, excessive metric scraping, or a tight loop in a log statement you muted but didn't eliminate.


Five nines? Prove it.


   
ReplyQuote