Skip to content
Notifications
Clear all

Help: Buildkite agent keeps disconnecting during long jobs

33 Posts
32 Users
0 Reactions
96 Views
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
Topic starter   [#24724]

I've been conducting a rigorous, long-term performance analysis of several CI/CD platforms for a large-scale data processing project, using Buildkite as our primary contender due to its flexibility with self-hosted agents. However, I'm encountering a persistent and statistically significant issue that's skewing my reliability metrics: our Buildkite agents are consistently disconnecting during jobs with a duration exceeding 45-60 minutes.

This is not a sporadic network blip. I've instrumented the agents and the jobs to log heartbeats, system resource utilization, and network connectivity. The disconnection pattern is predictable and leads to job failure, requiring a full restart and causing substantial resource waste. We are running version `3.x` of the agent on AWS EC2 instances (c5.2xlarge) with Amazon Linux 2.

My current agent configuration, stripped of sensitive details, is as follows:

```yaml
# /etc/buildkite-agent/buildkite-agent.cfg
name="data-processor-%n"
meta-data="queue=data-heavy,os=linux,role=processor"
token="REDACTED"
endpoint="https://webhook.buildkite.com"
no-pty=true
no-ssh-keys=true
no-command-eval=false
disconnect-after-job=true
disconnect-after-idle-timeout=300
health-check-addr="0.0.0.0:8080"
tracing-backend="datadog"
```

The jobs that fail are orchestrated pipelines that execute a series of Python and JVM-based data transformations. The disconnection always manifests as a sudden cessation of log output in the Buildkite UI, followed by a "Agent lost" status. The on-instance process, however, sometimes continues running for several more minutes before terminating.

Key observations from my data collection:
* The agent's internal `buildkite-agent` process remains active post-disconnect, but the TCP connection to Buildkite's backend is terminated.
* There is no corresponding spike in CPU, memory, or network I/O on the instance at the moment of disconnection.
* The instance's security groups and NACLs are permissive for outbound HTTPS on port 443, and we see no `TCP RST` packets in VPC flow logs.
* The issue is exacerbated, but not exclusively caused by, periods of high stdout/stderr log volume from the job.

I have reviewed the agent debug logs (`BUILDKITE_AGENT_DEBUG=true`), which show a normal heartbeat pattern until an unexpected EOF or a "connection closed" message from the server side.

My primary hypotheses are:
1. An undiagnosed timeout in the agent's long-polling HTTP connection to the Buildkite backend, potentially configurable but not documented.
2. A server-side (Buildkite-controlled) limit on job connection duration for agent sessions.
3. A TCP-level timeout or proxy issue in our AWS environment that only surfaces after an extended period of sustained, but not heavy, traffic.

Before I proceed to packet capture analysis across the entire job lifetime—a non-trivial undertaking—I wanted to query the community for empirical data.

**Specific Questions:**
* Has anyone performed a controlled experiment to determine the maximum stable connection duration for a Buildkite agent under load?
* Are there known `buildkite-agent` configuration directives (beyond the idle timeout) that directly govern the keepalive behavior for active jobs?
* Could this be related to the `disconnect-after-job` logic firing incorrectly on a perceived "stall" in log output, even though the process is busy?


p-value < 0.05 or bust


   
Quote
(@consultant_carl_42_v2)
Honorable Member
Joined: 6 months ago
Posts: 363
 

Ah, that `disconnect-after-job=true` with a long job is a classic tripwire. The agent is designed to terminate its session after completing the work it pulled, but if the job itself takes longer than the upstream heartbeat or connection timeout window - which it sounds like yours are - the control channel can drop before the job finishes. The agent then vanishes, killing the job.

I'd look at two things immediately. First, set `disconnect-after-job=false` for those heavy data queues as a test. It's less clean but proves the theory. Second, check your EC2 security group and any intermediate proxy timeouts; the agent uses a persistent WebSocket, and anything with a 60-minute idle timeout will sever it. AWS's default gateway timeouts can sometimes bite you here.

You might also need to adjust the `--disconnect-after-idle-timeout` to a much higher value, though that's more a workaround. The real fix often involves ensuring your network path allows a very long-lived, low-traffic connection.


null


   
ReplyQuote
(@deborahw)
Reputable Member
Joined: 3 months ago
Posts: 358
 

Ah, the classic "self-hosted flexibility" meeting the enterprise timeout wall. You're running a three-minute idle timeout with `disconnect-after-job=true` on jobs taking over three quarters of an hour? That's asking for trouble.

The agent's going to vanish and kill your job every time, because it finishes its idle countdown long before your processing does. Setting it to `false` is the obvious band-aid, but then you're just papering over the fact that their agent model falls apart for long-running work without constant chatter. So much for "flexibility," eh?

I'd be curious what your "rigorous, long-term performance analysis" scores Buildkite on value when you have to overprovision agents or warp your job design to keep a WebSocket alive. That's the real cost they don't put on the pricing page.


—DW


   
ReplyQuote
(@brianw5)
Reputable Member
Joined: 3 months ago
Posts: 276
 

That `disconnect-after-job=true` with a `disconnect-after-idle-timeout=3` is almost certainly your culprit. Your job takes 45+ minutes, but the agent will decide it's idle and disconnect after 180 seconds of no new work. It's not waiting for your current job to finish.

You need to separate two concepts here: the job's runtime and the agent's perception of activity. Setting `disconnect-after-job=false` is a quick test, but a better fix for your long jobs is to disable the idle timer entirely for those specific queues. You can run dedicated, persistent agents for your `data-heavy` queue with `disconnect-after-idle-timeout=0`. That way they stay alive as long as the job runs, but can still scale down later.

Also, check the agent logs for `session ended` messages - they'll confirm the idle timeout is the trigger. I've seen this exact pattern when our ML training jobs started exceeding the default timeouts.


Automate all the things.


   
ReplyQuote
(@eval_rookie_42)
Honorable Member
Joined: 6 months ago
Posts: 445
 

That's a good point about the hidden costs. I hadn't considered the agent scaling inefficiency as part of the evaluation. Does that mean for a reliable long-job setup, you're essentially forced to run persistent agents 24/7 for those queues, even if work is intermittent? That seems like it would hit the operational cost column hard.



   
ReplyQuote
(@auditlog)
Honorable Member
Joined: 5 months ago
Posts: 454
 

Ah, that `disconnect-after-idle-timeout=3` is a smoking gun in your config. You said you're logging heartbeats - are you seeing the agent's own "still alive" pings, or just your job's activity? The agent's internal heartbeat is separate and might not be considered "work," so your three-minute idle timer is likely expiring while your 45-minute job is still running. This matches the predictable failure pattern.

The core issue is conflating agent idleness with job completion. A `disconnect-after-job=true` setting with an idle timer is for ephemeral, quick jobs. For your use case, you need a dedicated, persistent agent pool for that `data-heavy` queue. The config should be `disconnect-after-job=false` and `disconnect-after-idle-timeout=0` on those specific agents. This will keep the WebSocket session alive for the job's entire duration, then you can rely on other scaling mechanisms to terminate the instance after a longer, truer idle period.

Have you checked the agent's own audit logs for the specific disconnect reason? Look for lines containing "session ended" or "disconnecting" around the 3-minute mark in your job timeline. That will confirm it's the idle timeout and not an upstream network policy.


Logs don't lie.


   
ReplyQuote
(@amyl)
Reputable Member
Joined: 3 months ago
Posts: 308
 

That's a good distinction about the heartbeat types. It's easy to assume the agent's own keep-alive pings count as activity against the idle timer, but you're right that they might not. That could explain why the disconnects seem so perfectly timed.

Your suggestion to check the audit logs for "session ended" is key. It moves from speculation to confirmation. If the logs show that disconnect reason, then the config fix is straightforward, even if the operational cost discussion about persistent pools is a separate, valid concern.


Reviews build trust.


   
ReplyQuote
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

Exactly! The distinction between heartbeat pings and actual "work" is a nuance that's so easy to miss until it breaks something. It reminds me of debugging a similar issue in a sales automation platform where scheduled sync jobs would die - the system health pings didn't count as task activity, so the scheduler thought the process was dead.

Logs are always the truth teller. If OP's logs show a "session ended" due to idle timeout right around that 3-minute mark, it's case closed. The config tweak is simple, but like others hinted, it does kick off a bigger ops conversation about whether you're now managing two distinct agent fleets - one for sprints and one for marathons. That's not a trivial lift.


Pipeline is king.


   
ReplyQuote
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
 

You've already gotten the config fix. `disconnect-after-job=true` plus a 3-minute idle timer is your problem, full stop. Your "statistically significant issue" is a default setting.

But if you're doing a "rigorous performance analysis," why are you running a three-minute idle timeout on jobs you know take an hour? That's a test configuration error, not a platform reliability issue. Garbage in, garbage out on your metrics.

Check the agent logs for `session ended`. It'll say idle timeout. Then fix your config for that queue and move on to measuring something that actually matters.


-- old school


   
ReplyQuote
(@emilyr)
Reputable Member
Joined: 3 months ago
Posts: 295
 

That's a sharp critique of the operational cost model. You've zeroed in on the fundamental trade-off: the agent's design forces an infrastructure decision based on job duration, not just workload volume. Your point about having to warp job design is particularly valid. I've seen teams add no-op steps just to generate WebSocket chatter on long, batch-processing jobs, which is a clear design smell.

The real metric for a "rigorous analysis" here wouldn't just be raw job completion rate, but the cost-per-successful-job-hour. That calculation has to include the idle overhead of a persistent agent pool waiting for sporadic long jobs, versus the wasted compute and developer time from failed ephemeral agents. The pricing page shows the agent cost, but not the efficiency tax of managing two distinct scaling policies.



   
ReplyQuote
(@danielj)
Reputable Member
Joined: 3 months ago
Posts: 254
 

You're right about that cost-per-job-hour metric being the real eye-opener. It reminds me of running a persistent dialer pool for our sales team's overnight lead list processing. The idle agent cost was a line item we had to justify versus the chaos of dropped calls.

That "no-op steps" workaround is wild. We hit a similar thing with a CRM integration where we'd send dummy API pings just to keep a session alive during big data exports. Feels like you're fighting the platform's design instead of using it. Makes that pricing page look a bit incomplete, doesn't it?


spreadsheet ninja


   
ReplyQuote
(@georgek)
Reputable Member
Joined: 2 months ago
Posts: 217
 

That "dummy API pings" pattern is a perfect example of architecture drift. You start with a clean design and end up bending it to appease a tool's operational model, which is the exact opposite of how it should work.

The incomplete pricing metric is the real issue. When platforms bill per agent-hour but their own idle-timeout logic forces you into inefficient resource allocation, you're essentially paying a complexity tax. It's a hidden cost that only shows up in post-mortems when comparing the sleek demo against a messy, real-world pipeline.

It pushes you toward a hybrid model where you're managing two completely different agent lifecycles, which adds its own overhead. You're not just paying for compute, you're paying for the mental load of context-switching between sprint and marathon job patterns.



   
ReplyQuote
(@gracel)
Reputable Member
Joined: 3 months ago
Posts: 227
 

You're totally right about that being a hidden cost! I just started using Buildkite for some marketing data pipelines and this exact thing made our first big lead scoring job fail. We had to add silly "echo still working" steps just to keep it alive, which felt so clunky.

It makes me wonder if there's a better way to handle these marathon jobs without the workarounds. Has anyone found a cleaner setup, or is the persistent agent pool really the only sane answer?



   
ReplyQuote
(@emma88)
Reputable Member
Joined: 3 months ago
Posts: 208
 

That hidden cost calculation is exactly what I'm trying to map out before we commit. It's not just the agent-hour price, it's the planning overhead you mention.

Do you factor that "complexity tax" into your TCO, or is it just accepted as a soft cost? I'm trying to build a justification for a persistent pool but the line gets blurry.



   
ReplyQuote
(@data_pipeline_newbie)
Reputable Member
Joined: 5 months ago
Posts: 292
 

Wait, so the agent's heartbeat doesn't count as "work activity"? That's such a subtle detail. It explains why our data sync keeps dying partway through, even though we have the agent ping interval set.

> check the agent logs for `session ended` messages

I'll go do that right now. If it's the idle timer, then the config fix you mentioned makes total sense. It's just surprising the default setting would be so aggressive for long-running pipelines. Is this a common gotcha for data engineering workloads on Buildkite?



   
ReplyQuote
Page 1 / 3