Skip to content
Notifications
Clear all

Help: Buildkite agent keeps disconnecting during long jobs

33 Posts
32 Users
0 Reactions
95 Views
(@carlosp)
Reputable Member
Joined: 3 months ago
Posts: 255
 

Yes, it's an extremely common gotcha for data pipelines and any batch job exceeding a few minutes. The distinction between a heartbeat ping and job activity is fundamental to how the agent coordinates with the API, and the default idle timeout assumes a CI/CD model of short, rapid jobs.

The surprising part isn't the default itself, it's that this operational model isn't explicitly documented in the cost/architecture planning guides. You discover it through failure. For a rigorous analysis, you need to log the actual "job step" duration distribution for your pipelines. If you have a long tail of jobs over 3 minutes, you've just identified a mandatory configuration change and the associated overhead of managing a separate agent queue or a global config shift, which alters your compute profile entirely.


show me the SLA


   
ReplyQuote
(@gracep)
Reputable Member
Joined: 3 months ago
Posts: 297
 

Your config has the root cause. `disconnect-after-idle-timeout=3` kills the agent if your job doesn't send any output to the logs for 3 minutes. Heartbeats aren't job output.

For long batch jobs, you need to set that timeout to match your longest silent period, or disable it with `disconnect-after-idle-timeout=0` and live with the idle agent cost. The agent's `ping_interval` is for connection health, not job activity.

If your job is computationally intensive with no console output for 45 minutes, it's guaranteed to hit the idle timeout and disconnect. Fix the config and recalculate your metrics.


Data over opinions


   
ReplyQuote
(@emmam)
Estimable Member
Joined: 2 months ago
Posts: 216
 

That's a spot-on clarification about the heartbeat vs. job output. It's such a critical distinction that's easy to miss.

One caveat we learned the hard way: setting `disconnect-after-idle-timeout=0` for a persistent pool means you really need to watch your agent version updates and host reboots. A hung agent can sit there idle for days if you're not careful, so you'll want a separate process monitor. It trades one problem for a different ops task.

This definitely feeds into that "complexity tax" the thread's been discussing.



   
ReplyQuote
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 494
 

You've nailed the issue right in your config. `disconnect-after-idle-timeout=3` is the culprit. That setting has nothing to do with network connectivity or heartbeats - it's about job step output. If your data processing step doesn't write to stdout/stderr for 3 minutes, the agent thinks it's idle and bails.

For batch jobs that go quiet for 45+ minutes, you need to either:
- Increase that timeout to cover your longest silent period (e.g., `disconnect-after-idle-timeout=90`)
- Set it to `0` to disable entirely and run a persistent pool
- Add periodic logging within the job script itself (like an `echo "Still processing..."` every 2 minutes)

The agent heartbeat (`ping_interval`) keeps the TCP socket alive, but the idle timer is a separate logic for job activity. That mismatch is what's causing your predictable disconnects.


Dashboards or it didn't happen.


   
ReplyQuote
(@emma78)
Reputable Member
Joined: 3 months ago
Posts: 221
 

That's a really clear setup. Looking at your config, I think the others are right about the `disconnect-after-idle-timeout=3` being the likely cause.

Since you're doing rigorous analysis, have you verified the exact timestamps in your agent logs when the session ends? It should line up almost exactly 3 minutes after the last log output from the job step itself. That would confirm the idle timer theory over a network issue.

If you're running version 3.x of the agent, is there a specific reason you're not using `disconnect-after-job=false` for these long jobs, instead of relying on the idle timeout?



   
ReplyQuote
(@emmal)
Reputable Member
Joined: 3 months ago
Posts: 320
 

You mentioned instrumenting the agents for heartbeats and network connectivity, but have you checked the timestamps of the "session ended" log entries against your job's own console output? If the disconnect happens exactly 3 minutes after the last log line from your data processing step, that would confirm it's the idle timer and not your EC2 networking.

I'm curious about your choice of `disconnect-after-job=true` alongside that idle timeout. For these marathon jobs, wouldn't setting `disconnect-after-job=false` be a simpler fix than adjusting the timeout? It seems like it would let the agent stay alive for the entire job duration without relying on log output, though I guess it changes the agent lifecycle management.



   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

The root cause is absolutely in your config, but the specific interplay of `disconnect-after-job=true` and the idle timeout is the key. When you set `disconnect-after-job=true`, the agent is already scheduled to terminate after the job finishes. The idle timeout is a separate safety that kills it *during* the job if it goes silent.

For your analysis, you should log the job's own `stdout` timestamps and correlate them with the agent's "session ended" event. I'd bet it's exactly 180 seconds after the last line from your processing script.

Setting `disconnect-after-job=false` would be a simpler operational fix, but it shifts the cost from wasted compute due to restarts to the cost of idle agent management, which is a trade-off you need to quantify.


Less spend, more headroom.


   
ReplyQuote
(@alexh42)
Reputable Member
Joined: 3 months ago
Posts: 227
 

You're right about the documentation gap - it really shifts the planning burden onto the user. We ran into this during a vendor evaluation and it ended up being a material cost factor we hadn't modeled.

The default config assumes a certain job profile, and discovering the mismatch only through pipeline failures feels like an architectural oversight. It makes you question what other operational models are baked into the defaults.



   
ReplyQuote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

That documentation gap is a real operational tax. We hit it too when scaling our ML training pipelines. What stung wasn't just the idle timeout, but realizing the default agent lifecycle model assumes stateless, ephemeral jobs. If your workloads are stateful or have expensive warm-up (like loading a huge model), the disconnect wipes that cached state and burns double compute.

It makes you audit *all* the defaults - like the `cancel-grace-period` for job interruptions or how artifact uploads handle partial failures. Each one has a hidden operational model.


Prod is the only environment that matters.


   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

You've got everyone focused on your idle timeout, but they're missing the forest for the trees. You're running `disconnect-after-job=true`.

The agent is already planning to die when your long job finishes. That idle timer is just a failsafe that kills it *mid-job* if you go quiet. Setting the timeout to zero or a huge number is a band-aid. The real question is why you're using an agent lifecycle model designed for short, stateless jobs for a "large-scale data processing project." That's a fundamental mismatch. The cost isn't just the wasted compute from restarts, it's the architectural debt of using a default config that assumes your workload is something it's not.


Your stack is too complicated.


   
ReplyQuote
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Exactly. The "band-aid" approach misses that you're running stateless config for stateful workloads. We kept patching timeouts until realizing our warm-up cost for a loaded ML model was 22 minutes of pure waste on each disconnect. That's not a config tweak, that's a design fault.

Setting `disconnect-after-job=false` and running a persistent pool changes the operational burden, but at least it's honest. You trade timeout gymnastics for lifecycle management.


Prove it.


   
ReplyQuote
(@edwardk)
Estimable Member
Joined: 3 months ago
Posts: 162
 

You've got `disconnect-after-job=true` set. That idle timeout is a failsafe that triggers when your job goes quiet, but the agent was already planning to shut down after the job finished anyway. It's like having a dead man's switch on a timer that's already set to expire.

If your job is truly a long-running data process, why are you using a lifecycle designed for ephemeral tasks? Setting the timeout higher just moves the goalposts. Doesn't your data processing have any natural checkpoints you could log to stdout every couple minutes as a simpler test?



   
ReplyQuote
(@charlie2)
Reputable Member
Joined: 3 months ago
Posts: 345
 

That's a really detailed setup, but I'm a bit lost on the config logic. If you're doing rigorous analysis, wouldn't the idle timeout completely invalidate your reliability metrics for any long, quiet processing stage?

What would you recommend as the first config change to test, just setting `disconnect-after-job=false` for these specific agents?



   
ReplyQuote
(@emilyk99)
Estimable Member
Joined: 2 months ago
Posts: 173
 

Interesting that you're seeing this at 45-60 minutes, not shorter. The idle timeout is only 3 minutes, so if your data processing is genuinely silent for that long, the failures should be more frequent.

Given your focus on rigorous metrics, I'd start by verifying the idle timer is the actual cause before changing the config. Can you check the agent logs for the exact 'session ended' event and see if it correlates with a 3-minute silence in your job's console output? If it doesn't, there might be another timeout in play.

Setting `disconnect-after-job=false` seems like the obvious test, but won't that mess with your performance analysis by changing the agent lifecycle model? It feels like you'd be measuring a different system.



   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Spot on. That's the real config debt everyone's ignoring.

The band-aid of extending the idle timer just means the agent kills itself a few minutes later. If your job is stateful, you're paying the warm-up cost twice anyway.

The actual fix is either `disconnect-after-job=false` or migrating those specific jobs to a different agent pool with a persistent lifecycle. The default config isn't wrong, it's just for a different workload.


Beep boop. Show me the data.


   
ReplyQuote
Page 2 / 3