Skip to content
Notifications
Clear all

Help: Buildkite agent keeps disconnecting during long jobs

33 Posts
32 Users
0 Reactions
94 Views
(@elenab)
Estimable Member
Joined: 2 months ago
Posts: 202
 

You've perfectly described the hidden tax of mismatched defaults. That 22 minute warm-up cost is the kind of figure that should be in bold on the vendor's pricing page, not something you discover through trial and error.

The audit point is critical. We found the same with `cancel-grace-period`. Its default assumes your job can be killed instantly without cost, which is fine for a linter but financially reckless for a GPU instance halfway through a render. Each of these parameters encodes a financial assumption about your workload.

It forces you to build a shadow configuration guide, because the vendor's defaults are a one-size-fits-none model optimized for their demo cases, not your actual production costs.


show me the tco


   
ReplyQuote
 danw
(@danw)
Reputable Member
Joined: 3 months ago
Posts: 387
 

The complexity tax is never soft. It's real engineering hours you could have spent elsewhere.

We build it into TCO as a multiplier: each mismatched default adds 20% to operational burden. Persistent pools don't just change cost, they change ownership. You're now responsible for uptime, patching, and scaling.

That's the justification. Is your team ready to own the pool, or are you just trying to avoid vendor timeouts?



   
ReplyQuote
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

That's a super clean config you've shared. I'm zeroing in on the `disconnect-after-idle-timeout=3`. That three-minute timer starts ticking during any quiet period in your job's output, not from the job's start. Your data processing likely has phases where it's crunching numbers silently for longer than 180 seconds.

The first thing I'd do is add a simple heartbeat to your job script, something like `echo "[$(date)] Still processing..."` every 60 seconds, just to test if the disconnections stop. It's a hack, but it proves the idle timer is the culprit before you overhaul your agent lifecycle.

If that works, then the real decision is between that perpetual heartbeat and switching to `disconnect-after-job=false` for a dedicated, persistent pool.


Pipeline is king.


   
ReplyQuote
Page 3 / 3