Skip to content
Notifications
Clear all

My results after enforcing a 10-minute timeout on all jobs

25 Posts
24 Users
0 Reactions
21 Views
(@benjislack)
Reputable Member
Joined: 2 months ago
Posts: 244
Topic starter   [#25489]

Everyone talks about optimizing compute minutes, but no one talks about the easiest lever: just killing jobs.

We forced a 10-minute hard timeout on every job in our GitHub Actions workflows. No exceptions for now.

Results from last month:
* 23% of all jobs were terminated.
* Saved ~$850 on our GitHub Actions bill.
* Main casualties: flaky integration tests, docker builds with bloated contexts, and "analysis" jobs that hung.

The noise was immediate. Engineers complained about "broken builds." Told them to fix their scripts. The flaky tests that used to pass by eventually finishing now fail fast. That's a win.

It's not elegant, but it forces efficiency. You quickly learn what's actually necessary. Next step is applying the same logic to our Atlassian Bamboo instances, but those contracts are a harder fight.


your mileage will vary


   
Quote
(@carols)
Estimable Member
Joined: 2 months ago
Posts: 142
 

The cost savings are compelling, but have you isolated the impact on developer productivity? There's a trade-off between the $850 saved and the cumulative hours engineers spent diagnosing timeouts versus actual work. A blanket policy can mask inefficiencies in the underlying runner specs or job parallelization.

>flaky tests that used to pass by eventually finishing now fail fast

This is the real benefit. It turns a reliability cost into a visible, actionable ticket. Consider logging the terminated jobs by type and owner. That data is crucial for justifying the policy long-term and identifying which jobs genuinely need an exception. Without that, you're just shifting cost from the cloud bill to engineering time.


Buy once, cry once.


   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 363
 

Interesting approach! The cost savings are impressive. I'm curious, did you have any pushback about security scans or code quality checks that might take longer? Those always seem to be the ones that creep up on us.

Also, logging the terminated jobs like user1393 mentioned seems like a smart next step. It would help show the team which scripts need love the most. Did you set up any kind of alert for the owners when their job times out?



   
ReplyQuote
(@gracep)
Reputable Member
Joined: 3 months ago
Posts: 297
 

Good. You've hit the core benefit: forcing a failure mode.

The next step is metrics. Log the job name, owner, and runtime at termination. Otherwise you're just creating churn without data.

We did this. The histogram of job runtimes (before the cap) showed 80% finished under 5 minutes. The long tail was the problem. It gave us the proof to carve out specific, justified exceptions for the 2% that were genuinely necessary, like a compliance scan. The rest got fixed or removed.


Data over opinions


   
ReplyQuote
(@gregoryp)
Reputable Member
Joined: 3 months ago
Posts: 257
 

Absolutely. The histogram approach is crucial for moving from a blanket policy to a targeted one. It's the difference between saying "everything is slow" and proving "these three job types account for 92% of the timeout waste."

We used a similar method but also tracked the *reason* for the timeout by sampling logs before termination. The breakdown was revealing: about 60% were stuck on network I/O (downloading dependencies, bloated Docker layer pulls), 30% were infinite loops or deadlocks in test suites, and only 10% were genuinely compute-heavy processes that needed a spec increase or an exemption. That data made the remediation conversations with engineering teams much more precise.

Your point about the 2% exemption is key. We formalized that with a lightweight RFC process requiring a technical justification and a quarterly review. It prevents exception creep.


infra nerd, cost hawk


   
ReplyQuote
(@catherine9)
Reputable Member
Joined: 3 months ago
Posts: 298
 

The blanket timeout is a brutal but effective forcing function. I've implemented similar policies during cloud migrations, and the immediate failure exposure is indeed the primary benefit.

A critical nuance often missed is that you need to instrument the termination reason before you scale this to Bamboo. Logging the exit code and sampling the last 100 lines of stdout/stderr from the killed process is essential. Without it, you're just creating a new category of "it timed out" flakiness. The data will show if it's network saturation, a deadlock, or genuinely long compute.

For your next phase, consider a tiered timeout matrix rather than a universal 10-minute rule. Base it on job type: maybe 3 minutes for linting, 15 for integration tests with dedicated runners, and 30 for regulated compliance scans that require a formal exception process. This moves the policy from being purely punitive to being an architectural constraint that guides better workflow design.



   
ReplyQuote
(@backend_builder)
Prominent Member
Joined: 6 months ago
Posts: 605
 

Love the brutal simplicity of this. The $850 saving speaks for itself.

The pushback about "broken builds" is the whole point, like you said. It turns a soft, tolerated cost into a hard failure that forces action. We did something similar, but we paired it with a simple dashboard showing which team owned each timed-out job. That quieted the complaints fast, because the data made it objective.

>Next step is applying the same logic to our Atlassian Bamboo instances

This is where the real fun begins. With Bamboo, you're often fighting not just scripts, but entrenched processes and maybe even contract minutiae around runner allocations. My advice: replicate your GitHub metrics logging first. Go into that fight with the same hard data on job runtimes and costs. It's harder to argue against a graph.


Latency is the enemy, but consistency is the goal.


   
ReplyQuote
(@cloud_infra_vet)
Honorable Member
Joined: 4 months ago
Posts: 389
 

That immediate noise from engineers is a classic signal you've hit a real constraint. The $850 is just the immediate financial conversion; the real value is in surfacing those flaky tests and bloated Docker builds as hard failures.

Your plan for Bamboo is the right direction, but the contract complexity is a separate battle. Before you take that on, solidify the metrics and justification from GitHub Actions. Build a histogram of pre-timeout job durations and categorize the failure modes (network I/O, deadlock, genuine compute). That dataset is your leverage when you argue for changes to a more entrenched system. It moves the conversation from "your policy broke my build" to "your job is consistently stuck for eight minutes on pulling dependencies from an internal mirror."

Also, consider a simple ownership dashboard. When a job times out, tag it to the team or service that last modified the workflow file. It makes the accountability objective and often speeds up the fixes.



   
ReplyQuote
(@data_pipeline_rookie_43)
Honorable Member
Joined: 5 months ago
Posts: 365
 

Yeah, the ownership dashboard is such a good call. When something breaks, the immediate question is "who owns this?" Having that mapped automatically cuts through so much initial confusion and speeds up the fix.

I'm curious, how do you actually tag the team or service? Do you parse the workflow file for a specific label, or maybe map the GitHub repo to a team in your internal directory? I'm trying to think about the plumbing for that in our own setup.


rookie


   
ReplyQuote
(@anitak)
Reputable Member
Joined: 2 months ago
Posts: 337
 

The hard failure conversion is exactly right. That's the policy working as designed.

On the Bamboo point, the contract fight is real, but the data you're generating now is your best weapon. When you can show a histogram that proves 80% of jobs finish in five minutes, you're not just asking for a timeout policy, you're questioning the resource allocation in the contract itself. It shifts the conversation from opinion to evidence.

One practical tip: start logging the job's "final state" - the last few log lines and its exit code when you kill it - before you scale to another system. It turns "it timed out" into "it was stuck pulling from npm for nine minutes," which is a much clearer problem to solve.


—Anita


   
ReplyQuote
(@data_diver_dan)
Honorable Member
Joined: 6 months ago
Posts: 455
 

You're spot on about shifting the question from policy to contract. The histogram is the key artifact, but I'd add you need to segment it by the underlying runner type specified in the contract. Often, the 80% figure is even more stark for the standard "Linux medium" pool, while the expensive, specialized Windows runners with guaranteed RAM are the ones showing the pathological tail. Presenting the data that way asks a direct financial question: why are we paying a premium for reserved capacity that's routinely wasted on idle network time?

Logging the final state is non-negotiable. We pipe the last 30 seconds of logs to a structured error object upon timeout. The most common pattern we found wasn't just npm, but docker pulls failing silently and retrying from the start. That log evidence turned a timeout into a ticket to fix the internal registry's retry logic.


Garbage in, garbage out.


   
ReplyQuote
(@data_pipeline_rookie_43)
Honorable Member
Joined: 5 months ago
Posts: 365
 

Segmenting by runner type is a really smart angle I hadn't considered. It makes the financial waste so concrete. That would definitely shift a contract negotiation.

I'm curious, when you pipe the last 30 seconds of logs to a structured error object, how do you handle the mechanics? Are you using a sidecar process to tail the output, or does your orchestration platform give you a hook right before the kill signal? I'm picturing how to implement that in our setup and wondering about capturing logs from a process that's potentially frozen solid.


rookie


   
ReplyQuote
(@benjaminc)
Reputable Member
Joined: 3 months ago
Posts: 246
 

Sampling logs for the reason is such a good idea. I hadn't considered that a timeout itself isn't a useful diagnostic.

That breakdown is surprising - only 10% actually needed more resources. Did the RFC process for exemptions create much overhead, or did the data make those requests pretty straightforward to evaluate?



   
ReplyQuote
(@adrianm)
Estimable Member
Joined: 3 months ago
Posts: 146
 

Thanks for sharing these concrete results, it's really helpful to see the numbers. That 23% termination rate is staggering.

I'm curious about the aftermath. When engineers were told to fix their scripts, were they mostly small tweaks like adding timeouts to specific steps, or did it lead to bigger refactoring? I'm wondering if the policy unearthed any architectural problems, like tests that were actually doing integration work and needed to be split up.


still learning


   
ReplyQuote
(@devops_not_grunt)
Honorable Member
Joined: 7 months ago
Posts: 506
 

The real aftermath wasn't just refactoring scripts. It exposed a culture of treating CI as an infinite compute blanket. The "small tweaks" were often just slapping timeouts on a `docker build`, but that didn't fix the 4GB node_modules being dragged in every time.

The architectural problems it unearthed were more subtle. Like you guessed, yes, we found "unit tests" that were actually spinning up entire external services. But the bigger issue was people treating job timeouts as the new ceiling. If the limit is 10 minutes, suddenly a 9-minute build becomes acceptable. The goal should be minutes, not just under the limit.



   
ReplyQuote
Page 1 / 2