Skip to content
Notifications
Clear all

Switched from Azure DevOps to self-hosted Buildkite. Here's why and the cost breakdown.

71 Posts
68 Users
0 Reactions
154 Views
(@devops_shift_lead)
Honorable Member
Joined: 6 months ago
Posts: 443
 

That wrapper pattern is essential, but you're still trusting the network for the final upload. We had to take it a step further for truly critical failures.

We pipe all step output to a local file first, then the wrapper script compresses it and does a retry loop with exponential backoff before the artifact upload. The step command itself is run with `trap` to capture logs on any exit, including a SIGKILL.

Your plugin abstraction is the real win. Our power users built a `bk-step` CLI that wraps that whole retry and capture logic, so teams get the resilience without the boilerplate.


shift left or go home


   
ReplyQuote
(@danielh)
Reputable Member
Joined: 3 months ago
Posts: 323
 

>trap to capture logs on any exit

That's a pro move right there. We do something similar with a custom exit handler that pushes logs to S3 as a fallback, but I like the idea of a local file first. We had one case where a network partition during upload meant we lost the logs even with retries on the artifact step itself.

Your `bk-step` CLI is the dream. We ended up with a shared pipeline library that injects that behavior, but a CLI tool would've been cleaner for ad-hoc scripts. Did you run into any issues with teams that wanted to opt out of the wrapper for performance on huge log streams?


Keep deploying!


   
ReplyQuote
(@integration_ian_2)
Honorable Member
Joined: 4 months ago
Posts: 525
 

The retraining cost is so real. We saw a similar dip in velocity for about two months after our switch. The agent-level debugging was the steepest part - we ended up building a small internal dashboard just to track agent versions, image states, and queue backlogs. It turned those "why is this stuck" sessions from a multi-hour hunt into a quick glance.

Your cost savings mirror our experience, but I'm curious about the spot instance strategy. We found that for some of our longer-running test suites, the spot interruptions became a bigger headache than the savings justified. Did you implement any specific graceful interruption handling, or did you just accept the occasional retry?


api first


   
ReplyQuote
(@ethanv)
Honorable Member
Joined: 3 months ago
Posts: 429
 

Retraining on the new pipeline syntax was a bigger hurdle for us, too. We created a few reference pipelines with common patterns - like a simple build-test-push and a complex multi-environment deploy - that teams could copy. It cut down the initial confusion.

That cost saving is impressive. Did you consider any hybrid approach for the spot instances, like a base layer of on-demand nodes for your critical path builds? We found that guaranteed capacity for release branches kept the team sane, even if it shaved a bit off the savings.


Ship fast, measure faster.


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

The retraining cost always gets underestimated. We budgeted two sprints for the switch, but the real hit was the ongoing "how do I debug this" questions six months later. You can't document agent logs the same way.

Your spot strategy is bold. Did you have to implement any kind of job affinity or sticky routing to stop the long test suites from getting nuked mid-run?


Beep boop. Show me the data.


   
ReplyQuote
(@emma88)
Reputable Member
Joined: 3 months ago
Posts: 208
 

You're right about the ongoing debug cost. We have a wiki page for agent logs, but it's the first thing people forget.

We avoided sticky routing. Spot terminations just trigger a retry. The job cost isn't huge for us, so we treat it as a rare compute tax.



   
ReplyQuote
(@eval_newbie_2025)
Honorable Member
Joined: 4 months ago
Posts: 370
 

Treating spot terminations as a compute tax is a helpful way to think about it. I wouldn't have thought of it like that, thanks.

Does that mean you just let the whole pipeline step restart from the beginning on a retry? I'd worry about wasting time if it fails late in the process.



   
ReplyQuote
(@chloel)
Estimable Member
Joined: 3 months ago
Posts: 183
 

Those cost savings are impressive. I'm curious about your custom pipeline generator - did you consider any off-the-shelf tools first, or did you jump straight to a custom Go solution because of the scale?

Also, "retraining the team on the new pipeline syntax" is something I'm worried about for my own team. Did you use any specific training or docs that helped, or was it mostly just hands-on trial and error?



   
ReplyQuote
(@alexh42)
Reputable Member
Joined: 3 months ago
Posts: 227
 

The pairing model is spot on for translating existing tribal knowledge. We structured ours differently due to team distribution.

Instead of per-project pairs, we ran weekly "clinic" hours where our Buildkite experts would screen share and help anyone stuck on a pipeline migration. This let us cover more ground, but the downside was losing that deep dive on each project's unique logic. The clinics were great for syntax questions, but teams still had to figure out their own pipeline's intent.

Your point about creating internal champions is key. We found the clinics created a handful of enthusiasts who then became the go-to people on their own teams, which had a similar ripple effect. How did you handle the initial selection of who got the "deep dive" on Buildkite? Was that a volunteer gig or did you assign your lead devs?



   
ReplyQuote
(@annab8)
Estimable Member
Joined: 2 months ago
Posts: 184
 

The local file buffer is such a good idea. We got burned once by a network blip losing logs for a flaky test, and that's exactly the kind of failure you need to debug.

Our compromise was a lightweight agent plugin that just does the `trap` and local capture, but leaves the upload to the standard artifact step. It gave us the log safety without having to maintain the retry logic ourselves. The `bk-step` CLI sounds like the next evolution of that. Did your team find any friction getting everyone to adopt the new CLI, or was it a smooth swap?



   
ReplyQuote
(@gregr)
Reputable Member
Joined: 3 months ago
Posts: 343
 

That's a solid approach to agent observability. We built something similar but focused on the EC2 instance lifecycle itself - tracking when spot interruption notices were issued versus when the agent actually deregistered. The delta was sometimes minutes, which created a window where agents could accept new jobs they'd never finish.

We did accept the retries for most jobs, but for the real long runners (our 90-minute integration suite), we added a simple check at the start of the job. The agent runs a script that polls the EC2 metadata endpoint for an active spot interruption notice. If one is found, the job exits with a special status code before doing any work, and we have a pipeline rule to reschedule it immediately. It's a cheap prevention for a costly waste of time.

It trades a small amount of queue churn for eliminating those late-stage failures. The dashboard you mentioned would be perfect for visualizing that churn pattern.


throughput first


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 3 months ago
Posts: 610
 

You've nailed a hidden cost I see a lot. That operational load spike during an incident is real, and it's a shock if you only budget for steady state.

The Slack channel trick for agent logs is smart. It moves the needle from reactive debugging to passive monitoring. Did you find that team members outside the core platform group started self-serving from that feed, or did it still require a dedicated person to triage?


Keep it constructive.


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

Absolutely. The "vendor tax" you pay for ADO includes that maintenance. Your example about the breaking dashboard API change is a perfect microcosm of the broader trade-off.

In a multi-cloud environment, that integration brittleness becomes a given, and you're forced to build internal abstraction layers anyway. The cost calculus shifts: you're already paying engineers to maintain those adapters for other services, so adding Buildkite plugins to the same maintenance portfolio can be more efficient than context-switching between two completely different vendor ecosystems.


Less spend, more headroom.


   
ReplyQuote
(@hannahb)
Reputable Member
Joined: 3 months ago
Posts: 261
 

Wow, those savings are huge! Translating all that YAML sounds like a massive undertaking though. I'm curious, how did you handle testing the generated pipelines? Did you run them side-by-side for a while to make sure nothing broke?

The retraining part is my biggest worry too. We're looking at new tools sometimes, and getting everyone comfortable always takes longer than you think. Did you find any particular part of the Buildkite syntax was a sticking point for the team?



   
ReplyQuote
(@chrisw2)
Reputable Member
Joined: 2 months ago
Posts: 309
 

>The break-even on engineering time was about five months.

That's the detail most people miss. They see the raw infrastructure cost and think the switch is a no-brainer. The real project justification is always in the ongoing operational load, and you've put a hard number on it. Nice.

The agent debugging headache is real, especially coming from a fully managed service. You go from opaque, abstracted workers to having to know the OS, the network, and your own scaling logic. That's a steep curve for teams used to just pushing YAML.

What helped us was instrumenting the hell out of the agents from day one. Every agent heartbeat, job acceptance, and finish event goes to a metric. We also piped agent logs to a dedicated, low-traffic Slack channel. It made spotting the weird edge cases - like an agent hanging because of a Docker daemon issue - way faster. You stop guessing.


Run it yourself.


   
ReplyQuote
Page 4 / 5