Skip to content
Notifications
Clear all

ELI5: Should I host my own CI/CD runners or use cloud?

3 Posts
3 Users
0 Reactions
20 Views
(@elliotk)
Reputable Member
Joined: 3 months ago
Posts: 323
Topic starter   [#15284]

Okay, so I’m in the middle of containerizing a pretty standard LLM-powered app—FastAPI backend, some RAG pipelines, a couple of finetuning scripts—and I’ve hit the classic infrastructure crossroads. Right now, my CI/CD is basically a free-tier cloud runner (you know the ones). It’s fine for tiny pushes, but my builds are starting to creep past 20 minutes because of the GPU steps for embedding model tests and the occasional small training run.

I’m trying to think this through like a proper tinkerer: do I bite the bullet and set up my own runners on a beefy machine, or do I just upgrade to a paid cloud plan with more powerful instances?

I need a proper, grounded comparison. Not just “self-hosted is cheaper” (is it, though, when you count my time?), but real trade-offs for a mid-sized, ML-heavy project. Think:
* A monorepo with ~5 services, some needing CUDA for certain CI steps.
* Maybe 15-20 pipeline runs per day across the team.
* We already have a decent NAS and a spare server in the rack.

The big things I’m wrestling with:

**For self-hosted runners:**
* **Upfront & Maintenance Time:** How many hours am I *really* signing up for per month? Setting up the runner agents, networking, security patches, dealing with runner updates, and the inevitable “why is the runner offline?” Slack messages.
* **Hardware Cost vs. Cloud:** If I get a used server with an A5000, how many months of cloud GPU runner time does that buy me? Is the capex worth it?
* **The “It Works on My Runner” Problem:** How do you ensure consistency between dev machines and the self-hosted runner environment? Do you just run everything in Docker anyway, making the host OS less critical?

**For cloud runners:**
* **Cost Predictability:** Are we talking “bill shock” territory if someone’s branch has a bug and runs a 2-hour GPU job 10 times? Are there good guardrails?
* **Cold Start & Performance:** For those big Docker builds with multiple LLM dependency layers, how fast are the cloud VMs compared to a local machine with a fast NVMe?
* **Vendor Lock-in:** How painful is it to move your pipeline definitions and secrets later if you need to?

What I’m *really* after are some benchmark-ish numbers or rules of thumb from people who’ve lived through this. Things like:
* “Our self-hosted runner on a Xeon with 64GB RAM handles a full build in 8 minutes, same build on Cloud Provider Y’s comparable ‘large’ instance took 9 minutes but costs $0.45 per run.”
* “The maintenance overhead for our 5 runners is about 2-3 hours every month for updates and troubleshooting.”
* “For GPU jobs, cloud became cheaper than our self-hosted rig after 6 months because of our sporadic usage pattern.”

Has anyone done a detailed side-by-side for a data/ML stack? I feel like the GPU component changes the math completely compared to a standard web app build.



   
Quote
(@heatherm)
Reputable Member
Joined: 3 months ago
Posts: 255
 

You've nailed the main trade-off: your time is the real currency here. For your scale and that GPU requirement, the maintenance hours can add up fast.

That spare server is a tempting anchor for a self-hosted setup, but you have to factor in patching, monitoring, and inevitable driver updates for those CUDA steps. I've seen teams get bogged down just keeping the runner images in sync with their cloud counterparts.

Honestly, with 15-20 daily runs needing GPU bursts, a paid cloud plan with powerful spot instances might surprise you on cost. It turns a capital expense and time sink into a predictable, variable one. Run a two-week cost simulation on both options before you commit any weekend hours to setup.


Ask me about my RFP template


   
ReplyQuote
(@code_panda)
Reputable Member
Joined: 5 months ago
Posts: 294
 

> A monorepo with ~5 services, some needing CUDA for certain CI steps.

This is the key detail. If only some steps need the GPU, you can actually hybridize. Use a cloud GPU runner for the heavy CUDA jobs (model tests, training), and keep your own, cheaper CPU runner for the rest of the build and deployment steps. Most platforms let you tag jobs and match them to specific runner types.

You're right to question the "self-hosted is cheaper" mantra. For GPU workloads, the cloud spot market can be insane value, and you skip the headache of managing driver compatibility across those 5 services. That spare server might be better used as a dedicated CPU runner to cut your baseline cloud costs, while you rent the GPUs only when you need them.


Spreadsheets > marketing slides.


   
ReplyQuote