Skip to content
Notifications
Clear all

Am I the only one who thinks CI should be a fixed cost, not variable?

22 Posts
20 Users
0 Reactions
56 Views
(@cloud_watcher_99)
Prominent Member
Joined: 3 months ago
Posts: 668
Topic starter   [#25439]

Okay, I need to get this off my chest. I've been deep in my AWS and Datadog bills this week, and the CI/CD line item is giving me anxiety. Every month it's a surprise. 1500 build minutes one month, 4500 the next. It feels like I'm being punished for my team's productivity or for fixing more bugs.

I get the "pay for what you use" model for things like S3 or Lambda. But CI feels different. It's core infrastructure. I want to be able to tell my finance person, "This is what it costs to ship software," and have that number be predictable. Right now, I'm constantly tweaking things just to shave minutesβ€”like aggressively pruning old Docker images or setting crazy short timeouts. It adds toil.

I ran the numbers for our mid-sized team (about 10 devs). Our managed CI (let's just say it's a popular cloud one) averages about $380/month, but it's swung from $220 to $550. I modeled a self-hosted runner cluster on EC2 Spot with a fallback, and the *infrastructure* cost is actually pretty predictable (~$200 fixed for the baseline + maybe $50 in variables). But then you add the hidden cost: my time to monitor, patch, and scale it. That's the real killer.

```yaml
# This is the kind of config optimization I'm talking about. Feels like micro-managing pennies.
jobs:
build:
timeout-minutes: 15 # Aggressive timeout to cap cost
steps:
- uses: docker/setup-buildx-action@v3
- name: Build and push
uses: docker/build-push-action@v5
with:
cache-from: type=gha
cache-to: type=gha,mode=max
# Caching helps, but it's another thing to manage
```

Am I crazy for wishing there was a sane middle ground? Like a committed-use discount for CI minutes, or a tier where you get a build "concurrency" pool for a flat fee, then pay a tiny variable on top? The current model makes FinOps for CI feel like a game where the house always wins.

What are you all doing? Eating the variable cost, self-hosting, or have you found a provider with a more predictable model?


cost first, then scale


   
Quote
(@data_diver_42)
Honorable Member
Joined: 7 months ago
Posts: 400
 

Totally feel you on the bill anxiety. That swing from $220 to $550 is rough for planning.

Have you looked at the newer commit-based plans some providers offer? They're still variable, but the cap makes it less spiky. We moved to one and it helped smooth things out, though you do lose some flexibility.

The hidden ops cost for self-hosted is so real. I ran a similar model and the break-even point for my time was way higher than I thought. It only made sense once we had a dedicated platform engineer to own it.


Data is the new oil - but it's usually crude.


   
ReplyQuote
(@cost_optimizer_elle)
Reputable Member
Joined: 4 months ago
Posts: 370
 

Commit-based plans are a decent bandage, but they're still just a pricing trick. You're not actually fixing the problem, you're just paying for a bigger bucket of minutes you might not use.

>break-even point for my time was way higher than I thought

This is the key everyone misses. The real cost isn't the EC2 bill, it's the Friday night pager alert because your self-hosted runner scale-down lambda choked. I've seen teams burn 40+ engineering hours a month babysitting their "cheap" setup. At that point, the $550 variable bill looks like a bargain.

If you go self-hosted, you must run it as a product with an SLA. Otherwise you're just trading a variable cloud bill for a wildly variable, and often higher, ops tax.


- elle


   
ReplyQuote
(@consultant_carl_42)
Reputable Member
Joined: 4 months ago
Posts: 381
 

You've put your finger on the exact tension. You're trying to budget for a core utility, but the vendor sells you a commodity resource. Of course they want it variable; their margins are better when your team is productive.

The real question your finance person should ask isn't "what does it cost to ship software," but "what's the maximum we're willing to pay for the privilege of shipping faster?" That $550 surprise month is the answer. It's the cost of a burst of productivity or, more likely, a month of broken builds and re-runs. That variability isn't a bug in the model, it's the feature.

Your self-hosted math is classic. Everyone models the EC2 bill. Nobody accurately values the cognitive load of being the person who gets to think about pruning Docker images *and* scaling lambdas. You just traded one variable line item for a massive, unplanned drain on your most finite resource: focused engineering time. Predictable cash outlay, wildly unpredictable productivity tax.


Test the migration.


   
ReplyQuote
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
 

You're absolutely right about the ops tax being the hidden multiplier. I've modeled this for clients.

The "run it as a product with an SLA" is the critical line. Most teams aren't structured to do that. They assign it to a devops-savvy engineer as "extra responsibility." The cost then becomes the opportunity cost of that engineer's *other* work, plus the burnout risk.

A reasonable middle ground I've seen is using a managed service for the control plane (orchestration, queue) with self-hosted, auto-scaling runners on spot instances. You cap the compute variable cost and offload the most brittle scaling logic to the vendor. You still own runner maintenance, but the pager alerts for "no capacity" disappear.


Data is the only truth.


   
ReplyQuote
(@brian)
Reputable Member
Joined: 3 months ago
Posts: 282
 

You're just trading one variable cost for another. Spot instances can disappear. Now you're on the hook for both the managed service fee *and* a different kind of scaling headache.

The middle ground often becomes a no-man's-land where you pay for management but still eat the operational risk.


Trust but verify.


   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

The problem you're describing is a budget planning failure, not a pricing model failure. Finance needs a forecast, not a fixed cost. You take your average, add a 20-30% buffer for those productive months, and call that your line item. If you're consistently blowing past that buffer, then your development process is the variable, and you need to address that.

You're optimizing for the wrong thing by pruning images and setting short timeouts. That creates friction and failure. Pay the $550 when it happens and treat it as data. If it's a month of bug fixes, that's a cost of quality. If it's inefficient pipelines, that's a cost of technical debt. The bill tells you which one it is. A fixed cost just hides that signal.


β€”AF


   
ReplyQuote
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
 

Forecasting with a buffer is still reacting to a variable. You're just pre-paying for the surprise.

>Pay the $550 when it happens and treat it as data.

The signal is useless if the cost of collecting it is your engineers wasting time on build-time optimizations instead of features. A fixed cost isn't about hiding the signal, it's about removing a distraction. My database doesn't charge me more for a complex join, and my CI shouldn't charge me more for a productive team.


SQL is enough


   
ReplyQuote
(@gracep)
Reputable Member
Joined: 2 months ago
Posts: 297
 

>The bill tells you which one it is.

It tells you far too late. By the time you see the $550 spike, the inefficient pipeline has already burned a month of developer time waiting for builds. That's the real cost, not the invoice.

A fixed cost doesn't hide the signal. It moves the signal from the finance spreadsheet to the engineering metrics you should be watching anyway: build duration, queue time, flaky test rate. You act on those leads, not on a cost lag.


Data over opinions


   
ReplyQuote
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 494
 

That config snippet hits home. I've been down that exact rabbit hole - trying to squeeze minutes by tweaking timeouts and caching layers, then wondering why our test flakiness went up 😅

You mentioned the hidden cost of your time for self-hosting. That's the real equation. I track mine in Datadog with a custom "ops toil" dashboard. Turns out I was spending 8-10 hours monthly just babysitting runners, which at my hourly rate made the $550 variable bill look cheap. The predictable cost wasn't the EC2 bill, it was my Friday nights.

Have you looked at setting up billing alerts on your CI provider? I have a simple one that fires at 75% of our average monthly spend. Doesn't fix the variability, but at least eliminates the surprise when finance comes asking.


Dashboards or it didn't happen.


   
ReplyQuote
(@cloud_watcher_99)
Prominent Member
Joined: 3 months ago
Posts: 668
Topic starter  

That config snippet hits home. I've been down that exact rabbit hole - trying to squeeze minutes by tweaking timeouts and caching layers, then wondering why our test flakiness went up 😅

You mentioned the hidden cost of your time for self-hosting. That's the real equation. I track mine in Datadog with a custom "ops toil" dashboard. Turns out I was spending 8-10 hours monthly just babysitting runners, which at my hourly rate made the $550 variable bill look cheap. The predictable cost wasn't the EC2 bill, it was my Friday nights.

Have you looked at setting up billing alerts on your CI provider? I have a simple one that fires at 75% of our average monthly spend. Doesn't fix the variability, but at least eliminates the surprise when finance comes asking.


cost first, then scale


   
ReplyQuote
(@felixr47)
Reputable Member
Joined: 2 months ago
Posts: 292
 

You've got a great analogy with the database, but I think it breaks down on a fundamental level. A database join is a pure compute operation - it's predictable. CI minutes are a mix of compute, storage, and human activity.

That $550 bill isn't the database charging for the join. It's the database charging you for the 50 developers all running exploratory queries at the same time, plus the storage for their temporary result sets. The fixed cost model works for the infrastructure *capability*, but not for the aggregated *consumption* of a team, unless you buy a massive fixed capacity you'll rarely use.

The distraction you mention is real, but I've found it's less about the cost being variable and more about the cost being *surprising*. A transparent, predictable variable cost - like a clear per-minute rate with a dashboard - feels very different from a black-box monthly invoice with a random spike.



   
ReplyQuote
(@elliotr)
Reputable Member
Joined: 2 months ago
Posts: 229
 

The core of your anxiety isn't the variable cost itself, it's the misalignment between the pricing metric and the business value. You're being measured on minutes, a low-level technical unit, when your goal is delivering a predictable capability to ship software. This misalignment forces the cost optimization behavior you described.

The comparison to a database is apt. You license the *capability* of a database, not its CPU cycles. For CI, the capability is "my team can merge and deploy with confidence." Some providers do offer fixed-price enterprise tiers for this exact reason, but they're often priced for much larger scale, leaving mid-sized teams in this variable-cost trap.

Your self-hosted analysis reveals the real trade-off. The predictable infrastructure cost comes with a highly variable operational burden, which is just a different kind of cost volatility. The question becomes whether you prefer to manage financial volatility or time volatility. For a team of your size, the financial volatility of a cloud CI is likely the lesser evil, but it requires treating the spend as a semi-variable with a communicated buffer, not a fixed line item.



   
ReplyQuote
 danf
(@danf)
Estimable Member
Joined: 2 months ago
Posts: 168
 

Exactly, and that misalignment is why I keep circling back to the database analogy. But the problem is, we're not licensing Oracle. We're renting AWS or GCP compute by the second. The promise of cloud CI was that it *wasn't* a licensed capability model, it was utility computing. The disappointment is that the pricing didn't follow through, it just swapped capital expense for a confusing operational expense that's still tied to the wrong unit.

The real issue is that the "capability" you're describing isn't fixed. A database license buys you a fixed performance envelope. The CI "capability" of a merging team is wildly variable - is it two devs patching or twenty-five shipping a new product? A fixed price for that is either a ripoff or a loss leader for the vendor. The enterprise tiers you mention are just pre-paying for a huge pool of those same minutes, and you're right, the break-even point is astronomical.

So we're stuck. The financial volatility is indeed the lesser evil, but calling it "semi-variable" is just a fancy way of saying you're still guessing. The buffer isn't a plan, it's an admission that the model doesn't fit.


Anecdotes aren't data.


   
ReplyQuote
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
 

You're right about the operational tax. The hidden cost isn't just time, it's context switching. Every minute spent tuning a scaling lambda or debugging a runner's provisioning is a minute pulled from the core product. That's why the ops tax has such a high multiplier.

Your SLA point is critical. Most teams treat self-hosted CI as infrastructure, but you're correct that it needs product-level reliability tracking. The moment you start defining actual SLOs for build queue time and runner availability, you realize the maintenance burden justifies a managed service's premium.

Even with a perfect SLA, the fixed infrastructure cost still doesn't align with a team's output value. You're paying for idle capacity during quiet periods, which is just the variable cost problem in reverse.


benchmark or bust


   
ReplyQuote
Page 1 / 2