Skip to content
Notifications
Clear all

GitLab CI after 12 months - real experience with scaling and invoices

5 Posts
5 Users
0 Reactions
16 Views
(@finnm)
Reputable Member
Joined: 2 months ago
Posts: 280
Topic starter   [#25023]

So, I've been running our GitLab CI for our small SaaS for a full year now. We started with a few hundred shared runner minutes, but after growing and adding more automated tests, our usage exploded.

Our last invoice was... a lot 😅. I'm trying to understand if this is normal. We're hitting about 3,000-4,000 compute minutes per month now. Anyone else scaled up and have real numbers to share? Specifically:

* What does your monthly invoice look like at this volume?
* Did you hit a point where moving to a self-hosted runner made financial sense?
* Any gotchas with managing your own runner that aren't obvious?

Just trying to figure out if we're on the right track or if we need a major rethink. Our budget for dev tools is getting squeezed.



   
Quote
(@cloud_infra_rookie)
Noble Member
Joined: 4 months ago
Posts: 552
 

Yeah, that monthly usage sounds familiar, we're in a similar spot. Our invoice jumped when we added integration tests. Have you looked at caching dependencies in your pipelines? That cut our minutes down a fair bit.

We switched to a self-hosted runner on a beefy AWS spot instance about six months ago. The break-even math worked for us pretty fast. The main hidden cost is the maintenance time, honestly. You have to patch the runner and monitor its disk space.

Did you consider just optimizing your current jobs first, before jumping to self-hosted? Maybe splitting tests to run in parallel?



   
ReplyQuote
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
 

Spot-on about the maintenance time, that's the killer they never put on the pricing page. Patching is one thing, but runners quietly filling their disk with old job logs and cache will murder you at 3 AM.

Parallelizing tests is the first lever to pull, but only if your test suite isn't a tangled mess of shared state. If it is, you'll spend more time fixing flaky tests than you ever saved. Cache dependencies aggressively, but also look at your job structure - are you rebuilding your Docker image from scratch in every pipeline stage? That's a common minute-burner.

The real math for self-hosted isn't just GitLab's per-minute rate versus an EC2 instance. It's that rate versus the instance *plus* your hourly rate for babysitting it. If your team already has a solid IaC setup, the overhead is lower. If not, that invoice might start looking reasonable again.


Speed up your build


   
ReplyQuote
(@dragonrider)
Honorable Member
Joined: 3 months ago
Posts: 367
 

Absolutely true about dependency caching being a first line of defense. I'd add that you need to audit your cache keys regularly. We once had a misconfigured key that basically never updated, so our runners were stuck using ancient, busted dependencies for weeks.

The jump to a beefy spot instance is smart for raw compute, but have you tracked the "human patching tax" against your actual saved minutes? I'm a data nerd, so I started logging every minute we spent on runner upkeep. Turns out, for our team size, the sweet spot was optimizing GitLab jobs *and* using a smaller, reserved instance for our longest-running tasks only. We kept shared runners for the quick stuff.

Spot instances are great until they get reclaimed mid-deploy. Did you build any automation to handle that, or just accept the occasional pipeline failure?


Try everything, keep what works.


   
ReplyQuote
(@grafana_knight_shift)
Reputable Member
Joined: 6 months ago
Posts: 324
 

Parallelization is a huge win, but only if you have good telemetry on your pipeline stages first. I set up a Grafana dashboard from our GitLab instance's metrics to spot which jobs are the real hogs. You'd be surprised how often one slow, sequential linter job burns more minutes than the whole test suite.

The patching tax is real. We started treating our self-hosted runners like cattle, using the GitLab runner helm chart on a spot node group. When a node gets reclaimed or needs a patch, the pod just reschedules. It adds a bit of startup latency, but it's cheaper than my 2 AM SSH sessions.



   
ReplyQuote