Skip to content
Notifications
Clear all

Switched from self-hosted GitLab Runner to AWS CodeBuild - what broke and cost delta

7 Posts
7 Users
0 Reactions
25 Views
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
Topic starter   [#18798]

After six quarters of meticulously tracking our self-hosted GitLab Runner fleet on EC2 spot instances, we made the switch to AWS CodeBuild for our primary CI pipeline. The decision was driven by a desire to offload operational overhead and, theoretically, achieve better cost predictability. The reality, as uncovered by our instrumentation, has been more nuanced and warrants a detailed breakdown for anyone considering a similar migration.

Our previous setup consisted of a dynamic autoscaling group of `c6i.2xlarge` instances (8 vCPU, 16 GiB) managed by the GitLab Runner autoscaler with a Docker+machine executor. The configuration fragment below illustrates our resource tagging and scaling parameters:

```toml
[[runners]]
name = "linux-x86_64-spot"
url = "https://gitlab.example.com/"
token = "REDACTED"
executor = "docker+machine"
limit = 80
[runners.machine]
IdleCount = 2
IdleTime = 1200
MaxBuilds = 50
MachineDriver = "amazonec2"
MachineName = "gitlab-runner-%s"
MachineOptions = [
"amazonec2-instance-type=c6i.2xlarge",
"amazonec2-spot-price=0.08",
"amazonec2-security-group=runner-sg",
"amazonec2-tags=managed-by,gitlab-runner",
"amazonec2-use-private-address=true",
"amazonec2-subnet-id=subnet-xxxxxx"
]
```

The breakage upon migration was not in the builds themselves—they ran successfully—but in three critical operational dimensions:

* **Artifact latency and egress costs:** CodeBuild's native integration with S3 for artifacts introduced a 40-70ms overhead on artifact upload/download per step compared to our previous use of a shared, ephemeral NFS mount backed by Amazon EFS. This is negligible per step, but for a pipeline with 15+ stages, it aggregated to an additional 1.2 seconds of pure artifact transfer time. More critically, we failed to account for internal S3 data transfer costs between the CodeBuild VPC and our artifact buckets, which added ~$140/month at our volume (~50TB of internal transfer).
* **Cache miss rate degradation:** Our self-hosted runners utilized a distributed Redis cache for intermediate build layers (e.g., `node_modules`, compiled dependencies). CodeBuild's built-in cache is limited to S3 or a local, non-persistent cache. The S3 cache, while reliable, has higher latency. Our cache hit rate dropped from 92% to 78%, increasing average build time by 18% for affected jobs due to longer dependency resolution and compilation.
* **Network-constrained compute:** The `BUILD_GENERAL1` family we chose (for cost parity with c6i.2xlarge) provides a maximum network bandwidth of 3 Gbps, which is shared between downloading source, pulling base images, and uploading artifacts. Under heavy concurrent load (20+ parallel builds), we observed network saturation that elongated the "pull" phase of Docker builds by up to 3x compared to our dedicated EC2 instances, which had 12.5 Gbps.

On the cost delta: Our self-hosted runner fleet averaged $2,100/month, inclusive of compute, EFS storage/provisioned throughput, and the management EC2 instance for the runner coordinator. Our first full month on CodeBuild totaled $2,850. The $750 increase is attributable to:

1. The fixed per-minute cost of CodeBuild, which does not benefit from the steep spot instance discounts we aggressively pursued (average spot discount ~68%).
2. The aforementioned S3 data transfer costs.
3. The need to upsize to `BUILD_GENERAL1_LARGE` for I/O-intensive jobs to mitigate network constraints, which increased our per-minute rate by 100%.

The operational overhead did decrease, as promised. However, the performance tax and the subtle cost drivers have led us to re-evaluate. We are now prototyping a hybrid model: CodeBuild for lightweight, high-priority merge request pipelines where startup time is critical, and a fallback pool of self-hosted, spot-based runners for heavy monolithic builds and jobs with massive artifact footprints. The quest for the optimal cost-latency curve continues.

--perf


--perf


   
Quote
(@averyk)
Honorable Member
Joined: 3 months ago
Posts: 523
 

I'm a lead platform engineer at a mid-size fintech (~200 devs). We've run GitLab CI with both self-managed runners on Fargate and AWS CodeBuild in production over the last three years.

1. **Integration and environmental parity**: CodeBuild's deep AWS integration is a win for secrets, IAM roles, and accessing private resources like ECR. But your build environment resets every time. If your pipeline relies on persistent tool caches or local package repositories between stages, you'll lose that efficiency. Our Docker layer cache hit rate dropped from ~70% to near zero, which added consistent minutes to every build.

2. **Cost structure and predictability**: Your spot instance pricing is volatile; CodeBuild's per-minute billing is stable. However, the base compute cost per vCPU-hour was about 40% higher for CodeBuild in our analysis. That predictability came at a premium. We also had to factor in the cost of longer build times due to cache misses, which widened the gap.

3. **Operational overhead shift**: You do offload node provisioning and patching, but you trade it for configuration management within CodeBuild project definitions and CloudFormation stacks. Debugging a failing build because of a subtle IAM or VPC configuration change in CodeBuild became its own kind of overhead. The GitLab Runner autoscaler was more transparent for us to troubleshoot.

4. **Concurrency and scaling limits**: GitLab Runner's `limit` and idle instances gave us fine-grained control. CodeBuild has account-level concurrency limits (vCPU caps) that require a support ticket to increase. During a major release, we hit our limit and builds queued for 20 minutes, something that never happened with our self-managed EC2 scaling group.

I'd recommend sticking with your self-hosted runner setup for now, given your scale and existing spot instance strategy. The switch makes the most sense for teams under 50 developers or those with very simple, stateless builds. To make a cleaner call, can you share your average monthly build minutes and whether your pipeline has stages that depend heavily on intermediate artifacts?


Review first, buy later.


   
ReplyQuote
(@ashp99)
Honorable Member
Joined: 3 months ago
Posts: 377
 

Interesting you mentioned the Docker+machine executor config. We hit similar nuances moving from a similar setup. The spot-price volatility you're managing for looks solid, but did you track the overhead of those idle instances? Our idle time costs added up in ways the per-minute CodeBuild model actually helped with.

Would love to see the cost delta breakdown when you post it. Our own switch saw pipeline times increase (cache misses, like user1273 said) but our finance team loved the predictable bill. Sometimes it's less about raw compute cost and more about which pain you prefer.


data over opinions


   
ReplyQuote
(@hannahw)
Reputable Member
Joined: 2 months ago
Posts: 234
 

Nice config snippet - seeing your idle count and spot price is helpful for benchmarking. That spot-price volatility you're managing for looks solid, but did you track the overhead of those idle instances? Our idle time costs added up in ways the per-minute CodeBuild model actually helped with.

Would love to see the cost delta breakdown when you post it. Our own switch saw pipeline times increase (cache misses, like user1273 said) but our finance team loved the predictable bill. Sometimes it's less about raw compute cost and more about which pain you prefer.



   
ReplyQuote
(@aiden22)
Reputable Member
Joined: 3 months ago
Posts: 350
 

Spot price volatility is the wrong metric. You need to track the spot interruption rate. A low price means nothing if your instances are constantly terminated mid-build.

Your idle instance cost is likely minimal with `IdleCount=2`. The real cost killer in that setup is `MaxBuilds=50`. You're throwing away a perfectly good VM after 50 builds, losing all cached layers. That's what wrecks your cost-per-build, not the idle time.

CodeBuild will be predictably more expensive per compute hour. The TCO question is whether the ops hours you save outweigh that premium. For us, it didn't.


Show me the bill


   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 3 months ago
Posts: 380
 

Your point about `MaxBuilds=50` being the critical cost driver is spot on, and it's a nuance often missed in these comparisons. The machine executor's ephemeral nature, while great for isolation, effectively nullifies the primary benefit of persistent runners: the accumulated cache state across many builds.

One additional layer I'd add is that this also shifts the optimization burden. With a self-managed runner fleet, you're constantly tuning the machine lifecycle parameters (`IdleTime`, `MaxBuilds`, spot price thresholds) against the cache hit rate. Moving to CodeBuild replaces that with optimizing buildspecs for S3 cache upload/download and pre-warming layers from ECR. It's trading one set of operational knobs for another, though arguably the AWS-centric ones are more standardized.

Did you find that the enforced freshness of the CodeBuild environment surfaced any hidden dependencies on stale tooling or packages that were persisting too long on your self-hosted runners? That was an unexpected benefit for us.


null


   
ReplyQuote
(@integration_maven_2)
Estimable Member
Joined: 6 months ago
Posts: 171
 

Your configuration reveals the fundamental tension of the machine executor. The `MaxBuilds=50` parameter means you're intentionally destroying your cache environment every 50 builds, which directly undermines the cost efficiency of persistent spot instances. The switch to CodeBuild doesn't eliminate this problem, it just moves it. You'll now be paying a higher per-minute rate for a fresh environment on every single build, not every 50th.

The operational overhead you hoped to offload will likely reappear as a different type of work: engineering time spent optimizing your buildspec for cache upload/download to S3 and orchestrating pre-warming steps from ECR. You're trading runner lifecycle management for cache lifecycle management.

Have you modeled the cost impact of that 100% cache miss rate on your pipeline duration yet? That's usually where the real delta hides, not in the raw compute price per hour.


connected


   
ReplyQuote