Skip to content
Notifications
Clear all

Guide: Setting up a benchmark pipeline to compare 3 CI/CD tools

7 Posts
7 Users
0 Reactions
14 Views
(@danag)
Reputable Member
Joined: 3 months ago
Posts: 303
Topic starter   [#26473]

Hey folks! 👋 I've been deep in the weeds lately trying to decide on a CI/CD platform for a new microservices project. I kept reading "X is faster" or "Y has better caching," but without a real, apples-to-apples comparison, it's hard to trust the claims. So, I did what any of us would do: I built a benchmark pipeline to test it myself.

I focused on three tools: GitHub Actions, GitLab CI, and CircleCI. The goal was to measure the real-world time and cost for a typical Python/FastAPI service pipeline: lint, test, build a Docker image, and push. No marketing fluff, just timings and configs. I used the same simple FastAPI app and pytest suite across all three.

Here's the core of the test pipeline I replicated in each tool's syntax. It's a straightforward four-stage flow:

```yaml
# This is the GitHub Actions structure; others were semantically similar.
jobs:
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
- run: pip install ruff
- run: ruff check .

test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
- run: pip install -r requirements.txt pytest
- run: pytest -v --cov=app

build:
runs-on: ubuntu-latest
needs: test
steps:
- uses: actions/checkout@v4
- run: docker build -t myapp:${{ github.sha }} .
- run: docker save myapp:${{ github.sha }} | gzip > image.tar.gz

push:
runs-on: ubuntu-latest
needs: build
if: github.ref == 'refs/heads/main'
steps:
- # Simulated push step for consistency
run: echo "Pushing to registry..."
```

The key was ensuring parity: same Ubuntu runner version, same Docker layer caching strategy (where possible), and the same order of operations. I ran each pipeline 10 times on the same code commit, cleared caches between tool switches, and recorded the total execution time and any cost implications.

I'll share the raw numbers and my observations in the next post, but I'm curious: has anyone else run a similar head-to-head? What were your deal-breakers? I'm particularly interested in the developer experience around debugging failing pipelines and local testing.



   
Quote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

This is exactly the kind of thing I need. I'm about to pick a tool for my team's first real pipeline. When you say you measured cost, did you factor in the caching setup time and config for each? I've heard GitHub's cache actions can be tricky to get right compared to GitLab's built-in cache keys, and that would eat into the time savings fast.



   
ReplyQuote
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
 

Caching config is the whole ballgame with these platforms. The benchmark likely just uses the vendor's default caching, which is useless. Setting up a correct, reusable cache key strategy in GitHub Actions takes more YAML than the entire pipeline. If they didn't count that config and debugging time, the results are just a toy test.

Cost comparisons are even worse. GitLab's shared runners are a different beast than GitHub's per-minute compute. Good luck getting a real cost per build without running it for a month and watching the credits evaporate.


SQL is enough


   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

Exactly. Caching config is where most benchmarks fall apart. People post "fastest" times using perfectly primed caches on a single run, which never reflects a team's actual workflow.

The real test is how it handles a cold cache on a branch merge, or when a dependency layer changes. GitLab's cache key syntax is at least predictable. GitHub's cache action feels like a separate product you have to debug.

And you're right on cost. The pricing models are intentionally opaque. CircleCI's credits, GitHub's minutes, GitLab's compute - comparing them from a toy pipeline is pointless. The real cost is the engineering hours burned tweaking configs to shave a minute off a build.


If it's not a retention curve, I don't care.


   
ReplyQuote
(@felixr47)
Reputable Member
Joined: 2 months ago
Posts: 292
 

This is a great approach, and I'm glad you're taking the time to do a hands-on test. The initial timings you get will be a solid baseline, but you're right that the devil is in the details.

To make your data truly valuable, I'd suggest adding a second phase to your benchmark. Run the same pipeline again, but *immediately after an intentional cache bust*. Change a single dependency hash in your `requirements.txt` or `pyproject.toml`. That'll simulate a dependency update and show you how each platform handles a partial cache miss on the `pip install` layer. You'll see a huge difference in behavior there, especially in how they handle the Docker build cache.

Also, factor in the *config complexity* for each caching method. GitHub's `actions/cache` requires explicit key and restore-key YAML, GitLab uses a `cache:key:` block, and CircleCI has its own `save_cache`/`restore_cache` steps. The lines of YAML and mental overhead for each is a real cost. Maybe note that down alongside your runtime numbers.

Looking forward to seeing your results, especially the cold-start pipeline time.



   
ReplyQuote
(@henryb)
Reputable Member
Joined: 2 months ago
Posts: 214
 

That's a good point about the real cost being engineering hours. I've spent what feels like weeks just trying to get GitHub's cache to work reliably across branches. The config gets so long it's hard to even track what it's supposed to do.

Is it fair to say that if a tool's default caching is good enough, that's a major time saver? Or do most teams always end up needing custom cache keys anyway?



   
ReplyQuote
(@crmsurfer_43)
Honorable Member
Joined: 7 months ago
Posts: 398
 

Your experience with GitHub's cache action is exactly why I switched my team's project to GitLab last year. That config sprawl isn't just annoying, it becomes a real source of bugs when someone copies a workflow and doesn't understand why the cache key isn't restoring.

To your question, if the default caching is good enough, it's a massive time saver... but in my experience, that rarely lasts. You start simple, then you need cache for branches, then you need to invalidate on lockfile changes, and suddenly you're right back in the weeds.

That's the hidden cost they never mention: the maintenance burden of those "custom keys" you inevitably need.



   
ReplyQuote