I wanted to share a data-driven case study from my team's recent migration. We were long-time Travis CI users, but with the changes to their pricing model and our desire for tighter GitHub integration, we switched to GitHub Actions. Our initial assumption was that the tighter workflow integration would lead to efficiency gains. However, our first full month's bill for compute minutes was roughly 2.1x our previous Travis CI invoice.
After a deep dive into our workflows and GitHub's pricing model, the reasons became clear. I'll break down the primary cost drivers we identified.
**Key Differences in Execution Model & Billing:**
* **Travis CI:** We were on a legacy plan with concurrent job limits. This acted as a natural throttle, causing builds to queue. This was frustrating for developers but capped our peak compute usage.
* **GitHub Actions:** The pricing is based on aggregate minute consumption across the entire org, with much higher default concurrent job limits. Our parallelized workflows, especially for testing matrices, now run fully in parallel, burning minutes aggressively.
**The Configuration Culprit:**
Our main cost came from a matrix build for testing across multiple PostgreSQL versions. In Travis, due to concurrency limits, these jobs often serialized. In GitHub Actions, they all fire at once.
```yaml
# Example of our expensive strategy
jobs:
test:
runs-on: ubuntu-latest
strategy:
matrix:
postgres-version: [12, 13, 14, 15]
steps:
- uses: actions/checkout@v4
- name: Start PostgreSQL
run: docker run -d -p 5432:5432 ...
# This runs 4 concurrent jobs, each consuming minutes independently
```
**Mitigation Strategies We're Implementing:**
* **Refining Job Concurrency:** Using the `max-parallel` keyword in matrix strategies and the `concurrency` group to limit simultaneous runs for non-critical workflows.
* **Caching Dependencies Aggressively:** We've overhauled our setup to use `actions/cache` for Go modules, compiled binaries, and even Docker layers, shaving 2-3 minutes off each job.
* **Rightsizing Runners:** We defaulted to `ubuntu-latest`, but are now evaluating if `ubuntu-22.04` is sufficient and testing `runs-on: [self-hosted, linux, x64]` for heavy integration suites.
* **Scheduled vs. PR Triggers:** Moving some exhaustive linting and security scans from pull request triggers to nightly scheduled runs.
The takeaway is that GitHub Actions provides more raw power and parallelism, but this directly translates to higher costs if your workflows aren't designed with its consumption model in mind. Optimization is less about raw speed and more about intelligent resource throttling and caching. We're now treating our CI minutes like database query costs—something to be monitored and optimized relentlessly.
Has anyone else performed a similar analysis? I'm particularly interested in strategies for efficiently managing matrix builds.
-- latency
sub-100ms or bust