Hey everyone, I've been wrestling with this for weeks and I'm hoping someone here has gone through the same fire and come out the other side. I'm in the middle of migrating a large, monolithic analytics backend (think billions of rows across MySQL and Postgres, with a MongoDB layer for real-time) to a new cloud-native architecture. Our CI/CD pipeline on CircleCI has become a major bottleneck.
The project is huge—the monorepo is about 25GB, and the build process involves extensive data validation scripts, schema migrations for all three database types, and generating a slew of artifacts. Lately, nearly every other build fails with various `insufficient resources` errors. The most common ones are:
```
ERROR: Insufficient resources to satisfy: requested memory: 4096M, available: 2048M
---
FAIL: Your container was terminated due to out of memory (OOM).
---
ERROR: No available runners matched the resource class: 'large'
```
We're on the "Performance" plan, but it feels like we're constantly hitting a ceiling. Our `.circleci/config.yml` specifies resource classes, but it seems like they're just not available when we need them. Here's a snippet of our current setup for the most resource-intensive job:
```yaml
jobs:
run-massive-migration-tests:
docker:
- image: cimg/python:3.10-node
resource_class: large
steps:
- checkout
- run:
name: Build and Seed Test Databases
command: |
./scripts/start_local_dbs.sh # Spins up local MySQL, Postgres, Mongo instances
./scripts/seed_test_data.py # This is the memory hog, loads ~10GB of test data
./scripts/run_migration_suite.py
```
We've tried:
* **Breaking the job into smaller pieces** – but the integration tests are stateful and need the full dataset.
* **Using `medium` plus memory tuning** – but the OOM killer gets triggered during the data seeding phase.
* **Caching the database snapshots** – which helped a bit, but the memory pressure during the actual migration simulation is still immense.
My gut tells me we're either:
1. Pushing CircleCI beyond its intended scale for a single job.
2. Needing a fundamentally different approach, like moving the heavy lifting to a self-hosted runner or a completely different CI/CD paradigm for this part of the pipeline.
Has anyone successfully managed **extremely resource-hungry CI jobs**—especially ones involving large-scale database migration simulations—on CircleCI? Did you find a configuration sweet spot, or did you have to offload to another platform/service for just that part?
I'm particularly interested in any benchmark numbers: actual job durations, success rates, and costs before/after any changes you made. I'll gladly share our full config and scripts if it helps! This migration is teaching me more about CI/CD limits than I ever wanted to know 😅
—B
Backup first.
That memory error is a classic sign of parallel job definitions competing for the same resource pool on their platform. The 'large' class availability issue, especially on the Performance plan, suggests you're hitting their concurrency limits for those higher-tier resources.
Have you audited your entire config for implicit default resource requests? Jobs without an explicit `resource_class` will still consume a standard allocation, which can fragment availability. You might need to implement a sequential workflow for the heaviest jobs, like the data validation and multi-database migrations, to avoid them all spinning up simultaneously and contending for the same 'large' instances.
For a monorepo of that size, are you using workspace persistence correctly? A 25GB checkout on every job, even with caching, will dominate I/O and contribute to runner exhaustion before your scripts even run.
The concurrency limit is real but not the only bottleneck. Their platform's actual instance availability fluctuates wildly. I've seen identical configs fail then pass minutes apart.
You're right about the default `resource_class`. People forget the `setup_remote_docker` command also reserves resources automatically, which can trigger the same error if you're on the edge.
For a 25GB monorepo, workspace persistence creates its own resource tax. The attach_workspace step can hang for minutes on large volumes, eating into your job time before any code runs. Sometimes you're better off with aggressive selective checkout and a package-level artifact strategy.
Benchmarks don't lie.
Sequencing heavy jobs is the right call, but you can't ignore the underlying cost. Running those large migrations back-to-back on `resource_class: large` will spike your CircleCI spend.
Check your plan's included credits and the overage rate. A `large` instance is 10x the credits of a `medium`. If your migrations take an hour each, you're burning through credits fast.
Sometimes the cheaper fix is to split the monolithic job. Run the memory-hungry validation in one `large` job, but move the schema generation to a `medium` with a longer timeout. You trade runtime for cost.
null