I’ve been running some massive security scanning jobs—think 50+ parallel jobs in a matrix, each with a hefty container spinning up Trivy, Semgrep, and a custom toolchain. The kind of workload you’d expect from a monorepo with 50+ microservices.
GitHub Actions is consistently timing out on these. The 6-hour job limit feels like a hard wall we’re hitting 2-3 times a week now. The matrix strategy seems to be the culprit, where the overhead of scheduling and coordinating all those runners just burns minutes before the real work even starts.
My questions for anyone else in this trench:
* Are you seeing this 6-hour limit as a fundamental scaling blocker, or are we just configuring things poorly?
* What’s the actual upper practical limit for matrix size before the orchestration overhead becomes punitive? Our 50-parallel-job setup is clearly past it.
* Concrete proof of scale: Has anyone successfully run a matrix build with 100+ jobs without hitting the timeout? If so, what was your runner configuration (self-hosted? GitHub-hosted larger runners?) and what was the *actual* total execution time vs. wall-clock time lost to setup?
The vendor docs are… optimistic. I need real numbers from teams pushing similar scale, especially in security/scanner workloads where container spin-up is non-negotiable. The alternative is stitching together a mess of custom runners and third-party orchestration, which defeats the point of a managed CI.
The 6-hour limit is absolutely a scaling blocker for intensive matrix workflows. You're correct that the overhead isn't negligible. With 50+ parallel jobs, even a 5-minute per-job setup/teardown for container pulls and coordination can consume over 4 hours of the budget before a single line of security tooling runs.
We've successfully run matrices exceeding 100 jobs, but only by abandoning the "one toolchain per matrix job" pattern. The practical limit for a monolithic job-per-service approach on standard GitHub-hosted runners is roughly 20-30 before orchestration drag dominates. Our solution was to shift the parallelism inward: we use a matrix of 5-10 runners, each a self-hosted beefy instance (16 vCPUs, 64GB RAM), and each runner internally multiplexes scanning across 10-15 services using a custom orchestration script. This reduces the scheduler overhead dramatically.
The total execution time for scanning 120 services is now around 3 hours wall-clock, with about 20 minutes lost to runner startup and checkout. The key metric is moving the coordination complexity from GitHub Actions' scheduler into your own job logic, where you have fine-grained control. Have you considered a batched approach, where each matrix job handles a group of services?
Shifting to self-hosted runners is the right move, but the cost of those "beefy instances" you're running 5-10 of will blindside you if you don't cap concurrency. That's a massive ongoing compute bill.
The real trick is using spot or preemptible instances from your cloud provider. Your custom orchestration script can live in a container, and the main job just fires up a disposable spot instance to run it. You pay a fraction of the cost for the same vCPUs.
You're solving the scheduler problem but creating a new one if those 16 vCPU runners sit idle between jobs.
show me the bill