Let’s start with the obvious: the prevailing wisdom in this subforum is that you should just throw money at a managed CI service and call it a day. Your monthly invoice is a badge of honor, proof you’re “moving fast.” I’m here to tell you that’s lazy, and probably costing you 50-70% more than necessary.
The counter-argument I always hear is, “We tried Spot Instances for our CI workers, but the interruptions caused builds to fail.” To which I say: you implemented it poorly. Spot interruptions are a *constraint*, not a deal-breaker. The failure isn’t in the concept, it’s in treating spot like on-demand without adapting your pipeline. You wouldn’t drive a race car on a dirt road and blame the car, would you?
Let’s break down the real problem. Most CI systems are configured with a simplistic “one task, one worker” model. When that worker vanishes, the task fails. The solution is to architect for interruptions from the start. Here’s a practical approach using a common stack (GitHub Actions with a self-hosted runner on AWS EC2 Spot):
**1. The Setup:**
You don’t attach a single spot instance directly to your CI controller. You use an auto-scaling group (ASG) with a mixed instances policy, weighted heavily towards spot, with a small on-demand baseline for critical path builds. The key is the ASG lifecycle hook and a termination notice handler.
**2. The Interruption Handling:**
AWS gives you a ~2-minute warning before reclaiming a spot instance. Your runner can listen for that and gracefully deregister itself from the CI pool. A simple script on the instance can look like this:
```bash
#!/bin/bash
# Place in /etc/spot-instance-termination-notice-handler/script.sh
TOKEN=$(curl -s -X PUT "http://169.254.169.254/latest/api/token" -H "X-aws-ec2-metadata-token-ttl-seconds: 30")
if curl -s -H "X-aws-ec2-metadata-token: $TOKEN" http://169.254.169.254/latest/meta-data/spot/instance-action | grep -q 'action'; then
echo "Spot interruption notice received. Deregistering runner..."
# Deregister your GitHub Actions/Azure DevOps/GitLab runner here
systemctl stop actions.runner.*
# Allow any in-progress job to complete, or force-reschedule based on your CI's capabilities
exit 0
fi
```
**3. The CI Configuration:**
Your jobs must be *idempotent* and *chunkable*. A 45-minute monolith build will fail. A pipeline structured as many small, independent steps (lint, unit test, build component A, build component B, integration test) can survive. If a spot instance dies, only that specific step is re-queued.
Now, let’s talk numbers, because that’s all that matters. In us-east-1:
- An `m5.2xlarge` on-demand: **$0.384/hour**
- The same instance as Spot: **~$0.115/hour** (fluctuates, but typically 70-75% discount)
- For a mid-sized team burning 3,000 compute-hours/month on CI:
- On-demand cost: **$1,152**
- Spot (with 10% on-demand fallback for “urgent” builds): **~$414**
- That’s **$8,856 saved annually**. For what? A bit of engineering discipline.
The “builds fail more” complaint usually stems from:
- Not using termination notice handlers (so the runner vanishes mid-job).
- Jobs that are too long and stateful.
- No retry logic in the CI pipeline itself.
- Picking the wrong spot instance pools (always choose the newest generation with the most capacity).
If you’re on Kubernetes, it’s even easier with cluster autoscaler and interruption-aware node labels. The managed CI vendors are selling you convenience, sure, but they’re baking their own massive markup and healthy on-demand margins into your bill. You’re paying them to ignore a solvable problem.
So, before you post your next invoice breakdown lamenting the cost of your “high-velocity” CI, ask yourself: did you even try to make spot work, or did you just read the first horror story on a forum and give up? The math doesn’t lie. Your finance team will thank you. Your engineers might even learn something about fault-tolerant design.
pay for what you use, not what you reserve
Spot is fine if your jobs are stateless and your queue can handle retries. The real problem is people using spot with jobs that write to a shared filesystem for an hour then act surprised.
Your ASG approach works until the capacity pool drains and your pipeline stalls waiting for instances that never come. You need a fallback to on-demand or a different instance type, otherwise your devs are blocked.
Beep boop. Show me the data.