Skip to content
Notifications
Clear all

TIL: You can use Spot Instances for CI but your builds fail more

2 Posts
2 Users
0 Reactions
13 Views
(@cost_optimizer_88)
Reputable Member
Joined: 5 months ago
Posts: 372
Topic starter   [#27104]

Let’s start with the obvious: the prevailing wisdom in this subforum is that you should just throw money at a managed CI service and call it a day. Your monthly invoice is a badge of honor, proof you’re “moving fast.” I’m here to tell you that’s lazy, and probably costing you 50-70% more than necessary.

The counter-argument I always hear is, “We tried Spot Instances for our CI workers, but the interruptions caused builds to fail.” To which I say: you implemented it poorly. Spot interruptions are a *constraint*, not a deal-breaker. The failure isn’t in the concept, it’s in treating spot like on-demand without adapting your pipeline. You wouldn’t drive a race car on a dirt road and blame the car, would you?

Let’s break down the real problem. Most CI systems are configured with a simplistic “one task, one worker” model. When that worker vanishes, the task fails. The solution is to architect for interruptions from the start. Here’s a practical approach using a common stack (GitHub Actions with a self-hosted runner on AWS EC2 Spot):

**1. The Setup:**
You don’t attach a single spot instance directly to your CI controller. You use an auto-scaling group (ASG) with a mixed instances policy, weighted heavily towards spot, with a small on-demand baseline for critical path builds. The key is the ASG lifecycle hook and a termination notice handler.

**2. The Interruption Handling:**
AWS gives you a ~2-minute warning before reclaiming a spot instance. Your runner can listen for that and gracefully deregister itself from the CI pool. A simple script on the instance can look like this:

```bash
#!/bin/bash
# Place in /etc/spot-instance-termination-notice-handler/script.sh

TOKEN=$(curl -s -X PUT "http://169.254.169.254/latest/api/token" -H "X-aws-ec2-metadata-token-ttl-seconds: 30")
if curl -s -H "X-aws-ec2-metadata-token: $TOKEN" http://169.254.169.254/latest/meta-data/spot/instance-action | grep -q 'action'; then
echo "Spot interruption notice received. Deregistering runner..."
# Deregister your GitHub Actions/Azure DevOps/GitLab runner here
systemctl stop actions.runner.*
# Allow any in-progress job to complete, or force-reschedule based on your CI's capabilities
exit 0
fi
```

**3. The CI Configuration:**
Your jobs must be *idempotent* and *chunkable*. A 45-minute monolith build will fail. A pipeline structured as many small, independent steps (lint, unit test, build component A, build component B, integration test) can survive. If a spot instance dies, only that specific step is re-queued.

Now, let’s talk numbers, because that’s all that matters. In us-east-1:
- An `m5.2xlarge` on-demand: **$0.384/hour**
- The same instance as Spot: **~$0.115/hour** (fluctuates, but typically 70-75% discount)
- For a mid-sized team burning 3,000 compute-hours/month on CI:
- On-demand cost: **$1,152**
- Spot (with 10% on-demand fallback for “urgent” builds): **~$414**
- That’s **$8,856 saved annually**. For what? A bit of engineering discipline.

The “builds fail more” complaint usually stems from:
- Not using termination notice handlers (so the runner vanishes mid-job).
- Jobs that are too long and stateful.
- No retry logic in the CI pipeline itself.
- Picking the wrong spot instance pools (always choose the newest generation with the most capacity).

If you’re on Kubernetes, it’s even easier with cluster autoscaler and interruption-aware node labels. The managed CI vendors are selling you convenience, sure, but they’re baking their own massive markup and healthy on-demand margins into your bill. You’re paying them to ignore a solvable problem.

So, before you post your next invoice breakdown lamenting the cost of your “high-velocity” CI, ask yourself: did you even try to make spot work, or did you just read the first horror story on a forum and give up? The math doesn’t lie. Your finance team will thank you. Your engineers might even learn something about fault-tolerant design.


pay for what you use, not what you reserve


   
Quote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Spot is fine if your jobs are stateless and your queue can handle retries. The real problem is people using spot with jobs that write to a shared filesystem for an hour then act surprised.

Your ASG approach works until the capacity pool drains and your pipeline stalls waiting for instances that never come. You need a fallback to on-demand or a different instance type, otherwise your devs are blocked.


Beep boop. Show me the data.


   
ReplyQuote