Just had a fantastic ops win this week and had to share with folks who'd appreciate it! 😊 We've been running some hefty batch processing workloads on `c5.4xlarge` on-demand instances, and the bill was starting to sting.
I knew about Spot Instances, but the fear of sudden termination and workflow disruption always held me back from going all-in. Then I dug into **Spot Fleets with a mixed instances policy and fallback**. The key is setting up your request so it *automatically* falls back to On-Demand if Spot capacity isn't available, maintaining your desired capacity. Our savings so far? A consistent **~40%** reduction.
Here's a simplified version of the CloudFormation snippet that made it work. The magic is in the `MixedInstancesPolicy`:
```yaml
SpotFleet:
Type: AWS::EC2::SpotFleet
Properties:
SpotFleetRequestConfigData:
IamFleetRole: !Sub "arn:aws:iam::${AWS::AccountId}:role/aws-ec2-spot-fleet-role"
TargetCapacity: 10
AllocationStrategy: lowestPrice
OnDemandTargetCapacity: 2 # This ensures at least 2 instances are On-Demand
OnDemandAllocationStrategy: lowestPrice
LaunchTemplateConfigs:
- LaunchTemplateSpecification:
LaunchTemplateId: !Ref MyLaunchTemplate
Version: !GetAtt MyLaunchTemplate.LatestVersionNumber
Overrides:
- InstanceType: c5.4xlarge
- InstanceType: m5.4xlarge
- InstanceType: r5.4xlarge
Type: maintain
InstanceInterruptionBehavior: terminate
```
**Why this setup rocks:**
* It automatically chooses the cheapest Spot pool from a list of instance types you define (c5, m5, r5 in our case).
* The `OnDemandTargetCapacity: 2` acts as a buffer. We always have at least 2 instances running on-demand, so if Spot capacity evaporates for all our types, the fleet maintains capacity by launching more On-Demand.
* `AllocationStrategy: lowestPrice` and the mixed overrides are crucial for max savings.
We paired this with a simple Prometheus alert on the `aws_ec2_spot_fleet_request_capacity` metrics to warn us if the On-Demand count spikes, which could indicate Spot market pressure. The Grafana dashboard for this is super satisfying.
Has anyone else tried a similar setup? Curious if you've tweaked the allocation strategies or used different fallback mechanisms with other cloud providers.
If it's not monitored, it's broken.
That's a solid implementation of the strategy, and your savings figure is right in line with what I've seen for stateless, interruptible workloads. The mixed policy with explicit `OnDemandTargetCapacity` is the correct safety net.
The crucial detail many teams miss is aligning the `AllocationStrategy` with their workload's flexibility. You've used `lowestPrice`, which is perfect for batch jobs. For more sensitive fleets where instance uniformity matters, like a distributed compute cluster, `capacityOptimized` is often a better choice despite potentially slightly lower savings, as it reduces the frequency of interruptions.
One caveat to consider for your own planning: that consistent 40% savings is dependent on your region and instance family's Spot market depth. I've observed it can fluctuate, sometimes dipping to 30% or jumping to 60% during periods of low regional demand for C5s. It's wise to build a simple dashboard tracking your effective Spot vs. On-Demand price ratio over time; it can alert you to a sustained market shift that might warrant reevaluating your instance family or even region for that workload.
Nice find! That spot fleet fallback is a game-changer for cost control on interruptible work. One thing I'd add from a security perspective: when you're launching instances dynamically like this, please make sure your `LaunchTemplate` isn't referencing any secrets directly in its user data or configs. It's tempting to bake in an API key, but that's a huge risk if the template gets exported or logged. Use something like AWS Secrets Manager with an IAM instance profile, so the secret is fetched at runtime and never sits in the template itself.
Encrypt all the things.
Good point about the secrets, but that's just shuffling deck chairs if your vendor pricing is the real problem.
You lock into their secret manager, their IAM, their whole ecosystem. That's the real API key they don't want you extracting. The runtime fetch just makes you more dependent.
Savings disappear fast when you add on the cost of Secrets Manager and the engineering hours to wire it up "correctly" for what's supposed to be a cheap batch job.