Skip to content
Notifications
Clear all

Guide: How to build a cutover checklist that actually works

2 Posts
2 Users
0 Reactions
30 Views
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
Topic starter   [#11349]

Alright, let's cut through the usual "here's a template from Notion" nonsense. Everyone loves a good checklist until it's 3 AM, DNS is propagating slower than a tectonic plate, and you realize step 14 depends on step 7 which you can't roll back because the checklist said "DESTROY LEGACY RESOURCES" right after the cutover. Brilliant.

A cutover checklist isn't a to-do list. It's a state machine with rollback triggers, and most teams treat it like a grocery list. The difference is the inclusion of *validation gates* and *rollback procedures for every single step*. If your checklist doesn't have those, you're just hoping.

Here's the structure I've been forced to adopt after a few spectacularly expensive migrations (that we'll politely call "learning experiences"). This is for a lift-and-shift of a stateful service, say, moving from an EC2-hosted database to RDS, but the principle applies to any cutover.

First, your checklist must exist in something that can be programmatically validated. A Confluence page is a narrative, not an operational tool. I use a markdown file in the repo that holds the Terraform, but the real magic is in the simple Python script that parses it and checks state. It looks for specific markers.

```python
# This isn't production code, it's a sanity enforcer.
import re
CHECKLIST_FILE = "CUTOVER.md"

def parse_checklist(filepath):
with open(filepath, 'r') as f:
content = f.read()
# Looks for patterns like [GATE], [ROLLBACK], [VALIDATE:some_command]
gates = re.findall(r'[GATE] (.*?)', content)
validations = re.findall(r'[VALIDATE:(.*?)]', content)
rollbacks = re.findall(r'[ROLLBACK] (.*?)', content)
return gates, validations, rollbacks
```

Now, the checklist itself. Every step is a triad.

**Pre-cutover: The Snapshot & Silence**
- [GATE] Confirm no write operations for last 24h exceed normal pattern (check CloudWatch metrics). If yes, abort.
- [VALIDATE:aws rds describe-db-snapshots --db-instance-identifier legacy-db --query 'DBSnapshots[0].Status'] Ensure final snapshot of legacy DB is `available`.
- [ROLLBACK] If snapshot fails, notify and investigate. No cutover today. Stand down.

**Cutover: The Moment of Truth**
- [GATE] Set application maintenance mode ON (API returns 503, health checks fail). Confirm all load balancer connections drain to zero.
- [VALIDATE:curl -s -o /dev/null -w "%{http_code}" https://api.internal/health ] Returns 503.
- [ROLLBACK] If cannot enable maintenance mode, abort. Disable maintenance mode. Checklist ends.
- [GATE] Update application configuration to point to new RDS endpoint. Deploy configuration only.
- [VALIDATE:aws rds describe-db-instances --db-instance-identifier new-rds --query 'DBInstances[0].DBInstanceStatus'] Returns `available`.
- [ROLLBACK] If new RDS is not available, revert config deployment to old endpoint. Disable maintenance mode.

**Post-cutover: The Unseen Chaos**
- [GATE] Disable application maintenance mode.
- [VALIDATE:run smoke tests suite ./smoke_tests.sh] Exit code must be 0.
- [ROLLBACK] If smoke tests fail > 3 times, re-enable maintenance mode. Revert config to old endpoint. Disable maintenance mode. This is a full rollback; you are now back to square one. Schedule post-mortem.
- [GATE] Monitor new RDS CPU and connections for 60 minutes. Threshold: not 2x baseline.
- [VALIDATE:check CloudWatch for CPUCreditBalance] No sustained depletion.
- [ROLLBACK] If metrics exceed threshold, initiate rollback procedure (as above). Yes, even now.

The transition for a moderately complex service using this method took us about 6 hours of actual execution time, spread over two weekends of dry runs. The dry runs are non-negotiable. They expose the "oh, that command needs this IAM role" issues you only find at 3 AM.

Without this level of paranoid, automated validation and explicit rollback for each step, you're just performing a hope-driven deployment. And hope is not a strategy; it's a liability.

-- cynical ops


Your k8s cluster is 40% idle.


   
Quote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

> "the real magic is in the simple Python script that parses it and checks sta"

I've done the markdown parser route. It works until someone puts a typo in a heading or a junior dev decides to "reformat" the checklist and breaks the regex. Then you're debugging a script at 2 AM instead of the actual cutover.

What I've switched to is a YAML file with explicit fields for each step, like:

```yaml
steps:
- id: disable_writes
action: "execute runbook/mysql_disable_writes.sh"
validate: "curl -s http://localhost:8080/health | jq -e '.writes_enabled == false'"
rollback: "execute runbook/mysql_enable_writes.sh"
rollback_validate: "curl -s http://localhost:8080/health | jq -e '.writes_enabled == true'"
depends_on: []
```

Then a simple Go binary (or Python if you hate yourself) walks the DAG, runs validate before each step, blocks on failure, and if you call `--rollback` it walks the dependency graph in reverse. That's forced the team to actually think about what "validate" means for each step.

The other thing nobody talks about is that validation gates should be *idempotent* health checks, not just "did the last command exit 0". I've seen a checklist that said "check DNS propagation" via a one-off dig, but the script didn't retry or timeout. That's not a gate, that's a gamble.


Automate everything. Twice.


   
ReplyQuote