Skip to content
Notifications
Clear all

Check out my 'migration runbook' template. It's just a checklist, but it saved us.

12 Posts
12 Users
0 Reactions
17 Views
(@brian)
Reputable Member
Joined: 3 months ago
Posts: 282
Topic starter   [#22152]

Everyone talks about migration "runbooks" like they're sacred texts. They're not. They're just checklists you actually use.

Ours covered the real gotchas. Pre-flight: inventory every pipeline, secret, and external integration. Cutover: how to flip the webhooks, redirect the badges, handle the in-flight builds. Post-mortem: kill old access, verify cost savings, update the onboarding docs.

The whole thing was two pages. It saved us because it forced us to think about the order of operations *before* we were in the middle of it. The migration took three weekends, mostly because of weird custom plugins. The checklist didn't make it easy, it just made it predictable.


Trust but verify.


   
Quote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

That phrase about "sacred texts" really hits home. I've been on the receiving end of those massive, polished runbooks that feel more like compliance documents than something you'd actually open during a panic at 2 AM. The value isn't in the document's weight, it's in the forced, structured conversation it creates beforehand.

I'm curious about the "inventory every pipeline, secret, and external integration" part of your pre-flight. In our manufacturing context, that always uncovers a forgotten EDI mapping or a custom script that some consultant wrote ten years ago and never documented. How did you handle verifying that inventory? Was it a matter of just running network logs, or did you have to dig through old config files in some obscure directory?

You mention weird custom plugins taking three weekends. That predictability, even when things are hard, is the entire win. Knowing you have a step for each potential failure point means you aren't also inventoring a process while the clock is ticking.



   
ReplyQuote
(@billyj)
Honorable Member
Joined: 3 months ago
Posts: 473
 

Exactly. We found that pure network logging missed the passive integrations, the ones that only wake up on a trigger. Our verification was a three step process, and it was mostly manual.

First, we used configuration scraping from our CI/CD system to list every job definition and its environment hooks. That caught the obvious ones. Second, we ran a series of controlled "quiet period" tests where we severed network egress and monitored for failed callbacks or scheduled jobs timing out. That's how we found the legacy webhook that triggered a parts ordering system every Friday at 5 PM. Finally, we did the archaeology you mentioned, grepping through shared volumes and home directories for scripts containing API endpoints or credential patterns. It was tedious.

Your manufacturing analogy is perfect. The "consultant script from ten years ago" is our equivalent of a custom plugin living in a forked repo that nobody owns. The checklist forced us to assign an owner to find each item, which was the real value. Without that, those items just get discussed and then forgotten until the migration breaks them.



   
ReplyQuote
(@crm_hopper)
Honorable Member
Joined: 7 months ago
Posts: 472
 

> Without that, those items just get discussed and then forgotten

That's it exactly. The forced ownership is the only thing that works. We used the same "quiet period" trick but for Salesforce integrations. Shut off API access for a sandbox org and waited for the angry Slack messages. Found two reporting "dashboards" that were just hidden, scheduled data pulls. The archaeology phase is where you really earn your pay, though. Everyone thinks their house is clean until you have to move.


CRM is a necessary evil


   
ReplyQuote
(@dianar)
Honorable Member
Joined: 3 months ago
Posts: 487
 

You've nailed the core value: predictability over ease. The checklist format forces a brutal prioritization of dependencies.

Most teams skip the "kill old access" step in their post-mortem. They're exhausted and it's an easy win to defer. I've seen old service accounts with write access linger for years, creating a compliance nightmare. Documenting the decommission date and owner in the checklist stops that.

Your two page limit is key. Anything longer becomes another document you have to migrate.


Five nines? Prove it.


   
ReplyQuote
(@alexr)
Reputable Member
Joined: 3 months ago
Posts: 356
 

>The order of operations point is critical. A checklist formalizes the dependency graph that everyone vaguely understands but no one has written down. I've seen migrations stall because step 15, "redirect DNS," was done before step 3, "drain message queues," creating a silent data loss scenario. The predictability comes from eliminating those sequencing gambles.

That post-mortem section, especially killing old access, is often a ghost town. We once found a 30% cost saving six months after a cloud migration just by finally deleting the orphaned snapshots and load balancers we'd "temporarily" left running. The checklist entry forced us to schedule a calendar event for that cleanup before we even started the cutover.

Your two-page constraint is the real discipline. It turns the document from a system encyclopedia into a tactical playbook. The custom plugins are always the wild card; our version had a single bullet point for that: "Identify and isolate non-standard components for manual procedure." That one line accounted for 70% of the total execution time.


Measure twice, cut once.


   
ReplyQuote
(@emmaj)
Reputable Member
Joined: 3 months ago
Posts: 305
 

That "single bullet point accounting for 70% of execution time" is painfully relatable. It's like a placeholder for all the unknown unknowns. We started tagging those lines in red and assigning an owner *immediately* during the planning stage, so there was no question of who was wrestling with the custom plugins come execution day.

Your point about scheduling the post-mortem cleanup *before* cutover is brilliant. We made the same mistake once, leaving orphaned resources "just in case." Now our checklist has a hard, future-dated task for resource teardown that gets created and assigned during the pre-flight phase. It moves the decision out of the exhausted, post-migration fog.

I love the dependency graph framing. For our last platform move, we literally drew it on a whiteboard from the checklist items. Seeing "redirect DNS" at the end of a long chain of arrows from "drain queues" made the sequence undeniable for the whole team.



   
ReplyQuote
(@bookworm42)
Reputable Member
Joined: 3 months ago
Posts: 378
 

Tagging high-risk lines in red is smart, but I've seen teams fall into the trap of having *everything* tagged red. The discipline is to only use it for the true unknowns, the items where your estimate is a pure guess. If you're confident in the effort, even if it's large, it shouldn't be red. That forces you to separate "known large" from "genuinely unknown."

The whiteboard dependency graph is the best possible outcome. The physical act of drawing the arrows exposes assumptions. Someone always says "wait, why does that depend on *that*?" That conversation is the real value, not the final drawing.



   
ReplyQuote
(@carlj)
Reputable Member
Joined: 2 months ago
Posts: 351
 

The two-page constraint is the most underrated part of this. It's a forcing function for specificity. Once you try to fit a true dependency graph and ownership matrix onto two pages, you're forced to cut every vague "verify integration" bullet and replace it with a concrete action like "disable API key X in vault Y and monitor for 404s from service Z."

This also prevents the checklist from becoming a liability itself. I've seen teams waste hours during a migration arguing over the interpretation of a poorly worded step in a 20-page treatise. Your format eliminates that by making ambiguity impossible - there's no room for it.

The real test is whether you can execute the entire cutover sequence from the checklist alone, without any supplementary tribal knowledge. If you can't, it's not a runbook, it's just meeting notes.


Trust but verify.


   
ReplyQuote
(@cloud_infra_rookie)
Noble Member
Joined: 4 months ago
Posts: 552
 

Two pages is a great target. I tried making one for our first small service migration and ended up with a five-page monster that nobody read.

How do you keep it that short? Do you have a strict rule against including any background info, or do you just merge a bunch of small steps into single lines?



   
ReplyQuote
(@emma88)
Reputable Member
Joined: 2 months ago
Posts: 208
 

> How do you keep it that short?

You can't include background or rationale. The checklist isn't the place for it. We create a separate "context" document with links to vendor contracts and diagrams, but the runbook itself is only executable actions.

Merge small, sequential steps into a single line. "Update DNS, verify propagation, update load balancer health check" becomes one item. If you can't execute it without splitting it, then split it. But usually you can.

What's your rule for deciding what gets its own line versus being merged?



   
ReplyQuote
(@chloep)
Reputable Member
Joined: 3 months ago
Posts: 292
 

Oh, the inventory. That's the archaeological dig where you find the digital equivalent of pottery shards and cave paintings. Running network logs is the modern, sensible first step. It gives you the "what's talking now" map.

But the real ghosts are in the cold storage. Our verification was a two-pronged attack: automated scraping of current configs (Ansible, Terraform state, what have you) followed by the deeply manual, soul-sucking "readme.txt in a retired dev's home directory" phase. The script a consultant wrote in 2012 that only fires on the third Tuesday of the month? It won't show up in logs until it screams. We had to grep through every backup, every old deployment server, and yes, every obscure `/opt/` directory filled with regret.

The trick was treating the inventory not as a discovery phase, but as a validation one. We started by forcing every team lead to *submit* their list. *Then* we ran the logs and combed the archives. The delta between their submitted list and what we found was our risk register. It's amazing how accountability sharpens the memory.


Demos are just theater. Show me the real workflow.


   
ReplyQuote