Skip to content
Notifications
Clear all

Step-by-step: building a runbook that your team will actually use

1 Posts
1 Users
0 Reactions
28 Views
(@charlotte2)
Reputable Member
Joined: 3 months ago
Posts: 337
Topic starter   [#20430]

Let’s be honest: most runbooks are aspirational fiction. They’re written in a vacuum, approved in a meeting, and then promptly ignored during the 2 AM firefight because they’re either too generic, painfully outdated, or read like a graduate thesis.

So how do you build one that engineers will actually *want* to use? You start by accepting a brutal truth: a runbook is a user interface for a stressed-out human. Its primary job isn't documentation—it's decision support.

First, ditch the monolithic "Step 1 through 47" format. Structure it around the alert title and the initial symptom, not the root cause. If the alert is "Database CPU > 95%", the runbook entry should start with the immediate triage actions to verify and stabilize, not a lecture on query optimization. Assume the person reading it has been woken up and can only process basic if-this-then-that logic.

Second, bake in the context they’re missing. Every step should answer "what should I see if this is working?" and "what's the escape hatch if it fails?". Instead of "Restart the service," write "Run `systemctl restart xyz` (takes ~45s, check dashboard for green health indicator. If not green after 90s, see 'Rollback to v2.1' below)." You’re pre-answering the next frantic question.

Finally, the only way to keep it alive is to make updating it a byproduct of the incident review. During the post-mortem, when you ask "What was unclear in the heat of the moment?", the answer becomes a direct edit to the runbook. If it’s not in the runbook, it didn’t happen. Treat it as a living artifact, not a compliance checkbox.

Otherwise, you’re just creating a very detailed paperweight.

Just stirring the pot


But what about the edge case?


   
Quote