Skip to content

Runbooks: what they are and why yours is already stale

Scribelet Team
11 min read

It's 3:40 a.m. and the page says the payment queue is backing up. You're half awake, you've never seen this alert before, and you do the correct thing: you open the runbook.

Step one, check the dashboard. The link 404s; that dashboard moved when the team migrated to the new Grafana instance. Step two, run the drain command. The flag it uses was renamed two releases ago and the command exits non-zero. Step four references a service that was folded into another service in March.

Somebody wrote this carefully. It was accurate the day it was written. Nobody lied to you, and nobody was lazy. The runbook just aged, quietly, in a repo where nothing forced anyone to look at it until the exact moment you needed it to be right.

This is the part of runbook documentation that almost nobody writes about. There's no shortage of guidance on what a runbook is and which sections belong in one. There's almost nothing on what happens to it over the following eighteen months. So: what a runbook is, how it differs from a playbook, the four specific ways it goes wrong, and how to build one that tells you when it has stopped being true.

What is a runbook?

A runbook is a documented, step-by-step procedure for a specific, repeatable operational task: restarting a stuck service, rotating a certificate, draining a queue, failing over a database. It names the trigger that starts the procedure, the exact commands or checks to run, the expected result at each step, and who to escalate to when the steps don't work. Writing one down at all is the point: it moves a procedure out of the head of whoever set it up, which makes a runbook one of the most direct ways to raise the bus factor on an operational task.

The value is not that it contains knowledge. It's that it contains knowledge in a form somebody can execute at 2 a.m. with no context and adrenaline in their bloodstream. Google's SRE team puts a number on it: consulting a playbook during an incident produces roughly a 3x improvement in mean time to repair compared with improvising. That's the whole case for writing them. It's also the case for keeping them accurate, which is the harder half, and the reason we built background verification into a notebook in the first place.

A useful operational runbook has five parts:

  1. Trigger. The alert, symptom, or schedule that means "open this document." Be specific enough that someone can match a page to a runbook without reading six of them.
  2. Preconditions. Access you need, systems that must be up, and anything you should check before touching production.
  3. Steps. Numbered, with the literal command and the expected output. Not "restart the consumer" but the command, and what a healthy response looks like.
  4. Verification. How you know it worked, stated as an observable: a metric returning to baseline, a queue depth falling, a health endpoint returning 200.
  5. Escalation and rollback. Who to wake, and how to undo what you just did if it made things worse.

AWS's Well-Architected Framework frames the target well: a runbook should carry the minimum information necessary to perform the procedure successfully. Everything past that minimum is surface area that can rot.

If you want to see what mature ones look like in the wild rather than in a template, GitLab publishes its production runbooks openly. It's an unusually honest look at the format, including how much of it is maintenance scaffolding.

Diagram of the five parts of a runbook: trigger, preconditions, steps, verification, and escalation, arranged as a vertical flow with a verified-on stamp

Runbook vs playbook vs SOP

These three get used interchangeably in most organizations, which is fine until someone asks you to "write a runbook" and hands you a scope you didn't expect. The distinction that actually matters is how much judgment the document assumes the reader will apply.

RunbookPlaybookSOP
ScopeOne specific task or alertA whole class of incident or eventA recurring business process
ReaderOn-call engineer, mid-incidentIncident commander, coordinatingAnyone performing the process
Judgment assumedAlmost none. Follow the steps.Substantial. Choose a branch.Some. Follow policy, apply context.
Typical lengthOne pageSeveral pagesVaries, often long
Failure mode when staleWrong command at 2 a.m.Wrong escalation pathCompliance gap

An incident runbook says "the payment queue is backing up, here is how to drain it." A playbook says "we're in a Sev1 involving payments, here is how we run the incident, who declares, who communicates, when we page the exec on call." An SOP says "here is how we onboard a vendor." The runbook is the tactical one, and it's the one that gets read fastest and checked least.

Runbook automation sits at the end of that spectrum: you take a procedure stable enough that no judgment is required, and turn its steps into a script or workflow. That's a real improvement, and it moves the decay problem rather than solving it. An automated script that calls a renamed endpoint fails just as hard as a human following a stale instruction, and it fails without a person there to notice something looked off.

Why runbooks rot faster than any other document

Every document in your organization decays. Runbooks decay faster, for three reasons that compound.

The subject changes constantly. A runbook describes a live production system, and that system ships. Every deploy, rename, migration, dependency bump, and dashboard reorganization is a chance for a step to fall out of sync. If you deploy weekly, the document has roughly fifty opportunities a year to become wrong. Design docs describe intentions, which stay fixed. Runbooks describe infrastructure, which does not. That's the same mechanism behind why engineering docs go stale generally, running at a much higher clock speed.

Almost nobody reads it. A good runbook covers a failure that's rare. Rare failure means the document sits untouched for months, so the normal correction mechanism (someone reading it and noticing an error) almost never fires. The documents that stay accurate are the ones people read constantly, and by definition your runbooks are not those.

The error surfaces at the worst moment. A wrong entry in your architecture notes costs you an awkward meeting. A wrong step in an on-call runbook costs you minutes in an outage, or worse, it makes you run a command that does something you didn't intend. The document with the highest cost of being wrong is also the one with the fewest natural opportunities to be corrected. That inversion is the entire problem, and it's a sharper version of knowledge decay than most notes ever face.

Chart showing a runbook accuracy line falling steadily over eighteen months while system changes accumulate, with the gap between them labeled as drift

The four ways a runbook actually goes wrong

Runbook rot isn't one failure. It's four, and they need different fixes.

Dead references are the most common and the least dangerous, because they announce themselves. A dashboard URL that 404s, a wiki page that moved, a Slack channel that was archived, a person who left in 2025 listed as the escalation contact. You lose time, you don't lose data.

Drifted commands are worse. The command still exists and still runs, but a flag was renamed, a default changed, or the tool moved from v1 to v2 with different semantics. The step no longer does what the document says it does, and sometimes it does something adjacent and harmful. This is the failure mode that turns a 15-minute incident into a postmortem.

Phantom steps survive their own purpose. A step exists because a service used to need a manual cache flush before restart. That service was rewritten and no longer does. The step is now a no-op at best, and the document is teaching every new engineer a piece of folklore about a system that doesn't exist. Nobody removes it, because deleting a step feels riskier than leaving it.

Then there are missing steps, which are invisible by construction. A new dependency was added to the deploy path six months ago. The runbook was never updated, so it doesn't mention it, and there's nothing in the document to indicate an absence. You only discover a missing step by executing the procedure and having it not work, which means you discover it during an incident. Reviews rarely catch these, because a reviewer reads what's on the page and evaluates whether it's correct. It is correct. It's just incomplete.

Why "review it quarterly" doesn't work

The standard advice is to assign an owner and review runbooks on a schedule. Every team that has tried this knows how it goes. The review lands in a sprint that's already full, one person skims fifteen documents in an hour, and each one gets a fresh last updated date that certifies nothing. Now the runbook is stale and it looks current, which is worse than stale and obviously old.

Calendar review fails because the calendar isn't what changed. Your runbook didn't become wrong on the first of the quarter. It became wrong the afternoon somebody merged a PR that renamed a service. The review interval and the decay event have no relationship to each other, so you're either reviewing documents that didn't need it or missing the ones that did.

The second failure is that last updated is doing work it can't do. Editing a typo bumps the date. So does rewriting every command. The field records that a human touched the file, not that anyone confirmed the contents still match production, which is the distinction behind last edited is not last verified. A runbook needs a separate verified-on field, changed only when someone actually ran or checked the procedure against the current system.

A runbook with a recent edit date and no verification date is a document that looks trustworthy and isn't. That combination is worse than no runbook, because it stops the on-call engineer from being appropriately suspicious.

A runbook that tells you when it's wrong

The fix isn't more discipline. It's building the expiry into the document and attaching the check to the thing that actually causes drift.

Start by writing down what would make each step wrong, next to the step. Not a vague "review if things change," but the specific dependency: this step breaks if the queue consumer is renamed or the drain endpoint moves. These are named drift triggers, the same technique that keeps architecture decision records honest, and they convert a periodic review into a targeted one. When you rename the consumer, you know which document to open, because the document told you in advance.

Give every runbook a verified-on date and a verification method. The method matters as much as the date: "read it" is not verification, "ran it in staging" is, and "executed during the last game day" is better. Google's incident response practice and Atlassian's incident response guidance both land on the same conclusion, which is that a procedure nobody has executed recently is a hypothesis rather than a procedure.

The strongest signal is execution. Run the runbook deliberately, on a schedule, in a controlled window: a game day, a failover drill, a staging exercise where somebody unfamiliar with the system follows the document literally and reports every place it lied. This is the only technique on the list that catches missing steps, because it's the only one that exercises the procedure rather than reading it.

For the rest, automate the checking. Scribelet's AI verification runs background checks on a schedule you set, comparing what a note claims against current sources and returning a diff: green for what's changed, red for what no longer holds, with a link to the source for each. For a runbook, the useful source is usually your own codebase, and the GitHub connector points verification at your repos instead of the open web, so a step that references a renamed service or a deleted flag gets flagged against the code that actually shipped. You review the diff and accept or dismiss it. Nothing is rewritten without you seeing it.

None of that removes the need to think about your runbooks. It removes the need to remember to think about them, which is the part that fails. If you'd rather start manual, the same instinct scales down into a 30-minute weekly review system that touches the documents most likely to have drifted rather than all of them. Point a background check at your first procedure and see what it finds; the free plan includes 10 verifications a month, which is enough to test the idea on the documents you trust least.

A runbook template with an expiry built in

Most runbook templates give you sections. This one gives you sections plus the two fields that determine whether the sections are still true.

# Runbook: [task or alert name]
 
**Trigger:** [the specific alert, symptom, or schedule]
**Owner:** [team, not a person]
**Verified on:** 2026-08-08 by [name], method: [executed in staging]
**Drift triggers:** [rename of X service, change to Y endpoint, Z dashboard migration]
 
## Preconditions
- Access required: [role, VPN, credentials]
- Confirm before starting: [system states to check]
 
## Steps
1. [Command]
   Expected: [observable result]
2. [Command]
   Expected: [observable result]
 
## Verification
- [Metric returning to baseline, queue depth, health endpoint]
 
## If this doesn't work
- Rollback: [how to undo the steps above]
- Escalate to: [team or rotation, not an individual]
 
## Known failure modes
- [Symptom] usually means [cause]. See [linked runbook].

Two rules make the template work. First, verified on only changes when someone confirms the procedure against the current system, never when someone fixes a typo. Second, drift triggers are written at the same time as the steps, while you still remember which parts of the system the procedure depends on. Trying to reconstruct them a year later is guesswork.

The document you only read when you can't afford it to be wrong

Runbooks are the highest-stakes writing most engineering teams do, and they're treated as the lowest-status. They get written during an incident retro, committed, and never opened again until the next incident, at which point their accuracy is a matter of luck and deploy velocity.

The reframe is small. A runbook isn't a description of your system; it's a claim about your system, and claims expire. Once you treat it that way, the maintenance work becomes obvious: date the claim, write down what would falsify it, and give something the job of checking. Keep your runbooks true before the next page.

Share this article

We use cookies for analytics to improve your experience. Learn more