A runbook is a written procedure that lets someone who did not build a system diagnose and mitigate one known problem on it, quickly and safely, usually while tired and under pressure. Most teams have runbooks; few have runbooks that work at 3 a.m. The usual failure is not missing information but the wrong kind: prose that explains how the system works instead of telling the reader what to type, what they should see and what to do when they see something else.

This guide is about the craft of writing one. It covers choosing which problems deserve a runbook, collecting the steps from real incidents rather than memory, a strict grammar for steps, a worked rewrite of a weak runbook into a strong one, testing it with someone who has never seen the system, and keeping it true with a linter. The broader questions of runbook architecture, diagnosis trees and automation are covered in Runbook architecture; here we stay at the level of one page and one author.

Advertisement

Decide what deserves a runbook

Start from alerts, not from components. A useful rule: every alert that can page a human links to exactly one runbook, and an alert that has no sensible runbook should not page anyone. That rule does two jobs. It gives responders a starting point for every page, and it exposes alerts that are really dashboards in disguise.

Beyond alerts, write runbooks for recurring manual tasks that are risky or rare: rotating a certificate, failing over a database, restoring from backup, draining a node. Rarity is the key signal. A task done weekly lives in someone's fingers; a task done once a year lives nowhere unless it is written down.

CandidateWrite a runbook?Why
Paging alert with a known set of causesYes, firstHighest stress, clearest scope
Annual or quarterly risky task (restore, failover)YesNobody remembers the steps
Task done daily by the same personAutomate insteadA script is a better runbook
Novel outage with no patternNoUse the incident process, then write one afterwards
Alert nobody acts onNo: delete or demote the alertA runbook cannot fix a bad signal

Rank candidates by how often the alert fired in the last quarter and how long its incidents took to mitigate. The alert that fired twelve times with a long median time to mitigate is your first runbook.

Gather the steps from evidence, not memory

Runbooks written from memory describe how the author thinks the system fails. Runbooks written from incidents describe how it actually fails. Before writing, collect the timelines of the last few incidents for this alert, the shell and query history of whoever fixed them, and the chat transcript of the incident channel. Every command someone actually ran is a candidate step; every dead end is a candidate branch.

Then interview the person who fixes this best, with the evidence in front of you. Ask what they look at first and why, what they would see that makes them rule a cause out, what they would never do, and when they would call someone else. The last two questions produce the most valuable lines in the runbook, because experts rarely volunteer what they avoid.

Advertisement

The anatomy of a page

  1. Header. Alert name, owning team, a last-tested date, and one sentence on user impact if nothing is done. The reader must know within seconds whether this is urgent.
  2. Stop conditions. The situations in which the reader must stop and escalate rather than continue. Put them at the top, not in a footnote.
  3. Confirm. A read-only check that the alert reflects a real problem.
  4. Diagnose. Read-only commands whose outputs split the problem into branches, each branch naming its next step.
  5. Mitigate. One section per branch, safest action first, destructive actions marked and gated.
  6. Verify. How to tell the mitigation worked, with a time window.
  7. Escalate and afterwards. Who to call with what information, and what to record so the page improves.

Notice what is missing: architecture overviews, history, and explanations of how replication works. Link to those. A runbook explains only as much as the reader needs to choose the next step.

Step grammar: action, expected result, branch

Every step should have the same three parts. The action is an exact, copy-pasteable command or a precise UI path. The expected result says what normal output looks like. The branch says what to do when the output is something else. A step without an expected result leaves the reader unable to tell success from failure; a step without a branch leaves them stuck the first time reality differs.

A few rules make steps safe to run while tired:

  • Read-only steps come before any step that changes state, and are labelled read-only.
  • Destructive steps carry a WARNING line directly above them that states the consequence and who must approve.
  • Placeholders are obvious, such as <SLOT_NAME>, and never a realistic-looking value that might be pasted unchanged.
  • Each command starts with $ so tools can find it, runs on its own, and states where to run it: which host, which cluster, which context.
  • Time is explicit: "recheck in 5 minutes", not "wait a while".

Worked example: rewriting a weak runbook

Here is a runbook of the kind found on most wikis, for a Postgres primary whose data volume is filling with write-ahead log (WAL):

## Postgres disk full
Check disk space. If WAL is big, clean it up. Restart Postgres if needed.
Ask the DB team if unsure.

Every line fails a 3 a.m. reader. "Check disk space" has no command and no threshold. "Clean it up" invites the single most damaging action available, deleting files in the WAL directory, which corrupts the cluster. "Restart if needed" gives no condition. "Ask the DB team" gives no channel and no information to bring.

Rewritten using the incident history (the last three incidents were an abandoned replication slot, a broken archive command, and one unexplained case), it becomes:

# PostgresWALVolumeFilling  (owner: data-platform, last_tested: 2026-09-14)
Alert: disk used on /var/lib/postgresql above 85% on a primary.
Impact if ignored: at 100% the primary stops accepting writes.
STOP and escalate to the data-platform on-call if: this is a replica that is
also lagging, or you are about to drop a replication slot.

## 1. Confirm (read-only, 2 minutes)
  $ df -h /var/lib/postgresql
  Expected: Use% matches the alert. If below 80%, it was a spike; go to step 5.

## 2. Find what is holding WAL (read-only)
  $ psql -c "SELECT slot_name, active,
           pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS retained
           FROM pg_replication_slots
           ORDER BY pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn) DESC;"
  $ psql -c "SELECT failed_count, last_failed_wal, last_failed_time FROM pg_stat_archiver;"
  - A slot with active = f retaining many GB      -> step 3a
  - failed_count rising, last_failed_time recent  -> step 3b
  - neither                                       -> step 3c

## 3a. Inactive replication slot
  Find the consumer named by the slot and restart it. Recheck retained in 5 minutes.
  WARNING (destructive): dropping the slot frees WAL but forces that consumer
  to resynchronise from scratch. Only with data-platform approval:
    $ psql -c "SELECT pg_drop_replication_slot('<SLOT_NAME>');"

## 3b. WAL archiving failing
  Read the Postgres log for the archive_command error (credentials, bucket, network).
  Fix the cause; archived segments are then recycled at the next checkpoint.

## 3c. Unknown cause
  Escalate. Do NOT delete files in pg_wal by hand: it corrupts the cluster.

## 4. Verify
  $ df -h /var/lib/postgresql   (usage falling over 10 minutes; the alert resolves)

## 5. Afterwards
  Note in the incident channel which branch you took and anything this page got wrong.

The diagnosis queries are real: pg_replication_slots with pg_wal_lsn_diff shows how much WAL each slot is holding back, and pg_stat_archiver exposes archiving failures. The destructive action, dropping a slot, is gated behind a warning that states its cost. The unknown branch ends in escalation together with an explicit prohibition, because the most dangerous moment in an incident is when the reader has run out of instructions.

Test it with someone who has never seen the system

The author cannot test a runbook, because the author fills every gap from memory without noticing. Give the page to an engineer from another team, in a staging environment where the fault has been injected, and watch without helping. Record every question they ask and every place they pause; each is a defect in the text. Twenty minutes of this finds more problems than any review.

For alerts that cannot be safely reproduced, do a tabletop run: read the page aloud, and for each step have the tester say exactly what they would type and what they expect to see. Then put the most important runbooks into a periodic game day, and record the date as last_tested in the header so the age is visible.

A runbook's life: written from evidence, exercised, and corrected by every usePaging alertlinks one runbookRunbooktriage, branch, actIncidentnotes what failedPostmortemrunbook action itemsEdit + reviewowner approvesNovice dry runor game dayLint in CIsections, stalenesslast_testedA runbook nobody has executed since it was written is a hypothesis, not a procedure.
Runbooks stay correct only when incidents, postmortems and dry runs feed edits back into them.

Keep it true: ownership, review and a linter

Runbooks rot because systems change and nobody is responsible for the page. Give each runbook an owning team, keep it next to the code or alert definition it describes so changes are reviewed together, and make "update the runbook" a standard postmortem action item whenever a responder had to improvise. Store runbooks as Markdown in the repository rather than in a wiki, so they get diffs, reviews and CI.

A small linter enforces the structural rules mechanically, leaving review time for content:

#!/usr/bin/env python3
# runbook_lint.py: fail CI when a runbook is missing structure or has gone stale.
import datetime, pathlib, re, sys

REQUIRED = ["## 1.", "Expected:", "STOP", "## 4. Verify"]
MAX_AGE_DAYS = 90
COMMAND = re.compile(r"^\s*\$ ")
DESTRUCTIVE = re.compile(r"(\b|_)(drop|delete|truncate|failover)(\b|_)|rm -rf|kill -9", re.I)

def lint(path):
    text = path.read_text(encoding="utf-8")
    errors = [f"missing '{s}'" for s in REQUIRED if s not in text]
    m = re.search(r"last_tested:\s*(\d{4}-\d{2}-\d{2})", text)
    if not m:
        errors.append("no last_tested date")
    else:
        age = (datetime.date.today() - datetime.date.fromisoformat(m.group(1))).days
        if age > MAX_AGE_DAYS:
            errors.append(f"last tested {age} days ago")
    lines = text.splitlines()
    for i, line in enumerate(lines):
        if COMMAND.match(line) and DESTRUCTIVE.search(line) and "WARNING" not in "\n".join(lines[max(0, i - 3):i + 1]):
            errors.append(f"line {i + 1}: destructive command without a WARNING within 3 lines above")
    return errors

failed = False
for path in sorted(pathlib.Path("runbooks").glob("*.md")):
    for err in lint(path):
        print(f"{path}: {err}")
        failed = True
sys.exit(1 if failed else 0)

The required strings match the template above; adjust them to yours. The staleness check is the most valuable line: it converts "we should retest runbooks" from an intention into a failing build.

Failure modes

  • The essay. Paragraphs of explanation with commands buried inside. Fix: move explanation to a linked page, and keep the runbook to numbered steps.
  • The happy path only. Steps assume each check passes. Fix: every diagnostic step needs an expected result and at least one branch.
  • Dangerous by omission. A runbook that never says what not to do. Fix: list prohibited actions explicitly, especially the obvious-looking ones.
  • Copy-paste traps. Commands with example hostnames, wrong contexts or line breaks that change their meaning. Fix: placeholders in angle brackets, and run every command during the dry run.
  • Orphaned pages. Alerts renamed, links broken, runbook untouched. Fix: alert definitions reference the runbook path, and CI fails when the target is missing.
  • Automation without a manual fallback. A one-click fix that fails silently. Fix: the runbook documents what the automation does, so a human can do it when the button breaks.

What to do next

  1. List the alerts that paged in the last quarter and mark those without a runbook; delete or demote any alert that no runbook could address.
  2. Pick the alert with the highest count times median mitigation time and pull its last three incident timelines and command histories.
  3. Interview the person who fixes it best, asking what they check first, what they never do and when they escalate.
  4. Write the page with stop conditions, confirm, diagnose, mitigate, verify and escalate sections, using action, expected result and branch for every step.
  5. Run a novice dry run in staging and fix every question the tester asked.
  6. Add the linter to CI, a last-tested date to every page, and runbook updates to your postmortem template.
  7. Keep learning with Runbook architecture, On-call runbooks, On-call architecture and How to run a blameless postmortem.
Key takeaway: A good runbook lets someone who did not build the system fix one known problem safely while under pressure. Write one for every paging alert and every rare, risky task, and build it from incident timelines, command histories and an expert interview rather than memory. Put stop conditions first and keep explanation out. Write every step as an exact action, its expected result, and what to do otherwise, with read-only checks before changes and explicit warnings on anything destructive. Then test the page by watching a novice run it, keep it in the repository with an owner and a last-tested date, lint it in CI, and feed every incident's corrections back into it.