A postmortem is the written account of an incident: what happened, why it made sense to the people involved at the time, which conditions let it happen, and what will change. Most organisations already hold the meeting. Fewer produce a document anyone reads afterwards, and fewer still learn anything from the hundredth postmortem that they did not learn from the first. The difference is not effort. It is how the document is written and what happens to it after publication.

This article treats the postmortem as an artefact. The meeting itself is covered in How to Run a Blameless Postmortem, and the review pipeline and factor taxonomy in Incident review architecture. Here the focus is on writing: language that keeps people honest without blaming them, causal analysis that finds more than one thing to fix, action items that actually reduce risk, and tooling that turns a folder of documents into evidence for investment. One incident runs through the whole article as a worked example.

Advertisement

Who the document is for

A postmortem has at least four readers, and they want different things. The responders want an accurate record that does not make them look careless for decisions that were reasonable at the time. The owning team wants a prioritised list of changes. Other engineering teams want to know whether the same conditions exist in their systems, which means they need the mechanism explained, not just the symptoms. Leadership wants impact, risk and the cost of fixing it, in one paragraph.

Write for the reader who knows least: an engineer from another team who has never seen this service. If the document works for them, it works for everyone. A useful structure:

SectionAnswers
SummaryWhat broke, for whom, for how long, is it fixed
ImpactUsers, requests, money, data, SLO budget consumed
ContextHow the system normally works, at the level needed
TimelineWhat happened and what people knew, with timestamps
AnalysisMechanism and contributing factors, as a graph
What went wellThings that limited damage, so they are kept
Action itemsOwner, due date, verification, strength of control

Local rationality: the timeline from inside

The most important idea in postmortem writing is local rationality: people do what makes sense to them given their goals, their knowledge and their attention at that moment. If an action looks foolish in hindsight, the useful question is what made it look reasonable then. The answer always points at something in the system: a dashboard that showed the wrong thing, a runbook that was out of date, a review tool that hid context, a deadline that made one risk look smaller than another.

Hindsight bias works against this: once you know the outcome, the signals that pointed to it look obvious. So for each decision point in the timeline, record the information available, the interpretation, and the alternatives considered. That turns a list of mistakes into a list of design problems.

Language gives hindsight away. Counterfactual phrases such as should have, failed to and did not notice describe a world that did not happen, and they quietly assign blame. Rewrite them as descriptions of what did happen and why.

Counterfactual or blamingDescriptive rewrite
The engineer should have run the load test before changing the TTL.The change was a one-line config edit; the team's process required load tests for code changes, not config changes.
On-call failed to notice the database alert.The database alert routed to the team that owned the service until July; the current owners were not paged.
The reviewer carelessly approved the change.The review tool showed a one-line diff with no information about traffic or cache hit rate.

Blameless does not mean unaccountable: teams still own the action items. It removes the incentive to hide information, which the next diagnosis depends on.

Advertisement

Causal graphs instead of five whys

The five whys technique asks why until you reach a root cause. It is easy to teach, and it has two structural problems. First, it produces a chain, but incidents have many causes that combine. Each why has several true answers, and the chain follows only one of them, usually the one the facilitator finds most familiar. Second, it tends to stop at a person: why did the TTL change? Because an engineer changed it. A chain that ends at a human gives you one action item, and it is a weak one.

A causal graph fixes both problems. Put the outcome at the top. Below it, list every condition that was necessary for the outcome or made it worse. For each condition, ask what made it possible, and keep going until you reach decisions about how the organisation designs, tests, deploys and operates systems. Edges mean contributed to. Several branches usually converge, and that is the point: each branch is a separate place to intervene.

Rules that keep the graph honest: every node must be a fact supported by the timeline or data, not a guess; a node may not name a person, only a condition or decision; and the graph is finished when every leaf is something the organisation can change. A graph for the worked example follows.

Checkout p99 over 8 s41 minutes, 09:02-09:43DB CPU saturatedcache misses x10Page went to wrong teamstale ownership mapRollback took 24 minchange not in deploy logCache TTL 300 s to 30 sconfig-only changeApplied at 09:00 Mondaypeak ramp, no canaryService moved teamsregistry not updatedConfig bypasses pipelineno diff, no one-click undoLoad test used warm cacheTTL risk invisibleReview saw a one-line diffno load context shownMany causes converge on one outcome; each box is a place to intervene, not a person to blame
Causal graph for the worked example: three conditions made the outcome worse, and each traces to a different organisational decision.

Worked example: a cache TTL change on a Monday morning

What happened. At 09:00 on a Monday, an engineer on the payments team shortened the product-price cache TTL from 300 seconds to 30 seconds to make price updates appear faster. The change was a configuration value, applied directly through the config service. Traffic was ramping toward its weekly peak. Cache misses rose roughly tenfold, the pricing database reached 100% CPU by 09:02, and checkout p99 latency climbed above eight seconds.

Detection and response. A latency alert fired at 09:06, but it paged the team that had owned the pricing service until a reorganisation in July. That team spent eleven minutes confirming nothing of theirs had changed before escalating. The current owners joined at 09:19. They suspected a code deploy, found none in the deploy history, and only at 09:37 found the TTL change in the config service audit log. They reverted it at 09:41; the cache warmed and latency recovered by 09:43.

Analysis. A five-whys chain would end at the engineer who changed the TTL, and produce an action item like remind engineers to load test config changes. The graph above finds three independent branches. The outcome was severe because the database could not absorb a tenfold miss rate, and the load test that should have shown this ran against a warm cache. It lasted long because paging used a stale ownership map. It took long to roll back because config changes did not appear where responders looked for changes. Each branch gets its own action item, and none of them depends on anyone being more careful.

What went well. The cache degraded to the database instead of failing closed, so checkout slowed but did not stop. Record that so a refactor does not remove it.

Action items: rank by strength of control

Not all fixes are equal. Safety engineering ranks controls by how much they depend on human attention. Adapted to software:

StrengthKind of controlExample from this incident
StrongestElimination: the hazard cannot occurPricing reads go through a request-coalescing layer, so a miss storm becomes one query per key
StrongEngineering: the system detects or blocks itConfig changes go through the deploy pipeline with canary analysis on p99 latency
StrongEngineering: recovery is fastDeploy history shows config changes, with one-click revert
MediumAutomated check in the workflowPaging targets are generated from the service catalogue and validated nightly
WeakAdministrative: process, training, remindersDocument that TTL changes affect database load

A postmortem whose action items are all administrative has not finished its analysis. Every action item needs an owner (a team, not a person who might leave), a due date, and a verification: the observable evidence that it worked. For the pipeline item, the verification is that a TTL change appears in deploy history and is stopped at canary when p99 regresses. Track items with other engineering work; an item deferred three times should be funded or closed as an accepted risk, with the reason written down.

Tie the priority to the error budget. If this incident consumed most of the quarter's checkout budget, the reliability work outranks feature work under the policy described in error budget policy. That is how a postmortem turns into scheduled engineering time rather than a wish list.

Make the document machine-readable

A postmortem is prose, but its key facts should be structured: timestamps, severity, impact, factors from a controlled vocabulary, and action items. Put them in front matter. Prose stays free-form; facts become queryable, and time to detect and mitigate fall out without a spreadsheet.

# postmortems/2026-09-14-checkout-latency.yaml  (front matter of the document)
id: PM-2026-091
title: Checkout latency after cache TTL change
severity: SEV2
started: 2026-09-14T09:02:00Z
detected: 2026-09-14T09:06:00Z
mitigated: 2026-09-14T09:43:00Z
customer_impact: "p99 checkout latency above 8 s for 41 min; 3.1% of checkouts abandoned"
factors:            # controlled vocabulary, reviewed quarterly
  - change.config_outside_pipeline
  - test.unrepresentative_load
  - alerting.routing_stale
  - rollback.not_discoverable
action_items:
  - id: AI-1
    kind: prevent          # prevent | detect | mitigate | process
    control: engineering   # elimination | engineering | administrative
    text: Route cache config through the deploy pipeline with canary analysis
    owner: payments-platform
    due: 2026-10-30
    verify: "a TTL change shows in deploy history and stops at canary on p99 regression"

A small linter run in review catches the common defects before the meeting: missing fields, counterfactual phrases in the narrative, action items without an owner or verification, and documents where every action item is administrative. It is a prompt for the author, not a gate on publication; a flagged phrase is sometimes the right one in context.

import re, sys, yaml, datetime as dt

BLAME = re.compile(r"\b(should have|failed to|careless|human error|forgot|neglected|obviously)\b", re.I)
REQUIRED = ["id", "severity", "started", "detected", "mitigated", "customer_impact", "factors", "action_items"]

def lint(path):
    raw = open(path, encoding="utf-8").read()
    front, _, body = raw.partition("\n---\n")
    doc = yaml.safe_load(front)
    problems = [f"missing field: {k}" for k in REQUIRED if not doc.get(k)]
    for n, line in enumerate(body.splitlines(), 1):
        for m in BLAME.finditer(line):
            problems.append(f"line {n}: counterfactual or blame phrase '{m.group(0)}'")
    for ai in doc.get("action_items", []):
        for k in ("owner", "due", "verify", "kind"):
            if not ai.get(k):
                problems.append(f"{ai.get('id', '?')}: action item has no {k}")
    kinds = {ai.get("control") for ai in doc.get("action_items", [])}
    if kinds and kinds <= {"administrative"}:
        problems.append("every action item is administrative (training, reminders); add a stronger control")
    return problems

if __name__ == "__main__":
    issues = lint(sys.argv[1])
    print("\n".join(issues) or "ok")
    sys.exit(1 if issues else 0)

From one postmortem to a corpus

One postmortem fixes the conditions behind one incident. A hundred postmortems, tagged with a shared factor vocabulary, tell you where the organisation keeps getting hurt. That is the most valuable output of the practice, and the one most teams never collect.

Incidenttimeline capturedDraftfacts, graph, factorsReviewlint + meetingPublishtagged, searchableAction itemsowner, due, verifyCorpusfactor tags by quarterThemesinvestment proposalsfewer, smaller incidentsA single postmortem fixes an incident; the corpus fixes the organisation
The learning loop: each document feeds a corpus, and themes across the corpus become funded projects.

Rank factors by impact rather than count. Ten small incidents tagged alerting.routing_stale may matter less than two large ones tagged change.config_outside_pipeline. A short script over the front matter is enough to start:

import collections, glob, yaml

def quarter(ts):            # "2026-09-14 09:02:00+00:00" -> "2026Q3"
    return f"{ts[:4]}Q{(int(ts[5:7]) - 1) // 3 + 1}"

counts = collections.Counter()
minutes = collections.Counter()
for path in glob.glob("postmortems/*.yaml"):
    doc = yaml.safe_load(open(path, encoding="utf-8").read().partition("\n---\n")[0])
    q = quarter(str(doc["started"]))
    impact = (doc["mitigated"] - doc["started"]).total_seconds() / 60
    for f in doc["factors"]:
        counts[(q, f)] += 1
        minutes[(q, f)] += impact

for (q, f), n in sorted(counts.items(), key=lambda kv: -minutes[kv[0]])[:15]:
    print(f"{q}  {f:40s} incidents={n:3d}  impact_minutes={minutes[(q, f)]:7.0f}")

Review the output each quarter with engineering leadership. A factor that appears across several teams is a platform investment, and the corpus is the evidence for funding it. Keep the vocabulary to twenty to forty tags and prune it, or it stops grouping anything.

Spread the learning too: a monthly digest of instructive postmortems, and links from the runbooks and code they changed so the next responder finds them.

Failure modes and trade-offs

  • Postmortem theatre. Documents are written because policy requires them and read by nobody. Symptom: action items repeat across quarters. Fix: track verification, not completion, and review the corpus.
  • Blame by another name. Names disappear but the narrative still centres one decision. Symptom: one-branch graphs. Fix: require at least one factor in each of change, detection and recovery, or state why none applied.
  • Delay. A postmortem written a month later is built from memory. Draft within two working days, using tool logs as described in incident response, and publish within a week or two.
  • Secrecy. Restricting documents to the owning team wastes the lesson. Default to organisation-wide visibility and redact sensitive details instead.

What to do next

  1. Adopt the section structure above as your template, with the summary and impact first.
  2. Add front matter with timestamps, severity, factors and action items, and start a factor vocabulary of twenty to forty tags.
  3. Run the linter on your last five postmortems and rewrite every counterfactual phrase it finds.
  4. Redo the analysis of one recent incident as a causal graph and compare its action items with the original.
  5. Classify open action items by strength of control; add an engineering control wherever only administrative ones exist.
  6. Give every action item a verification and review overdue ones in a standing meeting.
  7. Run the corpus script each quarter and bring the top three factors by impact to engineering leadership.
Key takeaway: A postmortem is only as useful as the document it produces and what happens to that document afterwards. Write for an engineer who has never seen the system. Describe decisions from inside, using local rationality instead of counterfactual blame. Replace five-whys chains with a causal graph whose leaves are organisational decisions, and rank action items by strength of control, each with an owner and a verification. Store the key facts as structured front matter so the corpus can be mined. Recurring factors across teams are the evidence for platform investment.