An incident ends when service is restored. The learning only starts there, and most organisations lose it: the people involved are tired, the chat scrollback is long, and a week later the memory has been rewritten into a story with a villain. A postmortem is the structured process that turns one bad day into changes that make the next one less likely or less damaging.

"Blameless" is the part people misunderstand. It does not mean nobody is accountable or that nothing was wrong. It means the review explains how the system, including its tools, defaults, pressures and information, made the actions taken look reasonable at the time, because that is the only explanation that leads to fixes. Blame produces a quieter organisation, not a safer one: people stop reporting near misses and start editing their own timelines. This guide is the procedure, from trigger to the last closed action item.

Advertisement

When to hold one

Write the triggers down before you need them, so holding a postmortem is never an accusation. Typical triggers are: any incident at or above an agreed severity; any user-visible data loss or security exposure; an outage that consumed a meaningful share of an error budget; an incident that needed more than one escalation; and any near miss that someone thinks could have been serious. Also allow anyone involved to request one. A team that only reviews big outages misses the cheap lessons in the small ones.

Timing matters. Assign the owner within a working day of resolution, gather evidence while logs are retained and memories are fresh, and aim to hold the review within about a week. Not every incident needs a meeting; small ones can get a short written review approved asynchronously. The severity scale and live roles that feed this process are described in Incident Response Architecture.

Roles

  • Owner. Writes the document and drives it to final. Usually someone close to the incident, often the incident commander, but not a manager judging the responders.
  • Facilitator. Runs the meeting and protects its rules. Ideally someone not involved in the incident, trained to redirect blame and to keep questions open.
  • Participants. Responders, owners of the systems involved, and someone who represents affected users or support.
  • Approver. A senior engineer or manager who confirms the analysis is complete and the action items are funded, and who is accountable for them being prioritised.
The postmortem pipeline: from trigger to closed action items, with an owner at every stepTriggerseverity, impact, near missAssignowner + facilitatorGather evidencelogs, chat, deploysMerged timelineUTC, sourcedDraft analysiscontributing factorsReview meetingfacilitated, 60 minAction itemsowner, date, priorityPublishsearchable, sharedTrackingweekly review until every item is closed or re-decidedProgramme metricstime to draft, item closure, repeatsFeeds back intorunbooks, alerts, design reviews
Each stage has an owner; the process ends when the action items close, not when the meeting ends.
Advertisement

Gather evidence and build a merged timeline

The timeline is the backbone, and building it is where most of the insight appears. Pull from every source that carries timestamps: alerts, deploy and configuration history, feature-flag changes, the incident chat channel, paging records, status page updates, support tickets and dashboards. Normalise everything to UTC; time zone confusion has produced more false causal stories than any other single error. Each line should say what happened and where the evidence lives.

A small script that merges exported sources is worth keeping in your incident tooling:

import csv, json, sys
from datetime import datetime, timezone

def parse_ts(s):
    dt = datetime.fromisoformat(s.replace("Z", "+00:00"))
    return (dt if dt.tzinfo else dt.replace(tzinfo=timezone.utc)).astimezone(timezone.utc)

def load(sources):
    rows = []
    for src, path, ts_key, text_key in sources:
        with open(path, encoding="utf-8") as f:
            items = json.load(f) if path.endswith(".json") else list(csv.DictReader(f))
        for it in items:
            rows.append({"ts": parse_ts(it[ts_key]), "source": src, "text": it[text_key].strip()})
    rows.sort(key=lambda r: r["ts"])
    return rows

SOURCES = [
    ("alert",  "alerts.json",        "fired_at",  "summary"),
    ("deploy", "deploys.csv",        "finished",  "change"),
    ("chat",   "incident_chat.json", "timestamp", "message"),
    ("status", "statuspage.csv",     "posted",    "body"),
]

if __name__ == "__main__":
    rows = load(SOURCES)
    t0 = rows[0]["ts"]
    for r in rows:
        mins = int((r["ts"] - t0).total_seconds() // 60)
        print(f'{r["ts"]:%H:%M:%S}Z  +{mins:>3}m  [{r["source"]:<6}] {r["text"][:110]}')

Then annotate the merged timeline by hand with four markers: when the problem started, when it was detected, when it was mitigated and when it was resolved. The gaps between them, time to detect and time to mitigate, frame the analysis. Add a column for what responders believed at each point and what information they had. The distance between what was true and what people could see is where the most useful fixes are.

Analyse contributing factors, not a root cause

Complex systems rarely fail for one reason. A bad configuration reaches production because a validation was missing, the rollout was not staged, the alert threshold was too loose and the rollback path was slow; remove any one of those and the incident is smaller or absent. Asking for the single root cause pushes the review to stop at the first plausible answer, usually the last human action. Ask instead for the conditions that made the incident possible, made it larger, and made it last longer.

Useful questions, roughly in order:

  • What did each person know, and how did they know it? What did the tools show them?
  • Why did the action that triggered the incident look safe at the time? What would have shown otherwise?
  • What stopped the problem from being caught earlier: in review, in testing, in staging, by an alert?
  • What made diagnosis slow: missing dashboards, misleading alerts, unclear ownership, unfamiliar runbooks?
  • What made mitigation slow: no rollback path, slow deploys, permissions, coordination?
  • Where did we get lucky? Luck is an unplanned control and deserves an action item of its own.

The five-whys technique can help if you branch at each step instead of following one chain, and stop when the answer becomes "because a person made a mistake": that is a symptom to explain, not an explanation. Replace "why did you" with "what made it reasonable to", which invites the context you need.

Run the review meeting

The meeting is for checking and deepening a draft, not for writing one from scratch. Circulate the draft and timeline a day ahead. A 60-minute agenda that works:

MinutesStepFacilitator's job
0-5Ground rulesState the blameless frame; we are explaining the system, names appear only as roles
5-20Walk the timelineCorrect facts, fill gaps, add what people believed at each point
20-40Contributing factorsAsk the questions above; branch, do not stop at the first answer
40-55Action itemsOne owner each, a date, a type; reject items without owners
55-60CloseConfirm open questions, who finalises the document, and when

The facilitator's main work is language. When someone says "she should have checked the dashboard", ask "what would have prompted anyone to check it, and what would it have shown?" When someone apologises, thank them and move to the conditions. Counterfactual phrases such as "should have", "failed to" and "just needed to" are signals to redirect. Keep the most senior person in the room speaking last on causes, so the group does not anchor on their view.

Write action items that close

Action items are the product. Each one needs a single owner, a due date, a priority, a type (prevent, detect, mitigate or process) and a ticket in the team's normal backlog, not in a postmortem-only list nobody reads. Prefer changes to systems over changes to people: an automated check beats a reminder, a safer default beats a training session. Limit the list to what will be done; a dozen items with no capacity behind them is a list of things that will not happen.

Balance the types. Teams naturally write prevention items for the trigger they just saw, but detection and mitigation items pay off across many future incidents with different triggers: an alert on the right symptom, a faster rollback, a runbook step. New runbooks should follow the structure in On-Call Runbooks. Review open items weekly in an existing team meeting until each is closed or explicitly re-decided with a reason; quietly letting items expire teaches everyone the process is theatre.

A document template

# Postmortem: <short description> (<incident id>)
Status: draft | in review | final      Owner: <name>      Facilitator: <name>
Severity: <level>    Detected: <UTC>    Mitigated: <UTC>    Resolved: <UTC>

## Summary            (3-5 sentences a non-engineer can follow)
## Impact             (who, how many, how long, what they could not do; data loss yes/no)
## Timeline           (UTC, each line sourced: alert / deploy / chat / status page)
## Detection          (how we found out, how long it took, what should have told us sooner)
## Response           (what helped, what slowed us, decisions and the information behind them)
## Contributing factors
   - <condition>: <how it made the incident possible, bigger or longer>
## What went well
## Where we got lucky
## Action items
   | # | Action | Type (prevent/detect/mitigate/process) | Owner | Due | Ticket |
## Open questions

The "where we got lucky" section is the one teams most often skip and most often regret. Keep the summary readable by people outside the team, because the document's second audience is everyone else who runs similar systems.

A worked example

A payment service returned errors for 47 minutes after a routine configuration change lowered a connection pool limit. The merged timeline showed the change deployed to all regions at once, error alerts firing nine minutes later, and responders spending twenty minutes investigating a database that looked overloaded on its dashboard. The configuration change was rolled back after someone noticed it in the deploy log.

A blame-oriented review would stop at "the engineer set the wrong value". The blameless review found five contributing factors: configuration changes skipped the staged rollout that code changes used; the value had no validation or documented safe range; the error alert measured a symptom far from the cause; the database dashboard was the first one in the runbook, which pointed responders the wrong way; and recent configuration changes were not shown on the service dashboard. The engineer's value had been copied from a staging file, which was a reasonable thing to do given what the tooling showed.

Action items: stage configuration rollouts like code (prevent, platform team, two weeks); add range validation for pool settings (prevent); add a pool saturation alert (detect); overlay configuration and deploy events on service dashboards (mitigate); reorder the runbook to check recent changes first (mitigate). Every factor became an action, and none of them was "be more careful".

Failure modes of the process

  • Blame by other means. Names in the causes section, or a manager's private follow-up. People stop reporting. Keep names to roles and keep reviews separate from performance management.
  • Root cause stops at human error. Nothing changes. Ask what made the action reasonable.
  • Postmortems only for big outages. Small and near-miss lessons are lost. Set low triggers and a light written format.
  • Action items in a separate list. They never get prioritised. Put them in the normal backlog with a priority.
  • Too many actions. None close. Fund a few, record the rest as rejected with a reason.
  • Documents nobody reads. The same incident repeats in another team. Publish in a searchable place and review notable ones across teams.

Measure the programme

Track a few numbers per quarter: time from resolution to draft, time to final, the share of action items closed by their due date, the share that are detection or mitigation rather than prevention, and incidents that repeat a known contributing factor. The last is the strongest signal that the process is not working. Look at error budget consumption alongside, as described in Error budget architecture, so reliability work competes for time on evidence. If detection gaps dominate, consider rehearsing failures deliberately; see chaos engineering.

What to do next

  1. Write your postmortem triggers and severity thresholds into your incident policy, including near misses and a right for anyone to request one.
  2. Adopt the template above and keep the timeline-merging script with your incident tooling.
  3. Train two or three facilitators and make facilitation a role distinct from ownership.
  4. Run your next review with the 60-minute agenda and the language rules, and ask what made each action reasonable.
  5. Put every action item in the team backlog with an owner, date, type and priority, and review them weekly until closed.
  6. Publish finished postmortems where other teams can search them, and track repeats of known contributing factors each quarter.
Key takeaway: A blameless postmortem explains how the system made the actions taken look reasonable, because that explanation is what leads to fixes. Trigger reviews by written policy, give each one an owner and a separate facilitator, build a sourced UTC timeline, analyse contributing factors rather than one root cause, run a facilitated meeting with clear language rules, and treat action items as the product: owned, dated, typed, in the real backlog and tracked until closed.