Most organisations can run a good review meeting for a big outage. Fewer can say, a year later, which kinds of failure keep recurring, which promised fixes were actually built, and whether those fixes worked. That gap is architectural. A review produces information, and if that information lives only in a document written once and never queried, every incident is learned from in isolation and the same contributing factors return under new names.
This article treats incident review as a system: a structured incident record, timelines assembled automatically from the tools responders already use, triage that decides how much review each incident deserves, a contributing-factor taxonomy that makes incidents comparable, action items tracked to verification, and analysis across incidents. How to facilitate a single review meeting is covered in How to run a blameless postmortem; here the subject is everything around it.
A review is a pipeline, not a meeting
Picture the flow. During the incident, the pager, the incident chat channel, the deploy system, feature-flag changes and the status page each record what happened in their own format. After mitigation, those records are pulled into one incident record with an impact measure and a timeline. Triage decides whether this incident gets a full review, a light written one or none. The review produces two durable outputs, contributing factors and action items, which go into stores that outlive the document. Analysis across those stores finds themes, and the findings change alerts, runbooks, designs and staffing.
Each stage has an owner and a definition of done. The incident commander owns the record until mitigation; a named review owner owns it from then until the review is closed; owning teams own their action items until they are verified, not merely marked done. Incident response architecture covers the response side that feeds this pipeline.
The incident record
Free-form documents cannot be queried, so the core of the system is a small relational model. An incident has a peak severity, timestamps for impact start, detection, mitigation and resolution, how it was detected, the services involved and an impact measure you can add up. Timeline events, contributing factors and action items hang off it. Keep the review document for narrative and link it from the record; do not try to put judgement into columns.
CREATE TABLE incident (
id text PRIMARY KEY, -- INC-2026-0412
title text NOT NULL,
severity_peak text NOT NULL, -- SEV1..SEV4, highest reached
impact_start timestamptz NOT NULL, -- first customer impact, not first page
detected_at timestamptz NOT NULL,
mitigated_at timestamptz,
resolved_at timestamptz,
detected_by text NOT NULL, -- alert | customer | employee | synthetic
services text[] NOT NULL,
impact_units bigint, -- e.g. failed requests or customer-minutes
review_depth text, -- none | light | full (set by triage)
review_doc text,
review_state text NOT NULL DEFAULT 'pending' -- pending | drafted | reviewed | closed
);
CREATE TABLE timeline_event (
incident_id text REFERENCES incident(id),
at timestamptz NOT NULL,
source text NOT NULL, -- pager | chat | deploy | flag | status | manual
actor text,
kind text NOT NULL, -- signal | decision | action | comms | change
summary text NOT NULL,
ref text, -- link back to the original record
clock_note text -- set when the source clock is suspect
);
CREATE TABLE factor (
incident_id text REFERENCES incident(id),
category text NOT NULL, -- from the taxonomy below
detail text NOT NULL
);
CREATE TABLE action_item (
id text PRIMARY KEY,
incident_id text REFERENCES incident(id),
ticket_key text NOT NULL, -- lives in the team's normal tracker
kind text NOT NULL, -- prevent | detect | mitigate | process
owner_team text NOT NULL,
due_at date NOT NULL,
state text NOT NULL, -- open | done | verified | dropped
verified_by text -- the test, alert or drill that proved it
);Three choices in this schema matter more than they look. impact_start is when customers were first affected, not when someone was paged; the difference is your detection gap, and it is often the largest single number in a review. detected_by separates incidents your monitoring found from those customers reported, which is one of the most useful trend lines you can keep. And impact_units is one additive measure, such as failed requests or customer-minutes, so you can rank themes by harm rather than by count. Make the impact measure match the SLOs you already run; running an SLO programme covers defining them.
Timelines from the tools, not from memory
Reconstructing a timeline by hand days later is slow and biased toward what people remember. Most of it already exists as timestamped records: pages and acknowledgements, chat messages, deploys, configuration and flag changes, status-page updates and ticket events. A timeline builder pulls events from each source for the affected services in a window around the incident, normalises them to one shape in UTC, merges, and drops duplicates such as the same alert relayed into chat.
from datetime import timedelta
def build_timeline(incident, sources, window=timedelta(hours=2)):
"""Pull events around the incident from each source, normalise them, merge and de-duplicate."""
lo, hi = incident.impact_start - window, (incident.resolved_at or incident.detected_at) + window
events = []
for src in sources: # pager, chat, deploys, flags, status page
for raw in src.fetch(lo, hi, services=incident.services):
ev = src.normalise(raw) # -> at (UTC), source, actor, kind, summary, ref
if src.clock_suspect:
ev.clock_note = f"{src.name} clock; ordering within ~{src.skew_s}s is uncertain"
events.append(ev)
events.sort(key=lambda e: (e.at, e.source))
merged = []
for ev in events: # same alert relayed by pager and chat
if merged and ev.ref and ev.ref == merged[-1].ref:
continue
merged.append(ev)
return merged # a draft: people annotate, correct and add contextTreat the output as a draft. People add what tools cannot see: what responders believed at each point, why they chose one action over another, and what they tried that left no trace. Flag sources whose clocks are suspect, because a few seconds of skew can reverse the apparent order of a deploy and the first error. Keep a ref back to each original record so anyone can check an entry, and keep the chat export with the incident under the same retention as the review itself.
Triage: how much review each incident gets
Reviewing everything in depth exhausts people, and reviewing only the biggest outages misses cheap lessons from small ones. Triage makes the decision explicit and consistent. A full review gets a written analysis, a meeting and tracked actions. A light review gets the generated timeline, factors and actions from one person, with no meeting. None means the record and its factors are still captured for trend analysis.
def review_depth(inc, recent_factors):
"""Decide how much review an incident gets. Tune the thresholds to your own impact units."""
if inc.severity_peak in ("SEV1", "SEV2"):
return "full"
if inc.detected_by == "customer": # our monitoring missed it
return "full"
if any(f in recent_factors for f in inc.suspected_factors):
return "full" # a repeat is worth more than a novelty
if inc.impact_units and inc.impact_units > IMPACT_FULL:
return "full"
if inc.severity_peak == "SEV3":
return "light" # timeline, factors, actions; no meeting
return "none" # logged for trend analysis onlyTwo triggers deserve full review regardless of size: incidents customers found before your monitoring did, and incidents whose suspected factors match recent ones. A repeat is evidence that an earlier fix did not work or was never built, which is more valuable to understand than a novel failure. Run triage within two working days of resolution and record who decided and why. Light reviews should be cheap enough that teams actually do them; a 30-minute target is reasonable.
A contributing-factor taxonomy
To compare incidents you need a shared vocabulary for why they happened. Complex failures rarely have one root cause, so record several contributing factors per incident, each with a category from a short, stable list and a sentence of detail. The list below is a starting point; keep it under about fifteen categories so people apply it consistently, and review it yearly rather than adding a category for every new incident.
| Category | Examples |
|---|---|
| Change | Code deploy, configuration or flag change, schema migration |
| Capacity | Load growth, quota or limit reached, noisy neighbour |
| Dependency | Third-party outage, internal service degradation, certificate expiry |
| Detection | Missing alert, alert too slow, alert ignored as noisy |
| Mitigation | No rollback path, slow failover, runbook wrong or missing |
| Design | Single point of failure, unsafe retries, missing isolation |
| Process | Unreviewed change, unclear ownership, handoff gap |
| Knowledge | Unfamiliar system, documentation missing or stale |
Categorise factors, not people. "Engineer ran the wrong command" is not a factor; "the command had no confirmation and production and staging differed only by a flag" is two, Design and Process. Tagging is done in the review by the group, and a second person checks categorisation consistency across reviews each month.
Action items that are verified
Action items fail in two ways: they are never done, or they are done and do not work. Put each action in the owning team's normal tracker rather than a separate review tool, so it competes for priority in the same place as other work, and link the ticket key back to the incident. Classify each as prevent, detect, mitigate or process, give it an owning team and a due date proportional to severity, and require evidence of verification: the test that now fails on the old behaviour, the alert that fired in a drill, the game-day that exercised the new failover.
-- Overdue and unverified action items, oldest first
SELECT a.id, a.ticket_key, a.owner_team, a.kind, i.severity_peak,
current_date - a.due_at AS days_overdue
FROM action_item a JOIN incident i ON i.id = a.incident_id
WHERE (a.state = 'open' AND a.due_at < current_date)
OR (a.state = 'done' AND a.verified_by IS NULL)
ORDER BY days_overdue DESC;
-- Contributing-factor themes over two quarters, weighted by impact rather than count
SELECT f.category,
count(DISTINCT f.incident_id) AS incidents,
sum(i.impact_units) AS impact
FROM factor f JOIN incident i ON i.id = f.incident_id
WHERE i.impact_start >= now() - interval '6 months'
GROUP BY f.category
ORDER BY impact DESC;The first query is the weekly ageing report each engineering manager sees for their teams. The second ranks factor categories by impact over two quarters, which tells you where investment buys the most reliability. Mean time to resolve is a tempting headline, but incident durations are heavily skewed and small samples swing widely, so report distributions and counts alongside any average. On-call architecture covers feeding detection and runbook findings back into rotations.
Worked example
An API platform team finishes a SEV2: a configuration change lowered a connection-pool limit, and error rates rose for 47 minutes before a customer reported it. The timeline builder assembles 212 events from the pager, chat, the deploy log and the flag service; the review owner removes duplicates, annotates three decision points and notes that the flag service's clock was 9 seconds behind. Triage marks it full, because it is SEV2 and customer-detected.
The review records three factors: Change (the limit was edited by hand without review), Detection (the error-rate alert used a 30-minute window) and Mitigation (rollback of configuration required a full deploy). The learning store then shows that Detection appeared in four other incidents in the last six months, two of them also customer-detected. That turns one action item into a programme: alert windows across the platform are reviewed against the error budgets, and the new alert is verified in a drill that injects the same pool exhaustion. Of five action items, four are verified within 30 days; the fifth, configuration rollback without deploy, is re-scoped and tracked on the monthly ageing report.
Failure modes and trade-offs
| Failure | What you see | Fix |
|---|---|---|
| Reviews as documents only | Same factors recur under new names | Structured factors and queries across incidents |
| Everything reviewed in depth | Late, shallow reviews; people avoid declaring incidents | Triage into full, light and none |
| Actions in a separate tool | Actions never prioritised | Tickets in the owning team's tracker, linked back |
| Done but not verified | Repeat incidents after fixes | Require verification evidence before closing |
| Taxonomy sprawl | Inconsistent, unanalysable tags | Short, stable list; monthly consistency check |
| Blame in the data | Factors name people; reporting drops | Categorise conditions, not individuals |
| Metrics as targets | Incidents downgraded to avoid review | Report distributions and repeats, not quotas |
The trade-offs are between rigour and cost. More structure makes analysis possible but adds work to every review, so keep the required fields few and generate what you can. Deeper triage thresholds catch more lessons and consume more engineering time. Automating timelines saves hours but can lend false confidence to an ordering that clock skew has scrambled, so keep the annotations human.
What to do next
- Create the incident, timeline, factor and action tables, and backfill the last six months of significant incidents.
- Write a timeline builder for your two richest sources, usually the pager and the incident chat, and add deploys and flags next.
- Adopt explicit triage rules, including customer detection and repeated factors as automatic full-review triggers.
- Agree a contributing-factor taxonomy of no more than fifteen categories and tag every new review with it.
- Move action items into owning teams' trackers with due dates and verification evidence, and publish a weekly ageing report.
- Review factor themes by impact every quarter and fund the top theme as reliability work.