Most organisations can run a good review meeting for a big outage. Fewer can say, a year later, which kinds of failure keep recurring, which promised fixes were actually built, and whether those fixes worked. That gap is architectural. A review produces information, and if that information lives only in a document written once and never queried, every incident is learned from in isolation and the same contributing factors return under new names.

This article treats incident review as a system: a structured incident record, timelines assembled automatically from the tools responders already use, triage that decides how much review each incident deserves, a contributing-factor taxonomy that makes incidents comparable, action items tracked to verification, and analysis across incidents. How to facilitate a single review meeting is covered in How to run a blameless postmortem; here the subject is everything around it.

Advertisement

A review is a pipeline, not a meeting

Incident review as a pipeline: records in, verified learning outPager and alertsIncident chatDeploys and flagsStatus page, ticketsIncident recordfacts, impact, timelineTriagereview depthReviewfactors, actionsAction trackertickets, ageing, verifyLearning storefactor taxonomyCross-incident analysisthemes by impact, overdue actions, repeatsfindings change alerts, runbooks and design
Sources feed a structured record; triage sets review depth; reviews produce factors and actions; the learning store makes incidents comparable and feeds changes back into operations.

Picture the flow. During the incident, the pager, the incident chat channel, the deploy system, feature-flag changes and the status page each record what happened in their own format. After mitigation, those records are pulled into one incident record with an impact measure and a timeline. Triage decides whether this incident gets a full review, a light written one or none. The review produces two durable outputs, contributing factors and action items, which go into stores that outlive the document. Analysis across those stores finds themes, and the findings change alerts, runbooks, designs and staffing.

Each stage has an owner and a definition of done. The incident commander owns the record until mitigation; a named review owner owns it from then until the review is closed; owning teams own their action items until they are verified, not merely marked done. Incident response architecture covers the response side that feeds this pipeline.

The incident record

Free-form documents cannot be queried, so the core of the system is a small relational model. An incident has a peak severity, timestamps for impact start, detection, mitigation and resolution, how it was detected, the services involved and an impact measure you can add up. Timeline events, contributing factors and action items hang off it. Keep the review document for narrative and link it from the record; do not try to put judgement into columns.

CREATE TABLE incident (
  id               text PRIMARY KEY,           -- INC-2026-0412
  title            text NOT NULL,
  severity_peak    text NOT NULL,              -- SEV1..SEV4, highest reached
  impact_start     timestamptz NOT NULL,       -- first customer impact, not first page
  detected_at      timestamptz NOT NULL,
  mitigated_at     timestamptz,
  resolved_at      timestamptz,
  detected_by      text NOT NULL,              -- alert | customer | employee | synthetic
  services         text[] NOT NULL,
  impact_units     bigint,                     -- e.g. failed requests or customer-minutes
  review_depth     text,                       -- none | light | full (set by triage)
  review_doc       text,
  review_state     text NOT NULL DEFAULT 'pending'  -- pending | drafted | reviewed | closed
);
CREATE TABLE timeline_event (
  incident_id text REFERENCES incident(id),
  at          timestamptz NOT NULL,
  source      text NOT NULL,                   -- pager | chat | deploy | flag | status | manual
  actor       text,
  kind        text NOT NULL,                   -- signal | decision | action | comms | change
  summary     text NOT NULL,
  ref         text,                            -- link back to the original record
  clock_note  text                             -- set when the source clock is suspect
);
CREATE TABLE factor (
  incident_id text REFERENCES incident(id),
  category    text NOT NULL,                   -- from the taxonomy below
  detail      text NOT NULL
);
CREATE TABLE action_item (
  id          text PRIMARY KEY,
  incident_id text REFERENCES incident(id),
  ticket_key  text NOT NULL,                   -- lives in the team's normal tracker
  kind        text NOT NULL,                   -- prevent | detect | mitigate | process
  owner_team  text NOT NULL,
  due_at      date NOT NULL,
  state       text NOT NULL,                   -- open | done | verified | dropped
  verified_by text                             -- the test, alert or drill that proved it
);

Three choices in this schema matter more than they look. impact_start is when customers were first affected, not when someone was paged; the difference is your detection gap, and it is often the largest single number in a review. detected_by separates incidents your monitoring found from those customers reported, which is one of the most useful trend lines you can keep. And impact_units is one additive measure, such as failed requests or customer-minutes, so you can rank themes by harm rather than by count. Make the impact measure match the SLOs you already run; running an SLO programme covers defining them.

Advertisement

Timelines from the tools, not from memory

Reconstructing a timeline by hand days later is slow and biased toward what people remember. Most of it already exists as timestamped records: pages and acknowledgements, chat messages, deploys, configuration and flag changes, status-page updates and ticket events. A timeline builder pulls events from each source for the affected services in a window around the incident, normalises them to one shape in UTC, merges, and drops duplicates such as the same alert relayed into chat.

from datetime import timedelta

def build_timeline(incident, sources, window=timedelta(hours=2)):
    """Pull events around the incident from each source, normalise them, merge and de-duplicate."""
    lo, hi = incident.impact_start - window, (incident.resolved_at or incident.detected_at) + window
    events = []
    for src in sources:                               # pager, chat, deploys, flags, status page
        for raw in src.fetch(lo, hi, services=incident.services):
            ev = src.normalise(raw)                   # -> at (UTC), source, actor, kind, summary, ref
            if src.clock_suspect:
                ev.clock_note = f"{src.name} clock; ordering within ~{src.skew_s}s is uncertain"
            events.append(ev)
    events.sort(key=lambda e: (e.at, e.source))
    merged = []
    for ev in events:                                 # same alert relayed by pager and chat
        if merged and ev.ref and ev.ref == merged[-1].ref:
            continue
        merged.append(ev)
    return merged                                     # a draft: people annotate, correct and add context

Treat the output as a draft. People add what tools cannot see: what responders believed at each point, why they chose one action over another, and what they tried that left no trace. Flag sources whose clocks are suspect, because a few seconds of skew can reverse the apparent order of a deploy and the first error. Keep a ref back to each original record so anyone can check an entry, and keep the chat export with the incident under the same retention as the review itself.

Triage: how much review each incident gets

Reviewing everything in depth exhausts people, and reviewing only the biggest outages misses cheap lessons from small ones. Triage makes the decision explicit and consistent. A full review gets a written analysis, a meeting and tracked actions. A light review gets the generated timeline, factors and actions from one person, with no meeting. None means the record and its factors are still captured for trend analysis.

def review_depth(inc, recent_factors):
    """Decide how much review an incident gets. Tune the thresholds to your own impact units."""
    if inc.severity_peak in ("SEV1", "SEV2"):
        return "full"
    if inc.detected_by == "customer":                 # our monitoring missed it
        return "full"
    if any(f in recent_factors for f in inc.suspected_factors):
        return "full"                                 # a repeat is worth more than a novelty
    if inc.impact_units and inc.impact_units > IMPACT_FULL:
        return "full"
    if inc.severity_peak == "SEV3":
        return "light"                                # timeline, factors, actions; no meeting
    return "none"                                     # logged for trend analysis only

Two triggers deserve full review regardless of size: incidents customers found before your monitoring did, and incidents whose suspected factors match recent ones. A repeat is evidence that an earlier fix did not work or was never built, which is more valuable to understand than a novel failure. Run triage within two working days of resolution and record who decided and why. Light reviews should be cheap enough that teams actually do them; a 30-minute target is reasonable.

A contributing-factor taxonomy

To compare incidents you need a shared vocabulary for why they happened. Complex failures rarely have one root cause, so record several contributing factors per incident, each with a category from a short, stable list and a sentence of detail. The list below is a starting point; keep it under about fifteen categories so people apply it consistently, and review it yearly rather than adding a category for every new incident.

CategoryExamples
ChangeCode deploy, configuration or flag change, schema migration
CapacityLoad growth, quota or limit reached, noisy neighbour
DependencyThird-party outage, internal service degradation, certificate expiry
DetectionMissing alert, alert too slow, alert ignored as noisy
MitigationNo rollback path, slow failover, runbook wrong or missing
DesignSingle point of failure, unsafe retries, missing isolation
ProcessUnreviewed change, unclear ownership, handoff gap
KnowledgeUnfamiliar system, documentation missing or stale

Categorise factors, not people. "Engineer ran the wrong command" is not a factor; "the command had no confirmation and production and staging differed only by a flag" is two, Design and Process. Tagging is done in the review by the group, and a second person checks categorisation consistency across reviews each month.

Action items that are verified

Action items fail in two ways: they are never done, or they are done and do not work. Put each action in the owning team's normal tracker rather than a separate review tool, so it competes for priority in the same place as other work, and link the ticket key back to the incident. Classify each as prevent, detect, mitigate or process, give it an owning team and a due date proportional to severity, and require evidence of verification: the test that now fails on the old behaviour, the alert that fired in a drill, the game-day that exercised the new failover.

-- Overdue and unverified action items, oldest first
SELECT a.id, a.ticket_key, a.owner_team, a.kind, i.severity_peak,
       current_date - a.due_at AS days_overdue
FROM action_item a JOIN incident i ON i.id = a.incident_id
WHERE (a.state = 'open' AND a.due_at < current_date)
   OR (a.state = 'done' AND a.verified_by IS NULL)
ORDER BY days_overdue DESC;

-- Contributing-factor themes over two quarters, weighted by impact rather than count
SELECT f.category,
       count(DISTINCT f.incident_id)  AS incidents,
       sum(i.impact_units)            AS impact
FROM factor f JOIN incident i ON i.id = f.incident_id
WHERE i.impact_start >= now() - interval '6 months'
GROUP BY f.category
ORDER BY impact DESC;

The first query is the weekly ageing report each engineering manager sees for their teams. The second ranks factor categories by impact over two quarters, which tells you where investment buys the most reliability. Mean time to resolve is a tempting headline, but incident durations are heavily skewed and small samples swing widely, so report distributions and counts alongside any average. On-call architecture covers feeding detection and runbook findings back into rotations.

Worked example

An API platform team finishes a SEV2: a configuration change lowered a connection-pool limit, and error rates rose for 47 minutes before a customer reported it. The timeline builder assembles 212 events from the pager, chat, the deploy log and the flag service; the review owner removes duplicates, annotates three decision points and notes that the flag service's clock was 9 seconds behind. Triage marks it full, because it is SEV2 and customer-detected.

The review records three factors: Change (the limit was edited by hand without review), Detection (the error-rate alert used a 30-minute window) and Mitigation (rollback of configuration required a full deploy). The learning store then shows that Detection appeared in four other incidents in the last six months, two of them also customer-detected. That turns one action item into a programme: alert windows across the platform are reviewed against the error budgets, and the new alert is verified in a drill that injects the same pool exhaustion. Of five action items, four are verified within 30 days; the fifth, configuration rollback without deploy, is re-scoped and tracked on the monthly ageing report.

Failure modes and trade-offs

FailureWhat you seeFix
Reviews as documents onlySame factors recur under new namesStructured factors and queries across incidents
Everything reviewed in depthLate, shallow reviews; people avoid declaring incidentsTriage into full, light and none
Actions in a separate toolActions never prioritisedTickets in the owning team's tracker, linked back
Done but not verifiedRepeat incidents after fixesRequire verification evidence before closing
Taxonomy sprawlInconsistent, unanalysable tagsShort, stable list; monthly consistency check
Blame in the dataFactors name people; reporting dropsCategorise conditions, not individuals
Metrics as targetsIncidents downgraded to avoid reviewReport distributions and repeats, not quotas

The trade-offs are between rigour and cost. More structure makes analysis possible but adds work to every review, so keep the required fields few and generate what you can. Deeper triage thresholds catch more lessons and consume more engineering time. Automating timelines saves hours but can lend false confidence to an ordering that clock skew has scrambled, so keep the annotations human.

What to do next

  1. Create the incident, timeline, factor and action tables, and backfill the last six months of significant incidents.
  2. Write a timeline builder for your two richest sources, usually the pager and the incident chat, and add deploys and flags next.
  3. Adopt explicit triage rules, including customer detection and repeated factors as automatic full-review triggers.
  4. Agree a contributing-factor taxonomy of no more than fifteen categories and tag every new review with it.
  5. Move action items into owning teams' trackers with due dates and verification evidence, and publish a weekly ageing report.
  6. Review factor themes by impact every quarter and fund the top theme as reliability work.
Key takeaway: Incident review is a system, not a meeting. Capture each incident as a structured record with impact, detection source and timestamps, build timelines from the tools responders already use, triage review depth deliberately, tag contributing factors from a short shared taxonomy, and track action items in the owning team's tracker until they are verified. Then query across incidents so recurring factors become funded work rather than repeat outages.