Every engineering organisation says its postmortems are blameless. Fewer can show it. The test is simple: after a bad outage, does the engineer who ran the command that triggered it write the most detailed account in the review, or do they go quiet and update their CV? The template is the same in both companies. The culture is not.

This article is about that culture: the organisational conditions that make people tell the truth about failure, and the habits that quietly destroy them. It does not teach how to facilitate the meeting or write the document; for those, read How to Run a Blameless Postmortem and Postmortems, in depth. Here we cover what blameless means precisely, why it is an engineering requirement rather than a kindness, how leaders and incentives shape it, how observability data supports it, and how to measure whether you have it.

Advertisement

What blameless means, and what it does not

Blameless means the review asks how did our system make this action seem reasonable instead of who made the mistake, and that nobody is punished for an honest account of what they did, saw and believed. The phrase was popularised for software teams by John Allspaw's 2012 Etsy essay on blameless postmortems and just culture, which borrowed from decades of safety work in aviation and healthcare.

It does not mean consequence-free. Two kinds of accountability are easy to confuse. Backward-looking accountability asks who should pay for the outcome. Forward-looking accountability asks people to give a full account of what happened and to own the improvements. A blameless culture demands much more of the second, not less. The engineer who deployed the bad change is accountable for explaining it in detail and for helping fix the deploy pipeline that let it through.

Blameless also does not mean anything goes. David Marx's just culture model separates three kinds of behaviour. Human error, an inadvertent slip, calls for consoling the person and fixing the system. At-risk behaviour, a shortcut whose risk was not recognised or was normalised, calls for coaching and removing the incentive to take it. Reckless behaviour, a conscious disregard of substantial and unjustifiable risk, can warrant sanction. Almost every production incident falls in the first two buckets. Reckless cases are rare, they are handled by management outside the postmortem, and they must never be used to justify making every review an investigation.

Why blame is an engineering problem

Reliability depends on information about how work is actually done: the workaround everyone uses, the alert everyone ignores, the near miss that nearly took down the database last Tuesday. That information lives in the heads of the people closest to the work, and it only flows upward if sharing it is safe.

The sociologist Ron Westrum classified organisations by how they handle information. In pathological cultures information is a source of power and messengers are shot. In bureaucratic cultures information follows channels and rules. In generative cultures information is actively sought, messengers are trained, and failure leads to inquiry. The DORA research programme adapted Westrum's typology into a survey and found that generative culture predicts better software delivery and operational performance.

Two feedback loops: what an organisation does with the news of a failureBlame loopIncidentsomeone is namedFewer reportsnear misses go unsaidThin reviewscause: human errorLatent risk stayssame failure, new personrepeatLearning loopIncidentthe system is examinedMore reportspeople share near missesRich reviewstelemetry + accountsControls improvethe trap is removedfewerBoth loops reinforce themselves. The incident count can fall in either one,so a falling count is not evidence of safety; the reporting rate is the better signal.
Blame and learning are both self-reinforcing loops. The visible incident count can fall in either loop; only the volume of honest reports distinguishes them.

The diagram shows why blame is dangerous rather than merely unpleasant. Punish the person attached to an incident and people stop reporting near misses, reviews conclude with human error, the underlying trap stays, and the next person falls into it. The organisation sees fewer reported problems and concludes it is getting safer, right up to the large incident.

Advertisement

Leadership behaviour sets the culture

Engineers learn what is safe from what senior people do in the first hours after an incident, not from the policy page. A single director asking who did this in a public channel undoes a year of blameless templates.

MomentBehaviour that builds safetyBehaviour that destroys it
During the incidentAsk what responders need; protect them from status pingsAsk who pushed the change
First review meetingSenior person speaks last; thanks the people closest to the trigger for their accountOpens with severity, cost and disappointment
QuestionsWhat did you see? What did you expect? What made sense at the time?Why did you not check? Why did you not follow the runbook?
Action itemsOwned by teams, aimed at the system, resourcedRetraining for one person, more approvals
AfterwardsSenior people share their own mistakes in reviewsIncident quietly appears in a performance review

Leaders also model the forward-looking accountability they ask for. When a leader's decision, such as a staffing cut or a deadline that forced a risky launch, contributed to an incident, saying so in the review is the strongest possible signal that the review examines the whole system, management included.

Incentives that quietly bring blame back

Most blame does not come from anyone's intent. It comes from measures and processes that make blame the rational move. Look for these:

  • Incident counts as team targets. Teams respond by downgrading severities and not declaring incidents. Track error budgets against service level objectives instead; see SLIs, SLOs and error budgets.
  • Postmortems feeding performance reviews. Once a review document can be quoted in a promotion case, every author writes defensively. Keep them separate, and say so in writing.
  • Names as subjects. Titles and summaries like Bob's deploy broke checkout attach the incident to a person for ever in search results. Use roles, such as the deploying engineer, and name the system in the title.
  • Mandatory retraining as the default action. It tells everyone the cause was a person, and it rarely changes anything.
  • On-call overload. Exhausted responders make more slips and then get reviewed for them. Alert hygiene is part of postmortem culture; see alerting that does not burn out on-call.

Evidence over memory: the role of telemetry

Blame thrives on reconstruction from memory, because memory is fallible, and people asked to remember under scrutiny defend themselves. Observability changes the conversation from whose story is right to what the system recorded. A strong review starts from evidence collected automatically:

  • Deploy and config-change events with timestamps, from the pipeline rather than from recollection.
  • Alert firings, acknowledgements and pages, with who was paged when.
  • The chat log of the incident channel, which records what responders knew and when.
  • Snapshots of the dashboards responders were looking at, because a graph re-rendered a week later with different retention or rollups can show signals that were invisible at the time.
  • Traces and logs around the trigger, preserved past normal retention for the incident window.

That evidence reconstructs what the operator could see, which is the heart of understanding why an action made sense. If the dashboard showed green while the error rate climbed, the review is about the dashboard. Teams that automate this collection, as described in Incident review architecture, spend review time on understanding rather than arguing about timelines.

A blame-language linter for drafts

Language shapes reviews. Phrases like should have, failed to and careless frame events with hindsight and point at a person. A lightweight check in the docs pipeline catches them before a review is published. It cannot judge intent, so it warns rather than blocks, and it suggests a reframing question for each hit.

import re, sys

RULES = [
    (r"\bshould have\b|\bshould've\b", "Counterfactual. Describe what happened and what made it reasonable."),
    (r"\bfailed to\b|\bneglected to\b|\bforgot to\b", "Frames an omission as a failing. What made the step easy to miss?"),
    (r"\bcareless\w*|\bsloppy\b|\blazy\b|\bincompeten\w*", "Character judgement. Remove it."),
    (r"\bhuman error\b|\boperator error\b", "A label, not a cause. What in the system allowed the error?"),
    (r"\bsimply\b|\bjust had to\b|\bobviously\b", "Hindsight. Was it obvious with the information available then?"),
]

def lint(text, names=()):
    hits = []
    for n, line in enumerate(text.splitlines(), 1):
        for pattern, advice in RULES:
            if re.search(pattern, line, re.I):
                hits.append((n, advice, line.strip()))
        for name in names:                       # roster of people involved, from the incident record
            if re.search(rf"\b{re.escape(name)}\b", line):
                hits.append((n, "Use a role instead of a name.", line.strip()))
    return hits

if __name__ == "__main__":
    doc = open(sys.argv[1], encoding="utf-8").read()
    for n, advice, line in lint(doc, names=sys.argv[2:]):
        print(f"line {n}: {advice}\n    {line}")

Worked example: rewriting a draft

A first draft for a 40-minute checkout outage read: Priya failed to check the canary dashboard and should have noticed the error spike before promoting the release. This was human error. Action: Priya to complete deployment training. The linter flags four lines. The facilitator goes back to the evidence: the pipeline log shows promotion was manual, the canary dashboard was not linked from the promotion page, the canary served 1 percent of traffic so the error spike was 6 requests per minute, and the error panel's alert threshold was set in absolute counts.

The published version reads: The deploying engineer promoted the release 9 minutes after canary start, after the documented 5-minute minimum had passed. The canary's errors were real but small in absolute terms, and nothing on the promotion page showed them. Actions: gate promotion automatically on canary error ratio (platform team); link canary health on the promotion page (deploy tooling); change the canary alert to a ratio of requests (service team). Nobody is named, the engineer gave the account that made the fix possible, and the same trap will not catch the next person.

Measuring whether the culture is real

You cannot see culture directly, but you can see its effects. Track these quarterly and look at trends, never at a single team's number in isolation:

SignalWhat a healthy trend looks like
Near misses reported per monthRising at first, then stable; a sudden drop is a warning
Share of reviews written by the people closest to the triggerHigh and rising
Median days from incident to published reviewShort and stable, for example under 10
Action items aimed at systems versus peopleOverwhelmingly systems
Action item closure rate at 90 daysHigh; open items erode trust in the process
Repeat incidents with the same contributing factorFalling
Survey: messengers are not punished for bad news (Westrum item)Agreement rising over time

Do not turn these into targets for individual teams, or you will recreate the incident-count problem one level up. Use them to find where the culture is weak and to start a conversation there.

Blameless theatre and other failure modes

Failure modeWhat it looks likeRemedy
Blameless theatreTemplate says blameless; the hallway conversation names namesLeaders model accounts of their own mistakes
Blame by omissionNo names, but the timeline makes it obvious who erredExplain context and system conditions at every step
Everything is a system problemReal reckless behaviour is avoided, which feels unfair to othersHandle it through the just culture path, outside the review
Review fatigueEvery minor incident gets a full meetingTriage depth by severity and novelty
ShelfwareAction items never closeTrack closure; review overdue items in operations meetings

What to do next

  1. Publish a one-paragraph statement of what blameless means in your organisation, including the just culture boundaries.
  2. Write down that postmortems are never input to performance reviews, signed by engineering leadership.
  3. Brief every manager on the questions to ask, and to avoid, in the first hours of an incident.
  4. Replace incident-count targets with error budgets against service level objectives.
  5. Automate evidence collection: deploy events, pages, chat logs and dashboard snapshots per incident.
  6. Add a blame-language check to the postmortem drafting workflow, as a warning.
  7. Start a lightweight near-miss report channel and celebrate the first reports publicly.
  8. Run the Westrum survey items once now and once in six months, and compare.
Key takeaway: Blameless is an engineering requirement: reliability depends on honest reports of how work is really done, and those only flow when telling the truth is safe. Blameless does not mean unaccountable. It replaces punishment with forward-looking accountability and keeps a separate just culture path for rare reckless behaviour. Leaders set the culture in the first hours after an incident, incentives such as incident-count targets quietly undo it, and telemetry-based evidence keeps reviews about the system. Measure it by reporting rates, not by incident counts.