Every engineering organisation says its postmortems are blameless. Fewer can show it. The test is simple: after a bad outage, does the engineer who ran the command that triggered it write the most detailed account in the review, or do they go quiet and update their CV? The template is the same in both companies. The culture is not.
This article is about that culture: the organisational conditions that make people tell the truth about failure, and the habits that quietly destroy them. It does not teach how to facilitate the meeting or write the document; for those, read How to Run a Blameless Postmortem and Postmortems, in depth. Here we cover what blameless means precisely, why it is an engineering requirement rather than a kindness, how leaders and incentives shape it, how observability data supports it, and how to measure whether you have it.
What blameless means, and what it does not
Blameless means the review asks how did our system make this action seem reasonable instead of who made the mistake, and that nobody is punished for an honest account of what they did, saw and believed. The phrase was popularised for software teams by John Allspaw's 2012 Etsy essay on blameless postmortems and just culture, which borrowed from decades of safety work in aviation and healthcare.
It does not mean consequence-free. Two kinds of accountability are easy to confuse. Backward-looking accountability asks who should pay for the outcome. Forward-looking accountability asks people to give a full account of what happened and to own the improvements. A blameless culture demands much more of the second, not less. The engineer who deployed the bad change is accountable for explaining it in detail and for helping fix the deploy pipeline that let it through.
Blameless also does not mean anything goes. David Marx's just culture model separates three kinds of behaviour. Human error, an inadvertent slip, calls for consoling the person and fixing the system. At-risk behaviour, a shortcut whose risk was not recognised or was normalised, calls for coaching and removing the incentive to take it. Reckless behaviour, a conscious disregard of substantial and unjustifiable risk, can warrant sanction. Almost every production incident falls in the first two buckets. Reckless cases are rare, they are handled by management outside the postmortem, and they must never be used to justify making every review an investigation.
Why blame is an engineering problem
Reliability depends on information about how work is actually done: the workaround everyone uses, the alert everyone ignores, the near miss that nearly took down the database last Tuesday. That information lives in the heads of the people closest to the work, and it only flows upward if sharing it is safe.
The sociologist Ron Westrum classified organisations by how they handle information. In pathological cultures information is a source of power and messengers are shot. In bureaucratic cultures information follows channels and rules. In generative cultures information is actively sought, messengers are trained, and failure leads to inquiry. The DORA research programme adapted Westrum's typology into a survey and found that generative culture predicts better software delivery and operational performance.
The diagram shows why blame is dangerous rather than merely unpleasant. Punish the person attached to an incident and people stop reporting near misses, reviews conclude with human error, the underlying trap stays, and the next person falls into it. The organisation sees fewer reported problems and concludes it is getting safer, right up to the large incident.
Leadership behaviour sets the culture
Engineers learn what is safe from what senior people do in the first hours after an incident, not from the policy page. A single director asking who did this in a public channel undoes a year of blameless templates.
| Moment | Behaviour that builds safety | Behaviour that destroys it |
|---|---|---|
| During the incident | Ask what responders need; protect them from status pings | Ask who pushed the change |
| First review meeting | Senior person speaks last; thanks the people closest to the trigger for their account | Opens with severity, cost and disappointment |
| Questions | What did you see? What did you expect? What made sense at the time? | Why did you not check? Why did you not follow the runbook? |
| Action items | Owned by teams, aimed at the system, resourced | Retraining for one person, more approvals |
| Afterwards | Senior people share their own mistakes in reviews | Incident quietly appears in a performance review |
Leaders also model the forward-looking accountability they ask for. When a leader's decision, such as a staffing cut or a deadline that forced a risky launch, contributed to an incident, saying so in the review is the strongest possible signal that the review examines the whole system, management included.
Incentives that quietly bring blame back
Most blame does not come from anyone's intent. It comes from measures and processes that make blame the rational move. Look for these:
- Incident counts as team targets. Teams respond by downgrading severities and not declaring incidents. Track error budgets against service level objectives instead; see SLIs, SLOs and error budgets.
- Postmortems feeding performance reviews. Once a review document can be quoted in a promotion case, every author writes defensively. Keep them separate, and say so in writing.
- Names as subjects. Titles and summaries like Bob's deploy broke checkout attach the incident to a person for ever in search results. Use roles, such as the deploying engineer, and name the system in the title.
- Mandatory retraining as the default action. It tells everyone the cause was a person, and it rarely changes anything.
- On-call overload. Exhausted responders make more slips and then get reviewed for them. Alert hygiene is part of postmortem culture; see alerting that does not burn out on-call.
Evidence over memory: the role of telemetry
Blame thrives on reconstruction from memory, because memory is fallible, and people asked to remember under scrutiny defend themselves. Observability changes the conversation from whose story is right to what the system recorded. A strong review starts from evidence collected automatically:
- Deploy and config-change events with timestamps, from the pipeline rather than from recollection.
- Alert firings, acknowledgements and pages, with who was paged when.
- The chat log of the incident channel, which records what responders knew and when.
- Snapshots of the dashboards responders were looking at, because a graph re-rendered a week later with different retention or rollups can show signals that were invisible at the time.
- Traces and logs around the trigger, preserved past normal retention for the incident window.
That evidence reconstructs what the operator could see, which is the heart of understanding why an action made sense. If the dashboard showed green while the error rate climbed, the review is about the dashboard. Teams that automate this collection, as described in Incident review architecture, spend review time on understanding rather than arguing about timelines.
A blame-language linter for drafts
Language shapes reviews. Phrases like should have, failed to and careless frame events with hindsight and point at a person. A lightweight check in the docs pipeline catches them before a review is published. It cannot judge intent, so it warns rather than blocks, and it suggests a reframing question for each hit.
import re, sys
RULES = [
(r"\bshould have\b|\bshould've\b", "Counterfactual. Describe what happened and what made it reasonable."),
(r"\bfailed to\b|\bneglected to\b|\bforgot to\b", "Frames an omission as a failing. What made the step easy to miss?"),
(r"\bcareless\w*|\bsloppy\b|\blazy\b|\bincompeten\w*", "Character judgement. Remove it."),
(r"\bhuman error\b|\boperator error\b", "A label, not a cause. What in the system allowed the error?"),
(r"\bsimply\b|\bjust had to\b|\bobviously\b", "Hindsight. Was it obvious with the information available then?"),
]
def lint(text, names=()):
hits = []
for n, line in enumerate(text.splitlines(), 1):
for pattern, advice in RULES:
if re.search(pattern, line, re.I):
hits.append((n, advice, line.strip()))
for name in names: # roster of people involved, from the incident record
if re.search(rf"\b{re.escape(name)}\b", line):
hits.append((n, "Use a role instead of a name.", line.strip()))
return hits
if __name__ == "__main__":
doc = open(sys.argv[1], encoding="utf-8").read()
for n, advice, line in lint(doc, names=sys.argv[2:]):
print(f"line {n}: {advice}\n {line}")
Worked example: rewriting a draft
A first draft for a 40-minute checkout outage read: Priya failed to check the canary dashboard and should have noticed the error spike before promoting the release. This was human error. Action: Priya to complete deployment training. The linter flags four lines. The facilitator goes back to the evidence: the pipeline log shows promotion was manual, the canary dashboard was not linked from the promotion page, the canary served 1 percent of traffic so the error spike was 6 requests per minute, and the error panel's alert threshold was set in absolute counts.
The published version reads: The deploying engineer promoted the release 9 minutes after canary start, after the documented 5-minute minimum had passed. The canary's errors were real but small in absolute terms, and nothing on the promotion page showed them. Actions: gate promotion automatically on canary error ratio (platform team); link canary health on the promotion page (deploy tooling); change the canary alert to a ratio of requests (service team). Nobody is named, the engineer gave the account that made the fix possible, and the same trap will not catch the next person.
Measuring whether the culture is real
You cannot see culture directly, but you can see its effects. Track these quarterly and look at trends, never at a single team's number in isolation:
| Signal | What a healthy trend looks like |
|---|---|
| Near misses reported per month | Rising at first, then stable; a sudden drop is a warning |
| Share of reviews written by the people closest to the trigger | High and rising |
| Median days from incident to published review | Short and stable, for example under 10 |
| Action items aimed at systems versus people | Overwhelmingly systems |
| Action item closure rate at 90 days | High; open items erode trust in the process |
| Repeat incidents with the same contributing factor | Falling |
| Survey: messengers are not punished for bad news (Westrum item) | Agreement rising over time |
Do not turn these into targets for individual teams, or you will recreate the incident-count problem one level up. Use them to find where the culture is weak and to start a conversation there.
Blameless theatre and other failure modes
| Failure mode | What it looks like | Remedy |
|---|---|---|
| Blameless theatre | Template says blameless; the hallway conversation names names | Leaders model accounts of their own mistakes |
| Blame by omission | No names, but the timeline makes it obvious who erred | Explain context and system conditions at every step |
| Everything is a system problem | Real reckless behaviour is avoided, which feels unfair to others | Handle it through the just culture path, outside the review |
| Review fatigue | Every minor incident gets a full meeting | Triage depth by severity and novelty |
| Shelfware | Action items never close | Track closure; review overdue items in operations meetings |
What to do next
- Publish a one-paragraph statement of what blameless means in your organisation, including the just culture boundaries.
- Write down that postmortems are never input to performance reviews, signed by engineering leadership.
- Brief every manager on the questions to ask, and to avoid, in the first hours of an incident.
- Replace incident-count targets with error budgets against service level objectives.
- Automate evidence collection: deploy events, pages, chat logs and dashboard snapshots per incident.
- Add a blame-language check to the postmortem drafting workflow, as a warning.
- Start a lightweight near-miss report channel and celebrate the first reports publicly.
- Run the Westrum survey items once now and once in six months, and compare.