An incident is any unplanned event that degrades a service enough that someone has to stop what they are doing and fix it. Every organization with production systems responds to incidents somehow. The difference between a team that recovers in twenty minutes and one that takes four hours is rarely technical skill. It is usually structure: who decides, who talks to whom, which options are tried first, and how everyone knows what is going on.
This article treats incident response as an architecture with components, interfaces and failure modes, just like the systems it protects. It covers how signals become incidents, how to define severity, the roles that keep a response coordinated, an incident state machine you can implement in a chat bot, how to communicate, how to choose a mitigation, and what goes wrong. It ends where the incident review begins. Related practices have their own pages: runbooks, rollback strategy and error budgets.
The goal: restore service, then learn
Incident response has one primary goal during the incident: reduce customer impact as quickly as it can be done safely. Finding the root cause is a secondary goal and often waits until later. That ordering sounds obvious, but it is the most commonly violated principle in practice. Engineers are trained to debug, and an unexplained failure pulls them towards understanding it when rolling back the last change would end the impact in five minutes.
The second goal is to stay coordinated as the response grows. A ten-person incident across three teams without structure produces duplicate work, conflicting changes and silence towards customers.
Signals and triage
Incidents start from signals: automated alerts, customer reports through support, and staff who notice something wrong. The best alerts are symptom-based. They fire on what users experience, such as error rate, latency against a service level objective or failed checkouts, rather than on causes such as high CPU, which may or may not matter. Cause-based alerts belong on dashboards that responders use during diagnosis, not on pagers.
Every page goes to a named on-call engineer who acknowledges it and triages within a few minutes. Triage answers three questions. Is it real? Is it affecting users or about to? Is it bigger than one person can handle quickly? If the answers are yes, yes and possibly, the engineer declares an incident. Make declaring cheap and blameless. An incident declared and closed as minor costs a few minutes, while one declared an hour late costs the hour. A useful rule is that if you are wondering whether to declare, declare.
Give non-alert signals a path too: support staff and anyone else should know how to page on-call, because many serious incidents are first noticed by a person.
Severity: a scale you can apply in a minute
Severity decides how many people are pulled in, how often stakeholders are updated and who must be told. It should be based on impact, what customers and the business are experiencing, not on cause or on how interesting the problem is. It must also be quick to apply, because the person applying it is under pressure. A typical scale:
| Severity | Criteria (any one) | Response |
|---|---|---|
| SEV1 | core function down for most users; data loss or security breach in progress; major revenue stopped | IC plus all roles, leadership informed, updates every 30 minutes, public status |
| SEV2 | core function degraded or down for a significant subset; workaround exists but is painful | IC plus operations lead, updates hourly, public status if customers notice |
| SEV3 | minor feature broken or degraded; small number of users affected | on-call owns it, updates at key milestones |
| SEV4 | no current user impact; risk of impact if ignored | ticket, handled in working hours |
Two rules keep the scale honest. First, severity can change during an incident, upwards or downwards, and changing it is normal. Second, when in doubt, pick the higher severity. Over-responding to a SEV3 wastes some time; under-responding to a SEV1 lets the damage run.
Roles: incident command
The role structure most teams use descends from the Incident Command System developed for emergency services. Its core idea is that the person coordinating the response is not the person typing commands. When the most senior engineer both debugs and coordinates, both jobs suffer, and nobody notices that the status page has not been updated for an hour.
- Incident commander (IC). Owns the incident. Sets priorities, assigns work, decides between options, keeps the timeline of decisions and declares resolution. The IC does not debug. For small incidents the IC may be the on-call engineer; once others join, the IC should hand hands-on work to them.
- Operations lead. Directs technical work: who investigates what, which changes are applied, in what order. In large incidents there may be one per affected system, reporting to the IC.
- Communications lead. Writes stakeholder updates and the public status page on a fixed cadence, and shields responders from questions.
- Scribe. Records the timeline: what was observed, what was tried, what was decided and when. This record is the raw material for the review, and reconstructing it afterwards from chat logs is slow and unreliable.
- Subject matter experts. Engineers pulled in for specific systems. They report findings to the operations lead and make changes only when asked.
Hand roles over explicitly in the channel, so nobody doubts who is IC, and when more than five or six people report to one person, split the work into streams with their own leads.
The incident state machine
An incident moves through a small number of states, and making them explicit, in a chat bot or in your incident tracker, pays off in consistency. Each transition is a decision someone makes and records. The code below is a minimal model that enforces the transitions, requires a commander before declaration and tracks the update cadence for each severity. It is illustrative and not tied to any product.
from dataclasses import dataclass, field
from datetime import datetime, timezone
from enum import Enum
class State(Enum):
TRIAGE = "triage"
ACTIVE = "active" # declared, IC assigned, mitigation in progress
MONITORING = "monitoring" # mitigation applied, watching for relapse
RESOLVED = "resolved"
REVIEW = "review"
ALLOWED = {
State.TRIAGE: {State.ACTIVE, State.RESOLVED}, # resolved here = false alarm
State.ACTIVE: {State.MONITORING},
State.MONITORING: {State.ACTIVE, State.RESOLVED}, # relapse goes back to ACTIVE
State.RESOLVED: {State.REVIEW},
State.REVIEW: set(),
}
@dataclass
class Incident:
title: str
severity: int # 1 is worst
state: State = State.TRIAGE
commander: str | None = None
log: list = field(default_factory=list)
def note(self, who, text):
self.log.append((datetime.now(timezone.utc).isoformat(), who, text))
def move(self, who, new, reason):
if new not in ALLOWED[self.state]:
raise ValueError(f"{self.state.value} -> {new.value} is not allowed")
if new is State.ACTIVE and not self.commander:
raise ValueError("assign an incident commander before declaring")
self.note(who, f"{self.state.value} -> {new.value}: {reason}")
self.state = new
def update_due(self, last_update, now):
"""Minutes until the next stakeholder update is due, by severity."""
cadence = {1: 30, 2: 60, 3: 240}.get(self.severity)
if cadence is None or self.state not in (State.ACTIVE, State.MONITORING):
return None
return cadence - (now - last_update).total_seconds() / 60The important transitions are the loop between active and monitoring, because first mitigations often fail or only partly work, and the separation of resolved from review. Resolution means customer impact has ended and the team can stand down. The review, with its timeline and follow-up actions, is a separate piece of work with its own owner and deadline.
Communication architecture
Communication has three audiences with different needs. Responders need a single place for the incident, usually a dedicated chat channel created when the incident is declared, with a pinned summary of state, severity, roles and current hypothesis. Internal stakeholders, including support, account managers and leadership, need regular, plain-language updates on impact and expected recovery. Customers need honest, timely status on a public page. Keep side conversations out of direct messages, since decisions made there never reach the log.
Updates work best on a fixed cadence with a fixed template. A cadence means stakeholders stop interrupting responders to ask for news. A template means the communications lead can write an update in two minutes and it always answers the same questions:
[SEV2] Checkout errors for some card payments - update 3 (14:40 UTC)
Impact: About 18% of card checkouts in EU are failing. Other regions and payment
methods are unaffected. No data loss.
Status: Mitigating. The payment-routing config change at 13:52 is being rolled back;
error rate is falling in the first two canary zones.
Next: Rollback completes about 15:00. Next update 15:10 or sooner if things change.
Contact: IC: A. Rivera. Questions in #inc-2026-10-01-checkout, not in DMs.Say what is known, what is not known and when the next update will come. Never promise a fix time you cannot back up; give the time of the next update instead. And send an update even when nothing has changed, because silence is read as either neglect or a worse problem.
Mitigate first: the mitigation ladder
Mitigation is anything that reduces customer impact, whether or not it fixes the underlying fault. Teams that recover quickly have a ranked list of mitigations they try before deep debugging, ordered by speed and safety:
- Roll back the most recent change to the affected system: deploy, configuration or schema. Most incidents follow a change, and rollback is usually the fastest safe action if it has been practiced.
- Turn off the feature with a feature flag or kill switch, which can be faster and narrower than a rollback.
- Move traffic away from the failing zone, region or instance group, if capacity elsewhere allows.
- Shed or throttle load to protect the core path, for example by disabling expensive non-essential endpoints.
- Add capacity when the fault is saturation, remembering that new capacity takes time to warm up.
- Fail over to a standby, following the disaster recovery procedure, when the primary cannot be restored quickly.
The IC's question at each step is: what is the fastest action that reduces impact without making things worse? Record each mitigation with its time and outcome, so that the team can tell which action actually helped.
Worked example: a bad configuration push
At 13:52 a payment-routing configuration change reaches production. At 14:01 a symptom alert fires on card checkout failures in the EU region and pages the on-call engineer, who acknowledges at 14:03, sees an error rate of 18 percent and declares a SEV2 at 14:06, naming themselves IC and paging the payments on-call as operations lead. A support engineer joins as communications lead and posts the first update at 14:12.
The operations lead offers two hypotheses, the config change or a degraded card processor; other regions use the processor without errors, so the IC chooses the first rung of the ladder and orders a rollback at 14:24. Error rates fall in the first zones at 14:38 and recover fully at 14:58. The IC moves the incident to monitoring, watches for thirty minutes and resolves it at 15:30, handing the review to the payments team with the scribe's timeline.
Two timings from that story are worth tracking: 9 minutes from change to alert, and 18 minutes from declaration to the rollback decision. The review will ask why the change was not caught by a canary, and why rollback took eighteen minutes to choose when the change was the obvious suspect. Both are process findings, not individual failures.
Tooling and its failure domain
Incident tooling has one requirement that normal tooling does not: it must work while your systems are broken. Teams discover mid-incident that the status page runs on the platform that is down, or that chat login depends on the failing single sign-on. Map each tool in the response path, paging, chat, status page, runbooks, dashboards, deployment and access, to the infrastructure it depends on, and make sure the critical ones sit outside the failure domain they are meant to help with.
Keep a fallback for each: a secondary chat or conference bridge, an externally hosted status page, a copy of critical runbooks outside the wiki, and break-glass credentials that do not depend on the identity provider. Test the fallbacks in game days, because an untested fallback usually fails in the same incident that needs it.
Metrics, failure modes and trade-offs
Track time to detect, acknowledge, declare and mitigate per incident, and read the distribution and outliers rather than a mean, which one long incident can dominate. Track pages per on-call shift too; exhausted responders are slow.
- Late declaration. Engineers debug quietly for an hour before anyone knows. Fix it with a low declaration threshold and a blameless culture around false alarms.
- The IC debugging. The commander disappears into a terminal and coordination stops. Hand the keyboard to someone else.
- Too many cooks. Fifteen people join the channel and several make changes at once. Only the operations lead authorizes changes.
- Root-cause fixation. The team investigates while an available rollback sits untried.
- Silent incidents. Support hears about the outage from customers. Enforce update cadence.
- Process weight. A heavy process for minor incidents makes people avoid declaring. Scale the process with severity; that is the trade-off every design has to strike.
What to do next
- Write a one-page severity table based on customer impact and publish it where on-call engineers can find it.
- Define the IC, operations, communications and scribe roles, and train at least three people to act as IC.
- Create an incident channel template with a pinned summary and an update template like the one above.
- List your top mitigations for each critical service, rollback first, and confirm each has been practiced in the last quarter.
- Map every incident tool to its dependencies and set up a fallback for any that share a failure domain with production.
- Run a game day that goes from alert to resolution, then review the response itself, not only the injected fault.