Incident response has two jobs that pull in different directions: stop the harm as quickly as possible, and learn enough to stop it happening again. Most teams lose time because they do the second job during the first. They debug while customers are failing, chase the root cause through logs while a one-line rollback would have ended the outage, and then declare victory the moment a graph turns green, only to be paged again twenty minutes later when the recovery itself overloads something.
This guide is about the first job: getting systems back up fast and keeping them up. It assumes you already have severity levels and roles; incident response architecture covers severity, incident command and communication. Here the focus is the responder's craft: where the minutes go, what to do in the first five of them, how to find the change that caused the problem, which mitigations to reach for and what each one risks, how recovery goes wrong, and how to prove the fix held.
Where the minutes go
Time to mitigate is the sum of five phases: detect (the problem starts until a signal fires), acknowledge (the signal fires until a human owns it), orient (understanding scope and what changed), mitigate (taking an action that stops the harm) and verify (confirming it worked and stayed working). Root-cause diagnosis is deliberately not on that list. It matters, but it can happen after the bleeding stops, without customers waiting.
Measuring the phases separately tells you where to invest. If detection takes twenty minutes, better alerting beats any runbook; on-call architecture covers burn-rate alerts that fire on symptoms users feel. If acknowledgement is slow, the problem is paging and escalation. In most mature teams the largest and most variable phase is orientation, which is where the techniques below pay off.
The first five minutes
The first minutes decide the shape of the whole incident. A responder who has a fixed opening sequence does not waste them deciding what to do. A good one:
- Declare early. Open the incident channel and say it is an incident. Downgrading a false alarm costs nothing; a late declaration costs the minutes nobody was coordinating.
- State the impact in user terms. "Checkout returns errors for about 30% of requests in eu-west since 14:02" is actionable. "Database CPU is high" is a symptom of a symptom.
- Bound the blast radius. Which regions, which zones, which customer tiers, which client versions, which endpoints? A boundary is often the biggest clue you will get, because a problem in one zone points at infrastructure while a problem in one client version points at a release.
- Ask what changed. Pull the change timeline before reading a single log line.
- Pick the cheapest safe mitigation and say it out loud. Even if you do not execute it yet, naming it focuses the room.
Change-first triage
The majority of production incidents follow a change: a deploy, a configuration push, a feature flag flip, a certificate rotation, a dependency upgrade, an infrastructure event, a traffic shift, or a change on a vendor's side. Asking what changed is faster than asking why it broke, because the list of changes in the last few hours is short and the space of possible failure mechanisms is enormous.
The obstacle is that changes live in different systems: the deploy tool, the flag service, the config repository, the cloud provider's event log, the vendor status pages. Responders lose ten minutes opening each one. Build a single merged timeline in advance, so that one command shows every change around the incident start, in order, with minutes relative to the start.
#!/usr/bin/env python3
"""what_changed.py: merge every change source into one timeline around an incident."""
import json, sys
from datetime import datetime, timedelta, timezone
def load(path, source):
# Each exporter writes JSON lines: {"ts": ISO-8601, "service": ..., "summary": ...}
with open(path) as f:
for line in f:
e = json.loads(line)
e["ts"] = datetime.fromisoformat(e["ts"])
e["source"] = source
yield e
def timeline(start, window_h=6, sources=()):
lo, hi = start - timedelta(hours=window_h), start + timedelta(minutes=30)
events = [e for path, src in sources for e in load(path, src) if lo <= e["ts"] <= hi]
for e in sorted(events, key=lambda e: e["ts"]):
delta = (e["ts"] - start).total_seconds() / 60
print(f"{delta:+7.1f} min {e['source']:<8} {e['service']:<20} {e['summary']}")
if __name__ == "__main__":
t0 = datetime.fromisoformat(sys.argv[1]).astimezone(timezone.utc)
timeline(t0, sources=[("deploys.jsonl", "deploy"), ("flags.jsonl", "flag"),
("config.jsonl", "config"), ("infra.jsonl", "infra"),
("vendor.jsonl", "vendor")])Read the output from the incident start backwards. A change minutes before the start, to a service on the failing request path, is your first hypothesis, and its reversal is your first mitigation candidate. Be careful with two traps. A change can take effect long after it ships: a cache that expires hours later, a cron job that runs nightly, a certificate valid until midnight. And correlation is not proof; the rollback is the experiment that confirms or rejects the hypothesis, which is why you choose a reversible one.
Mitigate before you understand
The core discipline is to separate stopping the harm from understanding it. If a deploy went out four minutes before errors began, roll it back now, even though nobody has found the bug. If the rollback fixes it, you have both a mitigation and strong evidence. If it does not, you have lost a few minutes and eliminated a hypothesis. The alternative, reading code under pressure to be sure first, routinely turns a ten-minute incident into a ninety-minute one.
Prefer mitigations that are reversible, fast and narrow, in that order. Each lever in the catalogue below should exist, be tested and be documented before you need it; runbook architecture and on-call runbooks cover how to keep that knowledge executable.
| Mitigation | Use when | Main risk |
|---|---|---|
| Roll back a deploy | A recent release is on the failing path | Schema or data migrations that the old version cannot read |
| Turn off a feature flag | The failing behaviour sits behind a flag | Flag dependencies; turning one off may break another |
| Revert configuration | A config push preceded the problem | Config that other services already adapted to |
| Drain a zone or region | Impact is confined to one location | The remaining capacity must absorb the shifted load |
| Fail over a database or dependency | A primary is unhealthy and a replica is current | Replication lag means lost or duplicate writes |
| Scale out | Saturation from a real load increase | Slow to take effect; can overload a shared dependency |
| Shed load or rate-limit | Demand exceeds safe capacity | Deliberately failing some users; pick which ones |
| Block a client or tenant | One caller generates the harmful traffic | Blocking a legitimate customer by mistake |
| Restart | A leak or stuck state, with no better lever | Hides the evidence and often only buys time |
Hypothesis discipline under pressure
When the obvious change is not the cause, incidents degrade into many people trying many things at once. Then nobody knows which action changed the graphs. Three rules keep the investigation coherent.
- Keep a written hypothesis log. In the incident channel, each hypothesis gets a line: what we think, what evidence would confirm or refute it, who is checking, and the result. It stops the room from re-examining a theory it already disproved.
- One change at a time, announced first. "I am draining zone b at 14:21" followed by "done" lets everyone read the graph against a known moment.
- Diagnose by comparison. Healthy and unhealthy things side by side are faster than either alone: the canary against the baseline, zone a against zone b, the old client version against the new, the failing tenant against a healthy one. Whatever differs between them is where to look.
Hand diagnosis to a separate person once there is a plausible mitigation path. The incident commander keeps the room on mitigation; the investigator works the root cause in parallel and reports back only what changes the plan.
Recovery has its own failure modes
Many second outages happen during recovery, because a system that was broken has built up pressure. Before you remove a mitigation or re-admit traffic, consider each of these:
- Retry storms. Clients have been retrying, often several times per request with too little backoff. When the service returns, it receives normal traffic plus the retry backlog, and falls over again.
- Thundering herds of reconnects. Thousands of clients with the same reconnect timer arrive in the same second. Jitter in clients helps; admitting traffic in steps helps immediately.
- Cold caches. A restarted cache tier serves almost nothing from memory, so every request goes to the database the cache was protecting.
- Queue backlogs. Work queued during the outage must drain while new work keeps arriving; the drain rate is the headroom, not the capacity.
- Autoscaler lag. Scaled-down capacity takes minutes to return, and scale-up decisions based on an abnormal period can be wrong in either direction.
The arithmetic for backlogs is simple and almost always surprises people. A queue drains at capacity minus arrival rate, and adding workers helps only until the next bottleneck caps throughput.
def drain_minutes(backlog, arrival_per_s, capacity_per_s):
"""Minutes to clear a backlog while new work keeps arriving."""
headroom = capacity_per_s - arrival_per_s
if headroom <= 0:
return float("inf") # the queue never drains; add capacity or shed
return backlog / headroom / 60
# 1.8 million queued jobs, 2,000/s still arriving, workers at 2,500/s
print(drain_minutes(1_800_000, 2_000, 2_500)) # 60.0 minutes
# double the workers but the database caps them at 3,000/s
print(drain_minutes(1_800_000, 2_000, 3_000)) # 30.0 minutes, not 12Doing this calculation in the channel turns "when will it be fixed?" into a number you can give to customers, and shows early when the plan cannot work and you need to shed or defer the oldest work instead.
Worked example: losing a cache tier
At 09:40 a maintenance script restarts the nodes of a shared cache cluster in quick succession instead of one at a time. Product pages, which rely on a cache hit rate of about 95%, start sending every miss to the database. Database CPU reaches 100% within two minutes, latency rises to seconds, and the burn-rate alert for the product API pages the on-call engineer at 09:44.
In the first five minutes the responder declares, states the impact (product pages slow or failing everywhere, checkout unaffected because it does not read through this cache), and pulls the change timeline, which shows the maintenance job at 09:39. There is nothing to roll back: the restart is done and the caches are empty. The mitigation must protect the database while the cache refills. The team turns on load shedding at the edge for anonymous traffic, which halves database load, and enables the existing flag that serves stale product data from the CDN for up to an hour. Error rate falls within three minutes.
Recovery is where it could go wrong. Removing shedding at once would send full traffic to a half-warm cache and repeat the overload. Instead the team re-admits anonymous traffic in 25% steps every five minutes while watching the hit rate and database CPU, and enables request coalescing so that many concurrent misses for the same key produce one database query. By 10:25 the hit rate is back above 90%, shedding is off, and stale serving is disabled at 10:40. The follow-up questions, why the script could restart all nodes at once and why one cache tier was a single point of failure for the database, go to the postmortem, not the incident.
Declaring resolution
A green graph is not resolution. Agree on explicit criteria and check each one before closing:
- The user-facing SLIs are back within their objectives and have stayed there for a defined period, typically 15 to 30 minutes, with mitigations in their final state.
- Every segment is healthy, not only the aggregate: each region, tier, client version and major tenant. Averages hide the 2% who are still failing.
- Backlogs have drained, delayed jobs have run, and data written during the incident has been checked for corruption or duplicates.
- Temporary mitigations are either removed or recorded with an owner and an expiry, because a forgotten rate limit or stale-serving flag becomes the next incident.
- Synthetic checks pass from outside your network, from where customers actually are.
When incident response itself fails
- Hero debugging. The most senior engineer reads code for forty minutes while a rollback sits untried. Ask at every update what the current mitigation candidate is.
- Untested levers. The failover that has never been exercised fails during the incident. Practise levers in game days.
- Too many cooks. Twenty people in the channel making changes. Keep one commander and announce every action.
- Fatigue. Decisions degrade after a few hours. Hand over with a written state summary: impact, mitigations in place, hypotheses ruled out, next steps.
- Closing too early. Recovery traps reopen the incident. Use the resolution criteria and keep watching through the next traffic peak.
After resolution, the learning job begins: writing the postmortem and running the incident review turn this incident into fewer future ones.
What to do next
- Measure detect, acknowledge, orient, mitigate and verify separately for your last ten incidents and find the slowest phase.
- Build a merged change timeline covering deploys, flags, config, infrastructure and vendor events, runnable with one command.
- List the mitigation levers for each critical service, and test each one in a game day this quarter.
- Write your first-five-minutes sequence into the on-call handbook and practise it in a drill.
- Add jitter and backoff to client retries and reconnects, and make sure traffic can be re-admitted in steps.
- Adopt written resolution criteria, including per-segment checks and an owner for every temporary mitigation.