Most incident guides describe what an incident is: severity levels, roles, a status template. Those matter, and they are covered on this site in the incident management guide. This page is about the verb. Running an incident is a job with its own technique, and it is the job that most often goes undone, because the most senior engineer in the room is busy reading logs and nobody is steering.

The technique below is written for the person holding the incident commander role, the IC, from the moment of taking command until the incident is closed and handed to the review. It assumes the triage habits in incident response belong to the responders, and that the IC's job is to make sure those responders are working on the right things, in the right order, with the rest of the company kept informed.

Advertisement

Running is a different job from fixing

An incident has two kinds of work. Technical work finds and applies mitigations. Coordination work decides what the technical work should be, keeps people from duplicating or contradicting each other, notices when an approach has stalled, and tells everyone else what is happening. When one person does both, coordination loses every time, because a log line is more absorbing than a status update that is fifteen minutes overdue.

So the first rule of running an incident is that the IC does not debug. In a small incident the on-call engineer may start as both, but as soon as a second responder joins, one of them takes command and stops typing commands. Command is a role, not a rank: good ICs know the system well enough to ask sharp questions and are disciplined enough not to answer them.

Taking command

Command starts with an explicit statement in the incident channel and on the call, because ambiguity about who is in charge is the most common early failure. A good opening has four parts: who is IC, the impact as currently understood in user terms, the first objective, and when the next update will come. For example: I am IC. Card checkout errors at 12 percent in eu-west since 14:02. Objective: stop the errors, cause second. Next update 14:30.

Then fill the two roles that protect the IC's attention. A scribe keeps the timeline so the IC does not have to, and a communications lead writes stakeholder updates so the IC does not get pulled into wording them. In a small incident one person can do both. If you cannot fill them, use a tool that captures the timeline from the channel, and accept that updates will be shorter.

Finally, set the first objective deliberately. Early in an incident the objective is almost always mitigation: reduce user impact by the cheapest safe means, even before the cause is known. Saying it aloud stops the natural drift of a room of engineers towards root cause analysis while customers are still failing.

Advertisement

The command loop

Once command is established, the IC's work is a loop, turned roughly every 10 to 15 minutes on a timer. Each turn has five steps.

The command loop: one turn every 10 to 15 minutesAssessimpact, trend, streamsDecideobjective, optionsAssignowner, task, check-inRecorddecisions, timelineCommunicateupdate on cadenceIncident commanderdoes not debugWorkstream leads report in the same shape every turnConditions: what we see now. Actions: what we are doing. Needs: what we are blocked on.Every task leaves the loop with an owner and a time to report back
The IC turns this loop on a timer, not when something happens. Most turns are short; the value is that none are skipped.
  1. Assess. What is the impact now, and is it getting better or worse? What has each workstream learned since the last turn?
  2. Decide. Is the current objective still right? Which options are on the table, and is any decision overdue?
  3. Assign. Every piece of work leaves the loop with one named owner and a time to report back. Unowned work does not happen; work without a check-in time is not checked.
  4. Record. Decisions, with their reason and undo condition, go into the log.
  5. Communicate. If an update is due, the comms lead sends it. Updates go out on cadence even when nothing has changed.

The timer matters because incidents have long quiet stretches in which everyone is heads-down and the shape of the incident changes unnoticed. A turn at 14:45 that asks whether the database hypothesis has produced anything in thirty minutes stops a team from spending two hours on the wrong idea.

Running the call

A busy incident call is a shared radio channel, and it needs radio discipline. Use names when assigning work, and ask for the assignment to be repeated back. This closed-loop pattern, borrowed from aviation and medicine, catches the common failure where two people each think the other is rolling back the deploy. Keep the call for coordination and push technical detail into threads or breakout rooms that report back through their lead.

Manage the room size. Ask newcomers to state their purpose, and let people leave when their question is answered. When a senior leader joins, give them a one-minute summary and, in larger incidents, a liaison who answers their questions outside the call. Their presence never changes who is in command; if a leader wants to take over, that is an explicit handover, not a takeover by volume.

Keep a parking lot. Good ideas that are not the current objective, such as a monitoring gap, go into the log for later, which lets people stop arguing for them.

Workstreams and span of control

When more than about five people are reporting directly to the IC, the IC stops being able to track them, and the loop slows. The Incident Command System used by emergency services treats span of control as a hard design constraint, commonly cited as three to seven direct reports with five as the target. The software version is to split the incident into workstreams, each with a lead who reports to the IC and runs their own small loop.

Typical streams are mitigation, diagnosis, customer impact and data repair. Each lead reports on every turn in the same shape: conditions, what they see now; actions, what they are doing; needs, what they are blocked on. A fixed shape makes reports short and lets the IC spot a stream that has reported the same conditions three times, which usually means it is stuck. If a stream needs a decision only the IC can make, the needs line is where it surfaces.

Decisions under uncertainty

Incident decisions are made with incomplete information, and waiting for certainty is itself a decision that costs users. Two questions make them tractable. Is it reversible? A traffic shift, a feature flag or a rollback can be undone in minutes, so take it on weaker evidence. A data migration or a failover that cannot be failed back needs more evidence and usually a second opinion. And what signal would make us undo it? Writing that down before acting turns a hopeful change into an experiment with a stop condition.

A small helper keeps the loop honest. It is not an incident platform, just a log the IC or scribe updates during each turn: tasks with owners and due times, decisions with undo conditions, an overdue check and a handover brief generator.

from dataclasses import dataclass, field
from datetime import datetime, timedelta

@dataclass
class Task:
    owner: str
    what: str
    due: datetime
    done: bool = False

@dataclass
class Decision:
    at: datetime
    what: str
    why: str
    reversible: bool
    undo_if: str            # the signal that makes us reverse it

@dataclass
class IncidentLog:
    title: str
    ic: str
    update_every: timedelta = timedelta(minutes=30)
    last_update: datetime | None = None
    tasks: list = field(default_factory=list)
    decisions: list = field(default_factory=list)

    def assign(self, owner, what, minutes):
        self.tasks.append(Task(owner, what, datetime.now() + timedelta(minutes=minutes)))

    def decide(self, what, why, reversible, undo_if):
        self.decisions.append(Decision(datetime.now(), what, why, reversible, undo_if))

    def sent_update(self):
        self.last_update = datetime.now()

    def done(self, owner, what):
        for t in self.tasks:
            if t.owner == owner and t.what == what:
                t.done = True

    def tick(self, now=None):
        """Run at every turn of the command loop: what is overdue?"""
        now = now or datetime.now()
        late = [t for t in self.tasks if not t.done and t.due < now]
        for t in late:
            print(f"OVERDUE {t.owner}: {t.what} (due {t.due:%H:%M})")
        if self.last_update is None or now - self.last_update >= self.update_every:
            print("STATUS UPDATE DUE")

    def handover(self):
        open_tasks = [f"- {t.owner}: {t.what} (check {t.due:%H:%M})" for t in self.tasks if not t.done]
        last = [f"- {d.at:%H:%M} {d.what}; undo if {d.undo_if}" for d in self.decisions[-5:]]
        return "\n".join([f"{self.title}: IC {self.ic}", "Open tasks:", *open_tasks,
                          "Recent decisions:", *last])

Timebox hypotheses the same way. When a workstream starts on a theory, agree when it will report back and what result would rule the theory out. Without the timebox, the most interesting theory absorbs the most people, whether or not it is right.

Handovers and long incidents

Nobody runs a good incident for eight hours. Plan handovers before fatigue forces them: set a maximum IC shift, often around four hours, and rotate leads and responders on a schedule as well. A handover is a short, written brief followed by a spoken confirmation in the channel, and the outgoing IC stays available for fifteen minutes in case the new one has questions.

INC-2026-10-02 checkout errors: IC M. Osei -> handing to J. Park at 18:00
Impact now: card checkout errors 4% in eu-west (peak 31% at 14:20), falling
Objective: hold error rate under 1% until the payments fix is deployed
Open tasks:
- Payments stream (lead R. Chen): canary of fix at 18:10, report at 18:25
- Data stream (lead A. Okafor): count orders stuck in PENDING, report at 18:30
Recent decisions:
- 15:05 routed 50% of EU card traffic to processor B; undo if B error rate > 1%
- 16:40 paused nightly batch settlement; undo by 23:00 or finance escalation
Next status update: 18:15 to #inc channel and status page

The brief answers what a new IC needs first: current impact and trend, the objective, who owns which stream, what was decided and how to undo it, and when the next update is due. If you cannot write it in five minutes, the incident is not being recorded well enough, which is worth fixing before the next shift. For incidents that run across time zones, hand over to the next region rather than keeping people up, and make sure on-call rotations have someone trained to receive command, as described in the on-call guide.

Worked example: a five hour payments incident

The timeline below compresses a realistic incident into its command decisions. The technical story is simple; what made it go well was the running.

TimeWhat happenedWhat the IC did
14:06Page: checkout error rate 12% in eu-westDeclared SEV2, took command, opened the channel, set first update for 14:30
14:12Errors at 25%, eight people in the callNamed a scribe and a comms lead, asked everyone else to state their purpose or leave
14:20Errors at 31%; processor A returns timeoutsSplit two streams: payments (find cause) and mitigation (shift traffic)
14:45Mitigation stream proposes routing to processor BDecision recorded: reversible, undo if B errors exceed 1%
15:05Traffic shift live; errors fall to 6%Kept the decision under watch, cancelled two speculative fixes
16:30Orders stuck in PENDING reportedOpened a data stream with its own lead
16:40Data stream finds duplicate settlement riskPaused the batch job, escalated to finance
18:00Errors at 4%; fix readyHanded command over with a written brief
19:30Fix deployed; errors under 0.2% for an hourDowngraded, set the monitoring window, scheduled the review

Notice what the IC did not do: read a single log line. The shift of traffic to a second processor was a reversible mitigation with a written undo condition, chosen while the cause was still unknown. The settlement risk surfaced ten minutes after the IC opened a data stream for the stuck orders. And the handover at 18:00 took four minutes because the log was already up to date.

Standing down

Incidents end in stages. Mitigated means users are no longer affected but the system is in an abnormal state, such as traffic on a backup processor or a feature flag off. Resolved means the system is back to normal operation and has stayed there for an agreed monitoring window, often an hour for a fast-moving service. Declare each stage explicitly, with the evidence, and downgrade severity as you go so people can leave the call.

Before closing, make sure every temporary change has an owner who will reverse it and a date, every parked item is in the tracker, and the review is scheduled with the timeline attached. The review itself is a separate process, described in how to run a postmortem. The final message in the channel states what happened, what is still temporary and when the review is.

How running an incident fails

  • The IC debugs: coordination stops, updates are late and nobody notices a stalled stream.
  • Work without owners: someone should check the CDN, so nobody does.
  • Silent stretches: no timer, so a wrong hypothesis absorbs an hour.
  • Command by volume: a senior voice starts directing work without a handover, and two plans run at once.
  • Decisions without undo conditions: a mitigation stays in place for weeks because nobody wrote down when to remove it.
  • No handover plan: the IC is still in command at hour nine and making poor decisions.
  • Premature resolution: the incident is closed at the first good graph and reopened an hour later.

What to do next

  1. Write a one-paragraph IC opening script and pin it in your incident channel template.
  2. Add a 15 minute loop timer and a 30 minute update reminder to your incident bot, or run the helper above by hand.
  3. Adopt the conditions, actions, needs report shape for workstream leads.
  4. Require every mitigation decision to be logged with an undo condition.
  5. Set a maximum IC shift and practise a written handover in your next game day.
  6. Train at least two ICs per on-call rotation, and include engineers who are not the most senior.
  7. Define mitigated and resolved for your service, including the monitoring window, and keep runbooks for common mitigations current, following how to write a runbook.
Key takeaway: Running an incident is coordination, and the incident commander does not debug. Take command explicitly, fill scribe and comms roles, and set mitigation as the first objective. Turn a 10 to 15 minute loop of assess, decide, assign, record and communicate, giving every task an owner and a check-in time. Split into workstreams before you pass about five direct reports, log every decision with its undo condition, hand over in writing before fatigue sets in, and stand down in stages with evidence.