Classic incident response assumes something is visibly broken: errors spike, latency climbs, a health check fails. Many AI incidents look nothing like that. The service answers every request with HTTP 200 while it recommends a refund policy that does not exist, reveals another customer's order history, follows instructions hidden in an uploaded PDF, or spends a month's budget in an afternoon because an agent loop never terminates. Nothing pages. A customer tweets. By the time an engineer looks, the on-call runbook they open covers disk space and database failover.

This article is about designing a runbook for AI features specifically: what classes of incident to plan for, which containment levers must exist before the first incident, how to capture evidence when the evidence is prompts and model outputs, how to find what changed in a system whose behaviour depends on model, prompt, retrieval and tool versions, and how to close an incident so it cannot recur silently. General runbook structure and upkeep are covered in Runbook architecture and On-Call Runbooks; this page adds what those assume away.

Advertisement

What makes AI incidents different

  • No stack trace. A wrong answer is not an exception. Detection depends on evaluation signals, user feedback, sampled review and guardrail counters, not on error logs.
  • Probabilistic reproduction. The same input may produce a good answer nine times and a harmful one on the tenth. "Could not reproduce" is not evidence of a fix.
  • Many independent versions. Behaviour is the product of model version, prompt template, system instructions, retrieval index snapshot, tool definitions, guardrail configuration and sampling settings. Any one can change without a code deploy, and a hosted model can change without any action on your side.
  • The evidence is sensitive. Prompts and outputs contain customer data. Pasting them into an incident channel can itself become a second incident.
  • The fix is usually a lever, not a patch. Rolling back a prompt, pinning a model version or disabling one tool typically contains an AI incident within minutes; code changes come later.
Detect, triage by class, pull a pre-built lever, then diagnose by version dimensionsSignalsSLO burn, evals, reportsTriageclass + severityEvidence snapshottraces, versionsEscalationsafety, privacy, legalsev1/2Containment levers (built before the incident)kill switch | model pin/rollback | prompt rollback | route switch | tool disable | spend cap | output filter | index rollbackcontainDiagnosewhat changed? bisect versionsRecoverfix, re-run evals, rampClosepostmortem, new eval casesEvery step writes to the incident timelinewho pulled which lever, when, and what the signals did afterwards
The runbook flow. Triage assigns a class and severity, evidence is snapshotted before anything changes, a pre-built lever contains the impact, and diagnosis bisects over the versions that shape model behaviour.

A taxonomy that drives the first move

A runbook is most useful when it turns the first ten minutes into a lookup. Classify every AI incident into one of a small number of classes, each with its typical signal, its first containment lever and the people who must be involved:

ClassTypical signalFirst leverMust involve
Quality regressionEval score drop, thumbs-down rate, support ticketsRoll back prompt or pin previous modelFeature owner
Unsafe or harmful outputGuardrail hits, user report, pressTighten output filter; disable feature for affected surfaceSafety lead
Prompt injection exploitedUnexpected tool calls, data sent to odd destinationsDisable affected tools; restrict to read-onlySecurity
Data exposureOutput contains another tenant's data or secretsKill switch for the feature; preserve evidencePrivacy, legal
Cost runawaySpend per hour, tokens per task, loop countersSpend cap; cap steps per agent runFeature owner, finance
Provider or model outageErrors, time to first token, breaker openRoute switch or degraded modePlatform on-call
Silent model changeShift in output length, refusal rate, format errorsPin explicit model version; re-run evalsFeature owner

The class is a starting hypothesis, not a verdict. A cost runaway is often a quality regression in disguise, for example a prompt change that makes an agent retry a failing tool forever, so the runbook should tell the responder to re-classify as evidence arrives.

Advertisement

Severity rules that escalate the right things

Severity usually scales with user impact: how many users, for how long, with what workaround. Keep that, but add two AI-specific rules. First, safety and privacy classes start at high severity regardless of volume: one confirmed cross-tenant data exposure is a serious incident even if it happened once, because notification obligations may apply and legal needs to hear about it within hours, not after the postmortem. Second, exposure beats frequency: a harmful answer on a public, shareable surface is more severe than the same answer in an internal tool, because a screenshot outlives the fix. Tie the remaining severities to your error budgets so that a slow quality burn becomes an incident when it threatens the budget, using the approach in Burn-Rate Alerting.

Build the levers before you need them

The single largest predictor of a short AI incident is whether the containment levers already exist and have been exercised. A runbook that says "disable the tool" is useless if disabling the tool requires a code change, a review and a deploy. Every AI feature should ship with a set of runtime controls, owned by a config service that on-call can change in seconds with an audit trail:

# ai-controls/support-assistant.yaml  (served by the config service, audited, hot-reloaded)
feature: support-assistant
kill_switch: false                 # true -> static fallback page / human handoff
model:
  pinned_version: "model-2026-07-15"   # explicit version string, never an alias like "latest"
  rollback_to: "model-2026-05-02"
prompt:
  active: "support-system@v41"
  rollback_to: "support-system@v40"
retrieval:
  index_snapshot: "helpcenter-2026-09-28"
  rollback_to: "helpcenter-2026-09-21"
tools:
  issue_refund: {enabled: true, max_amount_minor: 20000}
  send_email:   {enabled: true}
  web_fetch:    {enabled: false}
limits:
  max_agent_steps: 12
  max_tokens_per_task: 60000
  spend_cap_per_hour_usd: 400      # breach -> degrade to retrieval-only answers
guardrails:
  output_filter_profile: "standard"    # "strict" available as a lever
routing:
  mode: "primary"                  # primary | secondary | degraded

Two properties matter more than the exact list. Each lever must be independent, so you can disable one tool without disabling the feature. And each must be reversible and tested: rehearse flipping every lever in a staging environment on a schedule, because a rollback target that points at a deleted prompt version or an expired index snapshot fails exactly when you need it.

The runbook skeleton

Write one runbook per feature, with one page per incident class, using the same six steps each time so responders do not have to learn a new structure under pressure:

  1. Trigger. The signals and thresholds that open this page, with links to the dashboards and queries.
  2. Verify. How to confirm the incident is real in under five minutes: a saved query, a replay of three reported examples, a check of the guardrail counter. Include how to tell this class from its look-alikes.
  3. Snapshot evidence. What to capture before changing anything (next section).
  4. Contain. Which lever to pull first, who may pull it, what signal should move afterwards and within how long. If it does not move, the next lever.
  5. Diagnose and recover. The bisection procedure, the fix path, the eval run required before un-containing, and the ramp plan.
  6. Close. Exit criteria, communications, postmortem owner and the eval cases that must be added.

Keep the steps imperative and specific: "set tools.web_fetch.enabled to false in ai-controls" rather than "consider limiting tool access". The responder at 3 a.m. may not be the feature's author.

Evidence capture when the evidence is prompts

AI incidents are diagnosed from traces: the full input as the model saw it after templating and retrieval, the output, the tool calls with arguments and results, and the versions of everything involved. If your tracing samples at one percent or truncates long prompts, the evidence for the incident may already be gone. Record, for every request, at least: request ID, tenant, timestamp, model version as reported in the response, prompt template version, retrieval snapshot and document IDs, tool calls, guardrail decisions, token counts and route. Store full prompt and output bodies in a restricted store with a retention period, and keep only IDs and metadata in general-purpose logs.

At the start of an incident, snapshot the relevant traces into an incident-scoped location with access limited to responders, and place a retention hold so routine deletion does not remove them mid-investigation. In the incident channel, refer to examples by trace ID, never by pasting the customer's text. If the class is data exposure, involve privacy before copying anything anywhere.

Diagnose by asking what changed

Most AI incidents are caused by a change: a new prompt, a new model version, a re-built index, a new tool, a guardrail configuration, or a shift in the traffic mix. Keep a single change log that records every one of these with a timestamp, whether it was deployed by your team or detected, for example by noticing a new model version string in responses. Then diagnosis becomes a join: find when the signal moved, list the changes just before, and bisect by grouping outcomes by each version dimension.

-- Which version dimension explains the regression?  (sampled, judged traffic)
SELECT prompt_version, model_version, index_snapshot,
       COUNT(*)                                        AS judged,
       AVG(CASE WHEN verdict = 'bad' THEN 1.0 ELSE 0 END) AS bad_rate
FROM   judged_samples
WHERE  feature = 'support-assistant'
  AND  ts BETWEEN TIMESTAMP '2026-09-27 12:00' AND TIMESTAMP '2026-09-28 12:00'
GROUP  BY prompt_version, model_version, index_snapshot
HAVING COUNT(*) >= 50
ORDER  BY bad_rate DESC;

Confirm a hypothesis by replay: run the evaluation set, plus the specific failing examples, against the suspected old and new configurations side by side, as described in LLM Evaluation Harness and Regression Testing. Because output is probabilistic, run each example several times and compare rates, not single outputs.

Roles and communication

Keep the standard roles, an incident commander, an operations lead and a communications lead with a scribe, and add two for AI incidents. A model owner, someone who understands the prompt, retrieval and evaluation setup, joins every quality, safety and silent-change incident. A safety or privacy liaison joins every harmful-output, injection and data-exposure incident and decides what can be said externally. External statements should describe impact and actions without overclaiming: "we have disabled the feature while we investigate" is accurate; "the model will never do this again" is not something anyone can promise.

Worked example: a policy the assistant invented

On Monday at 11:40, support agents report that the assistant is telling customers they can return opened software within 60 days; the real policy is 14 days. Triage classifies it as a quality regression at medium severity: many users, financial exposure, no privacy impact. The responder snapshots 200 traces containing the word "return" from the last 24 hours into the incident store. The change log shows two changes: prompt v41 shipped Friday, and the help-centre index was rebuilt Sunday night.

Grouping judged samples by version shows bad answers only with index snapshot 2026-09-28, under both prompt versions. The responder pulls the index rollback lever at 12:05; the bad-answer rate on a fresh sample falls to baseline within 20 minutes. Diagnosis finds that the rebuild ingested a draft page from the content system containing a proposed 60-day policy that was never approved. Recovery adds a publication-status filter to ingestion, rebuilds, and re-runs the eval set with ten new return-policy questions before switching the index forward again. The postmortem adds a pre-publish eval gate for index rebuilds and a daily canary question set on policy facts.

Close the loop

An incident is closed when the lever is back in its normal position, the fix has passed evaluation, and the incident has produced lasting detection: new eval cases built from the real failing examples, with sensitive data removed, a new or tuned alert, and, when a lever was missing or slow, a new lever. Run game days where someone introduces a known-bad prompt or a poisoned document in staging and the on-call engineer works the runbook, and measure time to contain. Track error budget consumption per incident as in Error Budgets, so that repeated AI incidents change release policy rather than just filling a postmortem folder.

What to do next

  1. List the seven incident classes for each AI feature and write the first lever and required people for each.
  2. Build runtime controls: kill switch, explicit model pin and rollback, prompt and index rollback, per-tool disable, step and spend caps, and a stricter output filter profile.
  3. Rehearse every lever in staging monthly and verify rollback targets still exist.
  4. Trace every request with model, prompt, index and tool versions; keep bodies in a restricted store with retention holds.
  5. Keep one change log for prompts, models, indexes, tools and guardrails, including detected provider-side model changes.
  6. Write one runbook page per class using the six-step skeleton, with imperative, specific actions.
  7. Add a model owner and a safety or privacy liaison to the incident roles.
  8. Close every incident with new eval cases, a detector and a game-day scenario.
Key takeaway: AI incidents often return HTTP 200, depend on several versions at once, and carry sensitive evidence. A good runbook classifies the incident to pick the first move, escalates safety and privacy regardless of volume, relies on runtime levers that were built and rehearsed in advance, snapshots traces before touching anything, diagnoses by grouping outcomes by version, and closes by turning real failures into eval cases and detectors.