Classic incident response assumes something is visibly broken: errors spike, latency climbs, a health check fails. Many AI incidents look nothing like that. The service answers every request with HTTP 200 while it recommends a refund policy that does not exist, reveals another customer's order history, follows instructions hidden in an uploaded PDF, or spends a month's budget in an afternoon because an agent loop never terminates. Nothing pages. A customer tweets. By the time an engineer looks, the on-call runbook they open covers disk space and database failover.
This article is about designing a runbook for AI features specifically: what classes of incident to plan for, which containment levers must exist before the first incident, how to capture evidence when the evidence is prompts and model outputs, how to find what changed in a system whose behaviour depends on model, prompt, retrieval and tool versions, and how to close an incident so it cannot recur silently. General runbook structure and upkeep are covered in Runbook architecture and On-Call Runbooks; this page adds what those assume away.
What makes AI incidents different
- No stack trace. A wrong answer is not an exception. Detection depends on evaluation signals, user feedback, sampled review and guardrail counters, not on error logs.
- Probabilistic reproduction. The same input may produce a good answer nine times and a harmful one on the tenth. "Could not reproduce" is not evidence of a fix.
- Many independent versions. Behaviour is the product of model version, prompt template, system instructions, retrieval index snapshot, tool definitions, guardrail configuration and sampling settings. Any one can change without a code deploy, and a hosted model can change without any action on your side.
- The evidence is sensitive. Prompts and outputs contain customer data. Pasting them into an incident channel can itself become a second incident.
- The fix is usually a lever, not a patch. Rolling back a prompt, pinning a model version or disabling one tool typically contains an AI incident within minutes; code changes come later.
A taxonomy that drives the first move
A runbook is most useful when it turns the first ten minutes into a lookup. Classify every AI incident into one of a small number of classes, each with its typical signal, its first containment lever and the people who must be involved:
| Class | Typical signal | First lever | Must involve |
|---|---|---|---|
| Quality regression | Eval score drop, thumbs-down rate, support tickets | Roll back prompt or pin previous model | Feature owner |
| Unsafe or harmful output | Guardrail hits, user report, press | Tighten output filter; disable feature for affected surface | Safety lead |
| Prompt injection exploited | Unexpected tool calls, data sent to odd destinations | Disable affected tools; restrict to read-only | Security |
| Data exposure | Output contains another tenant's data or secrets | Kill switch for the feature; preserve evidence | Privacy, legal |
| Cost runaway | Spend per hour, tokens per task, loop counters | Spend cap; cap steps per agent run | Feature owner, finance |
| Provider or model outage | Errors, time to first token, breaker open | Route switch or degraded mode | Platform on-call |
| Silent model change | Shift in output length, refusal rate, format errors | Pin explicit model version; re-run evals | Feature owner |
The class is a starting hypothesis, not a verdict. A cost runaway is often a quality regression in disguise, for example a prompt change that makes an agent retry a failing tool forever, so the runbook should tell the responder to re-classify as evidence arrives.
Severity rules that escalate the right things
Severity usually scales with user impact: how many users, for how long, with what workaround. Keep that, but add two AI-specific rules. First, safety and privacy classes start at high severity regardless of volume: one confirmed cross-tenant data exposure is a serious incident even if it happened once, because notification obligations may apply and legal needs to hear about it within hours, not after the postmortem. Second, exposure beats frequency: a harmful answer on a public, shareable surface is more severe than the same answer in an internal tool, because a screenshot outlives the fix. Tie the remaining severities to your error budgets so that a slow quality burn becomes an incident when it threatens the budget, using the approach in Burn-Rate Alerting.
Build the levers before you need them
The single largest predictor of a short AI incident is whether the containment levers already exist and have been exercised. A runbook that says "disable the tool" is useless if disabling the tool requires a code change, a review and a deploy. Every AI feature should ship with a set of runtime controls, owned by a config service that on-call can change in seconds with an audit trail:
# ai-controls/support-assistant.yaml (served by the config service, audited, hot-reloaded)
feature: support-assistant
kill_switch: false # true -> static fallback page / human handoff
model:
pinned_version: "model-2026-07-15" # explicit version string, never an alias like "latest"
rollback_to: "model-2026-05-02"
prompt:
active: "support-system@v41"
rollback_to: "support-system@v40"
retrieval:
index_snapshot: "helpcenter-2026-09-28"
rollback_to: "helpcenter-2026-09-21"
tools:
issue_refund: {enabled: true, max_amount_minor: 20000}
send_email: {enabled: true}
web_fetch: {enabled: false}
limits:
max_agent_steps: 12
max_tokens_per_task: 60000
spend_cap_per_hour_usd: 400 # breach -> degrade to retrieval-only answers
guardrails:
output_filter_profile: "standard" # "strict" available as a lever
routing:
mode: "primary" # primary | secondary | degradedTwo properties matter more than the exact list. Each lever must be independent, so you can disable one tool without disabling the feature. And each must be reversible and tested: rehearse flipping every lever in a staging environment on a schedule, because a rollback target that points at a deleted prompt version or an expired index snapshot fails exactly when you need it.
The runbook skeleton
Write one runbook per feature, with one page per incident class, using the same six steps each time so responders do not have to learn a new structure under pressure:
- Trigger. The signals and thresholds that open this page, with links to the dashboards and queries.
- Verify. How to confirm the incident is real in under five minutes: a saved query, a replay of three reported examples, a check of the guardrail counter. Include how to tell this class from its look-alikes.
- Snapshot evidence. What to capture before changing anything (next section).
- Contain. Which lever to pull first, who may pull it, what signal should move afterwards and within how long. If it does not move, the next lever.
- Diagnose and recover. The bisection procedure, the fix path, the eval run required before un-containing, and the ramp plan.
- Close. Exit criteria, communications, postmortem owner and the eval cases that must be added.
Keep the steps imperative and specific: "set tools.web_fetch.enabled to false in ai-controls" rather than "consider limiting tool access". The responder at 3 a.m. may not be the feature's author.
Evidence capture when the evidence is prompts
AI incidents are diagnosed from traces: the full input as the model saw it after templating and retrieval, the output, the tool calls with arguments and results, and the versions of everything involved. If your tracing samples at one percent or truncates long prompts, the evidence for the incident may already be gone. Record, for every request, at least: request ID, tenant, timestamp, model version as reported in the response, prompt template version, retrieval snapshot and document IDs, tool calls, guardrail decisions, token counts and route. Store full prompt and output bodies in a restricted store with a retention period, and keep only IDs and metadata in general-purpose logs.
At the start of an incident, snapshot the relevant traces into an incident-scoped location with access limited to responders, and place a retention hold so routine deletion does not remove them mid-investigation. In the incident channel, refer to examples by trace ID, never by pasting the customer's text. If the class is data exposure, involve privacy before copying anything anywhere.
Diagnose by asking what changed
Most AI incidents are caused by a change: a new prompt, a new model version, a re-built index, a new tool, a guardrail configuration, or a shift in the traffic mix. Keep a single change log that records every one of these with a timestamp, whether it was deployed by your team or detected, for example by noticing a new model version string in responses. Then diagnosis becomes a join: find when the signal moved, list the changes just before, and bisect by grouping outcomes by each version dimension.
-- Which version dimension explains the regression? (sampled, judged traffic)
SELECT prompt_version, model_version, index_snapshot,
COUNT(*) AS judged,
AVG(CASE WHEN verdict = 'bad' THEN 1.0 ELSE 0 END) AS bad_rate
FROM judged_samples
WHERE feature = 'support-assistant'
AND ts BETWEEN TIMESTAMP '2026-09-27 12:00' AND TIMESTAMP '2026-09-28 12:00'
GROUP BY prompt_version, model_version, index_snapshot
HAVING COUNT(*) >= 50
ORDER BY bad_rate DESC;Confirm a hypothesis by replay: run the evaluation set, plus the specific failing examples, against the suspected old and new configurations side by side, as described in LLM Evaluation Harness and Regression Testing. Because output is probabilistic, run each example several times and compare rates, not single outputs.
Roles and communication
Keep the standard roles, an incident commander, an operations lead and a communications lead with a scribe, and add two for AI incidents. A model owner, someone who understands the prompt, retrieval and evaluation setup, joins every quality, safety and silent-change incident. A safety or privacy liaison joins every harmful-output, injection and data-exposure incident and decides what can be said externally. External statements should describe impact and actions without overclaiming: "we have disabled the feature while we investigate" is accurate; "the model will never do this again" is not something anyone can promise.
Worked example: a policy the assistant invented
On Monday at 11:40, support agents report that the assistant is telling customers they can return opened software within 60 days; the real policy is 14 days. Triage classifies it as a quality regression at medium severity: many users, financial exposure, no privacy impact. The responder snapshots 200 traces containing the word "return" from the last 24 hours into the incident store. The change log shows two changes: prompt v41 shipped Friday, and the help-centre index was rebuilt Sunday night.
Grouping judged samples by version shows bad answers only with index snapshot 2026-09-28, under both prompt versions. The responder pulls the index rollback lever at 12:05; the bad-answer rate on a fresh sample falls to baseline within 20 minutes. Diagnosis finds that the rebuild ingested a draft page from the content system containing a proposed 60-day policy that was never approved. Recovery adds a publication-status filter to ingestion, rebuilds, and re-runs the eval set with ten new return-policy questions before switching the index forward again. The postmortem adds a pre-publish eval gate for index rebuilds and a daily canary question set on policy facts.
Close the loop
An incident is closed when the lever is back in its normal position, the fix has passed evaluation, and the incident has produced lasting detection: new eval cases built from the real failing examples, with sensitive data removed, a new or tuned alert, and, when a lever was missing or slow, a new lever. Run game days where someone introduces a known-bad prompt or a poisoned document in staging and the on-call engineer works the runbook, and measure time to contain. Track error budget consumption per incident as in Error Budgets, so that repeated AI incidents change release policy rather than just filling a postmortem folder.
What to do next
- List the seven incident classes for each AI feature and write the first lever and required people for each.
- Build runtime controls: kill switch, explicit model pin and rollback, prompt and index rollback, per-tool disable, step and spend caps, and a stricter output filter profile.
- Rehearse every lever in staging monthly and verify rollback targets still exist.
- Trace every request with model, prompt, index and tool versions; keep bodies in a restricted store with retention holds.
- Keep one change log for prompts, models, indexes, tools and guardrails, including detected provider-side model changes.
- Write one runbook page per class using the six-step skeleton, with imperative, specific actions.
- Add a model owner and a safety or privacy liaison to the incident roles.
- Close every incident with new eval cases, a detector and a game-day scenario.