When an LLM feature misbehaves at 2 a.m., the on-call engineer has two problems. The first is technical: is the model leaking data, following an injected instruction, regressing after a prompt change, or just suffering from a provider outage? The second is authority: may I switch off retrieval for every tenant, roll back the system prompt, revoke an agent's credentials, or tell customers? A good playbook answers both before the incident starts. It says how to recognise this kind of incident, which actions the responder may take at once without asking, and which decisions need a named person.
This article is about the playbook as an artifact: what goes in it, how to store it as data so tools can lint, route and execute it, a catalogue of the AI-specific scenarios worth writing first, a worked example run against the clock, and how to stop playbooks going stale. It does not cover severity scales or the general containment ladder; those are in LLM Incident Response. Exercises that test playbooks are covered in AI Tabletop Exercises.
Plan, playbook, runbook: what a playbook is for
Three documents get confused. An incident response plan is organisation-wide: who the incident commander is, how incidents are declared, how severity is set, how regulators are notified. A runbook is a mechanical procedure for one task: how to roll back a model version, how to rotate a vector store's API key. A playbook sits between them. It is scenario-shaped: for one class of incident it ties detection signals to a sequence of steps, each of which may call runbooks, and to decision points owned by roles.
NIST's incident response guidance, SP 800-61 Revision 3 (finalised April 2025), now frames response as part of overall cyber risk management, organised around the CSF 2.0 functions rather than a stand-alone lifecycle. That framing helps here. A playbook is not only a respond-phase document. It has a detect part (entry criteria), a protect part (pre-built switches it depends on), and a recover part (exit criteria and follow-up evals). If the switch a playbook tells you to flip does not exist yet, the playbook is fiction. Writing playbooks is a cheap way to find missing controls before an attacker does.
AI systems need their own playbooks because the usual first moves do not map cleanly. Blocking an IP address does not stop an injection planted in a shared document, and rolling back code does not roll back a poisoned vector index or an agent's memory. The playbook must name the prompt version, retrieval sources, tool grants, memory stores and model route, and say how to change each quickly.
Anatomy of a playbook
Every playbook in a repository should have the same fields, so responders always know where to look and tools can check them. The table lists the fields we recommend and the question each one answers under pressure.
| Field | Answers | Common mistake |
|---|---|---|
| Entry criteria | Is this the right playbook? | Vague prose such as 'model behaves oddly'; write observable signals |
| Severity hints | How bad can this get? | Restating the org scale instead of saying which facts raise severity |
| Roles | Who owns each decision? | Naming people, who leave; name roles that resolve to the rota |
| Pre-authorised actions | What may I do now, alone? | Listing actions that need approval from someone asleep |
| Decision points | What needs a named approver, and on what evidence? | No default if the approver cannot be reached |
| Evidence to preserve | What must I capture before I change anything? | Rolling back first and losing the prompt and retrieval logs |
| Comms | Who is told, when, with which template? | No regulatory clock; legal hears on day three |
| Exit criteria | When is it over? | 'Looks fine now'; write a measurable check |
| Follow-ups | What becomes a test so it cannot recur silently? | Action items with no owner and no eval |
Two fields carry most of the value. Pre-authorised actions cut time to containment, because the responder does not need to wake anyone to switch retrieval to read-only. Decision points with a default stop the incident stalling. For example: 'if the product owner has not answered within 20 minutes, keep the feature disabled.' A default that fails safe is a decision made in advance by people who were calm when they made it.
Architecture: playbooks in the incident path
Playbooks as linted data
Store playbooks as structured files in a repository, not as wiki pages. Structured files can be reviewed like code, versioned, linted in CI and loaded by the incident tooling, so the incident record shows exactly which version the responder followed. Here is a cross-tenant retrieval leak playbook in YAML. The action names are your own internal API, not a vendor's.
id: rag-cross-tenant-leak
version: 7
owner_role: ml-platform-lead
entry_criteria:
- signal: canary_token_in_output # planted per-tenant canary seen in another tenant's answer
- signal: user_report
match: "another company's|not our data|someone else's"
severity_hints:
raise_if: [pii_in_leaked_text, more_than_one_tenant, external_customer_reported]
pre_authorised:
- action: retrieval.set_mode
args: {mode: tenant_strict_only} # drop shared and cross-tenant indexes
runbook: rb/retrieval-mode
- action: response_cache.flush
runbook: rb/cache-flush
evidence_first:
- prompt_and_context_manifests: last_24h
- retrieval_logs: last_7d
- index_acl_snapshot: now
decisions:
- id: disable_feature
approver_role: product-owner
evidence: [leak_confirmed_by_replay]
timeout_minutes: 20
default: disable
- id: notify_customers
approver_role: privacy-counsel
evidence: [affected_tenant_list]
timeout_minutes: 120
default: escalate_to_ciso
exit_criteria:
- check: canary_sweep_clean
window_hours: 24
- check: acl_filter_eval_pass_rate
min: 1.0
follow_ups: [add_acl_regression_eval, add_canary_for_new_tenants]The linter below fails the build when a playbook refers to an action no system implements, a runbook that was deleted, or a role no rota resolves. It also fails when a decision has no timeout and default. Those are the four ways playbooks quietly rot. The three loaders are stubs you connect to your on-call tool, runbook directory and control-plane API.
import sys
import yaml # PyYAML
REQUIRED = ["id", "version", "owner_role", "entry_criteria", "pre_authorised",
"evidence_first", "decisions", "exit_criteria", "follow_ups"]
def lint(pb, known_roles, known_runbooks, known_actions):
errs = [f"missing field: {f}" for f in REQUIRED if f not in pb]
for a in pb.get("pre_authorised", []):
if a["action"] not in known_actions:
errs.append(f"action {a['action']} has no implementation")
if a.get("runbook") not in known_runbooks:
errs.append(f"runbook {a.get('runbook')} does not exist")
for d in pb.get("decisions", []):
if d["approver_role"] not in known_roles:
errs.append(f"decision {d['id']}: unknown role {d['approver_role']}")
if "default" not in d or "timeout_minutes" not in d:
errs.append(f"decision {d['id']}: needs timeout_minutes and a default")
for e in pb.get("exit_criteria", []):
if "check" not in e:
errs.append("exit criterion without a machine check")
return errs
if __name__ == "__main__":
pb = yaml.safe_load(open(sys.argv[1], encoding="utf-8"))
problems = lint(pb, load_roles(), load_runbooks(), load_actions()) # from your registries
for msg in problems:
print(f"{pb.get('id', '?')}: {msg}")
sys.exit(1 if problems else 0)Run the linter on every change to the playbook repository. Also run it nightly, because the things a playbook depends on change elsewhere. When someone renames the retrieval.set_mode endpoint in the platform repository, the nightly lint should fail before an incident does.
A starting catalogue of AI scenarios
Do not write fifty playbooks. Write the five or six that cover most AI incidents, make them excellent, and route everything else to a generic playbook that starts with 'preserve evidence, then contain with the kill switch'. This is a reasonable starting catalogue for a product with retrieval and tool-using agents:
| Scenario | Entry signal | First pre-authorised action | Key decision | Exit check |
|---|---|---|---|---|
| Indirect prompt injection drives a tool call | Tool call to an unexpected domain; egress alert | Revoke the agent's tool grants to read-only | Restore write tools? (security lead) | Replay of the payload blocked; no new egress for 24h |
| Cross-tenant or PII leak via retrieval | Canary token crosses tenants; user report | Tenant-strict retrieval, flush response cache | Notify customers or regulator? (privacy counsel) | Canary sweep clean; ACL eval at 100% |
| Jailbreak circulating publicly | Spike of a known prompt shape; social media | Add the shape to the input filter in monitor-then-block mode | Public statement? (comms) | Attack success rate on the variant set below threshold |
| Quality or safety regression after a change | Post-deploy eval drop; complaint rate | Roll back prompt or model route to last good version | Roll forward with a fix or hold? (product owner) | Eval suite back at baseline on live traffic sample |
| Cost or abuse spike | Tokens per user far above normal; key used from new ASN | Per-key rate limit; suspend the key | Refund or block account? (trust and safety) | Spend back inside budget band for 6h |
| Upstream model outage or silent change | Provider error rate; output drift detector | Fail over to secondary model route with stricter filters | Stay on secondary? (ML lead) | Primary passes the canary eval set |
Each first action is reversible and narrow. Revoking write tools hurts the product but loses no data. Tenant-strict retrieval hurts answer quality but is easy to undo. Reversibility is what makes it safe to pre-authorise an action. Irreversible steps, such as deleting an index, wiping agent memory or telling customers, belong at decision points. For the stop mechanics behind several of these actions, see Agent Kill Switch.
Worked example: a cross-tenant leak, minute by minute
Here is how the cross-tenant leak playbook runs against the clock. Times are minutes from the first alert.
| T+ | Step | What the playbook made possible |
|---|---|---|
| 0 | Canary alert: tenant A's canary string appears in an answer served to tenant B. | The router matched an entry criterion and opened the incident with playbook v7 attached. |
| 3 | Responder snapshots context manifests, retrieval logs and the index ACL table. | Evidence came first, so the later rollback did not destroy the trail. |
| 6 | Retrieval set to tenant_strict_only; response cache flushed. | Both actions were pre-authorised. No one had to be woken. |
| 15 | Replay in staging of the logged request reproduces the leak; cause found: a shared 'help centre' index carried tenant documents after a bulk import. | Replay tooling and manifests existed, because the playbook demanded them. |
| 20 | Product owner unreachable; decision 'disable_feature' falls to the default: disable. | The default kept the incident moving. |
| 45 | Affected tenant list built from retrieval logs: two tenants, one document each, no PII. | The evidence scope was defined in advance (seven days of logs). |
| 110 | Privacy counsel decides no regulator notification is needed; customer notices go out. | Comms templates and the decision record were ready. |
| 24h | Canary sweep clean; ACL eval at 100%; feature re-enabled. | Exit criteria were machine checks, not opinions. |
The review then changed two things. The system got an ingestion check that rejects tenant-labelled documents in shared indexes. The playbook got a new evidence item, the bulk import log, and its version went to 8. For investigation techniques used at T+15, see AI Forensics.
Keeping playbooks alive
Playbooks rot in predictable ways: actions get renamed, roles get reorganised, new features add surfaces no playbook mentions, and responders learn shortcuts that never get written down. Four habits keep them alive.
- Lint continuously. The CI and nightly lint above catches broken references. Add a check that every production AI feature appears in at least one playbook's scope list.
- Drill the pre-authorised actions. Flip each switch in staging every month and in production at least once a quarter in a low-traffic window. A kill switch that has never been flipped is a hypothesis.
- Diff after every incident. Compare the steps actually taken, from the incident record, with the playbook steps. Any step responders skipped or added becomes a pull request.
- Expire them. Give each playbook a review-by date. The router warns when it opens an incident with an expired playbook, and the owner role gets a ticket.
Failure modes
- Wrong playbook, confident responder. A provider outage looks like a regression; a leak looks like a hallucination. Mitigation: entry criteria include a short discriminating check, and every playbook has a 'this is not me if' line.
- Containment destroys evidence. Flushing caches and rolling back prompts before snapshotting manifests leaves nothing to replay. Mitigation: evidence steps come first and the tooling refuses to run the action until the snapshot exists.
- Pre-authorised action that is too broad. 'Disable the AI feature globally' as a first move turns a one-tenant incident into an outage for everyone. Mitigation: scope the first action to the affected path or tenant set.
- Approver bottleneck. Decision points without defaults stall at night. Mitigation: timeouts with fail-safe defaults and a deputy role.
- Playbook drift. Steps refer to dashboards or endpoints that no longer exist. Mitigation: lint, drills and review-by dates.
- Communications run ahead of facts. A statement goes out before the scope is known. Mitigation: comms steps are gated on evidence items; see AI Press and Incident Response.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Structured YAML vs wiki prose | Lintable, versioned, executable | Less room for nuance; needs tooling |
| Many narrow playbooks vs few broad ones | Precise steps | Routing errors and maintenance load grow with the count |
| Broad pre-authorisation | Fast containment | Risk of over-containment and user impact |
| Automated execution (SOAR style) | Seconds instead of minutes | Automation bugs act at machine speed; keep a human gate for irreversible steps |
What to do next
- List your AI components that a responder may need to change: prompt versions, model routes, retrieval indexes, tool grants, memory stores, caches.
- For each, confirm a fast, reversible control exists and has been exercised. Build the missing ones first.
- Write the cross-tenant leak and injection-driven tool-call playbooks in the schema above, with defaults on every decision.
- Add the linter to CI and a nightly job, wired to your role, runbook and action registries.
- Plant per-tenant canary strings so the leak playbook has a reliable entry signal.
- Run one drill per playbook this quarter and diff what happened against the file.
- After each real incident, open a playbook pull request and an eval before closing the review.