A tabletop exercise is a rehearsal in a meeting room. A facilitator walks a group of people through a realistic incident, releasing new information in stages, and the group talks through what they would do. Nobody touches production. The point is to find out, cheaply and before it matters, whether the people who would handle a real incident know who decides what, have the access and evidence they need, and can meet the clocks that apply to them.
AI systems need their own tabletops because their incidents do not look like the ones existing runbooks were written for. Is a model saying something harmful a bug, an attack or a vendor change? Who can turn off a tool an agent is using? Are prompts and outputs, which are both evidence and personal data, retained at all? This article shows how to design, run and score a tabletop for an LLM or agent system, with a complete worked scenario, an inject schedule in code and a scoring script that turns a decision log into findings. It assumes you already have the basics from LLM incident response.
What a tabletop is, and what it is not
Exercise programmes, such as the one described in NIST SP 800-84 and the US Homeland Security Exercise and Evaluation Program (HSEEP), separate discussion-based exercises from operations-based ones. A tabletop is discussion-based: people state what they would do, and the facilitator probes. A functional exercise or drill has people actually perform steps, such as revoking a key or rolling back a model, in a test or production environment. A red team attacks the system itself. They answer different questions:
| Activity | Tests | Typical finding |
|---|---|---|
| Tabletop | Decisions, roles, communication, clocks | Nobody owns the call to disable the agent |
| Functional drill | Whether a documented step works | The kill switch takes 40 minutes to deploy |
| Red team | Whether the system can be attacked | Indirect prompt injection exfiltrates data |
| Disaster recovery test | Restore and failover | Vector index restore loses embeddings metadata |
A good programme chains them: a red-team finding becomes a tabletop scenario, the tabletop exposes a missing control, and a drill proves the control works. Continuity and recovery testing is covered in AI disaster recovery and continuity.
Why AI incidents need their own exercises
Six properties of AI incidents make generic tabletops miss the hard parts:
- Ambiguous cause. The same symptom, a harmful or leaking answer, can come from a prompt change, a retrieval document, a user attack, a vendor model update or plain randomness. The first hour is spent deciding which, and that triage is itself a skill to exercise.
- Non-reproducibility. Sampling means a replay may not show the problem. Teams need to know whether request IDs, full prompts, retrieved context and model version are logged.
- Evidence is personal data. Conversation logs prove what happened and also contain user data, so retention, access and sharing with a vendor are decisions, not defaults.
- More rollback levers. Model version, system prompt, retrieval index, tool permissions and guardrail thresholds can each be rolled back separately, with different owners.
- Third parties. A hosted model provider controls part of the stack and has its own support process and timelines.
- Public visibility. Screenshots of model output spread fast and are easy to fake, so verification before comment matters; see AI press response.
Roles in the room
Keep the room small enough to talk, usually six to twelve players. The roles:
| Role | Responsibility |
|---|---|
| Sponsor | Sets objectives, attends the hot wash, owns funding for fixes |
| Facilitator | Runs the clock, delivers injects, asks follow-up questions, keeps it blameless |
| Players | The real responders: incident commander, ML or platform engineer, security, legal or privacy, comms, product owner |
| Simulation cell | Plays everyone not in the room: vendor support, regulator, journalist, angry customer |
| Scribe | Logs every decision with time, owner and rationale, plus open questions |
| Observers | Watch for gaps against objectives; do not play |
The facilitator should not be the person whose process is being tested. If you only have one security team, swap facilitators with another team or bring in someone from outside.
Designing objectives and scenarios
Start from objectives, not from a dramatic story. An objective names a decision and a standard, for example: within 15 minutes of evidence of data leaving through an agent tool, the on-call engineer disables the tool without needing approval. Three to five objectives per exercise is plenty. Then write a scenario that forces those decisions, and injects that release the evidence in a realistic order: partial, noisy, sometimes contradictory.
A scenario library for LLM and agent systems, each tied to a decision it tests:
| Scenario | Decisions it forces |
|---|---|
| Indirect prompt injection in a retrieved document makes a support agent email customer data | Disable a tool, quarantine a document, scope affected users, notification |
| Hosted model provider ships an update; refusal rates drop and harmful outputs rise | Pin or roll back a model version, contact vendor, re-run safety evals |
| Training or fine-tuning data found to contain poisoned or licensed material | Pull a model version, trace lineage, legal hold |
| Coding agent runs a destructive command in a shared environment | Revoke credentials, restore, review tool permission scopes |
| Leaked system prompt and jailbreak trending publicly | Verify, patch prompt, holding statement, monitor abuse |
| Personal data found in a vector index that should not hold it | Delete and re-index, data subject requests, retention review |
Pull scenarios from your own threat model, see threat modeling for LLM applications, and from red-team findings. A scenario your team has never considered is more valuable than a polished one they have rehearsed before.
Worked scenario: a poisoned page and an emailing agent
The worked scenario is a support agent with retrieval and a send_email tool. An attacker has planted instructions in a product page that the agent retrieves, and the agent follows them. The exercise runs 90 minutes of exercise time, compressed so that minute 45 can represent the next morning. The injects and the decisions they should trigger live in code, so the same scenario can be rerun next quarter and compared:
from dataclasses import dataclass
@dataclass
class Inject:
at_min: int # exercise clock, minutes from start
source: str # who delivers it in the story
text: str
expects: list # decisions this inject should trigger
deadline_min: int # minutes allowed after delivery
INJECTS = [
Inject(0, "support lead", "Three customers say the assistant emailed them "
"another customer's order history.", ["declare_incident", "assign_ic"], 10),
Inject(15, "monitoring", "send_email tool calls to recipients not on the ticket, "
"first seen 02:10 UTC.", ["disable_tool", "preserve_logs"], 15),
Inject(30, "engineering", "A product page in the RAG index has hidden text telling "
"the agent to email order history.", ["quarantine_document", "scope_query"], 20),
Inject(45, "legal", "Up to 1,840 customers; names, addresses, order contents.",
["notification_decision"], 20),
Inject(60, "press", "Journalist asks for comment on a screenshot.", ["holding_statement"], 15),
Inject(75, "vendor", "Model provider sees no platform issue, asks for request IDs.",
["share_evidence_policy"], 15),
]
def score(injects, decisions):
"""decisions: (minute, name, owner) from the scribe log. First occurrence counts."""
first = {}
for minute, name, owner in sorted(decisions):
first.setdefault(name, (minute, owner))
rows = []
for inj in injects:
due = inj.at_min + inj.deadline_min
for want in inj.expects:
if want not in first:
rows.append((want, "missed", None, None))
continue
minute, owner = first[want]
rows.append((want, "met" if minute <= due else "late", minute - inj.at_min, owner))
return rowsFeeding the scribe log from the run into score() gives this table:
| Decision | Status | Minutes after inject | Owner |
|---|---|---|---|
| declare_incident | met | 6 | support lead |
| assign_ic | late | 12 | eng manager |
| disable_tool | met | 7 | on-call SRE |
| preserve_logs | late | 26 | security |
| quarantine_document | met | 8 | ML engineer |
| scope_query | missed | - | - |
| notification_decision | late | 25 | DPO |
| holding_statement | met | 8 | comms |
| share_evidence_policy | met | 7 | security |
The numbers are a prompt for discussion, not a grade. The miss is the interesting one: nobody asked which other customers the poisoned page had been served to, so the notification decision at minute 70 rested on legal's estimate rather than a query. Logs were preserved late because conversation retention was seven days and nobody knew who could extend it. Both become findings.
Running the session
Practical rules for running the session:
- Open with ground rules: no-fault, decisions are hypothetical, say what you would actually do rather than what the policy says.
- Announce the time compression and keep a visible exercise clock.
- Deliver each inject as an artefact: a fake alert, a log excerpt, an email from the vendor. Text on a slide is less convincing.
- Ask follow-ups that force specifics: who exactly, with what access, how long does that take, how would you know it worked.
- Keep a parking lot for good questions that would derail the flow, and assign them at the end.
- Let the simulation cell improvise within limits, but do not rescue players who are stuck; being stuck is the finding.
- End with a hot wash of 20 to 30 minutes while memory is fresh: what went well, what was unclear, what scared you.
Notification clocks as injects
Regulatory and contractual clocks belong in the scenario because they force decisions under uncertainty. Under GDPR Article 33, a personal data breach must be notified to the supervisory authority without undue delay and, where feasible, within 72 hours of becoming aware of it. The EU AI Act, Article 73, adds serious-incident reporting duties for providers of high-risk AI systems, with deadlines that depend on the type of incident. Customer contracts and sector rules add their own. Rather than putting figures from memory into the scenario, ask counsel to supply the current deadlines that apply to your system and include them as an inject from legal. The exercise question is not the number, it is when the clock started and who decided whether it applies.
After-action report and fixes
The after-action report turns a conversation into change. Write it within a week, with one row per finding: what happened, why it matters, the fix, an owner, a due date and how the fix will be verified. Fixes for AI tabletops usually land in four places:
- Runbooks: who may disable a tool, roll back a model or quarantine a document without approval. Link them from alerts.
- Detections: the alert that should have fired at minute 0, such as outbound email to recipients not on the ticket; see the SOC playbook.
- Controls: kill switches per tool, log retention for incident holds, a query that lists users served a given document.
- Evals: the poisoned document becomes a regression test case for the agent.
Track findings like bugs, in the same system as engineering work. The next exercise opens by retesting the previous findings; an unverified fix is still an open finding.
Failure modes
| Failure mode | What it looks like | Countermeasure |
|---|---|---|
| Theatre | Scripted answers, everyone reads the policy | Ask what they would do at 3 a.m. with only their laptop |
| Too easy | All objectives met, nothing learned | Add ambiguity and conflicting injects; raise difficulty each round |
| Wrong room | Managers play, responders absent | Invite the on-call engineers who would actually be paged |
| Blame | People defend themselves instead of exploring | Facilitator enforces no-fault; sponsor states it at the start |
| Orphan findings | Same gaps next quarter | Owner and date per finding, retest at next exercise |
| Generic scenario | Ransomware with an AI label | Base scenarios on your architecture and red-team results |
Cadence and maturity
Run a tabletop at least twice a year per critical AI system, and additionally after major changes: a new model provider, a new tool with write access, a new data source in retrieval, or a real incident elsewhere that resembles your system. Grow difficulty over time: start with a single team and a clear cause, move to cross-functional exercises with ambiguous cause and third parties, then pair the tabletop with a functional drill where the kill switch is actually pulled in staging and timed. Keep scenarios, inject files and scores in version control so trends are visible.
What to do next
- Pick one AI system with tool access or customer data and write three objectives as decisions with deadlines.
- Choose a scenario from your threat model or last red-team report and write five or six injects as code.
- Book a 2-hour slot with the real responders, a facilitator from outside the team and a scribe.
- Before the session, confirm with counsel which notification clocks apply and turn them into a legal inject.
- Run it, log every decision with minute and owner, and score it with the script.
- Write the after-action report within a week with owner, due date and verification for each finding.
- Schedule the next exercise and open it by retesting this one's findings.