A tabletop exercise is a rehearsal in a meeting room. A facilitator walks a group of people through a realistic incident, releasing new information in stages, and the group talks through what they would do. Nobody touches production. The point is to find out, cheaply and before it matters, whether the people who would handle a real incident know who decides what, have the access and evidence they need, and can meet the clocks that apply to them.

AI systems need their own tabletops because their incidents do not look like the ones existing runbooks were written for. Is a model saying something harmful a bug, an attack or a vendor change? Who can turn off a tool an agent is using? Are prompts and outputs, which are both evidence and personal data, retained at all? This article shows how to design, run and score a tabletop for an LLM or agent system, with a complete worked scenario, an inject schedule in code and a scoring script that turns a decision log into findings. It assumes you already have the basics from LLM incident response.

What a tabletop is, and what it is not

Exercise programmes, such as the one described in NIST SP 800-84 and the US Homeland Security Exercise and Evaluation Program (HSEEP), separate discussion-based exercises from operations-based ones. A tabletop is discussion-based: people state what they would do, and the facilitator probes. A functional exercise or drill has people actually perform steps, such as revoking a key or rolling back a model, in a test or production environment. A red team attacks the system itself. They answer different questions:

ActivityTestsTypical finding
TabletopDecisions, roles, communication, clocksNobody owns the call to disable the agent
Functional drillWhether a documented step worksThe kill switch takes 40 minutes to deploy
Red teamWhether the system can be attackedIndirect prompt injection exfiltrates data
Disaster recovery testRestore and failoverVector index restore loses embeddings metadata

A good programme chains them: a red-team finding becomes a tabletop scenario, the tabletop exposes a missing control, and a drill proves the control works. Continuity and recovery testing is covered in AI disaster recovery and continuity.

Why AI incidents need their own exercises

Six properties of AI incidents make generic tabletops miss the hard parts:

  • Ambiguous cause. The same symptom, a harmful or leaking answer, can come from a prompt change, a retrieval document, a user attack, a vendor model update or plain randomness. The first hour is spent deciding which, and that triage is itself a skill to exercise.
  • Non-reproducibility. Sampling means a replay may not show the problem. Teams need to know whether request IDs, full prompts, retrieved context and model version are logged.
  • Evidence is personal data. Conversation logs prove what happened and also contain user data, so retention, access and sharing with a vendor are decisions, not defaults.
  • More rollback levers. Model version, system prompt, retrieval index, tool permissions and guardrail thresholds can each be rolled back separately, with different owners.
  • Third parties. A hosted model provider controls part of the stack and has its own support process and timelines.
  • Public visibility. Screenshots of model output spread fast and are easy to fake, so verification before comment matters; see AI press response.

Roles in the room

Keep the room small enough to talk, usually six to twelve players. The roles:

RoleResponsibility
SponsorSets objectives, attends the hot wash, owns funding for fixes
FacilitatorRuns the clock, delivers injects, asks follow-up questions, keeps it blameless
PlayersThe real responders: incident commander, ML or platform engineer, security, legal or privacy, comms, product owner
Simulation cellPlays everyone not in the room: vendor support, regulator, journalist, angry customer
ScribeLogs every decision with time, owner and rationale, plus open questions
ObserversWatch for gaps against objectives; do not play

The facilitator should not be the person whose process is being tested. If you only have one security team, swap facilitators with another team or bring in someone from outside.

Designing objectives and scenarios

Start from objectives, not from a dramatic story. An objective names a decision and a standard, for example: within 15 minutes of evidence of data leaving through an agent tool, the on-call engineer disables the tool without needing approval. Three to five objectives per exercise is plenty. Then write a scenario that forces those decisions, and injects that release the evidence in a realistic order: partial, noisy, sometimes contradictory.

A scenario library for LLM and agent systems, each tied to a decision it tests:

ScenarioDecisions it forces
Indirect prompt injection in a retrieved document makes a support agent email customer dataDisable a tool, quarantine a document, scope affected users, notification
Hosted model provider ships an update; refusal rates drop and harmful outputs risePin or roll back a model version, contact vendor, re-run safety evals
Training or fine-tuning data found to contain poisoned or licensed materialPull a model version, trace lineage, legal hold
Coding agent runs a destructive command in a shared environmentRevoke credentials, restore, review tool permission scopes
Leaked system prompt and jailbreak trending publiclyVerify, patch prompt, holding statement, monitor abuse
Personal data found in a vector index that should not hold itDelete and re-index, data subject requests, retention review

Pull scenarios from your own threat model, see threat modeling for LLM applications, and from red-team findings. A scenario your team has never considered is more valuable than a polished one they have rehearsed before.

Worked scenario: a poisoned page and an emailing agent

The worked scenario is a support agent with retrieval and a send_email tool. An attacker has planted instructions in a product page that the agent retrieves, and the agent follows them. The exercise runs 90 minutes of exercise time, compressed so that minute 45 can represent the next morning. The injects and the decisions they should trigger live in code, so the same scenario can be rerun next quarter and compared:

from dataclasses import dataclass

@dataclass
class Inject:
    at_min: int          # exercise clock, minutes from start
    source: str          # who delivers it in the story
    text: str
    expects: list        # decisions this inject should trigger
    deadline_min: int    # minutes allowed after delivery

INJECTS = [
    Inject(0, "support lead", "Three customers say the assistant emailed them "
           "another customer's order history.", ["declare_incident", "assign_ic"], 10),
    Inject(15, "monitoring", "send_email tool calls to recipients not on the ticket, "
           "first seen 02:10 UTC.", ["disable_tool", "preserve_logs"], 15),
    Inject(30, "engineering", "A product page in the RAG index has hidden text telling "
           "the agent to email order history.", ["quarantine_document", "scope_query"], 20),
    Inject(45, "legal", "Up to 1,840 customers; names, addresses, order contents.",
           ["notification_decision"], 20),
    Inject(60, "press", "Journalist asks for comment on a screenshot.", ["holding_statement"], 15),
    Inject(75, "vendor", "Model provider sees no platform issue, asks for request IDs.",
           ["share_evidence_policy"], 15),
]

def score(injects, decisions):
    """decisions: (minute, name, owner) from the scribe log. First occurrence counts."""
    first = {}
    for minute, name, owner in sorted(decisions):
        first.setdefault(name, (minute, owner))
    rows = []
    for inj in injects:
        due = inj.at_min + inj.deadline_min
        for want in inj.expects:
            if want not in first:
                rows.append((want, "missed", None, None))
                continue
            minute, owner = first[want]
            rows.append((want, "met" if minute <= due else "late", minute - inj.at_min, owner))
    return rows
Worked scenario: 90-minute timeline, injects above, decisions below0153045607590customer reportstool-call logspoisoned doclegal: 1,840 peoplepress queryvendor replydeclareIC namedtool offquarantinelogs heldholding stmtnotify callevidence policyGreen: within deadline. Amber: late. Never made: scope the affected-customer query (missed).
The run as the scribe logged it. Each inject expects decisions within a deadline; the scoring script classifies each one.

Feeding the scribe log from the run into score() gives this table:

DecisionStatusMinutes after injectOwner
declare_incidentmet6support lead
assign_iclate12eng manager
disable_toolmet7on-call SRE
preserve_logslate26security
quarantine_documentmet8ML engineer
scope_querymissed--
notification_decisionlate25DPO
holding_statementmet8comms
share_evidence_policymet7security

The numbers are a prompt for discussion, not a grade. The miss is the interesting one: nobody asked which other customers the poisoned page had been served to, so the notification decision at minute 70 rested on legal's estimate rather than a query. Logs were preserved late because conversation retention was seven days and nobody knew who could extend it. Both become findings.

Running the session

Practical rules for running the session:

  • Open with ground rules: no-fault, decisions are hypothetical, say what you would actually do rather than what the policy says.
  • Announce the time compression and keep a visible exercise clock.
  • Deliver each inject as an artefact: a fake alert, a log excerpt, an email from the vendor. Text on a slide is less convincing.
  • Ask follow-ups that force specifics: who exactly, with what access, how long does that take, how would you know it worked.
  • Keep a parking lot for good questions that would derail the flow, and assign them at the end.
  • Let the simulation cell improvise within limits, but do not rescue players who are stuck; being stuck is the finding.
  • End with a hot wash of 20 to 30 minutes while memory is fresh: what went well, what was unclear, what scared you.

Notification clocks as injects

Regulatory and contractual clocks belong in the scenario because they force decisions under uncertainty. Under GDPR Article 33, a personal data breach must be notified to the supervisory authority without undue delay and, where feasible, within 72 hours of becoming aware of it. The EU AI Act, Article 73, adds serious-incident reporting duties for providers of high-risk AI systems, with deadlines that depend on the type of incident. Customer contracts and sector rules add their own. Rather than putting figures from memory into the scenario, ask counsel to supply the current deadlines that apply to your system and include them as an inject from legal. The exercise question is not the number, it is when the clock started and who decided whether it applies.

After-action report and fixes

The after-action report turns a conversation into change. Write it within a week, with one row per finding: what happened, why it matters, the fix, an owner, a due date and how the fix will be verified. Fixes for AI tabletops usually land in four places:

  • Runbooks: who may disable a tool, roll back a model or quarantine a document without approval. Link them from alerts.
  • Detections: the alert that should have fired at minute 0, such as outbound email to recipients not on the ticket; see the SOC playbook.
  • Controls: kill switches per tool, log retention for incident holds, a query that lists users served a given document.
  • Evals: the poisoned document becomes a regression test case for the agent.

Track findings like bugs, in the same system as engineering work. The next exercise opens by retesting the previous findings; an unverified fix is still an open finding.

The exercise cycle: objectives in, tracked fixes outObjectivesdecisions to testScenarioplus injectsRun60 to 120 minHot washsame dayAARfindings, ownersFixesrunbooks, detections, kill switches, eval casesNext exerciseretest last findings, raise difficultyverifiedDecision log from the run feeds the scoring script; findings without an owner and a date do not count
A tabletop is one turn of a loop. The output that matters is verified fixes, not the meeting.

Failure modes

Failure modeWhat it looks likeCountermeasure
TheatreScripted answers, everyone reads the policyAsk what they would do at 3 a.m. with only their laptop
Too easyAll objectives met, nothing learnedAdd ambiguity and conflicting injects; raise difficulty each round
Wrong roomManagers play, responders absentInvite the on-call engineers who would actually be paged
BlamePeople defend themselves instead of exploringFacilitator enforces no-fault; sponsor states it at the start
Orphan findingsSame gaps next quarterOwner and date per finding, retest at next exercise
Generic scenarioRansomware with an AI labelBase scenarios on your architecture and red-team results

Cadence and maturity

Run a tabletop at least twice a year per critical AI system, and additionally after major changes: a new model provider, a new tool with write access, a new data source in retrieval, or a real incident elsewhere that resembles your system. Grow difficulty over time: start with a single team and a clear cause, move to cross-functional exercises with ambiguous cause and third parties, then pair the tabletop with a functional drill where the kill switch is actually pulled in staging and timed. Keep scenarios, inject files and scores in version control so trends are visible.

What to do next

  1. Pick one AI system with tool access or customer data and write three objectives as decisions with deadlines.
  2. Choose a scenario from your threat model or last red-team report and write five or six injects as code.
  3. Book a 2-hour slot with the real responders, a facilitator from outside the team and a scribe.
  4. Before the session, confirm with counsel which notification clocks apply and turn them into a legal inject.
  5. Run it, log every decision with minute and owner, and score it with the script.
  6. Write the after-action report within a week with owner, due date and verification for each finding.
  7. Schedule the next exercise and open it by retesting this one's findings.
Key takeaway: A tabletop tests decisions, not systems: who acts, with what access and evidence, within which deadline. For AI systems, design scenarios around ambiguous cause, non-reproducible outputs, evidence that is personal data, separate rollback levers and third-party models. Keep injects and objectives in code, score the decision log, and measure success by findings that are fixed and retested at the next exercise.