When an AI system misbehaves in public, the first sign is often not an alert but a screenshot: a chatbot inventing a policy, leaking what looks like another customer's data, or producing something offensive. Within an hour it may be shared widely and a journalist may be asking for comment. The technical incident response, which is covered in LLM Incident Response, in depth, runs in parallel with a second workstream: deciding what to say, to whom, and when.
This article treats that second workstream as an engineering problem. It explains why AI incidents are unusually hard to communicate about, how to verify a viral artefact against your own logs, how a fact register stops statements from getting ahead of evidence, how to track the legal notification clocks that may start, and how to sequence messages to the press, users, regulators and customers. It includes code for a statement gate and an artefact matcher, a worked example and a checklist. It is not legal advice: notification duties depend on jurisdiction, sector and contract, and counsel must confirm them.
The architecture
Why AI incidents are hard to talk about
Three properties make AI incidents harder to talk about than an outage. Nondeterminism. The same prompt may not reproduce the same output, so "we could not reproduce it" is weak evidence that it did not happen. Fabrication is cheap. A convincing chat screenshot takes minutes to fake with browser developer tools, so a viral image may be real, edited or invented. Attribution is unclear to the public. People treat what a chatbot says as the company speaking. In Moffatt v. Air Canada (2024), a Canadian tribunal rejected the airline's argument that its website chatbot was responsible for its own statements and held the airline to what the chatbot had told a customer. Plan on the assumption that your assistant's words are your words.
Those properties produce two opposite failure modes. Teams either deny too early ("this screenshot is fake") and are contradicted by their own logs, or concede too early ("our AI leaked customer data") before learning that the data came from the user's own earlier message. Both damage trust more than a short, accurate holding statement would have. The cure for both is the same: separate what you know from what you suspect, and only say what you know.
Roles and the communications workstream
Run communications as a named workstream inside the incident, not as a parallel effort that learns facts from chat. The roles that work:
- Incident commander owns the technical facts and the fact register. Nothing becomes "verified" without their sign-off.
- Communications lead owns timing, channels and drafts, and is the only path to press and social media. They attend the incident bridge.
- Legal and privacy decide which notification duties apply, approve external wording and own regulator contact.
- Customer lead handles enterprise accounts whose contracts carry their own notice terms.
- Spokesperson, chosen and trained in advance, says only what the approved statement says.
Write these names into the runbook before an incident. The general mechanics of running an incident, including the commander role, are in How to run an incident.
Verifying a viral artefact
The first technical question is whether the artefact is real. Treat it as evidence: save the original post, image and URL with timestamps before it is deleted, then search your own logs. Narrow candidates by product surface, time window, locale and any identifier visible in the image, then compare the screenshot's text with logged model outputs.
import re
def shingles(text, k=5):
words = re.findall(r"[a-z0-9]+", text.lower())
return {" ".join(words[i:i + k]) for i in range(max(1, len(words) - k + 1))}
def match_artifact(ocr_text, candidate_turns, threshold=0.6):
"""Compare text read from a screenshot with logged model outputs.
candidate_turns: (conversation_id, turn_id, text) narrowed first by time window,
surface and any visible identifiers. Returns the best match and a verdict.
"""
target = shingles(ocr_text)
best = (0.0, None)
for conv_id, turn_id, text in candidate_turns:
logged = shingles(text)
overlap = len(target & logged) / max(1, len(target)) # share of screenshot found
if overlap > best[0]:
best = (overlap, (conv_id, turn_id))
score, where = best
if score >= threshold:
return "found", where, score
if score > 0.2:
return "partial: possibly edited or a different turn", where, score
return "not found: fabricated, outside retention, or another surface", None, scoreMeasuring the share of the screenshot's word shingles found in a logged turn is robust to OCR noise and cropping. The three outcomes mean different things, and only one of them supports a public claim. Found: the output is real; next establish what preceded it. Partial: inspect by hand; edits often change only the damaging sentence. Not found is not proof of fabrication. Your retention window may have expired, the user may have used another surface or a partner's deployment, or logging may have sampled the turn out. Never say "fake" on the strength of a missing log line.
Once a turn is found, replay it with the logged context, model version and settings, and use context ablation to see which inputs caused the output. AI Forensics, in depth covers evidence preservation, replay under nondeterminism and ablation in detail.
The fact register and the statement gate
A fact register is a short, timestamped list of statements about the incident, each with a status (verified, suspected or retracted), the evidence behind it and an owner. Every external message is drafted as a list of claims that cite register entries, and a gate refuses drafts that cite anything not verified or not recently confirmed.
from dataclasses import dataclass
from datetime import datetime, timedelta
@dataclass
class Fact:
id: str
text: str
status: str # "verified" | "suspected" | "retracted"
evidence: list # log query ids, ticket ids, replay ids
owner: str
as_of: datetime
@dataclass
class Claim:
text: str
fact_ids: list
def gate(claims, facts, now, max_age=timedelta(hours=4)):
"""Block a draft statement unless every claim rests on fresh, verified facts."""
problems = []
for c in claims:
if not c.fact_ids:
problems.append(f"unsupported claim: {c.text!r}")
for fid in c.fact_ids:
f = facts.get(fid)
if f is None:
problems.append(f"unknown fact {fid}")
elif f.status != "verified":
problems.append(f"{fid} is {f.status}: {c.text!r}")
elif now - f.as_of > max_age:
problems.append(f"{fid} last confirmed {f.as_of:%H:%M}; re-confirm")
return problemsThis is deliberately mechanical. Under pressure people round up: "a small number of users" becomes "one user" because that was the first case found. The gate forces each number and each causal statement back to an entry the commander signed. It also makes retractions visible: if a fact flips to retracted, every published statement that cited it can be listed and corrected. Keep the register in the incident channel's tooling, not in a slide, so it has history.
Good holding statements need very few facts: that you are aware, what you have done so far (for example, disabled a feature), what you do not yet know, and when you will update. Avoid speculation about cause, numbers you have not counted and promises about outcomes.
Notification clocks
Some incidents start legal clocks, and the moment of awareness belongs in the timeline. The table lists common ones; which apply depends on your role, sector and jurisdiction.
| Regime | Trigger | Deadline | To whom |
|---|---|---|---|
| GDPR Art. 33 | Personal data breach | Without undue delay, where feasible within 72 hours of awareness | Supervisory authority |
| GDPR Art. 34 | Breach likely to cause high risk to people | Without undue delay | Affected individuals |
| NIS2 Art. 23 | Significant incident at an essential or important entity | Early warning within 24 hours, notification within 72 hours, final report within one month | CSIRT or competent authority |
| EU AI Act Art. 73 | Serious incident involving a high-risk AI system | No later than 15 days; 2 days for a widespread infringement or serious critical-infrastructure disruption; 10 days for a death | Market surveillance authority |
| SEC Form 8-K Item 1.05 | Material cybersecurity incident at a US-listed company | Four business days after determining materiality | Investors, via filing |
| Contracts | Defined in each agreement | Often shorter than statute | Enterprise customers |
Two notes on the AI Act. Article 73 binds providers of high-risk systems, and the 2026 Digital Omnibus moved when high-risk obligations apply (2 December 2027 for Annex III systems, 2 August 2028 for Annex I products); see EU AI Act, in depth. Separately, providers of general-purpose models with systemic risk have their own serious-incident duty under Article 55. Track each clock as an explicit timer with an owner, and remember that a regulator notification is not a press release: wording that is appropriate for one may be wrong for the other.
Sequencing the messages
Order matters. A workable default for a confirmed incident affecting users:
- Contain first. Disable the feature, tool or data source involved. A statement that says "we have turned it off" is far stronger than one that says "we are investigating".
- Holding statement within the first hour or two if the issue is public: aware, action taken, next update time. Post it where the conversation is happening and on the status page.
- Affected users before the press, where you can identify them. People should not learn from a news story that their data was involved.
- Regulators and contractual notices on their own clocks, drafted by legal from the same register.
- Updates on the promised schedule, even if the update is "no change". Missing a promised update reads as hiding something.
- Closing statement after recovery: what happened, who was affected, what changed. Publish the substance of the postmortem when you can.
Prepare the reusable parts in advance: statement templates per incident class (harmful output, data exposure, unauthorised action, outage), an FAQ skeleton, a status-page component per AI feature, and approval paths that work at night. Pre-approved language is what lets a holding statement go out in an hour rather than a day. The public promises you make in calmer times, such as how data is used, are the ones journalists will check; AI Customer Trust, in depth covers keeping those promises enforced and verifiable.
Worked example: the address screenshot
At 09:10 a post spreads showing a support assistant replying with a customer's full home address, captioned as a data leak. By 09:25 a reporter has asked for comment.
The commander opens the register. F1 (suspected): "assistant output contained a home address". The comms lead posts nothing yet. At 09:40 the matcher finds the turn: a 0.9 share of shingles in one conversation from 08:52; F1 becomes verified. Replay with the logged context shows the address came from the user's own message three turns earlier, which the assistant repeated back while confirming a delivery change. F2 (verified): "the address shown was supplied by the same user in that conversation". F3 (suspected): "no other customer's data was exposed"; that needs a wider search of outputs for addresses not present in the same conversation's inputs.
The draft holding statement that says "no customer data was leaked" fails the gate, because it cites F3, which is only suspected. The approved version says the company has reviewed the conversation, the address shown had been provided by that user earlier in the same chat, a broader review is under way, and an update will follow by 14:00. At 13:20 the sweep completes with no cross-customer matches, F3 is verified, and the update goes out. The team also ships a fix so the assistant masks addresses when echoing them, because the screenshot showed that echoing personal data, even the user's own, looks like a leak to everyone else.
Had the team said "fake" at 09:15, or "we are investigating a data leak", they would have been wrong either way.
Failure modes
- Denying on a missing log line. Absence of a match is not evidence of fabrication.
- Numbers that only grow. Early counts are floors; say "at least" or omit them.
- Blaming the model. "The AI made a mistake" reads as evasion; you deployed it.
- Inconsistent channels. Support agents improvising answers that contradict the statement. Give them the same approved text.
- Missed clocks. Awareness time not recorded, so nobody knows when 72 hours ends.
- Over-sharing exploit detail. Publishing the exact injection string before the fix invites copycats. Describe the class, not the payload, until it is closed.
- No retraction path. A corrected fact with no list of the statements that cited it.
Trade-offs
Speed versus accuracy. Silence lets others define the story; a fast wrong statement is worse. A holding statement with few, verified facts resolves most of the tension. Transparency versus security. Detailed disclosure builds trust and helps others defend, but premature detail helps attackers; disclose fully after the fix. Logging versus privacy. Verifying artefacts needs retained outputs; retaining them is itself a privacy commitment. Choose a retention window you can defend for both reasons and write it down.
What to do next
- Name the commander, comms lead, legal owner and spokesperson for AI incidents, with deputies.
- Write holding-statement templates for each AI incident class and get them pre-approved.
- Build the artefact matcher against your output logs and test it on real and edited screenshots.
- Stand up a fact register with statuses and a gate that blocks unsupported claims.
- List the notification regimes and contract terms that apply to you, with clock owners.
- Add a status-page component for each AI feature and a kill switch the statement can cite.
- Run a tabletop exercise with a fake viral screenshot, and time the first holding statement.