Coordinated disclosure assumes you have time. A reporter tells you privately, you fix the flaw, and you publish together when users are protected. Sometimes the time is gone. The flaw is being exploited against your users, or its details are already public, and silence now protects attackers more than users. Emergency disclosure is the procedure for that case: telling people what they need to protect themselves before a complete fix exists.
AI systems hit this case more often than traditional software. A jailbreak or injection string spreads on social media within hours. An exfiltration trick in one agent framework transfers to others. A server-side mitigation can protect everyone in minutes, but the underlying model weakness may take weeks to train out. This article covers what triggers an emergency, who may declare one, a decision function you can test, the interim mitigations available to AI services, how to write an advisory without handing over an exploit, and the order in which to notify people. It works through a full example and ends with a checklist. For normal embargo practice, read Responsible Disclosure for LLM Vulnerabilities first.
What makes a disclosure an emergency
An emergency is declared on evidence, not on worry. Write the triggers down so the on-call engineer does not have to argue the case at 3 a.m.:
| Trigger | Typical evidence in AI systems |
|---|---|
| Exploitation in the wild | Logs show the injection pattern against real tenants; unusual tool calls or outbound requests matching the report |
| Details public | A proof of concept posted, a conference talk moved up, a jailbreak prompt trending |
| Embargo broken | A coordinating vendor ships its fix early, or a journalist asks about the flaw |
| Ongoing harm without exploitation | A safety bypass producing harmful output for ordinary users, or a data exposure through normal use |
| Third-party compromise | An upstream model API, connector or library you depend on discloses an exploited flaw |
Each trigger needs a named source of truth. Exploitation evidence comes from detection queries you can re-run. Public-details evidence is a captured URL with a timestamp. Rumours start an investigation, not an emergency. Over-declaring costs credibility, and it burns reporters who trusted your embargo.
Break-glass authority
The worst emergency disclosures are late because the person who could approve them was asleep. Delegate the authority in advance, in writing, and use a two-person rule: an incident lead and a security lead may declare an emergency and publish from pre-approved templates together. Legal and communications are consulted with a time box. If they do not respond within 60 minutes, the pair proceeds and records the attempt. A template approved in peacetime is what makes this safe, because the wording has already been reviewed.
| Role | Owns | Must not |
|---|---|---|
| Incident lead | Timeline, decision record, fan-out order | Write technical detail into public text |
| Security lead | Trigger evidence, mitigation, what to withhold | Delay mitigation for wording |
| Communications | Customer and public text, channels | Promise a fix date nobody owns |
| Legal and privacy | Regulatory triggers, contract notice duties | Block a mitigation advisory without a reason |
| Reporter liaison | Keeping the reporter informed and credited | Publish before the reporter is told |
The decision: fix first, advise now or acknowledge
The central decision is the mode: fix first and then publish, publish an advisory about a mitigation now, or acknowledge the issue with a holding statement. Encode it so the rule can be tested and reviewed in calm conditions:
from dataclasses import dataclass
@dataclass
class Case:
exploited_in_wild: bool # evidence of use against real users or tenants
details_public: bool # embargo leaked, PoC posted, third party published
user_action_reduces_risk: bool # customers can usefully act right now
mitigation_deployed: bool # a server-side interim control is live
fix_eta_hours: float # best estimate for a complete fix
def decide(c: Case) -> dict:
if not (c.exploited_in_wild or c.details_public):
return {"mode": "coordinated", "audiences": []} # normal embargo process
if (not c.details_public and c.fix_eta_hours <= 24
and not c.user_action_reduces_risk):
mode = "fix-first" # private exploitation, fix within a day
elif c.mitigation_deployed or c.user_action_reduces_risk:
mode = "mitigation-advisory" # publish now: what to do, not how it works
else:
mode = "holding-statement" # acknowledge, give the next update time
audiences = ["regulators-if-triggered", "affected-customers"]
if mode != "fix-first":
audiences += ["all-customers", "public"]
return {"mode": mode, "audiences": audiences}The rule encodes three judgments. First, if exploitation is private and a complete fix lands within a day with nothing for users to do, publishing first only advertises the flaw. Fix it, tell affected customers, and publish the full advisory afterwards. Second, once details are public, attackers already know, so silence only keeps defenders uninformed. Third, an advisory needs something to say. With no mitigation and no user action, a holding statement that acknowledges the issue and names the next update time is more honest than a vague advisory. Directly affected customers are always told, because they may have their own breach-notification duties.
Interim mitigations for AI services
Hosted AI services have more interim levers than shipped software, and most can be pulled in minutes. Build and test them before you need them. The stop mechanisms for agents are covered in Agent Kill Switch, in depth.
| Lever | Time to deploy | Blast radius | What customers notice |
|---|---|---|---|
| Feature flag off (a tool, connector, markdown rendering, browsing) | Minutes | One capability | A feature disappears |
| Egress allowlist tightened | Minutes | Outbound calls | Some integrations fail |
| Input or output filter hotfix | Under an hour | Matching requests | Some false refusals |
| Model version rollback | Under an hour if pre-staged | All traffic on that model | Quality or behaviour shift |
| Token and key revocation | Minutes to hours | Affected credentials | Re-authentication |
| Read-only or offline mode | Minutes | Whole product | Outage |
Pick the narrowest lever that stops the exploitation path, and confirm it with the detection query that triggered the case. A filter that blocks the published string but not trivial paraphrases is a mitigation in name only. Test it with variants before you tell customers they are protected.
Writing the emergency advisory
An emergency advisory answers four questions: what is affected, what we have done, what you should do, and when we will update you. It leaves out the payload, exact prompts, the bypass technique and any detail that lets a reader rebuild the exploit. The detail ladder in AI Vulnerability Disclosure Norms helps when deciding how much to say once a fix ships. A skeleton:
SECURITY ADVISORY [ID] version 1 published 2026-10-06 14:20 UTC
Status: mitigated, fix in progress Next update: by 2026-10-07 02:00 UTC
Affected: assistants with the web-fetch connector enabled, all regions.
Impact: crafted web content could cause the assistant to send conversation
data to an external address.
Evidence of exploitation: yes, limited; affected customers are being contacted.
What we did: disabled automatic image and link rendering from connector output
at 13:05 UTC; added outbound URL allowlisting.
What you should do: rotate any API keys pasted into conversations since
[first exploitation date]; review the audit-log query linked below.
What we will not publish yet: technical details, until a fix is complete.
Credit: reported by [name], with consent.Publish at a stable URL and version it. Never edit silently. Put times in UTC, and give the next update time even if the update will be "no change". Missing a promised update does more damage than the original advisory.
Staged fan-out
Fan-out is a sequence, not a broadcast. Mitigate first wherever you can, so the announcement does not open a window for attackers. Then tell the reporter and any vendors in the coordination group, so nobody learns about the case from your status page. Then contact affected customers directly, through the security contacts in their contracts, not a marketing list. Then post the public advisory and status page entry, and add an in-product notice for administrators.
Regulatory clocks run in parallel from the moment you become aware, not from your publication. Which regimes apply, and how to compute the deadlines, is covered in Regulatory Reporting for AI Incidents. The practical point here is that the emergency decision record doubles as the awareness timestamp regulators will ask for. Keep the security-contact registry current with a quarterly bounce test. A stale registry is the most common reason affected customers first hear about an emergency from the press.
Worked example: an exploited browsing connector
Worked example. On day 3 of a 90-day embargo, a researcher reports that a web page fetched by an agent's browsing connector can instruct it to render a markdown image whose URL carries conversation text, which leaks data when the client loads the image. The fix, which separates connector content from instructions and adds a rendering policy, is estimated at three days of work plus a model-side change over several weeks.
On day 9 a variant is posted publicly. Within the hour the detection query finds outbound image requests with encoded payloads from 14 tenants since day 1. Both triggers now hold. The security lead disables image rendering from connector output, a five-minute flag change, and re-runs the query to confirm that new hits stop. The case is then:
decide(Case(exploited_in_wild=True, details_public=True,
user_action_reduces_risk=True, # rotate keys pasted into chats
mitigation_deployed=True, fix_eta_hours=72))
# {'mode': 'mitigation-advisory',
# 'audiences': ['regulators-if-triggered', 'affected-customers',
# 'all-customers', 'public']}The reporter is told first, and agrees to credit and to holding technical detail until the fix. Privacy counsel assesses the 14 tenants for personal-data exposure, which starts its own clocks. Those tenants get direct calls with their own log extracts. The public advisory above goes out two hours after the trigger, with updates every 12 hours. The complete fix ships on day 12, and a version 2 advisory adds the technical write-up. For contrast, if the exploitation had been seen in logs with nothing public and a one-line server fix ready and nothing for users to do, the function returns fix-first, and only the affected tenants hear before the full advisory.
Closing the emergency
An emergency ends when the complete fix is deployed and verified, not when the mitigation goes live. Then publish the final advisory version with technical detail at the level the norms support. Request identifiers where a versioned artifact exists. Close the coordination with partner vendors. Hold a review of the procedure itself as well as the flaw: how long the declaration took, whether the templates fitted, which contacts bounced. Containment and the postmortem format belong to LLM Incident Response; feed the timing numbers from that review back into the runbook.
Failure modes
- Waiting for a fix that is days away. Attackers use the public details while customers have no instructions.
- An advisory that is an exploit. Quoting the working prompt so readers can test hands it to everyone else.
- A mitigation nobody verified. The filter blocks the posted string, and a paraphrase gets through the same day.
- Reporter blindsided. The researcher reads about their own finding on a status page and stops reporting to you.
- No authority at night. The only approver is unreachable, and the advisory waits eight hours.
- Silent edits. The advisory changes without a version bump, and customers cannot tell what is new.
Trade-offs
Speed or completeness. An early advisory will contain uncertainty, and a complete one will be late. Publish what you know, label the unknowns, and use the update schedule to correct course.
Narrow or broad mitigation. Disabling a whole connector is certain and disruptive. A targeted filter preserves the product and may leak. When exploitation is confirmed, take the broad lever first and narrow it once a tested fix exists.
Transparency or attacker uplift. Every sentence about the mechanism helps defenders reason and helps copycats. In the emergency phase, describe the impact and the actions, and keep the mechanism for the final version.
What to do next
- Write the trigger table with a named evidence source for each trigger.
- Delegate emergency authority in writing to a two-person pair, with a 60-minute consultation time box.
- Encode the decision rule, review it, and add test cases for each mode.
- Build and drill the mitigation levers, including a verification query for each one.
- Pre-approve advisory, holding-statement and customer-email templates.
- Bounce-test the security-contact registry every quarter.
- Run a tabletop on a leaked jailbreak and an exploited connector, and time every step.