Coordinated disclosure assumes you have time. A reporter tells you privately, you fix the flaw, and you publish together when users are protected. Sometimes the time is gone. The flaw is being exploited against your users, or its details are already public, and silence now protects attackers more than users. Emergency disclosure is the procedure for that case: telling people what they need to protect themselves before a complete fix exists.

AI systems hit this case more often than traditional software. A jailbreak or injection string spreads on social media within hours. An exfiltration trick in one agent framework transfers to others. A server-side mitigation can protect everyone in minutes, but the underlying model weakness may take weeks to train out. This article covers what triggers an emergency, who may declare one, a decision function you can test, the interim mitigations available to AI services, how to write an advisory without handing over an exploit, and the order in which to notify people. It works through a full example and ends with a checklist. For normal embargo practice, read Responsible Disclosure for LLM Vulnerabilities first.

What makes a disclosure an emergency

An emergency is declared on evidence, not on worry. Write the triggers down so the on-call engineer does not have to argue the case at 3 a.m.:

TriggerTypical evidence in AI systems
Exploitation in the wildLogs show the injection pattern against real tenants; unusual tool calls or outbound requests matching the report
Details publicA proof of concept posted, a conference talk moved up, a jailbreak prompt trending
Embargo brokenA coordinating vendor ships its fix early, or a journalist asks about the flaw
Ongoing harm without exploitationA safety bypass producing harmful output for ordinary users, or a data exposure through normal use
Third-party compromiseAn upstream model API, connector or library you depend on discloses an exploited flaw

Each trigger needs a named source of truth. Exploitation evidence comes from detection queries you can re-run. Public-details evidence is a captured URL with a timestamp. Rumours start an investigation, not an emergency. Over-declaring costs credibility, and it burns reporters who trusted your embargo.

Break-glass authority

The worst emergency disclosures are late because the person who could approve them was asleep. Delegate the authority in advance, in writing, and use a two-person rule: an incident lead and a security lead may declare an emergency and publish from pre-approved templates together. Legal and communications are consulted with a time box. If they do not respond within 60 minutes, the pair proceeds and records the attempt. A template approved in peacetime is what makes this safe, because the wording has already been reviewed.

RoleOwnsMust not
Incident leadTimeline, decision record, fan-out orderWrite technical detail into public text
Security leadTrigger evidence, mitigation, what to withholdDelay mitigation for wording
CommunicationsCustomer and public text, channelsPromise a fix date nobody owns
Legal and privacyRegulatory triggers, contract notice dutiesBlock a mitigation advisory without a reason
Reporter liaisonKeeping the reporter informed and creditedPublish before the reporter is told

The decision: fix first, advise now or acknowledge

The central decision is the mode: fix first and then publish, publish an advisory about a mitigation now, or acknowledge the issue with a holding statement. Encode it so the rule can be tested and reviewed in calm conditions:

from dataclasses import dataclass

@dataclass
class Case:
    exploited_in_wild: bool         # evidence of use against real users or tenants
    details_public: bool            # embargo leaked, PoC posted, third party published
    user_action_reduces_risk: bool  # customers can usefully act right now
    mitigation_deployed: bool       # a server-side interim control is live
    fix_eta_hours: float            # best estimate for a complete fix

def decide(c: Case) -> dict:
    if not (c.exploited_in_wild or c.details_public):
        return {"mode": "coordinated", "audiences": []}      # normal embargo process
    if (not c.details_public and c.fix_eta_hours <= 24
            and not c.user_action_reduces_risk):
        mode = "fix-first"           # private exploitation, fix within a day
    elif c.mitigation_deployed or c.user_action_reduces_risk:
        mode = "mitigation-advisory" # publish now: what to do, not how it works
    else:
        mode = "holding-statement"   # acknowledge, give the next update time
    audiences = ["regulators-if-triggered", "affected-customers"]
    if mode != "fix-first":
        audiences += ["all-customers", "public"]
    return {"mode": mode, "audiences": audiences}

The rule encodes three judgments. First, if exploitation is private and a complete fix lands within a day with nothing for users to do, publishing first only advertises the flaw. Fix it, tell affected customers, and publish the full advisory afterwards. Second, once details are public, attackers already know, so silence only keeps defenders uninformed. Third, an advisory needs something to say. With no mitigation and no user action, a holding statement that acknowledges the issue and names the next update time is more honest than a vague advisory. Directly affected customers are always told, because they may have their own breach-notification duties.

Interim mitigations for AI services

Hosted AI services have more interim levers than shipped software, and most can be pulled in minutes. Build and test them before you need them. The stop mechanisms for agents are covered in Agent Kill Switch, in depth.

LeverTime to deployBlast radiusWhat customers notice
Feature flag off (a tool, connector, markdown rendering, browsing)MinutesOne capabilityA feature disappears
Egress allowlist tightenedMinutesOutbound callsSome integrations fail
Input or output filter hotfixUnder an hourMatching requestsSome false refusals
Model version rollbackUnder an hour if pre-stagedAll traffic on that modelQuality or behaviour shift
Token and key revocationMinutes to hoursAffected credentialsRe-authentication
Read-only or offline modeMinutesWhole productOutage

Pick the narrowest lever that stops the exploitation path, and confirm it with the detection query that triggered the case. A filter that blocks the published string but not trivial paraphrases is a mitigation in name only. Test it with variants before you tell customers they are protected.

Writing the emergency advisory

An emergency advisory answers four questions: what is affected, what we have done, what you should do, and when we will update you. It leaves out the payload, exact prompts, the bypass technique and any detail that lets a reader rebuild the exploit. The detail ladder in AI Vulnerability Disclosure Norms helps when deciding how much to say once a fix ships. A skeleton:

SECURITY ADVISORY  [ID]  version 1   published 2026-10-06 14:20 UTC
Status: mitigated, fix in progress           Next update: by 2026-10-07 02:00 UTC

Affected: assistants with the web-fetch connector enabled, all regions.
Impact: crafted web content could cause the assistant to send conversation
        data to an external address.
Evidence of exploitation: yes, limited; affected customers are being contacted.
What we did: disabled automatic image and link rendering from connector output
             at 13:05 UTC; added outbound URL allowlisting.
What you should do: rotate any API keys pasted into conversations since
                    [first exploitation date]; review the audit-log query linked below.
What we will not publish yet: technical details, until a fix is complete.
Credit: reported by [name], with consent.

Publish at a stable URL and version it. Never edit silently. Put times in UTC, and give the next update time even if the update will be "no change". Missing a promised update does more damage than the original advisory.

Staged fan-out

Emergency fan-out: the first 24 hours after the triggerT+0T+1hT+2hT+4hT+8hT+24hTrigger confirmedcase lead namedDecisionmode and audiencesMitigation liveflag, filter, revokeAffected tolddirect contactPublic advisorystatus pageReporter and partner vendorstold before publicationRegulatory clocksrun in parallel from awarenessScheduled updates until the fix ships
The order matters: mitigation before announcement, affected parties before the public, and the reporter before anyone reads about their finding.

Fan-out is a sequence, not a broadcast. Mitigate first wherever you can, so the announcement does not open a window for attackers. Then tell the reporter and any vendors in the coordination group, so nobody learns about the case from your status page. Then contact affected customers directly, through the security contacts in their contracts, not a marketing list. Then post the public advisory and status page entry, and add an in-product notice for administrators.

Regulatory clocks run in parallel from the moment you become aware, not from your publication. Which regimes apply, and how to compute the deadlines, is covered in Regulatory Reporting for AI Incidents. The practical point here is that the emergency decision record doubles as the awareness timestamp regulators will ask for. Keep the security-contact registry current with a quarterly bounce test. A stale registry is the most common reason affected customers first hear about an emergency from the press.

Worked example: an exploited browsing connector

Worked example. On day 3 of a 90-day embargo, a researcher reports that a web page fetched by an agent's browsing connector can instruct it to render a markdown image whose URL carries conversation text, which leaks data when the client loads the image. The fix, which separates connector content from instructions and adds a rendering policy, is estimated at three days of work plus a model-side change over several weeks.

On day 9 a variant is posted publicly. Within the hour the detection query finds outbound image requests with encoded payloads from 14 tenants since day 1. Both triggers now hold. The security lead disables image rendering from connector output, a five-minute flag change, and re-runs the query to confirm that new hits stop. The case is then:

decide(Case(exploited_in_wild=True, details_public=True,
            user_action_reduces_risk=True,   # rotate keys pasted into chats
            mitigation_deployed=True, fix_eta_hours=72))
# {'mode': 'mitigation-advisory',
#  'audiences': ['regulators-if-triggered', 'affected-customers',
#                'all-customers', 'public']}

The reporter is told first, and agrees to credit and to holding technical detail until the fix. Privacy counsel assesses the 14 tenants for personal-data exposure, which starts its own clocks. Those tenants get direct calls with their own log extracts. The public advisory above goes out two hours after the trigger, with updates every 12 hours. The complete fix ships on day 12, and a version 2 advisory adds the technical write-up. For contrast, if the exploitation had been seen in logs with nothing public and a one-line server fix ready and nothing for users to do, the function returns fix-first, and only the affected tenants hear before the full advisory.

Closing the emergency

An emergency ends when the complete fix is deployed and verified, not when the mitigation goes live. Then publish the final advisory version with technical detail at the level the norms support. Request identifiers where a versioned artifact exists. Close the coordination with partner vendors. Hold a review of the procedure itself as well as the flaw: how long the declaration took, whether the templates fitted, which contacts bounced. Containment and the postmortem format belong to LLM Incident Response; feed the timing numbers from that review back into the runbook.

Failure modes

  • Waiting for a fix that is days away. Attackers use the public details while customers have no instructions.
  • An advisory that is an exploit. Quoting the working prompt so readers can test hands it to everyone else.
  • A mitigation nobody verified. The filter blocks the posted string, and a paraphrase gets through the same day.
  • Reporter blindsided. The researcher reads about their own finding on a status page and stops reporting to you.
  • No authority at night. The only approver is unreachable, and the advisory waits eight hours.
  • Silent edits. The advisory changes without a version bump, and customers cannot tell what is new.

Trade-offs

Speed or completeness. An early advisory will contain uncertainty, and a complete one will be late. Publish what you know, label the unknowns, and use the update schedule to correct course.

Narrow or broad mitigation. Disabling a whole connector is certain and disruptive. A targeted filter preserves the product and may leak. When exploitation is confirmed, take the broad lever first and narrow it once a tested fix exists.

Transparency or attacker uplift. Every sentence about the mechanism helps defenders reason and helps copycats. In the emergency phase, describe the impact and the actions, and keep the mechanism for the final version.

What to do next

  1. Write the trigger table with a named evidence source for each trigger.
  2. Delegate emergency authority in writing to a two-person pair, with a 60-minute consultation time box.
  3. Encode the decision rule, review it, and add test cases for each mode.
  4. Build and drill the mitigation levers, including a verification query for each one.
  5. Pre-approve advisory, holding-statement and customer-email templates.
  6. Bounce-test the security-contact registry every quarter.
  7. Run a tabletop on a leaked jailbreak and an exploited connector, and time every step.
Key takeaway: Emergency disclosure applies when a flaw is exploited or public before a fix exists. Declare it on written triggers and evidence, with authority delegated in advance. Decide the mode with a tested rule, pull the narrowest verified mitigation, and publish an advisory that gives actions and impact without the exploit. Then notify in order: reporter and partners, affected customers, everyone, with regulatory clocks running in parallel from awareness.