Disclosure norms are the unwritten and written rules for what happens between the moment someone finds a flaw and the moment the world hears about it: who is told, how long they get, what gets published, and whether the finder is safe from legal threat. Software security spent three decades settling those rules. AI systems inherited them and then broke several of their assumptions at once.

This article is about the norms themselves, not the vendor-side machinery. The intake, routing, embargo and advisory process is covered in Responsible Disclosure for LLM Vulnerabilities, and paying for findings in LLM Bug Bounty. Here we look at where the norms came from, which assumptions AI flaws violate, the emerging answers for timelines, publication detail, cross-vendor notice, open weights and safe harbour, and a small planner in code that turns those answers into a concrete publication plan. A worked example runs the planner on a transferable jailbreak.

Where the norms came from

Early security culture swung between silent disclosure, where vendors could sit on reports for years, and full disclosure, which forced fixes but armed attackers before users could patch. Coordinated vulnerability disclosure (CVD) is the compromise: tell the vendor privately, give a fixed period, then publish whether or not a fix exists. The deadline is what makes the vendor move; the private period is what protects users.

The deadline numbers became norms by repetition. CERT/CC has long used a 45-day default. Google Project Zero popularised 90 days, later refined to 90+30: if the vendor fixes within 90 days, details wait a further 30 days so users can install the patch. In July 2025 Project Zero added a reporting-transparency trial, publishing within about a week that a report exists, naming the vendor, product and deadline but no technical detail, to shorten the gap between an upstream fix and downstream products shipping it. ISO/IEC 29147 and ISO/IEC 30111 codify the vendor side. None of these numbers is law; they are defaults that both sides can point to.

Which assumptions AI flaws break

Every CVD norm rests on assumptions that held for ordinary software. AI systems violate most of them.

AssumptionOrdinary softwareAI systems
A fix existsA patch removes the bugJailbreak classes are reduced, rarely eliminated; a retrain is slow and partial
One vendor owns itThe flaw is in one codebaseA technique often transfers across models from many vendors
Patching ends itUsers update and the window closesOpen-weight copies cannot be recalled or updated
Reproduction is binaryIt crashes or it does notSuccess is a rate across samples and model versions
Detail and harm are linked only through exploitationExploit code attacks the bugA working prompt may itself be the harmful payload
Testing is clearly boundedScope rules cover systemsTesting a model means producing the content its policy forbids

The last two rows are new in kind, not degree. For a buffer overflow, publishing a proof of concept helps defenders write detections. For a jailbreak that elicits serious uplift, the published prompt is the uplift. And a researcher testing whether a model will help with something dangerous must, by construction, try to make it do so.

Five questions every AI flaw disclosure has to answerFlaw foundreproduced, rate measured1. Who must know?vendor, other vendors, hosts2. When to publish?deadline, fix or no fix3. How much detail?class ... full prompt4. Can it be recalled?hosted vs open weights5. Was testing safe?rules of engagementPublication plandate, audience, detailExploited?shorten every clockClassic CVD answers these once per vendor. AI flaws often transfer across vendors,may have no patch at all, and can live forever in downloaded weights, so each answer changes.
Figure 1. The five questions a disclosure plan has to answer for an AI flaw. Active exploitation shortens every clock.

Timelines when there may be no fix

The deadline exists to force a fix. When no complete fix is possible, a pure deadline either publishes a live attack or extends forever. The emerging norm splits the clock in two. The disclosure clock still runs, usually 90 days, and at its end something is published. What is published depends on the mitigation state: if the vendor shipped a mitigation that measurably cuts the success rate, details can follow the usual 30-day adoption window; if not, publication is limited to the class and the measured rates, and full detail waits.

Report the rate, not an anecdote. A claim that a prompt works means little; a claim that it succeeded in 41 of 50 samples at temperature 0.7 on a named model version, and in 3 of 50 after the mitigation, gives the vendor a target and lets anyone check the fix later. Agree the measurement with the vendor at the start, because the deadline conversation becomes a conversation about whether that number moved.

Flaws that cross a boundary, such as an injection that exfiltrates another user's data, are ordinary vulnerabilities and keep the ordinary clock; the split applies to model-behaviour flaws.

How much to publish: the detail ladder

Think of publication as a ladder with five rungs, each adding detail: the class of attack; the mechanism, meaning why it works; metrics across models and versions; a redacted sample that shows the shape without the payload; and the full working prompt. Most of the defensive value sits on the first three rungs. Defenders need to know the class exists, why it works, and how widely, so they can build evaluations and filters. Attackers mostly need the top rung.

Two factors set the ceiling. The first is uplift: how much harm the detail itself enables. A jailbreak whose only effect is rude language can go to the top rung once mitigated; one that elicits meaningful help toward mass-casualty harm should stop at the mechanism permanently, whatever the fix status. The second is recall: whether every affected copy can be fixed. Many-shot jailbreaking, which Anthropic published in April 2024, is an instructive example. The paper explained the mechanism and the scaling of success with the number of shots in great detail, because that is what other developers needed, and Anthropic briefed other AI developers before publication.

Transferable flaws and cross-vendor notice

Transfer is the property that most separates AI flaws from software bugs. A technique that works against one chat model often works, at some rate, against others trained on similar data with similar methods. That turns a bilateral disclosure into a multi-party one, and multi-party disclosure needs someone to own the timeline.

The norm forming around this has three parts. Test transfer before reporting, on whatever models you can reach within their terms, and include the rates. Ask the first vendor to agree a shared date and to permit forwarding to other affected providers, or notify them yourself in parallel under the same date. And for broad, high-uplift classes, prefer a coordinator, because no single lab can be trusted by the others to set the clock. The position paper by Longpre and colleagues, In-House Evaluation Is Not Enough (2025), argues for exactly this: standardised AI flaw reports, broadly scoped flaw disclosure programmes with safe harbour, and shared infrastructure to route a report to every affected provider.

Open weights cannot be recalled

When weights are public, the publisher cannot patch every copy. A flaw in a hosted model ends when the host changes its model or filter; a flaw in released weights lives as long as someone hosts or runs those weights. Disclosure for open weights therefore targets the people who can act: the major hosting providers, inference platforms and downstream fine-tuners. The advisory carries mitigations they can apply, such as a new safety system prompt, an input classifier, a guard model, or a recommendation to move to a revised checkpoint, instead of a patch.

With no recall, the full-prompt rung stays closed, while the class, mechanism and an evaluation set are published so deployers can test their own stacks.

Safe harbour and rules of engagement

Software researchers rely on scope rules and, in the US, the Department of Justice's 2022 charging policy that good-faith security research should not be prosecuted under the Computer Fraud and Abuse Act. AI testing adds a contractual layer: terms of service forbid generating the very content a safety researcher must try to elicit, and account suspension is the usual enforcement. The 2024 open letter A Safe Harbor for AI Evaluation and Red Teaming, led by Longpre and colleagues, asked developers to commit to legal and technical safe harbour for good-faith research. Policies vary by vendor, so read the current text before testing rather than relying on a summary.

Useful safe-harbour language does four things. It defines good faith by conduct: test only your own accounts, stop at proof, do not use or share elicited harmful content beyond the report. It promises no legal action and no account termination for conduct inside those rules. It covers content-policy testing explicitly, not only security testing. And it gives a channel for model-behaviour reports that is not a bug bounty, since several vendors exclude content-only jailbreaks from paid scope and route them to feedback channels instead.

A disclosure planner in code

The planner below encodes the positions above as rules. It is a starting policy, not a standard: change the thresholds, but keep the separation between who is told, when, and how much.

from dataclasses import dataclass, field
from datetime import date, timedelta

LADDER = ["class", "mechanism", "metrics", "redacted_sample", "full_prompt"]

@dataclass
class Finding:
    uplift: str                # "none", "low", "high": harm the detail itself enables
    open_weights: bool         # affected model is downloadable
    exploited: bool            # evidence of use in the wild
    transfers_to: list = field(default_factory=list)   # other providers, with rates
    rate_before: float = 0.0   # measured success rate on the reported version
    rate_after: float | None = None   # after the vendor's mitigation, if any

def mitigated(f: Finding) -> bool:
    # "Fixed" for a behaviour flaw means a measured, large drop, not a closed ticket.
    return f.rate_after is not None and f.rate_after <= 0.2 * f.rate_before

def plan(f: Finding, reported: date) -> dict:
    audience = ["reporting vendor"] + [f"provider: {v}" for v in f.transfers_to]
    if f.open_weights:
        audience.append("major hosts and inference platforms")
    if len(f.transfers_to) >= 2 and f.uplift == "high":
        audience.append("coordinator")

    clock = 7 if f.exploited else 90
    publish_on = reported + timedelta(days=clock)

    if f.uplift == "high":
        ceiling = "mechanism"                    # never a working recipe
    elif not mitigated(f):
        ceiling = "metrics"
    elif f.open_weights or f.uplift == "low":
        ceiling = "redacted_sample"              # unpatched copies remain
    else:
        ceiling = "full_prompt"

    # On publish_on release up to "metrics"; anything above waits 30 more days.
    return {"notify": audience, "publish_on": publish_on, "ceiling": ceiling,
            "rest_after": publish_on + timedelta(days=30)}

mitigated() defines a fix by measured rate, and the ceiling is computed independently of the date, so a deadline can never by itself push a high-uplift recipe into public.

Worked example: a transferable jailbreak

A researcher finds that wrapping a request in a fictional tool-call transcript makes a hosted assistant ignore refusals for a category of moderately harmful content. Measured on the vendor's current version it succeeds in 37 of 50 samples. Within the vendor's terms, it also reproduces on two other providers' hosted models at 22 and 9 of 50, and on one popular open-weight model at 41 of 50. Uplift is judged low: the content is objectionable but freely available. There is no sign of use in the wild.

The planner says: notify the reporting vendor, both other providers and the major hosts of the open-weight model; publish on day 90; start at metrics. Sixty days in, the first vendor ships a classifier change and the researcher measures 3 of 50, so that vendor counts as mitigated. One other provider has not responded. On day 90 the researcher publishes the class, mechanism and rates for every model tested, including the unmitigated ones, which is the pressure the deadline exists to apply. The open-weight model cannot be recalled and is still at 41 of 50, so the post includes an evaluation set for deployers but no working prompt; a redacted sample follows 30 days later.

Failure modes

  • Anecdote reports. A single screenshot gives the vendor nothing to fix against and no way to show the fix worked. Report rates and versions.
  • Silent model swaps. The vendor ships a new model and closes the report, but the class still works. Re-measure on the new version before agreeing it is fixed.
  • Deadline publishes the payload. A team applying software norms literally posts the full prompt on day 90. Separate the detail ceiling from the date.
  • Single-vendor disclosure of a transferable flaw. One lab fixes it, the paper goes out, and every other provider learns from the press. Test transfer first.
  • Researcher banned for testing. Safe harbour that covers only security bugs, not content testing, chills exactly the reports vendors need.
  • Open-weight advisory with no audience. Publishing a notice on the model page reaches nobody who serves it. Name hosts and platforms explicitly.

Trade-offs

ChoiceGainCost
Fixed 90-day clock for behaviour flawsPredictable pressure on vendorsPublication may precede any real mitigation
Detail ceiling by upliftLimits harm from publicationLess reproducible science; harder independent verification
Cross-vendor notice before publicationProtects users of every affected modelMore parties, more leak risk, slower agreement
Coordinator for broad classesNeutral timeline ownerAdded latency and process; few coordinators handle AI today
Rate-based fix definitionFixes are verifiableMeasurement cost; sampling noise near thresholds

What to do next

  1. If you run an AI product, publish a policy that says how model-behaviour reports are handled, what clock applies and what safe harbour covers, including content testing.
  2. Adopt a standard flaw-report template with model version, sampling settings, sample count and success rate.
  3. Write your detail ladder and uplift tiers down now, and apply them to the next report rather than inventing them under pressure.
  4. As a researcher, measure transfer to other providers before reporting, and ask the first vendor to agree a shared date.
  5. For open-weight releases, keep a contact list of major hosts and platforms so an advisory reaches people who can act.
  6. Track every report in a register; see AI vulnerability databases for record formats, and feed confirmed classes into your red-team programme as regression tests.
Key takeaway: Coordinated disclosure assumed a single vendor, a real patch and a closed window. AI flaws often have none of these. Keep the deadline but split what it releases: publish class, mechanism and measured rates on time, hold working prompts until a measured mitigation exists, and never publish a high-uplift recipe. Test transfer and notify every affected provider, address open-weight advisories to hosts, define fixes by success rate, and give researchers safe harbour that covers content testing.