Responsible disclosure is the agreement, usually implicit, that whoever finds a security flaw tells the party who can fix it first, gives them a reasonable time to do so, and then makes the issue public so users can protect themselves. The industry term today is coordinated vulnerability disclosure (CVD), and the standards behind it are ISO/IEC 29147 for disclosure and ISO/IEC 30111 for the vendor's internal handling. LLM products strain the model in three ways: findings are often nondeterministic, the boundary between a security bug and a model behaviour problem is blurry, and a single flaw can sit in a model, a framework, a connector and a hundred deployments at once.
This article is about the process, from both sides: how a vendor builds intake, triage, clocks and advisories for an AI product, and how a researcher reports an LLM finding without causing harm. Paying for findings is a separate decision covered in LLM Bug Bounty; disclosure has to work whether or not you pay.
Is it a vulnerability? Routing AI reports
The first question in every AI report is whether it is a vulnerability at all. A useful test: does it let someone cross a boundary the system promised to enforce, such as another tenant's data, an action the user did not authorize, or a secret the operator configured? If yes, it goes through the security process with clocks and advisories. If it is a model saying something harmful or wrong with no boundary crossed, it is a safety or quality issue and goes to a different queue, with different timelines and usually no embargo.
| Report | Route | Why |
|---|---|---|
| Indirect prompt injection makes an agent email out another user's files | Security, high | Crosses tenant and authorization boundaries |
| System prompt revealed by asking politely | Security, low, unless it holds secrets | Prompts are not a secrecy boundary; credentials in them are |
| Tool call executes with the server's identity instead of the user's | Security, critical | Confused deputy, privilege escalation |
| Jailbreak produces disallowed content for the requesting user only | Safety queue | No other party harmed; fix is model or filter tuning |
| Training data extraction of personal information | Security and privacy | Disclosure of data the user had no right to |
| Model gives wrong medical advice | Quality or safety | Harmful, but not a vulnerability in the CVD sense |
Publish this routing in your policy. Reporters who know a jailbreak will go to the safety team, with no CVE and no bounty, stop arguing about severity and send the evidence you need.
The parties behind one fix
A report about your chatbot may really be about the model you call, the agent framework you embed, or a retrieval connector shared by many products. The vendor who receives it keeps ownership of the reporter relationship and the clock, even when the fix lives upstream. Forward the report to the upstream security contact with the reporter's permission, agree a single public date, and tell the reporter who else knows. If several vendors are involved, or upstream goes silent, bring in a coordinator such as CERT/CC or a national CSIRT. Their job is to keep one timeline across parties who do not trust each other.
Intake: security.txt and a policy reporters trust
Researchers look for a security contact before anything else. Publish one in machine-readable form at /.well-known/security.txt as defined by RFC 9116. Only Contact and Expires are required; the rest help reporters do the right thing.
Contact: mailto:security@example.com
Contact: https://example.com/security/report
Expires: 2027-10-01T00:00:00.000Z
Encryption: https://example.com/.well-known/pgp-key.txt
Policy: https://example.com/security/disclosure-policy
Acknowledgments: https://example.com/security/thanks
Preferred-Languages: en
Canonical: https://example.com/.well-known/security.txtThe policy page carries the commitments: what is in scope (name the AI features and their tool integrations), what testing is allowed (your own accounts only, no real users' data, no denial of service, rate limits), a safe harbour statement that you will not pursue legal action for good-faith research within those rules, your response targets, and your disclosure timeline. For AI products add one line many policies forget: what a reporter should do if a test surfaces another person's data. The right answer is stop, do not keep a copy, and report immediately.
Ask for evidence that survives nondeterminism: model and version if visible, full transcripts, the exact documents or web pages used for indirect injection, timestamps so you can find server-side traces, and a success rate over repeated attempts rather than one screenshot. How to reproduce such findings reliably is covered in AI Bug Bounties.
Cases with clocks
Treat every report as a case with explicit states and deadlines, so nothing depends on someone remembering. The sketch below is the core of a tracker; the timings are examples to adapt, not a standard.
from datetime import datetime, timedelta
from enum import Enum
class State(Enum):
RECEIVED = 1; TRIAGED = 2; CONFIRMED = 3; FIX_READY = 4; DEPLOYED = 5; PUBLISHED = 6; CLOSED = 7
ACK_WITHIN = timedelta(days=2)
TRIAGE_WITHIN = timedelta(days=7)
DISCLOSE_AFTER = timedelta(days=90)
DISCLOSE_IF_EXPLOITED = timedelta(days=7)
class Case:
def __init__(self, case_id, received: datetime, route: str):
self.id, self.received, self.route = case_id, received, route # "security" or "safety"
self.state, self.exploited, self.history = State.RECEIVED, False, []
self.parties = set() # upstream vendors notified
def disclosure_date(self):
return self.received + (DISCLOSE_IF_EXPLOITED if self.exploited else DISCLOSE_AFTER)
def move(self, new: State, note: str, now: datetime):
self.history.append((now, self.state, new, note)) # auditable trail for the reporter and regulators
self.state = new
def overdue(self, now: datetime):
alerts = []
if self.state == State.RECEIVED and now - self.received > ACK_WITHIN:
alerts.append("acknowledge the reporter")
if self.state in (State.RECEIVED, State.TRIAGED) and now - self.received > TRIAGE_WITHIN:
alerts.append("triage decision overdue")
if self.state.value < State.DEPLOYED.value and now > self.disclosure_date() - timedelta(days=14):
alerts.append("disclosure date within 14 days and no fix deployed")
return alertsRun overdue daily and page the owner, not a shared inbox. The history list matters more than it looks: when a reporter disputes the timeline, or a regulator asks when you knew, the case log is your answer.
Timelines, embargoes and regulatory clocks
There is no single legal deadline for disclosure, but norms have converged. Google Project Zero gives vendors 90 days, plus 30 days for users to patch before full technical details are published, and 7 days when a bug is being exploited in the wild. In a trial announced in July 2025 it also began publishing, within about a week of reporting, the fact that a report exists (vendor, product and deadline, without technical detail). CERT/CC has long used 45 days as its default. Most vendor policies ask for 90 days and promise to publish sooner if fixed.
Regulatory clocks now sit on top of those norms for some products. Since 11 September 2026 the EU Cyber Resilience Act requires manufacturers of products with digital elements to report actively exploited vulnerabilities through ENISA's single reporting platform: an early warning within 24 hours of becoming aware, a notification within 72 hours, and a final report within 14 days after a fix or mitigation is available. Software you ship (an SDK, a desktop assistant, an on-device model runtime) is the clear case; a purely hosted service is largely outside the Act, so check with counsel where your product falls. Separately, if a vulnerability was exploited against personal data, privacy breach rules may start their own clock; that hand-off belongs in LLM Incident Response.
Hosted AI changes what fixed means. When the fix is a server-side change to a prompt template, filter or tool permission, every user is protected the moment you deploy, so the case for a long embargo is weak. When the fix is a library release or a model version that customers must adopt, the clock has to include their upgrade time.
Advisories and identifiers
Publish an advisory for every confirmed security issue, even when no customer action is needed. It should say what was affected, which versions or dates, the impact in plain words, what was changed, what customers should do, and who reported it (with their consent). Machine-readable advisories in CSAF 2.0 let customers ingest them automatically; generate the JSON from the case record and validate it against the published schema.
Request a CVE when there is a versioned artifact customers run: a library, an agent framework, a model serving runtime. Practice varies for fully hosted fixes; some CNAs assign them, many vendors publish a notes entry instead. The CVE program's own funding and governance were contested in 2025, so confirm its current status rather than assuming it. How CVE, GHSA and OSV treat ML code and how classes such as CWE-1427 map to prompt injection are covered in AI Vulnerability Databases.
Disclosing as a researcher
If you find the flaw, the same norms bind you. Test on your own accounts and data. When an injection or extraction works, prove it with a canary you planted, not a real user's record; if real data appears anyway, stop, keep nothing beyond what proves the issue, and say so in the report. Do not pivot from a working injection into further systems to show impact; describe what the access would allow. Send the report to the published contact, encrypted if offered, with transcripts, inputs, versions, timestamps and a success rate. Agree the disclosure date in writing, and if the vendor stops answering, go to a coordinator rather than publishing early.
Worked example: an injection in a shared connector
A researcher finds that a document connector used by an enterprise assistant passes the hidden text of shared documents into the model's context. A shared document containing instructions causes the assistant to send summaries of the victim's other files to an external address through its email tool. It works in 7 of 10 attempts with the researcher's own two accounts. The mechanics of this class are in Prompt Injection via RAG Retrieval.
Day 0, the assistant vendor receives the report and acknowledges within a day. Day 3, triage routes it as security, high: it crosses a user boundary and uses a tool without consent. The connector is an open-source library used by other products, so the vendor asks the reporter's permission and notifies the library's maintainers; disclosure is set for day 90. Day 10, the vendor ships a server-side mitigation: outbound email from the assistant now requires user confirmation, and hidden text is stripped before retrieval. Its own users are protected. Day 41, the library releases a fixed version and requests a CVE. Day 60, with downstream deployers given time to upgrade, all parties publish together: the library advisory with the CVE, the vendor's advisory describing its mitigation, and credit to the researcher. No case went silent, and nobody learned about the flaw from a blog post.
Failure modes
- No contact. Without security.txt the report goes to support, is closed as a feature request, and reappears as a public post.
- Treating jailbreaks as vulnerabilities, or vice versa. Either floods the security queue or misses a real boundary crossing.
- Fixing silently. A server-side fix without an advisory leaves customers unable to judge past exposure or check their logs.
- Losing the upstream thread. The vendor patches its own app and never tells the framework, so every other deployer stays exposed.
- Unclear safe harbour. Researchers who fear legal action publish anonymously instead of reporting.
Trade-offs
Short embargoes protect users who are being attacked and pressure slow vendors, but they can publish details before downstream deployers upgrade. Long embargoes give time for coordination but leave users unaware. Broad scope invites more reports and more noise; narrow scope keeps triage cheap but pushes researchers toward public disclosure of whatever you excluded. Publishing advisories for hosted fixes costs some reputation in the short term and buys trust from customers who need to know what they were exposed to.
What to do next
- Publish /.well-known/security.txt with a monitored contact and an Expires date, and diary its renewal.
- Write a disclosure policy that names AI features in scope, testing rules, safe harbour and what to do on finding real user data.
- Publish the routing between security and safety reports, with examples like the table above.
- Run every report as a case with states, owners and dated deadlines, and alert on overdue cases daily.
- List the upstream vendors your AI features depend on and record each one's security contact.
- Decide whether your products fall under the Cyber Resilience Act and, if so, rehearse a 24-hour early warning.
- Template your advisory and CSAF output from the case record so publication takes minutes, not days.