Most writing about AI bug bounties is for hunters: what pays and how to write the report. This article is for the team that runs the program. It assumes you ship a product built on a language model, such as an assistant, an agent with tools or a retrieval system, and you are deciding whether and how to pay outside researchers to find its flaws. The researcher's side, including what major programs pay for, a sample scope file, severity and report quality, is covered in AI bug bounties. This article covers what that page does not: readiness, launch phases, legal safe harbour, a sandbox that makes findings reproducible, reward budgets, an intake pipeline that survives a flood of AI-generated reports, and the metrics that show whether the money is well spent.
The stakes are real. In January 2026, the curl project announced it would end its HackerOne bounty at the end of that month. The maintainers said a rising volume of low-quality, AI-generated reports was overloading the small security team, and that removing the payout removed the incentive. Press reports put the program's lifetime record at about 86,000 dollars paid and 78 confirmed vulnerabilities over six years. A program that did good work was shut down by intake cost. An LLM product attracts even more speculative reports, so plan intake first.
Readiness before rewards
A bounty pays strangers to find bugs. It only helps if you can confirm, fix and verify what they find. Check five things before you launch.
- Reproducible state. For every request, you log the model identifier, system prompt version, tool definitions, retrieved documents and sampling settings. Without this, the honest answer to most reports is that you cannot tell.
- A fix owner. A named team that fixes accepted findings, with an agreed time budget. Findings that sit unfixed are worse than no program, because researchers will eventually publish them.
- Basic hygiene done. If an internal red team or a penetration test has not looked at the product, it will find the easy bugs for a fraction of the cost. Paying bounties for them wastes the budget.
- Kill switches. You can disable a tool, a connector or a feature quickly when a critical report arrives.
- A disclosure policy. A public page that says how to report, what is authorised and what happens next. That is a vulnerability disclosure policy, or VDP, and you need one even with no bounty.
Launch in phases
Launch in three phases, and move to the next only when the metrics say you are ready.
| Phase | Who can report | Pays | Exit criteria |
|---|---|---|---|
| 1. VDP | Anyone | Thanks and credit | Intake handles volume; first response within your SLA |
| 2. Private bounty | A few dozen invited researchers | Yes | Valid-report rate is stable; fix backlog under control |
| 3. Public bounty | Anyone on the platform | Yes | Replay automation and triage staff can absorb many times the volume |
ISO/IEC 29147 describes how to receive and publish vulnerability information, and ISO/IEC 30111 describes how to handle it internally. They are useful checklists for phase 1. For safe harbour, adapt a published template such as those from the disclose.io project rather than drafting from scratch, and have counsel review it. In the United States, the Department of Justice said in May 2022 that it would not charge good-faith security research under the federal computer crime law. That policy does not bind private parties or other countries, so your own written authorisation still matters.
AI products need two extra clauses. The first says what researchers may do with harmful generated content, for example record it only as needed to prove the finding and never publish it. The second covers automation. Researchers will run thousands of prompts, so set rate limits for the sandbox and say that running scripted attacks inside them is allowed.
A sandbox that makes findings provable
The best investment for an LLM bounty is a dedicated tenant that runs the same model, prompts and tools as production, with fake data inside. Researchers attack it instead of real customers. Three features make findings provable.
- Canary documents. Seed each researcher's tenant, and a second victim tenant, with documents that contain unique random strings. If a string from the victim tenant appears in the attacker's output, or leaves through a tool call, cross-tenant exfiltration is proven with no argument about what the model meant.
- Stub tools. Email, payment and HTTP tools that record what they were asked to do, without doing it. A prompt injection that makes the agent send email shows up as a recorded call with the arguments.
- Trace identifiers. Every response carries a trace id that points to the full record: model version, prompt, retrieved chunks, tool calls. Reports must cite one.
Intake that survives a flood
The curl story shows the failure mode. Large language models make it cheap to produce a report that reads well and is wrong. A program for an LLM product receives more of these, because model behaviour is easy to describe confidently and hard to check. Do not try to detect AI-written text. Instead, require evidence a machine can verify, and verify it before a person spends time.
REQUIRED = ("trace_ids", "asset", "class", "impact_claim", "steps")
def intake(report, traces, replay, in_scope):
missing = [f for f in REQUIRED if not report.get(f)]
if missing:
return ("needs_info", f"missing fields: {missing}")
if not in_scope(report["asset"], report["class"]):
return ("out_of_scope", "see program scope")
recorded = [traces.get(t) for t in report["trace_ids"]]
if not all(recorded):
return ("needs_info", "trace ids not found in the bounty tenant")
# Re-run the attack from the recorded state, many times, on the pinned model.
result = replay(recorded[0], trials=20)
if result["canary_leaked"] or result["stub_tool_called"]:
return ("triage", f"reproduced {result['successes']}/20")
if result["successes"] == 0:
return ("needs_info", "could not reproduce from the recorded trace")
return ("triage_low", f"behaviour reproduced {result['successes']}/20, impact unproven")Three rules make this work. First, a report without a trace id is not rejected, but it is not triaged until one arrives. That alone removes most invented findings, because no attack took place. Second, the replay runs on the model version in the trace, not on whatever is live today, so a model update does not erase a real finding. Third, the replay checks impact with canaries and stub-tool logs, so a reply that merely sounds alarming does not count.
Add reputation as a secondary signal. Platforms track each researcher's ratio of valid to invalid reports. You can also cap the number of open reports per researcher, and require a track record to join a private phase. Use these to set queue priority, not to reject anyone outright.
Rewards and budget
Set rewards by impact and by the cost to fix, then estimate the budget. A simple model multiplies the expected count of valid findings in each severity by the reward for that severity.
REWARDS = {"critical": 20000, "high": 6000, "medium": 1500, "low": 300}
def expected_payout(valid_per_quarter, rewards=REWARDS, bonus_rate=0.10):
base = sum(valid_per_quarter[s] * rewards[s] for s in rewards)
return round(base * (1 + bonus_rate))
# Private phase guess from the pentest and red-team hit rates on the same product:
print(expected_payout({"critical": 0.5, "high": 3, "medium": 10, "low": 20})) # 53900The dollar values above are placeholders for the arithmetic. Set yours by reading the reward tables of programs that test similar products, and by asking what an incident would cost you. Then add the cost of running the program, which is often larger: platform fees, triage staff, the sandbox, and engineering time for fixes. For AI products, decide in writing three questions that otherwise cause disputes. Do you pay per root cause or per prompt? Per root cause, or you will pay for paraphrases. Do you pay at triage or at fix? Paying at triage keeps researchers engaged. And what is the payout for a finding that works 30 percent of the time? It should be close to full, because attackers retry.
Response targets and closure
Publish response targets and meet them; missed targets are the most common cause of researcher anger. Typical targets are a first response within three business days, triage within ten, and a bounty decision soon after triage. Disclosure timelines in the industry often use 90 days from report to public disclosure, a convention popularised by Google's Project Zero. Agree on an extension when a fix needs a model retrain.
Every accepted finding should also become a regression test that replays the attack on each model upgrade or prompt change. The sibling article shows that test. Here the operational point is ownership: the program owner checks that the test exists before the report is closed.
Metrics
| Metric | Why it matters | Warning sign |
|---|---|---|
| Valid-report rate | Signal-to-noise of intake | Falling below one in ten |
| Share auto-closed by replay | Triage time saved | Near zero: the replay is not being used |
| Time to first response | Researcher trust | Over the published target |
| Median time to fix, by severity | Real risk reduction | Critical fixes beyond weeks |
| Duplicate rate by root cause | Whether fixes hold | Same root cause reported after a fix |
| Cost per valid finding | Compare with pentest and red team | Rising while finding counts stay flat |
Worked example: the first quarter
Here is a worked first quarter. A company runs a support agent with email and refund tools. It publishes a VDP and builds the bounty tenant with canaries and stub tools. In the first month the VDP receives about forty reports. Most are system prompt disclosures or offensive outputs with no security impact, and the scope page routes them to product feedback. Two cite traces and reproduce. One is an indirect prompt injection in a support ticket that makes the agent call the refund tool, and the stub log shows the call.
The team fixes it by requiring user confirmation for refunds and by scoping the tool to the ticket's own customer. It adds a regression test, then opens a private bounty to twenty researchers. In the private phase, the replay harness auto-closes about half of the reports. The rest yield a cross-tenant retrieval bug, proven by a canary leaking from the victim tenant. Spend for the quarter comes in under the estimate. The team delays the public phase until the replay harness handles attachments, which a third of the reports needed.
Failure modes
- Launching before readiness: no traces, so every report becomes an argument.
- Testing in production: researchers touch real customer data, and a finding becomes an incident.
- Intake by hand only: the team drowns, as curl's did, and closes the program.
- Unclear content rules: jailbreak screenshots flood the queue because nobody said they are out of scope.
- Paying for prompts: rewards per prompt instead of per root cause drain the budget on paraphrases.
- Fixing with prompts alone: a system prompt patch regresses on the next model update.
Trade-offs
A bounty finds things internal teams miss, because many people with different habits look at the product. But it is noisy, slower to schedule and more expensive per finding than a pentest for known bug classes. Hackbots and automated scanners lower the cost of finding easy bugs, and the same tools raise the volume of low-quality reports, as described in autonomous hackbots. The usual balance is internal red teaming and pentests first, a VDP always, and a private bounty for the remaining hard classes: tool misuse, cross-tenant leaks and injection through data.
What to do next
- Publish a VDP with a safe-harbour clause based on a reviewed template.
- Log model id, prompt version, tools and retrieved chunks for every request, and expose a trace id.
- Build a bounty tenant with canary documents, a victim tenant and stub tools.
- Write the replay harness and require trace ids in reports.
- Estimate the budget from your red-team hit rates and decide per-root-cause and pay-at-triage rules.
- Run a private phase for one quarter and track the six metrics before going public.