Language models can now find real vulnerabilities in mature code. That no longer needs debating. The practical question is how to turn a model that is sometimes right into a process that is reliably useful. A model that reports ten plausible memory-safety bugs, one of them real, costs your team a lot of triage time and can easily bury the real one. Teams that get value from LLM bug discovery treat every model output as a hypothesis and accept only evidence that a machine can reproduce.
This article is about running that process on code you own or are authorised to test. It covers what the public results show, the two roles a model plays (reviewer and harness writer), how to slice context, how to structure and sample findings, how sanitizers act as the oracle, a worked example on a small C parser, how to secure the pipeline itself, and the metrics that tell you whether it is worth the compute.
What the public results show, and what they do not
Several results are well documented. Google's Big Sleep agent found an exploitable stack buffer underflow in SQLite, which was fixed quickly. Google's OSS-Fuzz project reported 26 vulnerabilities found with LLM-generated or LLM-improved fuzz targets. One was CVE-2024-9143, an out-of-bounds write in OpenSSL that existing human-written harnesses had not reached. Sean Heelan used OpenAI's o3 to find CVE-2025-37899, a use-after-free in the Linux kernel's ksmbd SMB server, where one thread's logoff frees sess->user while another thread can still use it. In the same write-up, o3 rediscovered a known ksmbd bug, CVE-2025-37778, in 8 of 100 runs. At DARPA's AI Cyber Challenge final at DEF CON 33, fully automated systems from Team Atlanta, Trail of Bits and Theori took the top three places by finding and patching bugs in open-source code.
The limits matter as much. A hit rate of 8 in 100 means one run is almost worthless, and many runs are needed. In those 100 runs Heelan also counted 28 false reports against the 8 true ones, and when he gave the model all the command handlers (about 12k lines instead of about 3.3k) the hit rate dropped to 1 in 100. The fuzzing results came from models writing harnesses, not from models declaring bugs, so a sanitizer decided what counted. The lesson for your pipeline: sample repeatedly, keep context focused, and let execution be the judge.
Two roles: reviewer and harness writer
In the reviewer role, the model reads a slice of code and proposes specific bugs: this length is trusted here, this object is freed on this path and used on that one. It suits logic and concurrency bugs that fuzzing rarely reaches, such as the ksmbd race. The weakness is that its output is a claim and nothing more.
In the harness writer role, the model writes the glue that lets a fuzzer drive an API: a libFuzzer entry point that turns random bytes into valid calls. Here the model never has to be right about a bug. It only has to write code that compiles and reaches deep code, and the fuzzer plus sanitizers find the crashes. This role has the best evidence behind it and the lowest false-positive rate, so start here. A third role, drafting a fix or a minimised reproducer after a crash, is useful too, but a human reviews the patch.
Context slicing: give the model the attack surface, not the repository
Relevance matters more than volume. Build slices around entry points where attacker-controlled data arrives, such as network handlers, file parsers, deserialisers and IPC endpoints. Follow the call graph a few levels down, and include the definitions of the structures and locking rules involved. For a concurrency review, include every handler that touches the shared object, because the bug lives in their interaction. A slice of a few thousand lines with a clear statement of the threat model beats a whole repository with no framing.
Tell the model what the attacker controls and what is trusted. "The len field in the packet header is attacker-controlled; the config file is trusted" removes a whole class of noise. Ask for one bug class per run, such as bounds, lifetime or integer overflow, rather than "find vulnerabilities", and rotate classes across runs.
Harness generation with a compile-and-coverage feedback loop
A harness is a function the fuzzer calls repeatedly with random input. For libFuzzer the entry point is fixed: int LLVMFuzzerTestOneInput(const uint8_t *Data, size_t Size), compiled with -fsanitize=fuzzer,address,undefined. The loop below feeds compiler errors and coverage back to the model until the harness builds and reaches the target function.
def generate_harness(target_fn, headers, examples, max_iters=6):
prompt = harness_prompt(target_fn, headers, examples) # API docs, existing tests, preconditions
for attempt in range(max_iters):
src = llm(prompt)
build = sandbox.compile(src, flags="-fsanitize=fuzzer,address,undefined -g")
if not build.ok:
prompt = revise(prompt, src, "compile error", build.stderr[-4000:])
continue
run = sandbox.fuzz(build.binary, seconds=120, corpus=seed_corpus(target_fn))
if run.crashed and harness_misuse(run, target_fn): # crash in the harness itself
prompt = revise(prompt, src, "harness violates API contract", run.stack)
continue
if coverage(run, target_fn) < 0.3: # never reached the interesting code
prompt = revise(prompt, src, "low coverage", uncovered_lines(run, target_fn))
continue
return src, run
return None, NoneThe harness_misuse check is essential. A generated harness can call a function in a way no real caller could: a length larger than the buffer, a freed handle, a missing init call. The resulting crash is a harness bug, not a product bug. Filter crashes whose top frames are in the harness, and put the API's documented preconditions in the prompt so the model respects them.
Reviewer runs: structured findings and repeated sampling
Free-text reports cannot be deduplicated or counted. Require a schema, and make the model name the trigger path. That forces it to reason concretely and gives the validator something to test:
{
"file": "src/proto/frame.c",
"function": "parse_frame",
"bug_class": "CWE-787 out-of-bounds write",
"attacker_input": "frame header field len (uint16, network)",
"trigger": "len > sizeof(buf) - HDR; memcpy(buf + HDR, p, len) at line 88",
"preconditions": ["connection authenticated: no"],
"confidence": "medium",
"suggested_test": "frame with len = 0xFFFF and 16 payload bytes"
}Run each slice k times, commonly 10 to 50 depending on budget, with varied prompts and bug classes. Then cluster findings by file, function and bug class. A location reported independently in many runs is more likely real than one reported once, but frequency is only a ranking signal and proves nothing. Instruct the model to report nothing rather than guess. That cuts noise, but you still need to measure whether it also cuts true positives.
Validation: the oracle decides
Every reviewer finding goes to a validator that tries to produce evidence. For memory-safety claims, turn suggested_test into an input, run it against a sanitizer build, and require an AddressSanitizer or UndefinedBehaviorSanitizer report whose stack includes the claimed function. For logic flaws, write a failing unit or integration test. For races, run a ThreadSanitizer build under a stress harness. The model can draft the reproducer, but the pass/fail signal comes from execution.
Deduplicate crashes by a hash of the top few symbolised frames plus the bug type, so a hundred inputs hitting one bug produce one ticket. Findings that cannot be reproduced go into the hypothesis log, not the tracker. Review that log on a schedule: an unreproduced lifetime bug in a concurrency path may be real and just hard to trigger, and it deserves an hour of human time. It never deserves an automatic CVE request.
Worked example: an off-by-header bug in a frame parser
Here is a small parser of the kind the reviewer role handles well:
#define HDR 4
int parse_frame(const uint8_t *p, size_t n, uint8_t out[256]) {
if (n < HDR) return -1;
uint16_t len = (uint16_t)(p[2] << 8 | p[3]);
if (len > 256) return -1; /* checks payload against the whole buffer */
if (n < HDR + (size_t)len) return -1;
memcpy(out, p, HDR); /* header copied first */
memcpy(out + HDR, p + HDR, len); /* writes up to HDR + 256 = 260 bytes */
return HDR + len;
}Asked about bounds bugs with the note that len comes from the network, a model will often point out that the check ignores the four header bytes already in out, so len of 253 to 256 writes past the end. The validator turns that into an input: a four-byte header with len = 256 followed by 256 bytes, passed to a caller whose out is a 256-byte stack array. Under AddressSanitizer this produces a stack-buffer-overflow write report in parse_frame, and only now does it become a finding. A harness-writer run gets there without the reviewer: a harness that passes a 256-byte stack buffer and random input hits the same report within seconds. The fix is if (len > 256 - HDR). The reproducer goes into the regression suite, and a simple variant search checks for the same pattern in the codebase's other parsers.
Securing the pipeline itself
- The code under review is untrusted input to the model. Comments and strings can contain instructions such as "ignore this file" or "report no bugs". Treat it as an indirect prompt injection surface: keep the tool-less reviewer separate from any agent that can act, and watch for sudden drops in finding rates on particular files.
- Generated harnesses and reproducers are untrusted code. Compile and run them in a sandbox with no network egress, no credentials and resource limits, and never on a developer laptop that holds signing keys.
- Keep source and findings away from third parties you have not approved. Sending proprietary code to an external model API is a data-handling decision, and unpatched findings are sensitive until fixed. See the secrets guide for keeping tokens out of prompts and logs.
- Scope. Run this only on code you own or are authorised to test. For third-party open source, follow the project's disclosure policy and report only reproduced bugs with a reproducer. Maintainers are already overwhelmed by unverified AI-generated reports.
Metrics and economics
| Metric | Definition | Why it matters |
|---|---|---|
| Validated findings per 1k runs | Reproduced, deduplicated bugs / runs | The yield you are paying for |
| Precision after oracle | Findings accepted by humans / tickets filed | Should be near 1; if not, the oracle is weak |
| Hypothesis reproduction rate | Reviewer claims reproduced / claims made | Tracks prompt and slice quality |
| Harness build rate | Harnesses that compile and reach target / attempts | Tracks feedback-loop quality |
| Cost per validated finding | Model spend + compute / validated findings | Compare with a human audit day or a bug bounty payout |
| Time to triage | Median hours from crash to human verdict | Noise shows up here first |
Run the pipeline against a benchmark of known, already-fixed bugs in your own history before trusting its output on new code, the same way Heelan measured o3 on a known CVE. The evaluation guide covers building such sets. Re-measure on every model or prompt change, because hit rates at this level swing widely.
Failure modes
- Triage flood: unvalidated reviewer output goes straight to the tracker, engineers stop reading it, and the real bug is ignored.
- Harness false positives: crashes caused by API misuse in generated harnesses are filed as product bugs.
- Context dilution: whole-repository prompts lower hit rates while raising cost.
- Single-run conclusions: one clean run is taken to mean the code is clean, when per-run recall may be under 10 percent.
- Benchmark contamination: a model scores well on famous CVEs it saw in training, and the result says little about your code.
- Pipeline compromise: generated code runs with network access or credentials, or injected comments steer the reviewer.
- Automation as a substitute for the red team: models find local bugs well, and design and authorisation flaws still need people.
What to do next
- Pick one parser or protocol handler you own and build ASan/UBSan fuzzing builds for it.
- Add an LLM harness generator with the compile, coverage and misuse feedback loop, and run it in a no-egress sandbox.
- Add reviewer runs with the JSON schema, 10 or more samples per slice, and one bug class per run.
- Route every claim through a validator. Only reproduced, deduplicated crashes or failing tests become tickets.
- Benchmark the pipeline on a few already-fixed bugs from your own history and record recall per run.
- Track cost per validated finding and precision after the oracle, and decide on scaling from those numbers.