Most organisations that build on large language models have run a red team exercise: a few people spend two weeks trying to make the model misbehave, write a report and move on. A red team program is different. It is a standing function that decides what to test from a risk register, measures results in a way that survives sampling noise, has the authority to hold a release, and leaves an evidence trail an auditor or regulator can follow. The difference matters because model behaviour is stochastic and keeps changing: a new system prompt, a model upgrade or a new tool can undo last quarter's fixes without anyone noticing.
This article is about running that program. The mechanics of an individual engagement (rules of engagement, the attack corpus, triage and independence) are covered in LLM Red Team Architecture, so they appear here only where the program depends on them. You will learn the maturity stages a program moves through, how to derive a coverage matrix from risk, how to compute attack success rates with honest confidence intervals, how to wire those numbers into a release gate, and how to turn the output into evidence. A worked example walks a customer support agent through two gate runs.
From exercise to program: maturity stages
Programs grow in recognisable stages. Knowing which stage you are in tells you what to build next and stops you buying an automation platform before you have anything worth automating.
| Stage | What exists | What is missing | Next investment |
|---|---|---|---|
| 0. Ad hoc | Occasional exercises by volunteers | Scope, repeatability, authority | A written charter and one owner |
| 1. Pre-launch review | A red team pass before major launches | Coverage of small changes, metrics | A coverage matrix tied to the risk register |
| 2. Release gate | Automated suites run on every model or prompt change, with thresholds | Novel attack discovery | Dedicated manual research time |
| 3. Continuous and adaptive | Gate plus scheduled manual campaigns plus production signal | Usually nothing structural; tune cost | External exercises and bounty intake |
The jump from stage 1 to stage 2 is the one that changes outcomes. Before it, a finding is a document; after it, a finding becomes a regression test that runs forever. Most of this article is about making that jump without the gate turning into either theatre or an obstacle that teams learn to route around.
A charter for stage 2 needs only a page: the systems in scope, who owns the gate, who may override it and how an override is recorded, how findings are triaged, and how results are retained. The hard part is the override rule. If any product manager can wave a release through, the gate is advisory; if nobody can, it will be bypassed informally. A named risk owner who signs an acceptance with an expiry date is the usual compromise.
Scoping from risk: the coverage matrix
Coverage starts from risk, not from a list of jailbreak techniques. Take each harm in the AI risk register and cross it with each surface where an attacker can deliver input: the chat box, uploaded files, retrieved web pages, tool outputs, and memory. Each cell of that matrix is a test objective with an owner and a threshold. Taxonomies such as the OWASP Top 10 for LLM Applications and MITRE ATLAS are useful checklists for filling cells, but they are inputs to the matrix, not the matrix itself.
Keep the matrix as data, so coverage is computed rather than claimed:
from dataclasses import dataclass
@dataclass
class Cell:
harm: str # e.g. "unauthorised refund"
surface: str # e.g. "retrieved web page"
severity: int # 1..4 from the risk register
max_asr: float # gate threshold on the 95% upper bound
min_trials: int # sample size needed to resolve that threshold
def coverage(cells, results):
"""results maps (harm, surface) -> (successes, trials)."""
missing, thin = [], []
for c in cells:
r = results.get((c.harm, c.surface))
if r is None:
missing.append(c)
elif r[1] < c.min_trials:
thin.append((c, r[1]))
return missing, thinA cell with no results is a known gap; a cell with fewer trials than it needs is a hidden one, because a clean result on 20 attempts proves very little. The next section shows why.
Measuring attack success honestly
A model given the same attack twice can refuse once and comply once. An attack success rate (ASR) is therefore an estimate of a probability, and it needs an interval. Three rules keep the numbers honest.
- Sample at production settings. Run the same model version, system prompt, tools and temperature that users get. Testing at temperature zero measures a system nobody uses.
- Report an interval, gate on the upper bound. The Wilson score interval behaves well for small counts and rates near zero, which is exactly where gates live.
- Measure the grader. Automated judges miss successful attacks. Have humans re-grade a random sample each run and track the judge's false negative rate; a judge that misses 30 percent of successes makes every ASR look 30 percent better than it is.
from math import sqrt
def wilson(successes, n, z=1.96):
"""95% Wilson score interval for a binomial proportion."""
if n == 0:
return 0.0, 1.0
phat = successes / n
denom = 1 + z * z / n
centre = (phat + z * z / (2 * n)) / denom
half = z * sqrt(phat * (1 - phat) / n + z * z / (4 * n * n)) / denom
return max(0.0, centre - half), min(1.0, centre + half)
def gate(cell, successes, n):
lo, hi = wilson(successes, n)
if n < cell.min_trials:
return "INSUFFICIENT", lo, hi
return ("PASS" if hi <= cell.max_asr else "FAIL"), lo, hiTwo useful facts follow. First, the rule of three: if you observe zero successes in n independent trials, the one-sided 95 percent upper bound on the true rate is roughly 3/n. To claim a rate below 1 percent with no observed failures you need about 300 trials per cell, not 30. Second, trials are only independent if the attacks differ. Running one prompt 300 times measures sampling variance for that prompt; running 300 distinct attacks drawn from the cell's corpus measures the cell. Use both, but gate on the second.
The release gate
The gate runs in the deployment pipeline whenever anything that changes behaviour changes: model version, system prompt, tool definitions, retrieval sources, or guardrail configuration. It reads the coverage matrix, runs each cell's suite, grades results, and emits PASS, FAIL or INSUFFICIENT per cell. Any FAIL on a severity 3 or 4 cell blocks the release unless a risk owner records an acceptance.
Keep the suite affordable. A full gate across 40 cells at 300 trials each is 12,000 conversations, many multi-turn, plus judge calls. Tier it: a fast smoke subset on every prompt change, the full suite on model or tool changes, and manual campaigns on a calendar. Cache nothing across model versions; the point is to detect that behaviour moved. The broader evaluation harness is described in LLM Security Evaluations.
Worked example: a support agent with a refund tool
A support agent can read order history, issue refunds up to a limit, and send email. The highest severity cell is unauthorised refund via indirect injection: instructions hidden in a product review or an email the agent reads. The threshold is an upper bound of 5 percent, with at least 200 trials.
Run 1. 200 distinct injection attacks, graded by a judge with a 10 percent human audit. Fourteen succeed. Wilson gives 4.2 to 11.4 percent. The gate fails; the team does not argue about whether the true rate is 7 percent or 5 percent, because the upper bound settles it.
Fix. Refunds now require the order ID to appear in the user's own message, not in retrieved content, and the refund tool rejects amounts above the order total. These are controls that hold regardless of what the model is persuaded to do.
Run 2. 300 distinct attacks, none reused from run 1, including 100 aimed at the new control (for example, coaxing the user into pasting an order ID). Four succeed: 0.5 to 3.4 percent. The original fourteen successes are re-run separately as regression checks and all must fail. The gate passes. All eighteen successful attacks go into the regression suite, and the manual team logs 'social engineering the user into supplying the trigger' as a new technique for the next campaign.
Note what the program did beyond the exercise: the threshold existed before the test, the sample size was chosen to resolve it, the fix was verified with fresh attacks rather than the ones it was tuned against, and the evidence is reproducible.
Turning results into evidence
Red team output increasingly doubles as compliance evidence, so record it in a form someone else can audit. Under the EU AI Act, providers of general-purpose AI models with systemic risk must evaluate their models, including conducting and documenting adversarial testing (Article 55). High-risk systems carry accuracy, robustness and cybersecurity duties (Article 15). In the United States, NIST AI 600-1, the Generative AI Profile of the AI Risk Management Framework, lists red teaming among its suggested actions. None of these prescribe a method, which is why the record matters more than the tooling.
A useful evidence pack per release contains: the coverage matrix with thresholds as they stood before testing; model, prompt and tool versions under test; per-cell counts and intervals; the judge audit rate and measured judge error; open findings with owners and dates; and any risk acceptances with expiry. Store raw transcripts with access controls, since successful attacks are themselves sensitive.
Cadence and program metrics
Cadence keeps the program alive between launches. A workable calendar is: the automated gate on every change; a two-week manual campaign each quarter focused on the newest capabilities; an external exercise yearly or before a major launch; and continuous intake from the bug bounty and production incidents. Pair with defenders: the AI blue team should turn each finding into a detection as well as a fix.
Measure the program, not just the model:
- Coverage: fraction of matrix cells with sufficient trials in the last 30 days.
- Time to fix: median days from confirmed finding to passing gate, by severity.
- Regression rate: how often an old finding reappears after a model or prompt change.
- Novelty rate: share of manual findings not caught by the automated suite. If it falls to zero, either the suite is excellent or the manual team has stopped exploring.
- Gate overrides: count and age of open risk acceptances.
Failure modes
- Testing a different system. Gate runs against a staging prompt or temperature that differs from production, so passing results describe a system users never see.
- Tiny samples read as clean. Zero successes in 25 trials is reported as 'resolved' when the upper bound is still about 13 percent.
- Overfitting the fix. Re-running only the attacks that found the bug shows the patch memorised them, not that the weakness is gone.
- Unmeasured judges. A grader model is updated and its miss rate doubles; ASR improves overnight with no change to the product.
- Permanent exceptions. Risk acceptances without expiry accumulate until the gate blocks nothing.
- Findings without owners. Reports land in a shared folder and nobody is accountable for the fix; route them through the same process as security bugs and, when exploited, through incident response.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Strict thresholds on upper bounds | Real protection, defensible evidence | Large sample sizes, slower releases |
| Automated judges | Scale to thousands of trials | Judge error must be audited continuously |
| Internal team | Context, speed, continuity | Blind spots and familiarity with the system |
| External exercises | Fresh technique, independence | Cost, ramp-up time, less context |
| Blocking gate | Changes behaviour | Pressure to override; needs an exception process |
What to do next
- Write a one-page charter naming the gate owner, the override authority and how acceptances expire.
- Build a coverage matrix from your risk register: harms crossed with input surfaces, each with severity and threshold.
- Compute minimum trials per cell from the threshold using the rule of three, and stop reporting cells below it.
- Add the Wilson interval to your result reports and gate on the upper bound.
- Audit your judge: human-grade a random 10 percent each run and track its miss rate.
- Wire a smoke suite into the deployment pipeline for prompt changes and a full suite for model and tool changes.
- Convert every confirmed finding into a regression test, and verify fixes with fresh attacks.
- Assemble the evidence pack for your next release and check that someone outside the team can follow it.