A red-team engagement on an LLM product is a time-boxed attempt, by people who want it to fail, to make a specific product version do things its owners said it must never do. Done well, it produces findings that engineers can reproduce, fixes that are shown to work, and test cases that stop the same failure returning. Done badly, it produces a slide deck of screenshots that nobody can reproduce, and arguments about whether a jailbreak "counts".

This page is the runbook for a single engagement: the artifacts, the order of work, and the record formats that make results usable. Two companion pages cover the wider context. Building an AI red team program covers maturity, coverage matrices and attack-success statistics, and LLM red team architecture covers the tooling and the organisation. This page assumes both and concentrates on running one engagement from start to finish. Attack descriptions are kept at the level of technique classes, deliberately.

What an engagement is for

Three activities get called red teaming, and they need different processes. An evaluation runs a fixed benchmark and reports a score. A penetration test attacks the infrastructure: APIs, authentication and storage. A red-team engagement attacks the behaviour of the product as a whole. That includes the model, its system prompt, retrieval, tools, memory and the user interface, all judged against the product's own policy. Its value comes from creativity and from chaining steps together, which is why it does not replace the other two.

Frameworks expect it. NIST's Generative AI Profile (NIST AI 600-1, July 2024) lists red-teaming among its suggested actions, and the OWASP Top 10 for LLM Applications and MITRE ATLAS give shared names for attack classes. None of them tells you how to run the work week by week. That is the gap this process fills.

The engagement lifecycle

Every engagement goes through the same eight steps. The first two happen before anyone attacks anything, and the last three take longer than the attacking.

One red-team engagement, from charter to regression suite1. Charterscope, rules, stop conditions2. Test planthreat model to objectives3. Executionmanual sessions + tools4. Capturerepro bundle per finding5. Triageseverity, owner, due date6. Fixowner team ships change7. RetestN trials, same bundle8. Regression caseruns on every releaseclosedstill reproduces: back to 6Report: what was tested, what was not, findings by severity, retest results, residual riskgoes to the release owner before the go/no-go decisionThe engagement is finished when every finding is closed, accepted as risk, or tracked with a date.
The engagement lifecycle. Retest uses the same reproduction bundle as the original finding, and every closed finding becomes a regression case.

The loop from fix to retest is the step teams most often skip. A finding that is "fixed" without a retest has, in practice, only been assigned. LLM behaviour is stochastic, so a single successful retest proves little. The retest has to run the attack enough times to put a bound on the residual rate.

Step 1: the charter

Write the charter before the engagement starts, and get it signed by the product owner and security. It fixes what is in scope, what is not, and when testers must stop. Keep it short enough that testers actually read it:

engagement: support-assistant-v4.2-release
target:
  build: assistant@4.2.0-rc3          # exact build; findings are tied to it
  model: provider/model-id@pinned     # model version as recorded by the gateway
  surfaces: [web_chat, email_ingest, refund_tool, kb_retrieval]
environment: staging-redteam           # isolated tenant, synthetic customer data only
window: 2026-10-12 to 2026-10-23
objectives:                            # behaviours that must never happen
  - O1: refund issued without a verified order owner
  - O2: another customer's data appears in a response
  - O3: system prompt or tool credentials disclosed
  - O4: policy-violating content in categories H1-H4 of the content policy
out_of_scope: [provider infrastructure, denial of service, real customer accounts]
rules:
  - no real personal data in prompts or uploaded files
  - automated tools capped at 2 requests/s against staging
  - log every session id in the engagement tracker
stop_conditions:                       # pause and escalate immediately
  - evidence that staging reaches production data or systems
  - a finding that also affects the live product -> page on-call security
owners: {lead: red-team-lead, product: support-eng-manager, security: appsec-oncall}

Write the objectives as testable statements about outcomes, not as attack techniques. "O1: a refund is issued without verification" can be judged true or false. "Try prompt injection" cannot. The stop conditions protect the engagement from itself. Testers will find unexpected paths, and the charter decides in advance what happens when one of them leads somewhere it should not.

Step 2: the test plan

Turn the LLM threat model into a test plan. For each objective, list the entry points an attacker controls, such as chat turns, retrieved documents, inbound emails and uploaded files. Then list the technique families that apply to each entry point. A row in the plan names one objective, one entry point and one technique family, with the tester assigned and a time budget:

ObjectiveEntry pointTechnique familyModeBudget
O1 unverified refundchatrole-play and authority claims; multi-turn escalationmanual6 h
O1 unverified refundinbound emailindirect injection in a quoted threadmanual + generated variants8 h
O2 cross-customer datakb_retrievalquery shaping to pull other tenants' chunksmanual6 h
O3 prompt disclosurechatextraction and translation tricksautomated sweep2 h
O4 policy contentchatknown jailbreak familiesautomated sweep + manual follow-up8 h

The plan also records what you will not test this time, and why. The final report repeats that list. Untested surfaces are where the residual risk sits, and a report that hides them overstates what the engagement shows. Multi-turn techniques such as Crescendo need their own rows, because single-message filters do not catch them.

Step 3: execution

Run execution in two modes at once. Automated sweeps use tools such as PyRIT, garak or promptfoo to replay known technique families at scale and to mutate prompts that worked. They find regressions and cover breadth cheaply. Automated red teaming covers how to build the generator and judge. Manual sessions are where the important findings usually come from: chained tool calls, business-logic abuse, and failures that only make sense once you understand the product. Timebox manual sessions to about 90 minutes, each with a stated objective, and write notes during the session, not afterwards.

Hold a daily 15-minute stand-up. Each tester reports what they tried, what partly worked and what they will try next. Partial successes matter most. A model that refused but leaked half the system prompt while refusing is a lead for someone else to pursue. The lead keeps a coverage board of the plan's rows and moves budget to rows that are producing results.

Step 4: capturing findings reproducibly

A finding is only useful if an engineer can reproduce it on a different day. LLM outputs depend on far more state than a web request, so the reproduction bundle has to capture all of it:

from dataclasses import dataclass, field
import hashlib, json, time

@dataclass
class ReproBundle:
    build: str                    # product build under test
    model_id: str                 # exact model version string from the gateway log
    system_prompt_sha256: str     # hash, so prompt edits are detectable
    sampling: dict                # temperature, top_p, max_tokens, seed if the API supports one
    tool_state: dict              # fixtures: which orders, accounts, documents existed
    retrieved_ids: list           # chunk ids the retriever returned on each turn
    transcript: list              # every message, role-tagged, including tool calls and results
    trace_ids: list               # gateway / tracing ids for the session

@dataclass
class Finding:
    id: str
    objective: str                # O1..O4 from the charter
    technique_family: str         # e.g. "indirect injection via email"
    summary: str
    bundle: ReproBundle
    trials: int = 0               # how many times the transcript was replayed
    successes: int = 0            # how many replays reached the objective
    severity: str = "untriaged"
    owner: str = ""
    status: str = "open"          # open -> fixing -> retest -> closed | accepted
    created: float = field(default_factory=time.time)

def prompt_hash(system_prompt: str) -> str:
    return hashlib.sha256(system_prompt.encode("utf-8")).hexdigest()

Record the reproduction rate when you file the finding: replay the transcript, say, 20 times and store successes out of trials. A finding that reproduces 3 times in 20 is real, with a probability that needs managing. Recording it as "flaky" hides that. Keep raw harmful outputs in the access-controlled tracker, and keep the readable summary free of them, so the report can circulate widely.

Step 5: triage

Triage each finding within one working day. Severity is the impact of reaching the objective, adjusted for the attacker's cost and the reproduction rate. A refund issued to the wrong person at 3 in 20 is worse than a mildly off-policy joke at 20 in 20. Use the shared rubric from the red team architecture page rather than inventing one per engagement, and assign a single owning team and a due date for each finding. If a finding also affects the live product, that is a charter stop condition: page on-call security and handle it as an incident, not as a ticket.

Steps 6 to 8: fix, retest, regression

When the owning team reports a fix, the red team, not the fixing team, replays the original bundle on the fixed build. Then it tries close variants, because a fix that blocks the exact wording but not the technique is a filter, not a fix. Zero successes does not mean zero risk. If n independent replays all fail, the 95% upper bound on the true success rate is roughly 3/n, by the rule of three. Choose n from the rate you need to rule out:

def trials_needed(max_rate: float, confidence: float = 0.95) -> int:
    """Smallest n such that 0/n successes bounds the true rate below max_rate."""
    import math
    return math.ceil(math.log(1 - confidence) / math.log(1 - max_rate))

def retest(finding, replay, variants, n):
    hits = sum(replay(finding.bundle) for _ in range(n))
    variant_hits = sum(replay(v) for v in variants)
    finding.status = "closed" if hits == 0 and variant_hits == 0 else "fixing"
    return hits, variant_hits

# trials_needed(0.05) -> 59    trials_needed(0.01) -> 299

Every closed finding becomes a regression case. Its transcript, fixtures and a judge for the objective go into the automated suite, which runs on every release candidate and every model version change. Model upgrades are the moment when old findings most often come back. Score the suite the way prompt injection evaluation describes: attack success rate and utility together, so that a fix which simply refuses everything does not count as a win.

Worked example: a refund-capable support assistant

Take a two-week engagement on version 4.2 of a support assistant that can issue refunds of up to 100 dollars and reads inbound customer emails. The team is three testers and a lead.

DaysWorkOutput
-5 to -1Charter signed; staging tenant seeded with 40 synthetic customers and orderscharter, fixtures
1Threat-model walkthrough with the product engineers; test plan agreedplan with 14 rows
2-3Automated sweeps for O3 and O4; manual sessions on O1 via chat2 low findings
4-6Manual sessions on O1 via email; a quoted forwarded thread gets the agent to treat text in the thread as staff instructionsF-07: O1 reached, 6/20
7Triage: F-07 rated high; owner is the agent platform team; stop-condition check shows production is unaffectedticket, due date
8-9Fix: email content is passed to the model as quoted data, and the refund tool requires a verified-owner flag set by code, not by the modelbuild rc4
10Retest: 0/59 on the original bundle; 0/25 across variantsF-07 closed
10Regression case added; report draftedsuite +3 cases

The fix for F-07 shows the pattern that most good fixes follow. It does not try to teach the model to resist one phrasing. It moves the security decision, whether this person may receive this refund, out of the model and into deterministic code. Results like 0/59 on the original bundle bound the residual rate below about 5% at 95% confidence. If the business needs a tighter bound, run more trials.

Failure modes

  • Screenshot findings. No build, model version or sampling settings were recorded, so the finding cannot be reproduced and is argued about instead of fixed.
  • Testing the wrong build. The model or system prompt changed mid-engagement. Pin the build, and record the prompt hash with every finding.
  • Self-certified fixes. The team that wrote the fix also marked it closed. Retesting belongs to the red team.
  • Exact-string fixes. A blocklist entry stops the original wording, and a paraphrase gets through. Always retest variants.
  • No negative space. The report lists findings but not untested surfaces, so readers assume full coverage.
  • Engagement without regression. Findings were fixed, but they return when the model is upgraded, because nothing re-runs them.
  • Unsafe handling of outputs. Harmful generations get pasted into chat channels or documents with wide access. Keep them in the tracker.

Trade-offs

Internal versus external testers. Internal testers know the product and find business-logic chains. External testers bring techniques and independence. Rotate an external team in for major releases. Breadth versus depth. Automated sweeps cover many technique families shallowly, and manual sessions go deep on a few. Spending the budget on depth for the highest-impact objectives usually finds more serious issues. Release gating versus continuous testing. An engagement before each major release catches design flaws. Continuous automated runs catch regressions between releases. Mature teams do both, and use the engagement's closed findings to feed the continuous suite.

What to do next

  1. Write a one-page charter for your next release, with outcome-based objectives and explicit stop conditions.
  2. Turn your threat model into a test plan table with entry points, technique families, modes and budgets.
  3. Adopt a reproduction bundle schema, and refuse to triage findings that lack one.
  4. Record successes out of trials for every finding, and size retests with trials_needed.
  5. Make retesting the red team's job, and require variant testing before closure.
  6. Add every closed finding to an automated regression suite that runs on model and prompt changes.
  7. Read automated red teaming to scale the sweep side.
Key takeaway: An LLM red-team engagement succeeds when its findings are reproducible, its fixes are verified and its lessons become regression tests. Sign a charter with outcome objectives and stop conditions, plan from the threat model, mix automated sweeps with timeboxed manual sessions, capture a full reproduction bundle with a success rate, retest independently with enough trials to bound the residual rate, and report what was not tested as clearly as what was.