Most LLM deployment checklists are a page of reasonable bullets: test for prompt injection, add content filters, monitor outputs. Teams tick them, ship, and find out in the first incident that nobody knew what passing meant, who had checked, or how to turn the feature off. A checklist only protects you when each item is a claim backed by evidence, with an owner, a pass criterion and an expiry date, checked by a machine where possible and re-run whenever the thing it describes changes.

This article is that kind of checklist for releasing an LLM-backed feature: the gates, what evidence each needs, a checklist-as-code format with a validator you can run in CI, the statistics behind evaluation thresholds, staged rollout with a kill switch and rollback, a worked example, and the events that should send a feature back through the gate. Infrastructure hardening (containers, networks, secrets, rate limiting) is covered in LLM deployment hardening; this page assumes it and focuses on the release decision.

Why checklists fail, and what an item needs

An LLM feature is not one artefact. It is a bundle: the model identifier and version, the system prompt, the tool definitions and their permissions, the retrieval corpus and its ingestion rules, the input and output filters with their thresholds, and the decoding settings. Change any part and the behaviour changes, sometimes dramatically. A safe release process therefore versions the bundle as a unit and attaches the evidence to that version.

Each checklist item should have five fields. The claim ("retrieved documents cannot trigger tool calls"), the owner who signs it, the evidence (a test report, a config file, a design review), the criterion that makes the evidence a pass, and the expiry: a date or a bundle hash after which the evidence no longer counts. Items without a criterion are opinions; items without expiry go stale silently.

The four gates

Group the items into four gates. The table lists the items most teams need; add your own domain items.

GateItemEvidencePass criterion
Scope and threatsIntended use and prohibited uses writtenProduct spec sectionApproved by product and security owners
Scope and threatsThreat model covers direct and indirect injection, data leakage, unsafe output, excessive agencyThreat model doc mapped to OWASP LLM Top 10Every high threat has a named control
ControlsTools follow least privilege; side-effects need confirmationTool manifest with scopesNo tool holds write scope it does not need
ControlsInput, document and output filters configuredPolicy table with thresholdsThresholds calibrated on labelled data, report attached
ControlsSecrets and PII never enter prompts or logs unredactedRedaction testsZero canary leaks in test set
Eval evidenceTask quality on the golden setEval report for this bundle hashMeets product threshold
Eval evidenceAttack success rate per familyScanner and red-team reportsUpper 95% bound below budget
Eval evidenceOver-refusal on benign promptsBenign set reportBelow agreed rate
Ops readinessKill switch testedDrill logFeature off within one minute
Ops readinessRollback to previous bundle testedDrill logPrevious bundle restored and verified
Ops readinessMonitoring and incident runbookDashboards, alert rules, runbook linkOn-call has acknowledged
Release gate for an LLM featureChangemodel, prompt, tool, dataGate: every item has owner, evidence, criterion, expiryvalidator reads release.yaml + evidence filesScope + threatsControlsEval evidenceOps readinessall passany failBlockedfix or signed exceptionStaged rolloutinternal -> 1% -> 10% -> 50% -> 100%watch SLOsKill switchflag off in secondsRollbackpinned bundle N-1Re-gate on: model version, system prompt, new tool, new data source, provider or policy change
Every change to the bundle passes the same gate; failures block or need a signed, expiring exception; passing releases roll out in stages behind a kill switch.

Checklist as code

Put the checklist in the repository next to the feature, as data, and validate it in CI. The validator does not judge quality; it enforces that every item has evidence that exists, is fresh, is tied to the current bundle and meets its numeric criterion.

# release.yaml
bundle:
  model: provider/model-2026-08-01
  system_prompt: prompts/support_v14.txt
  tools: tools/manifest.json
  filters: policy/filters.yaml
items:
  - id: threat-model
    owner: security@acme
    evidence: docs/threat_model.md
    max_age_days: 90
  - id: injection-asr
    owner: ml-safety@acme
    evidence: reports/injection_eval.json
    metric: {failures: failures, trials: trials, max_upper_bound: 0.02}
    bound_to_bundle: true
  - id: kill-switch-drill
    owner: sre@acme
    evidence: drills/kill_switch.json
    max_age_days: 30
import hashlib, json, math, sys, time, yaml
from pathlib import Path

def bundle_hash(bundle):
    h = hashlib.sha256()
    for key in sorted(bundle):
        value = bundle[key]
        data = Path(value).read_bytes() if Path(value).exists() else value.encode()
        h.update(key.encode() + b"\0" + data + b"\0")
    return h.hexdigest()[:16]

def wilson_upper(failures, trials, z=1.96):
    if trials == 0:
        return 1.0
    p = failures / trials
    denom = 1 + z * z / trials
    centre = p + z * z / (2 * trials)
    margin = z * math.sqrt(p * (1 - p) / trials + z * z / (4 * trials * trials))
    return (centre + margin) / denom

def check(path="release.yaml"):
    spec = yaml.safe_load(Path(path).read_text())
    current = bundle_hash(spec["bundle"])
    problems = []
    for item in spec["items"]:
        ev = Path(item["evidence"])
        if not ev.exists():
            problems.append(f"{item['id']}: missing {ev}"); continue
        age = (time.time() - ev.stat().st_mtime) / 86400
        if age > item.get("max_age_days", 10**6):
            problems.append(f"{item['id']}: evidence is {age:.0f} days old")
        if item.get("bound_to_bundle") or "metric" in item:
            data = json.loads(ev.read_text())
            if item.get("bound_to_bundle") and data.get("bundle_hash") != current:
                problems.append(f"{item['id']}: evidence is for bundle {data.get('bundle_hash')}, not {current}")
            if "metric" in item:
                m = item["metric"]
                ub = wilson_upper(data[m["failures"]], data[m["trials"]])
                if ub > m["max_upper_bound"]:
                    problems.append(f"{item['id']}: upper bound {ub:.3f} > {m['max_upper_bound']}")
    for line in problems:
        print("GATE FAIL", line)
    return 1 if problems else 0

if __name__ == "__main__":
    sys.exit(check(*sys.argv[1:]))

File modification time is a weak freshness signal in a fresh CI checkout; in practice store a generated_at field inside each evidence file and read that. The important property is the bundle hash: evaluation evidence produced for system prompt v13 cannot pass the gate for v14.

Statistics for evaluation gates

Evaluation gates are where checklists most often lie. "Zero jailbreaks in 50 attempts" sounds like a pass, but with 50 trials and no failures the 95% upper bound on the true failure rate is about 3/50, or 6%, by the rule of three. If your budget is 2%, you need at least 150 clean trials by that rule, and 189 with the Wilson bound the validator uses, before the data can support the claim at all. The validator above uses the Wilson score upper bound, which behaves sensibly near zero and is what the criterion should be written against.

Three more rules keep the numbers honest. Count trials per attack family, not pooled, because a 1% average can hide a 20% family. Sample several generations per prompt at your production temperature, because a failure that happens one time in five is real. And measure over-refusal on a benign set alongside attack success, otherwise the cheapest way to pass is to make the feature useless. Tools such as a continuous red team pipeline generate this evidence on every change; this gate consumes it.

Staged rollout, kill switch and rollback

Passing the gate earns a staged rollout, not a launch. Ship behind a feature flag that can disable the LLM path without a deploy and route users to a fallback (a search page, a human queue, the old flow). Roll out to internal users, then 1%, 10%, 50% and 100%, holding each stage long enough to see real traffic, with explicit stop conditions: output-filter block rate above its baseline band, user reports per thousand sessions, tool-call error rate, cost per session, and any severity-one safety report.

Rollback must restore the whole bundle. Pin the model version rather than an alias that the provider can move, store prompts and tool manifests in version control, and deploy them together with a bundle identifier that appears in every log line. Then a rollback is a flag flip to bundle N-1, and an incident report can say exactly which configuration produced a bad output. Drill both the kill switch and the rollback before launch and record the drill as evidence; an untested kill switch is a hope.

Logging needs the same care. Record the bundle identifier, filter decisions and tool calls for every turn; store prompts and outputs only under your retention and redaction policy; and make sure the on-call engineer can find a conversation by report identifier without needing database access granted during the incident. Output guardrail behaviour is covered in output guardrails.

Worked example: a support assistant with tools

A team is releasing a retrieval-augmented support assistant that can look up orders and open refund tickets. The figures in this example are illustrative. Scope and threats pass: prohibited uses are listed, and the threat model names indirect injection through help-centre articles and customer emails as the top risk. Controls mostly pass, but the tool review finds that the ticket tool can also close tickets, which the assistant never needs. The scope is reduced before anything else happens.

Evaluation evidence is where the gate bites. Direct injection: 4 failures in 600 trials, Wilson upper bound about 1.7%, under the 2% budget. Indirect injection through retrieved documents: 3 failures in 120 trials, upper bound about 7%, well over budget, and the sample is too small to pass even if it had been clean. The team adds document-level injection classification at ingestion, fences retrieved text as data in the prompt, and requires user confirmation before any refund ticket. They regenerate the evaluation on the new bundle with 400 trials: 1 failure, upper bound about 1.4%. Over-refusal on 500 benign questions rose from 1.2% to 2.0%, inside the agreed 3%.

Ops readiness fails once: the kill-switch drill took four minutes because the flag cache had a long TTL. The TTL is cut and the drill repeated in 40 seconds. The release goes to internal users for a week, then 1% for three days, where the output-filter block rate sits inside its baseline band, then onward. Two weeks later the provider announces a new model version; the bundle hash changes and the gate goes red until the evaluation is rerun, which is the system working as intended.

When to re-run the gate

Re-run the gate, or the affected items, when any of these change: the model version or provider; the system prompt; a tool definition, permission or new tool; a new retrieval source or ingestion rule; filter thresholds; decoding settings; the user population (for example opening to minors or a new market); or applicable regulation. Also re-run on a calendar, quarterly at least, because attacks evolve even when your system does not. Use a red-teaming playbook for the human portion of each cycle.

Failure modes

  • Checkbox theatre. Items marked done with no evidence link. The validator should reject them.
  • Stale evidence. An evaluation from three prompts ago attached to today's release. Bind evidence to the bundle hash.
  • Small-sample passes. Zero failures in 30 trials proves little; write criteria as upper bounds.
  • Floating model aliases. The provider updates the model behind a name and your evidence silently stops applying.
  • Exceptions that never expire. Every waiver needs an owner and a date.
  • Gate without rollout. A perfect pre-release gate does not catch distribution shift; staged rollout and stop conditions do.

Trade-offs

A strict gate slows releases, and teams under pressure route around slow processes. Keep the gate cheap: automate evidence generation, make the validator fast, scale rigour with risk (an internal summariser with no tools does not need the same injection budget as an agent that issues refunds), and allow signed, expiring exceptions so the gate bends visibly rather than breaking quietly. The goal is not zero risk; it is that every release decision is made deliberately, by a named person, on current evidence.

What to do next

  1. Define the bundle for each LLM feature and compute a hash over it in CI.
  2. Write release.yaml with owner, evidence, criterion and expiry for every item in the four gates.
  3. Add the validator to CI so a missing, stale or off-bundle item blocks the merge.
  4. Rewrite evaluation criteria as upper confidence bounds per attack family, with a benign over-refusal criterion alongside.
  5. Pin model versions and deploy prompts and tool manifests together with a bundle id in every log line.
  6. Put the feature behind a flag, drill the kill switch and rollback, and record both.
  7. Define rollout stages and stop conditions before the first user sees the feature.
  8. List re-gating triggers and add a quarterly calendar re-run.
Key takeaway: A deployment checklist protects you only when each item is evidence with an owner, a criterion and an expiry, tied to the exact bundle being released. Validate it in CI, write evaluation criteria as confidence bounds, roll out in stages behind a tested kill switch, and send the feature back through the gate whenever the model, prompt, tools or data change.