Most LLM deployment checklists are a page of reasonable bullets: test for prompt injection, add content filters, monitor outputs. Teams tick them, ship, and find out in the first incident that nobody knew what passing meant, who had checked, or how to turn the feature off. A checklist only protects you when each item is a claim backed by evidence, with an owner, a pass criterion and an expiry date, checked by a machine where possible and re-run whenever the thing it describes changes.
This article is that kind of checklist for releasing an LLM-backed feature: the gates, what evidence each needs, a checklist-as-code format with a validator you can run in CI, the statistics behind evaluation thresholds, staged rollout with a kill switch and rollback, a worked example, and the events that should send a feature back through the gate. Infrastructure hardening (containers, networks, secrets, rate limiting) is covered in LLM deployment hardening; this page assumes it and focuses on the release decision.
Why checklists fail, and what an item needs
An LLM feature is not one artefact. It is a bundle: the model identifier and version, the system prompt, the tool definitions and their permissions, the retrieval corpus and its ingestion rules, the input and output filters with their thresholds, and the decoding settings. Change any part and the behaviour changes, sometimes dramatically. A safe release process therefore versions the bundle as a unit and attaches the evidence to that version.
Each checklist item should have five fields. The claim ("retrieved documents cannot trigger tool calls"), the owner who signs it, the evidence (a test report, a config file, a design review), the criterion that makes the evidence a pass, and the expiry: a date or a bundle hash after which the evidence no longer counts. Items without a criterion are opinions; items without expiry go stale silently.
The four gates
Group the items into four gates. The table lists the items most teams need; add your own domain items.
| Gate | Item | Evidence | Pass criterion |
|---|---|---|---|
| Scope and threats | Intended use and prohibited uses written | Product spec section | Approved by product and security owners |
| Scope and threats | Threat model covers direct and indirect injection, data leakage, unsafe output, excessive agency | Threat model doc mapped to OWASP LLM Top 10 | Every high threat has a named control |
| Controls | Tools follow least privilege; side-effects need confirmation | Tool manifest with scopes | No tool holds write scope it does not need |
| Controls | Input, document and output filters configured | Policy table with thresholds | Thresholds calibrated on labelled data, report attached |
| Controls | Secrets and PII never enter prompts or logs unredacted | Redaction tests | Zero canary leaks in test set |
| Eval evidence | Task quality on the golden set | Eval report for this bundle hash | Meets product threshold |
| Eval evidence | Attack success rate per family | Scanner and red-team reports | Upper 95% bound below budget |
| Eval evidence | Over-refusal on benign prompts | Benign set report | Below agreed rate |
| Ops readiness | Kill switch tested | Drill log | Feature off within one minute |
| Ops readiness | Rollback to previous bundle tested | Drill log | Previous bundle restored and verified |
| Ops readiness | Monitoring and incident runbook | Dashboards, alert rules, runbook link | On-call has acknowledged |
Checklist as code
Put the checklist in the repository next to the feature, as data, and validate it in CI. The validator does not judge quality; it enforces that every item has evidence that exists, is fresh, is tied to the current bundle and meets its numeric criterion.
# release.yaml
bundle:
model: provider/model-2026-08-01
system_prompt: prompts/support_v14.txt
tools: tools/manifest.json
filters: policy/filters.yaml
items:
- id: threat-model
owner: security@acme
evidence: docs/threat_model.md
max_age_days: 90
- id: injection-asr
owner: ml-safety@acme
evidence: reports/injection_eval.json
metric: {failures: failures, trials: trials, max_upper_bound: 0.02}
bound_to_bundle: true
- id: kill-switch-drill
owner: sre@acme
evidence: drills/kill_switch.json
max_age_days: 30import hashlib, json, math, sys, time, yaml
from pathlib import Path
def bundle_hash(bundle):
h = hashlib.sha256()
for key in sorted(bundle):
value = bundle[key]
data = Path(value).read_bytes() if Path(value).exists() else value.encode()
h.update(key.encode() + b"\0" + data + b"\0")
return h.hexdigest()[:16]
def wilson_upper(failures, trials, z=1.96):
if trials == 0:
return 1.0
p = failures / trials
denom = 1 + z * z / trials
centre = p + z * z / (2 * trials)
margin = z * math.sqrt(p * (1 - p) / trials + z * z / (4 * trials * trials))
return (centre + margin) / denom
def check(path="release.yaml"):
spec = yaml.safe_load(Path(path).read_text())
current = bundle_hash(spec["bundle"])
problems = []
for item in spec["items"]:
ev = Path(item["evidence"])
if not ev.exists():
problems.append(f"{item['id']}: missing {ev}"); continue
age = (time.time() - ev.stat().st_mtime) / 86400
if age > item.get("max_age_days", 10**6):
problems.append(f"{item['id']}: evidence is {age:.0f} days old")
if item.get("bound_to_bundle") or "metric" in item:
data = json.loads(ev.read_text())
if item.get("bound_to_bundle") and data.get("bundle_hash") != current:
problems.append(f"{item['id']}: evidence is for bundle {data.get('bundle_hash')}, not {current}")
if "metric" in item:
m = item["metric"]
ub = wilson_upper(data[m["failures"]], data[m["trials"]])
if ub > m["max_upper_bound"]:
problems.append(f"{item['id']}: upper bound {ub:.3f} > {m['max_upper_bound']}")
for line in problems:
print("GATE FAIL", line)
return 1 if problems else 0
if __name__ == "__main__":
sys.exit(check(*sys.argv[1:]))File modification time is a weak freshness signal in a fresh CI checkout; in practice store a generated_at field inside each evidence file and read that. The important property is the bundle hash: evaluation evidence produced for system prompt v13 cannot pass the gate for v14.
Statistics for evaluation gates
Evaluation gates are where checklists most often lie. "Zero jailbreaks in 50 attempts" sounds like a pass, but with 50 trials and no failures the 95% upper bound on the true failure rate is about 3/50, or 6%, by the rule of three. If your budget is 2%, you need at least 150 clean trials by that rule, and 189 with the Wilson bound the validator uses, before the data can support the claim at all. The validator above uses the Wilson score upper bound, which behaves sensibly near zero and is what the criterion should be written against.
Three more rules keep the numbers honest. Count trials per attack family, not pooled, because a 1% average can hide a 20% family. Sample several generations per prompt at your production temperature, because a failure that happens one time in five is real. And measure over-refusal on a benign set alongside attack success, otherwise the cheapest way to pass is to make the feature useless. Tools such as a continuous red team pipeline generate this evidence on every change; this gate consumes it.
Staged rollout, kill switch and rollback
Passing the gate earns a staged rollout, not a launch. Ship behind a feature flag that can disable the LLM path without a deploy and route users to a fallback (a search page, a human queue, the old flow). Roll out to internal users, then 1%, 10%, 50% and 100%, holding each stage long enough to see real traffic, with explicit stop conditions: output-filter block rate above its baseline band, user reports per thousand sessions, tool-call error rate, cost per session, and any severity-one safety report.
Rollback must restore the whole bundle. Pin the model version rather than an alias that the provider can move, store prompts and tool manifests in version control, and deploy them together with a bundle identifier that appears in every log line. Then a rollback is a flag flip to bundle N-1, and an incident report can say exactly which configuration produced a bad output. Drill both the kill switch and the rollback before launch and record the drill as evidence; an untested kill switch is a hope.
Logging needs the same care. Record the bundle identifier, filter decisions and tool calls for every turn; store prompts and outputs only under your retention and redaction policy; and make sure the on-call engineer can find a conversation by report identifier without needing database access granted during the incident. Output guardrail behaviour is covered in output guardrails.
Worked example: a support assistant with tools
A team is releasing a retrieval-augmented support assistant that can look up orders and open refund tickets. The figures in this example are illustrative. Scope and threats pass: prohibited uses are listed, and the threat model names indirect injection through help-centre articles and customer emails as the top risk. Controls mostly pass, but the tool review finds that the ticket tool can also close tickets, which the assistant never needs. The scope is reduced before anything else happens.
Evaluation evidence is where the gate bites. Direct injection: 4 failures in 600 trials, Wilson upper bound about 1.7%, under the 2% budget. Indirect injection through retrieved documents: 3 failures in 120 trials, upper bound about 7%, well over budget, and the sample is too small to pass even if it had been clean. The team adds document-level injection classification at ingestion, fences retrieved text as data in the prompt, and requires user confirmation before any refund ticket. They regenerate the evaluation on the new bundle with 400 trials: 1 failure, upper bound about 1.4%. Over-refusal on 500 benign questions rose from 1.2% to 2.0%, inside the agreed 3%.
Ops readiness fails once: the kill-switch drill took four minutes because the flag cache had a long TTL. The TTL is cut and the drill repeated in 40 seconds. The release goes to internal users for a week, then 1% for three days, where the output-filter block rate sits inside its baseline band, then onward. Two weeks later the provider announces a new model version; the bundle hash changes and the gate goes red until the evaluation is rerun, which is the system working as intended.
When to re-run the gate
Re-run the gate, or the affected items, when any of these change: the model version or provider; the system prompt; a tool definition, permission or new tool; a new retrieval source or ingestion rule; filter thresholds; decoding settings; the user population (for example opening to minors or a new market); or applicable regulation. Also re-run on a calendar, quarterly at least, because attacks evolve even when your system does not. Use a red-teaming playbook for the human portion of each cycle.
Failure modes
- Checkbox theatre. Items marked done with no evidence link. The validator should reject them.
- Stale evidence. An evaluation from three prompts ago attached to today's release. Bind evidence to the bundle hash.
- Small-sample passes. Zero failures in 30 trials proves little; write criteria as upper bounds.
- Floating model aliases. The provider updates the model behind a name and your evidence silently stops applying.
- Exceptions that never expire. Every waiver needs an owner and a date.
- Gate without rollout. A perfect pre-release gate does not catch distribution shift; staged rollout and stop conditions do.
Trade-offs
A strict gate slows releases, and teams under pressure route around slow processes. Keep the gate cheap: automate evidence generation, make the validator fast, scale rigour with risk (an internal summariser with no tools does not need the same injection budget as an agent that issues refunds), and allow signed, expiring exceptions so the gate bends visibly rather than breaking quietly. The goal is not zero risk; it is that every release decision is made deliberately, by a named person, on current evidence.
What to do next
- Define the bundle for each LLM feature and compute a hash over it in CI.
- Write release.yaml with owner, evidence, criterion and expiry for every item in the four gates.
- Add the validator to CI so a missing, stale or off-bundle item blocks the merge.
- Rewrite evaluation criteria as upper confidence bounds per attack family, with a benign over-refusal criterion alongside.
- Pin model versions and deploy prompts and tool manifests together with a bundle id in every log line.
- Put the feature behind a flag, drill the kill switch and rollback, and record both.
- Define rollout stages and stop conditions before the first user sees the feature.
- List re-gating triggers and add a quarterly calendar re-run.