An output guardrail is any check that runs on what a language model produced before that output reaches a person, a browser, a database or another tool. Input guardrails try to stop bad requests; output guardrails exist because some will get through anyway, and because a perfectly benign request can still produce a harmful, wrong or leaky answer. The model is a probabilistic component, and its output should be treated the way a web server treats form input: untrusted until checked against the place it is going.
This overview is about the whole pipeline rather than any one detector. It covers what to check, in what order, what to do when a check fails, how to handle streaming, and how to measure whether the guardrails help. The individual detectors each have deep dives on this site, linked where they come up. The goal is that you can design a pipeline for your own application, defend its latency cost, and know its blind spots.
Start from where the output goes
Start from the sink, not the model. The same sentence is harmless in a chat bubble and dangerous inside an HTML template, a SQL string or a shell command. So the first question for any LLM feature is: where does this output go, and what can it do there? Common sinks and the failure each one enables:
| Sink | What goes wrong | Primary check |
|---|---|---|
| Rendered HTML or markdown | script injection, image-URL exfiltration | safe rendering, link allow-list |
| Parsed JSON / tool arguments | schema violations, injected fields | strict schema validation |
| Shown to an end user | toxicity, self-harm content, wrong facts | content classifiers, grounding |
| Logged or stored | PII and secrets persisted | PII and secret detectors |
| Fed to another agent | propagated prompt injection | treat as untrusted input downstream |
The general rules for handling output as untrusted data, including encoding, are covered in LLM output handling. Guardrails are the content layer on top: they decide whether the output should be released at all.
A staged pipeline
A practical pipeline has three stages, ordered by cost. Stage 1 is deterministic and nearly free: JSON schema validation, length limits, regular expressions for credentials and card numbers, URL allow-lists, and canary strings planted in the system prompt. Stage 2 runs classifiers: a toxicity or policy model, a PII recogniser, perhaps a safety model such as Llama Guard that scores the response against a harm taxonomy. Stage 3 checks the output against context: is each claim supported by the retrieved documents (grounding verification), does the answer reveal data this user may not see?
Each stage emits verdicts, not actions. A separate policy engine combines verdicts with severity and context and chooses one action. Keeping detection and decision apart is the single most useful structural choice: you can tune a threshold, or change what happens to a medium-severity PII hit, without touching detector code, and you can run a new detector in shadow mode where it is logged but cannot act.
From verdicts to actions
There are five useful actions, and choosing between them is product design as much as security. Allow. Redact: remove or mask the offending spans, which works for PII and secrets because the detector knows where they are, and fails for anything semantic. Regenerate: ask the model again with a stricter instruction, which fixes many grounding and format failures but doubles latency and cost, so bound it to one retry. Block: return a fixed safe message and a reason code. Flag: release the output but queue it for human review, used where the cost of a false block is high and the harm is low.
from dataclasses import dataclass
@dataclass
class Verdict:
check: str
passed: bool
severity: str # "low" | "medium" | "high"
score: float = 0.0
spans: list = None # character ranges for redaction
def run_pipeline(output, ctx, stages):
verdicts = []
for stage in stages: # ordered cheap -> expensive
for check in stage:
v = check(output, ctx)
verdicts.append(v)
if any(not v.passed and v.severity == "high" for v in verdicts):
break # short-circuit: no need for stage 3
return verdicts
def decide(verdicts, attempt, max_regen=1):
failed = [v for v in verdicts if not v.passed]
if not failed:
return "allow"
if any(v.severity == "high" for v in failed):
return "block"
if all(v.check in ("pii", "secret") for v in failed):
return "redact" # spans are known, remove them
if attempt < max_regen and any(v.check == "grounding" for v in failed):
return "regenerate" # retry with a stricter instruction
return "flag" # release, but queue for reviewTwo details matter in production. First, blocking must fail closed: if a classifier times out or errors, the policy treats it as a failed check for high-risk sinks, otherwise an attacker only needs to make the detector slow. Second, every decision is logged with the verdicts, scores and request ID, because you will tune thresholds from those logs. Libraries such as Guardrails AI and NeMo Guardrails implement the same verdict and on-fail shape; the design questions are the same whichever you choose.
Streaming output
Streaming breaks the simple model, because tokens reach the user before the output is complete. There are three options. Buffer the whole response and check it, which is safe and removes the benefit of streaming. Check incrementally and retract, which shows harmful text briefly and then removes it, acceptable for mild categories only. Or hold back a window: release text only up to a watermark a few tokens or a sentence behind generation, run fast checks on each new chunk, and abort the stream if a check fails before the unsafe part is released. The holdback approach is described in detail in streaming moderation and, for secrets and exfiltration links, in egress filtering.
Whatever you choose, run slow stage 3 checks after the stream ends and act on the stored answer: mark it, hide it from history, or alert. A gateway that labels a failed check on a stream without stopping it is giving you telemetry, not protection.
Worked example: a billing support answer
A customer support assistant answers billing questions from retrieved account documents. The user asks why they were charged twice. The model produces an answer that (a) explains the duplicate charge correctly, (b) quotes the last four digits of the card and also the full 16-digit number that appeared in a retrieved invoice, and (c) claims a refund was issued on a date that does not appear in any document.
Stage 1: the JSON envelope validates; the card-number regex with a Luhn check matches the 16-digit number, severity medium, span known. Stage 2: the toxicity score is low and the policy classifier passes. Stage 3: the grounding check finds the refund claim unsupported, severity medium. The policy sees two medium failures of different kinds. It redacts the card number first, then regenerates once with the instruction to state only facts in the provided documents. The second answer passes grounding, still mentions the last four digits (allowed by policy), and is released. Total added latency: one fast stage, one classifier call, one grounding call, then the same again for the regeneration, which is why the regeneration budget is one.
Measuring and choosing thresholds
A guardrail you have not measured is a guess. Build a labelled set of real outputs, including deliberately adversarial ones, and measure each detector's recall (unsafe outputs caught) and precision (blocks that were justified). Then choose thresholds from the trade-off, not from a default. An illustrative run on 2,000 labelled responses of which 100 are unsafe (invented numbers that show the arithmetic):
| Threshold | Recall | Precision | Missed unsafe | False blocks | Block rate |
|---|---|---|---|---|---|
| 0.3 | 94% | 37% | 6 | 160 | 12.7% |
| 0.5 | 88% | 59% | 12 | 60 | 7.4% |
| 0.7 | 79% | 78% | 21 | 22 | 5.1% |
| 0.9 | 61% | 91% | 39 | 6 | 3.4% |
At 0.3 the detector misses only 6 unsafe answers but wrongly blocks 160 good ones, 8% of all traffic: users will notice and route around the product. At 0.9 it is precise but misses 39 of 100. A reasonable choice here is 0.7 to block and 0.5 to flag for review. Re-run the evaluation whenever the model, prompt or detector changes, and sample production blocks weekly to keep the labels honest.
Where the guardrails run
Guardrails can run in three places, and most mature systems use more than one. In the application, as a library called after the model returns: the checks see full context such as the user's permissions and the retrieved documents, which grounding and access checks need, but every team must remember to call them. In an AI gateway between applications and model providers: checks apply to all traffic by default and are configured centrally, but the gateway usually sees only the request and response, not business context, and some gateways only label a failed check on a streamed response rather than stopping it. At the sink, inside the renderer, the tool executor or the database layer: the last line, and the only one that knows exactly what the output can do there.
A sensible split is: deterministic and classifier checks in the gateway, so no route ships without them; context checks such as grounding and data-access rules in the application; encoding and allow-lists at the sink. Record which layer made each decision, so an incident review can tell a missed detection from a missing layer.
Failure modes
- Detectors that only see the final text. Output split across tool calls, or encoded in base64 or a URL, passes a per-message classifier. Decode and check every channel.
- Fail-open on errors. A timeout returning allow turns a slow detector into a bypass.
- Regeneration loops. Unbounded retries burn cost and can be triggered on purpose.
- Over-blocking. High false positives teach users to rephrase until something passes, which trains them to evade the system.
- Language and format gaps. Classifiers trained on English prose underperform on other languages, code and tables.
- Leaky refusal messages. A block message that names the detector and threshold helps attackers tune around it. Return a reason code, log the detail.
- Guardrail drift after a model upgrade. A new model version phrases things differently, and a classifier tuned on the old outputs shifts its score distribution. Re-run the labelled evaluation before switching models, not after the first incident.
- Checking the wrong copy. The pipeline checks the text, then a formatter, translator or template step changes it before release. Run the final checks on the exact bytes that reach the sink.
Trade-offs
Every stage adds latency and cost. Deterministic checks are microseconds; a small classifier on a GPU or CPU adds tens of milliseconds; an LLM-based judge or grounding check can add as much as the generation itself. Order by cost, short-circuit on hard failures, run independent checks in parallel, and reserve stage 3 for high-risk routes. Buying a hosted moderation API is quick but sends output to another party and fixes the taxonomy; self-hosted classifiers give control and data locality at the price of evaluation work. Neither replaces safe handling at the sink: guardrails reduce what is released, and output encoding limits what a released string can do.
What to do next
- List every sink your LLM output reaches and the failure each one enables.
- Add stage 1 checks for each sink: schema validation, secret and PII regexes, URL allow-list.
- Pick one classifier for your highest-harm category and measure it on a labelled set.
- Separate detection from decision with a small policy function and log every verdict.
- Make every detector fail closed on high-risk routes, and bound regeneration to one retry.
- Choose a streaming strategy deliberately: buffer, holdback or retract.
- Run new detectors in shadow mode for a week before letting them block.
- Review a sample of blocks and passes weekly and re-tune thresholds.