Most teams add their first LLM guardrail after an incident. A user pastes a support transcript full of card numbers, or a chatbot quotes a refund policy that does not exist, or an agent emails the wrong customer. The fix is usually one more check bolted onto the request path. Months later there are nine checks with nine thresholds and log formats, and nobody can say which risk each one stops or what happens when one times out.
This article treats guardrails as a production control system rather than a list of filters. You start from the risks you actually carry, map each one to a control that can be tested, route every check through one decision engine that produces a single verdict and a single record, and change policy the way you change code: behind a regression suite and a staged rollout. Serving the guard models themselves (GPU placement, batching, timeouts) is covered in LLM Guardrails on GPU serving paths; the content taxonomy and human review queue are covered in LLM moderation architecture. Here the focus is the wiring, the evidence and the operations.
Start from a risk register, not a product
A guardrail is only useful if you can name the harm it prevents and show it working. So the first artefact is not code but a risk register: one row per risk, with who is harmed, how it would happen in your product, and which control owns it. It stops two common mistakes: assuming a broad content classifier covers data leakage, and stacking three classifiers on one risk while another, such as an agent deleting records, has no control at all.
| Risk | How it happens | Primary control | Evidence it works |
|---|---|---|---|
| Personal data leaks into prompts or logs | Users paste records; retrieval pulls unredacted rows | Deterministic PII detection and redaction before the model and before logging | Recall on a labelled PII set per locale |
| Prompt injection triggers an unwanted action | Instructions hidden in a retrieved page or email | Action guard: authorization in code, argument validation, approval for risky tools | Injection suite replayed against every tool |
| Ungrounded answers about policy or price | Model fills gaps from training data | Grounding check against cited sources, abstain when unsupported | Faithfulness rate on a fixed question set |
| Harmful or off-policy content | Adversarial or careless prompts | Input and output classifiers with per-category thresholds | Precision and recall per category |
| Malformed output breaks a downstream system | Model drifts from the JSON contract | Schema validation with one bounded repair attempt | Parse failure rate in production |
Notice that only one of the five rows is primarily a classifier. In most production systems the controls that prevent real damage are deterministic: redaction, authorization and schema validation. Classifiers handle the fuzzy categories where no rule can be written.
Three kinds of check
Every check falls into one of three kinds, and they behave differently under load and attack.
Deterministic checks are rules and parsers: regular expressions with checksum validation for card numbers, allow-lists of domains, JSON schema validation, length and rate limits. They are fast, explainable and do not drift, but only catch what you anticipated.
Model checks are classifiers or a second LLM acting as a judge: toxicity, self-harm, jailbreak likelihood, topic relevance, faithfulness. They generalise to unanticipated inputs, but they produce scores rather than facts, have false positives, and can be probed by an attacker. Every threshold needs a labelled dataset behind it.
Authorization checks decide whether an action may happen at all, given who the user is and what the request is trying to do. They must never depend on the LLM's own judgement, because the LLM is exactly the component an injection controls. If a user cannot refund an order through the normal UI, the agent acting for them cannot either, whatever the prompt says.
Give every check the same interface so the engine can treat them uniformly:
from dataclasses import dataclass, field
from typing import Literal
Action = Literal["allow", "redact", "block", "escalate"]
@dataclass
class Verdict:
check: str # "pii_regex", "jailbreak_clf", "refund_authz"
version: str # version of the check or its model
action: Action
score: float | None = None # model checks only
reason: str = "" # stable reason code, never free text from the LLM
redactions: list[tuple[int, int]] = field(default_factory=list)
latency_ms: float = 0.0
error: str | None = None # set when the check failed to runThe error field matters. A check that crashed has not said "allow". The engine, not the check, decides what an error means for that risk.
One decision engine, one verdict, one record
The decision engine is a small, boring piece of code with a big job: it runs the checks configured for a stage, combines their verdicts using a versioned policy, and writes one decision record. Checks never block or rewrite traffic on their own. That single point of decision is what makes the system auditable and testable.
SEVERITY = {"allow": 0, "redact": 1, "escalate": 2, "block": 3}
def decide(stage, request, checks, policy):
verdicts = []
for check in checks[stage]:
try:
v = check.run(request, timeout_ms=policy.timeout_ms[check.name])
except Exception as exc: # timeout or crash
v = Verdict(check.name, check.version, "allow", error=repr(exc))
if v.error:
# Per-risk failure mode lives in the policy, not in the check.
v.action = policy.on_error[check.name] # "block" for authz, maybe "allow" for tone
verdicts.append(v)
final = max(verdicts, key=lambda v: SEVERITY[v.action], default=None)
action = final.action if final else "allow"
record = DecisionRecord(
request_id=request.id, tenant=request.tenant, stage=stage,
policy_version=policy.version, action=action,
reason=final.reason if final else "",
verdicts=verdicts,
)
decision_log.append(record)
return action, recordThe most severe verdict wins, so adding a check can never make the system more permissive. Each check's failure mode is a versioned policy decision, so "what happens when the PII service is down" has a written answer. And one record per request means block rates and per-check latency come from one log.
Guarding actions, not just text
In agentic products the expensive failures are actions, not text. An action guard sits between the model's proposed tool call and its execution, and it checks four things in code: is this tool allowed for this user and tenant, are the arguments valid and within limits, do the arguments refer only to resources this user can reach, and does this action need a human to approve it.
def guard_tool_call(user, call, policy):
spec = TOOLS[call.name]
if call.name not in policy.allowed_tools(user.role):
return Verdict("tool_authz", spec.version, "block", reason="tool_not_allowed")
args = spec.schema.validate(call.arguments) # raises on bad types or extra keys
if not resources.user_can_access(user, spec.resource_ids(args)):
return Verdict("tool_authz", spec.version, "block", reason="resource_not_owned")
if spec.risk == "irreversible" or spec.amount(args) > policy.approval_limit(user):
return Verdict("tool_authz", spec.version, "escalate", reason="needs_approval")
return Verdict("tool_authz", spec.version, "allow")The resource check is the one teams skip. A model that has been told, through an injected document, to "refund order 991" will produce a perfectly valid tool call. Only a check that order 991 belongs to the current user stops it. Execute the tool with the user's own credentials or a token scoped to them, never with a service account that can reach every tenant. Broader agent-level limits, such as step budgets and kill switches, are covered in Agent guardrails.
Output checks: schema and grounding
Two output checks catch most of the failures users actually notice.
Schema validation. When output feeds code, validate it against a schema and allow at most one repair attempt. If the repaired output also fails, return a typed error; unbounded retries turn one bad answer into ten paid calls.
Grounding. For retrieval-backed answers about facts you own, such as prices, policies or account state, check that each claim is supported by a retrieved source. The cheapest version requires the model to cite source ids and verifies that each cited id was actually retrieved and that quoted spans appear in it. A judge model over each claim and source is stronger. Either way, the safe failure is to abstain ("I can't confirm that; here is the policy page") rather than to block silently. Techniques for constraining and verifying answers are covered in hallucination guardrails.
The decision log
The decision log is the evidence for tuning, audits, incidents and the eval set. Each record should hold the request id, tenant, stage, policy version, final action and reason code, plus every check's name, version, action, score, latency and error. Hold redacted text, or a pointer to the raw text in a store with tighter access and a short retention period, but never raw personal data in the main log.
Two metrics from the log deserve dashboards. The first is block rate per check and per tenant, because a sudden change almost always means a deploy, a model update or an attack. The second is error rate per check, because a guard that fails open silently is a guard you do not have. Sample blocked and allowed records into a daily labelling queue; the labels measure precision and recall and become regression cases.
Policy changes go through an eval suite
Change guardrail policy the way you change code. Keep a versioned eval set with three slices: known-bad inputs that must be blocked (jailbreaks, injections, PII samples, harmful requests), known-good inputs that must pass (including borderline but legitimate ones like medical questions), and recent production samples labelled by reviewers. Every policy or model change runs the full set in CI and fails the build if recall drops on the bad slice or the false-positive rate rises on the good slice.
def test_policy_candidate(candidate, baseline, eval_set):
res_new = run_eval(candidate, eval_set)
res_old = run_eval(baseline, eval_set)
for category in eval_set.categories:
assert res_new.recall(category) >= res_old.recall(category) - 0.01, category
assert res_new.false_positive_rate(category) <= res_old.false_positive_rate(category) + 0.005, category
assert res_new.p95_latency_ms <= LATENCY_BUDGET_MSThe tolerances are illustrative; size them to each slice. Scanners for injection specifically, and how to evaluate them, are covered in prompt injection scanners.
Worked example: a support assistant
Take a support assistant for an online shop. It answers questions from a help-centre index, can look up the user's orders and can issue refunds up to a limit. Walk one request through it: "My card 4111 1111 1111 1111 was charged twice for order 5521, please refund one of them."
- Input stage. The PII check finds a 16-digit sequence that passes the Luhn checksum and returns redact with the span. The jailbreak classifier scores 0.03, below its threshold, and returns allow. The engine takes the most severe verdict, redacts, and the model sees "My card [CARD] was charged twice...". The log records both verdicts under policy version 14.
- Model call. The model proposes
lookup_orders(user_id=current)and thenrefund(order_id=5521, amount=39.00). - Action stage. The action guard confirms the refund tool is allowed for customers, the arguments validate, order 5521 belongs to this user, and 39.00 is under the 50.00 self-service limit. It returns allow, and the tool runs with a token scoped to this user.
- Output stage. The reply cites help-centre article 210 on duplicate charges. The grounding check confirms article 210 was retrieved and contains the quoted refund window. Schema validation passes on the structured part of the reply.
- Record. One decision record ties the four stages together. If the same request had asked for a refund on order 7730, which belongs to another account, the action guard would have blocked with reason
resource_not_ownedwhatever the model believed.
Failure modes
- Fail-open by accident. A check's client library swallows a timeout and returns an empty result that the caller reads as "no findings". Make errors explicit in the verdict and decide per risk in the policy.
- One global threshold. A single score cut-off across categories and locales over-blocks some and under-blocks others. Calibrate per category, and per language where volume allows.
- Guarding the prompt but not the context. Input checks run on the user's message but not on retrieved documents or tool results, which is where indirect injection arrives. Run injection and PII checks on every untrusted text that enters the context.
- Silent policy drift. A vendor updates a hosted classifier and block rates move overnight. Pin model versions where you can, and alert on block-rate changes per check.
Rolling out and operating guardrails
Roll out every new check in shadow mode first. The engine runs it and logs its verdict, but excludes it from the final decision. A week of shadow data shows the real block rate, the latency cost and a sample of what it would have blocked, which reviewers can label before you turn it on. Then enforce it for one tenant or a small traffic share and widen gradually.
Keep policy as versioned data, not code: per-tenant thresholds, enabled checks and approval limits, with the version stamped in every decision record, so a bad policy rolls back in seconds without a deploy.
Write the incident runbook early: how to force a check into shadow or block-all mode, and how to find every decision made under a bad policy version. For tooling that wraps validators into reusable guards, see Guardrails AI in depth.
What to do next
- Write a risk register for your product with five to ten rows, and name one primary control and one piece of evidence for each.
- Give every existing check the same verdict interface, including an explicit error field, and route them all through a single decision engine.
- Decide and record the error behaviour for each check: authorization and PII fail closed; tone or style checks may fail open.
- Add an action guard that checks tool permission, argument schema, resource ownership and approval limits in code before any tool runs.
- Emit one decision record per request with the policy version, and build dashboards for block rate and error rate per check.
- Assemble a versioned eval set with bad, good and recent production slices, and make it a required CI gate for policy changes.
- Ship the next new check in shadow mode for a week, label a sample of its would-be blocks, and only then enforce it.