Most teams add their first LLM guardrail after an incident. A user pastes a support transcript full of card numbers, or a chatbot quotes a refund policy that does not exist, or an agent emails the wrong customer. The fix is usually one more check bolted onto the request path. Months later there are nine checks with nine thresholds and log formats, and nobody can say which risk each one stops or what happens when one times out.

This article treats guardrails as a production control system rather than a list of filters. You start from the risks you actually carry, map each one to a control that can be tested, route every check through one decision engine that produces a single verdict and a single record, and change policy the way you change code: behind a regression suite and a staged rollout. Serving the guard models themselves (GPU placement, batching, timeouts) is covered in LLM Guardrails on GPU serving paths; the content taxonomy and human review queue are covered in LLM moderation architecture. Here the focus is the wiring, the evidence and the operations.

Start from a risk register, not a product

A guardrail is only useful if you can name the harm it prevents and show it working. So the first artefact is not code but a risk register: one row per risk, with who is harmed, how it would happen in your product, and which control owns it. It stops two common mistakes: assuming a broad content classifier covers data leakage, and stacking three classifiers on one risk while another, such as an agent deleting records, has no control at all.

RiskHow it happensPrimary controlEvidence it works
Personal data leaks into prompts or logsUsers paste records; retrieval pulls unredacted rowsDeterministic PII detection and redaction before the model and before loggingRecall on a labelled PII set per locale
Prompt injection triggers an unwanted actionInstructions hidden in a retrieved page or emailAction guard: authorization in code, argument validation, approval for risky toolsInjection suite replayed against every tool
Ungrounded answers about policy or priceModel fills gaps from training dataGrounding check against cited sources, abstain when unsupportedFaithfulness rate on a fixed question set
Harmful or off-policy contentAdversarial or careless promptsInput and output classifiers with per-category thresholdsPrecision and recall per category
Malformed output breaks a downstream systemModel drifts from the JSON contractSchema validation with one bounded repair attemptParse failure rate in production

Notice that only one of the five rows is primarily a classifier. In most production systems the controls that prevent real damage are deterministic: redaction, authorization and schema validation. Classifiers handle the fuzzy categories where no rule can be written.

Three kinds of check

Every check falls into one of three kinds, and they behave differently under load and attack.

Deterministic checks are rules and parsers: regular expressions with checksum validation for card numbers, allow-lists of domains, JSON schema validation, length and rate limits. They are fast, explainable and do not drift, but only catch what you anticipated.

Model checks are classifiers or a second LLM acting as a judge: toxicity, self-harm, jailbreak likelihood, topic relevance, faithfulness. They generalise to unanticipated inputs, but they produce scores rather than facts, have false positives, and can be probed by an attacker. Every threshold needs a labelled dataset behind it.

Authorization checks decide whether an action may happen at all, given who the user is and what the request is trying to do. They must never depend on the LLM's own judgement, because the LLM is exactly the component an injection controls. If a user cannot refund an order through the normal UI, the agent acting for them cannot either, whatever the prompt says.

Give every check the same interface so the engine can treat them uniformly:

from dataclasses import dataclass, field
from typing import Literal

Action = Literal["allow", "redact", "block", "escalate"]

@dataclass
class Verdict:
    check: str                 # "pii_regex", "jailbreak_clf", "refund_authz"
    version: str               # version of the check or its model
    action: Action
    score: float | None = None # model checks only
    reason: str = ""           # stable reason code, never free text from the LLM
    redactions: list[tuple[int, int]] = field(default_factory=list)
    latency_ms: float = 0.0
    error: str | None = None   # set when the check failed to run

The error field matters. A check that crashed has not said "allow". The engine, not the check, decides what an error means for that risk.

One decision engine, one verdict, one record

The decision engine is a small, boring piece of code with a big job: it runs the checks configured for a stage, combines their verdicts using a versioned policy, and writes one decision record. Checks never block or rewrite traffic on their own. That single point of decision is what makes the system auditable and testable.

SEVERITY = {"allow": 0, "redact": 1, "escalate": 2, "block": 3}

def decide(stage, request, checks, policy):
    verdicts = []
    for check in checks[stage]:
        try:
            v = check.run(request, timeout_ms=policy.timeout_ms[check.name])
        except Exception as exc:          # timeout or crash
            v = Verdict(check.name, check.version, "allow", error=repr(exc))
        if v.error:
            # Per-risk failure mode lives in the policy, not in the check.
            v.action = policy.on_error[check.name]   # "block" for authz, maybe "allow" for tone
        verdicts.append(v)

    final = max(verdicts, key=lambda v: SEVERITY[v.action], default=None)
    action = final.action if final else "allow"
    record = DecisionRecord(
        request_id=request.id, tenant=request.tenant, stage=stage,
        policy_version=policy.version, action=action,
        reason=final.reason if final else "",
        verdicts=verdicts,
    )
    decision_log.append(record)
    return action, record
One request through a production guardrail pipelineClient requestuser + tenantInput checksrules, classifiersLLM callprompt + contextOutput checksschema, groundingAction guardauthorize tool calltool callTool / APIruns with user's rightsallowedDecision enginemerge verdicts by policy version: allow / redact / block / escalateResponse to useror refusal + reason codeDecision logone record per requestReview + eval setsampled, labelled, replayedsampleEvery check reports to the engine; only the engine decides. The log feeds the eval set that gates the next policy change.
Input checks, the model call, output checks and the action guard all report to one decision engine, which writes the decision log that feeds review and evaluation.

The most severe verdict wins, so adding a check can never make the system more permissive. Each check's failure mode is a versioned policy decision, so "what happens when the PII service is down" has a written answer. And one record per request means block rates and per-check latency come from one log.

Guarding actions, not just text

In agentic products the expensive failures are actions, not text. An action guard sits between the model's proposed tool call and its execution, and it checks four things in code: is this tool allowed for this user and tenant, are the arguments valid and within limits, do the arguments refer only to resources this user can reach, and does this action need a human to approve it.

def guard_tool_call(user, call, policy):
    spec = TOOLS[call.name]
    if call.name not in policy.allowed_tools(user.role):
        return Verdict("tool_authz", spec.version, "block", reason="tool_not_allowed")
    args = spec.schema.validate(call.arguments)          # raises on bad types or extra keys
    if not resources.user_can_access(user, spec.resource_ids(args)):
        return Verdict("tool_authz", spec.version, "block", reason="resource_not_owned")
    if spec.risk == "irreversible" or spec.amount(args) > policy.approval_limit(user):
        return Verdict("tool_authz", spec.version, "escalate", reason="needs_approval")
    return Verdict("tool_authz", spec.version, "allow")

The resource check is the one teams skip. A model that has been told, through an injected document, to "refund order 991" will produce a perfectly valid tool call. Only a check that order 991 belongs to the current user stops it. Execute the tool with the user's own credentials or a token scoped to them, never with a service account that can reach every tenant. Broader agent-level limits, such as step budgets and kill switches, are covered in Agent guardrails.

Output checks: schema and grounding

Two output checks catch most of the failures users actually notice.

Schema validation. When output feeds code, validate it against a schema and allow at most one repair attempt. If the repaired output also fails, return a typed error; unbounded retries turn one bad answer into ten paid calls.

Grounding. For retrieval-backed answers about facts you own, such as prices, policies or account state, check that each claim is supported by a retrieved source. The cheapest version requires the model to cite source ids and verifies that each cited id was actually retrieved and that quoted spans appear in it. A judge model over each claim and source is stronger. Either way, the safe failure is to abstain ("I can't confirm that; here is the policy page") rather than to block silently. Techniques for constraining and verifying answers are covered in hallucination guardrails.

The decision log

The decision log is the evidence for tuning, audits, incidents and the eval set. Each record should hold the request id, tenant, stage, policy version, final action and reason code, plus every check's name, version, action, score, latency and error. Hold redacted text, or a pointer to the raw text in a store with tighter access and a short retention period, but never raw personal data in the main log.

Two metrics from the log deserve dashboards. The first is block rate per check and per tenant, because a sudden change almost always means a deploy, a model update or an attack. The second is error rate per check, because a guard that fails open silently is a guard you do not have. Sample blocked and allowed records into a daily labelling queue; the labels measure precision and recall and become regression cases.

Policy changes go through an eval suite

Change guardrail policy the way you change code. Keep a versioned eval set with three slices: known-bad inputs that must be blocked (jailbreaks, injections, PII samples, harmful requests), known-good inputs that must pass (including borderline but legitimate ones like medical questions), and recent production samples labelled by reviewers. Every policy or model change runs the full set in CI and fails the build if recall drops on the bad slice or the false-positive rate rises on the good slice.

def test_policy_candidate(candidate, baseline, eval_set):
    res_new = run_eval(candidate, eval_set)
    res_old = run_eval(baseline, eval_set)
    for category in eval_set.categories:
        assert res_new.recall(category) >= res_old.recall(category) - 0.01, category
        assert res_new.false_positive_rate(category) <= res_old.false_positive_rate(category) + 0.005, category
    assert res_new.p95_latency_ms <= LATENCY_BUDGET_MS

The tolerances are illustrative; size them to each slice. Scanners for injection specifically, and how to evaluate them, are covered in prompt injection scanners.

Worked example: a support assistant

Take a support assistant for an online shop. It answers questions from a help-centre index, can look up the user's orders and can issue refunds up to a limit. Walk one request through it: "My card 4111 1111 1111 1111 was charged twice for order 5521, please refund one of them."

  1. Input stage. The PII check finds a 16-digit sequence that passes the Luhn checksum and returns redact with the span. The jailbreak classifier scores 0.03, below its threshold, and returns allow. The engine takes the most severe verdict, redacts, and the model sees "My card [CARD] was charged twice...". The log records both verdicts under policy version 14.
  2. Model call. The model proposes lookup_orders(user_id=current) and then refund(order_id=5521, amount=39.00).
  3. Action stage. The action guard confirms the refund tool is allowed for customers, the arguments validate, order 5521 belongs to this user, and 39.00 is under the 50.00 self-service limit. It returns allow, and the tool runs with a token scoped to this user.
  4. Output stage. The reply cites help-centre article 210 on duplicate charges. The grounding check confirms article 210 was retrieved and contains the quoted refund window. Schema validation passes on the structured part of the reply.
  5. Record. One decision record ties the four stages together. If the same request had asked for a refund on order 7730, which belongs to another account, the action guard would have blocked with reason resource_not_owned whatever the model believed.

Failure modes

  • Fail-open by accident. A check's client library swallows a timeout and returns an empty result that the caller reads as "no findings". Make errors explicit in the verdict and decide per risk in the policy.
  • One global threshold. A single score cut-off across categories and locales over-blocks some and under-blocks others. Calibrate per category, and per language where volume allows.
  • Guarding the prompt but not the context. Input checks run on the user's message but not on retrieved documents or tool results, which is where indirect injection arrives. Run injection and PII checks on every untrusted text that enters the context.
  • Silent policy drift. A vendor updates a hosted classifier and block rates move overnight. Pin model versions where you can, and alert on block-rate changes per check.

Rolling out and operating guardrails

Roll out every new check in shadow mode first. The engine runs it and logs its verdict, but excludes it from the final decision. A week of shadow data shows the real block rate, the latency cost and a sample of what it would have blocked, which reviewers can label before you turn it on. Then enforce it for one tenant or a small traffic share and widen gradually.

Keep policy as versioned data, not code: per-tenant thresholds, enabled checks and approval limits, with the version stamped in every decision record, so a bad policy rolls back in seconds without a deploy.

Write the incident runbook early: how to force a check into shadow or block-all mode, and how to find every decision made under a bad policy version. For tooling that wraps validators into reusable guards, see Guardrails AI in depth.

What to do next

  1. Write a risk register for your product with five to ten rows, and name one primary control and one piece of evidence for each.
  2. Give every existing check the same verdict interface, including an explicit error field, and route them all through a single decision engine.
  3. Decide and record the error behaviour for each check: authorization and PII fail closed; tone or style checks may fail open.
  4. Add an action guard that checks tool permission, argument schema, resource ownership and approval limits in code before any tool runs.
  5. Emit one decision record per request with the policy version, and build dashboards for block rate and error rate per check.
  6. Assemble a versioned eval set with bad, good and recent production slices, and make it a required CI gate for policy changes.
  7. Ship the next new check in shadow mode for a week, label a sample of its would-be blocks, and only then enforce it.
Key takeaway: Production guardrails work when they are a control system rather than a pile of filters: every risk maps to a testable control, every check reports to one decision engine that writes one record, actions are authorized in code rather than by the model, and policy changes ship behind an eval suite and shadow mode like any other code change.