Guardrails AI is an open-source Python framework (the guardrails-ai package, version 0.11.0 on PyPI as of August 2026) that wraps a language model call in a validation loop. You declare validators for the prompt and for the answer, and for each validator you choose what happens when it fails: raise, fix the value, drop it, return nothing, or ask the model again. The result comes back as a typed outcome object that says whether validation passed and what the final, possibly repaired, output is.
That framing matters for security work. Guardrails AI is not a classifier and not a firewall; it is an orchestration layer for checks. Its protection is only as good as the validators you install, the failure actions you pick, and whether your code actually reads the outcome. This article explains the loop from first principles, builds a custom validator, walks through a worked support-bot example, and lists the failure modes that turn a guard into decoration. API details were checked against the project's repository and documentation on 2026-10-03; where versions differ, the article says so.
What the framework is and is not
Three objects carry the whole design. A Guard holds a map from a target to a list of validators and an execution policy such as the reask budget. A Validator is a class whose _validate(value, metadata) method returns either PassResult() or a FailResult carrying an error message and, optionally, a fix_value. A ValidationOutcome is what you get back: raw_llm_output, validated_output, validation_passed, and, when the reask budget ran out, the reask that would have been sent next.
Targets are named by the on argument of Guard.use. According to the docstring on the main branch the valid options are "output" (the default), "messages" for the input conversation, or a JSON path starting with $. for one field of structured output. That gives you input guards and output guards from one object, which is the main reason teams pick the framework over hand-rolled checks.
It is worth being precise about what it is not. It does not constrain token sampling the way grammar-based structured output does; it validates after generation and repairs or retries. It does not ship detection models in the core package; detectors such as PII or toxicity live in separately installed validators, many of which run their own models. And it is a different project from Protect AI's LLM Guard, which is a scanner chain with a different API.
The validation loop
A guarded call runs in a fixed order. Input validators run on the messages first; a failure here with an exception policy stops the call before any tokens are spent. The model is called, through LiteLLM when you pass a model name or through a callable you supply. The response is parsed as a string or against a schema. Output validators run, and each failure is resolved by that validator's on_fail action. If any action is a reask and budget remains, Guardrails builds a new prompt that includes the validation errors and calls the model again; for structured guards the source defaults to re-requesting the whole schema, and the reask budget bounds how often this happens.
Every iteration is recorded in guard.history, an in-memory stack whose default length in the source is 10 calls. guard.history.last.iterations lets you count reasks, which is how you measure what the guard is costing you.
Installing and composing validators
Install the core package with pip install guardrails-ai and run guardrails configure once. The current README installs each validator as its own pip package and imports it from a guardrails_ai namespace; older releases used guardrails hub install hub://guardrails/<name> and imported from guardrails.hub. Check which form your pinned version expects instead of mixing them.
# pip install guardrails-ai guardrails-ai-regex-match guardrails-ai-detect-pii
from guardrails import Guard, OnFailAction
from guardrails_ai.regex_match import RegexMatch
from guardrails_ai.detect_pii import DetectPII
guard = (
Guard()
.use(DetectPII(pii_entities="pii", on_fail=OnFailAction.EXCEPTION), on="messages")
.use(RegexMatch(regex=r"^[A-Z]{3}-\d{6}$", on_fail=OnFailAction.EXCEPTION))
)
outcome = guard.validate("ABC-123456") # validate a string you already have
print(outcome.validation_passed, outcome.validated_output)Note the instance form: on the main branch, use takes validator instances positionally plus validators= and on= keywords. The README still shows a class-plus-keyword form; passing instances works on both and keeps every argument visible. guard.validate is for text you already have, such as a cached answer or a tool result; calling the guard itself with model and messages runs the full loop.
Choosing on_fail actions
The failure action is the security decision, so choose it per validator, not per guard. The enum in the source has eight values:
| Action | What happens | Use it when |
|---|---|---|
exception | Raises; nothing is returned | Input policy violations, secrets, anything that must never pass |
fix | Returns the validator's fix_value | Deterministic repairs: redaction, trimming, case |
filter | Drops only the failing field | Optional fields in structured output |
refrain | Returns None for the output | Unsafe answers where a fallback message is better than any text |
reask | Calls the model again with the errors | Format or content errors the model can plausibly fix |
fix_reask | Fixes, revalidates, reasks only if still failing | Cheap repair first, retry as backup |
noop | Records the failure, returns the output unchanged | Shadow mode while you calibrate |
custom | Your function on_fail(value, fail_result) decides | Logging, routing to review, custom fallbacks |
The default is noop. A validator you add without an on_fail argument flags the problem and lets the output through. That is fine for shadow testing and a serious bug in production, and it is the first thing to audit in any existing deployment.
Writing a custom validator
Most real policies need at least one validator of your own. The interface is small: subclass Validator, register it with a name and data type, and return a pass or fail result. This one blocks links to domains outside an allowlist and offers a fix that strips them:
import re
from typing import Any, Dict, Optional, Callable
from guardrails.validator_base import (
FailResult, PassResult, ValidationResult, Validator, register_validator,
)
URL = re.compile(r"https?://([^/\s]+)[^\s]*", re.I)
@register_validator(name="acme/allowed_domains", data_type="string")
class AllowedDomains(Validator):
def __init__(self, allowed: list[str], on_fail: Optional[Callable] = None, **kwargs):
super().__init__(on_fail=on_fail, allowed=allowed, **kwargs)
self.allowed = {d.lower() for d in allowed}
def _ok(self, host: str) -> bool:
host = host.lower().split(":")[0]
return any(host == d or host.endswith("." + d) for d in self.allowed)
def _validate(self, value: Any, metadata: Dict[str, Any]) -> ValidationResult:
bad = [m.group(0) for m in URL.finditer(value) if not self._ok(m.group(1))]
if not bad:
return PassResult()
cleaned = URL.sub(lambda m: m.group(0) if self._ok(m.group(1)) else "[link removed]", value)
return FailResult(
error_message=f"Reply contains {len(bad)} link(s) outside the allowed domains; remove them.",
fix_value=cleaned,
)Two design rules make custom validators trustworthy. First, make fix_value strictly safer than the input, never a guess at what the model meant. Second, write the error message for the model as well as for humans: on a reask it is pasted into the next prompt, so it should describe the problem without echoing the offending output: a count of disallowed links teaches the model what to change, while a bare failed does not, and quoting the URLs would feed attacker-controlled text back in. Unit-test _validate directly with adversarial strings, including uppercase schemes, ports and look-alike subdomains such as acme.com.evil.io, which the suffix check above correctly rejects.
Structured output guards
Guard.for_pydantic(output_class=Model) builds a guard whose output must parse into a Pydantic model. Parsing failures and field validators both feed the same loop, and field validators attach with a JSON path such as on="$.reply". When a reask happens on a structured guard, the source defaults to re-requesting the whole schema; for long objects that is expensive, and you can measure the cost from the iteration history before deciding.
If your provider supports native structured output or grammar-constrained decoding, use it for shape and keep Guardrails for content. Constrained decoding makes malformed JSON impossible at generation time; Guardrails then only has to judge meaning, which is where a validator earns its latency.
Worked example: a support-reply guard
Take a customer-support assistant that drafts replies from ticket text. The policy: customers' personal data must not be sent to the model provider; replies must not link outside the company's domains; replies must carry a ticket id in a fixed format; and no answer is better than an unsafe one.
from pydantic import BaseModel, Field
from guardrails import Guard, OnFailAction
from guardrails_ai.regex_match import RegexMatch
from guardrails_ai.detect_pii import DetectPII
# AllowedDomains is the custom validator above; ESCALATE_TO_HUMAN and
# log_iterations stand in for your fallback reply and metrics call.
class Reply(BaseModel):
ticket_id: str = Field(description="Ticket id like ABC-123456")
reply: str = Field(description="Reply to the customer, plain text")
guard = (
Guard.for_pydantic(output_class=Reply)
.use(DetectPII(pii_entities="pii", on_fail=OnFailAction.EXCEPTION), on="messages")
.use(RegexMatch(regex=r"^[A-Z]{3}-\d{6}$", on_fail=OnFailAction.REASK), on="$.ticket_id")
.use(AllowedDomains(["acme.com"], on_fail=OnFailAction.FIX), on="$.reply")
)
def draft(ticket_text: str, redacted: str) -> str:
try:
out = guard(model="gpt-4o-mini", num_reasks=1,
messages=[{"role": "user", "content": redacted}])
except Exception:
return ESCALATE_TO_HUMAN # input guard fired, or reask failed hard
if not out.validation_passed or out.validated_output is None:
return ESCALATE_TO_HUMAN
log_iterations(len(guard.history.last.iterations))
return out.validated_output["reply"]Trace three tickets. A ticket where redaction missed an email address trips the input validator: an exception, no provider call, escalation. A clean ticket where the model writes ticket 4471 instead of ABC-004471 fails the regex; one reask with the error usually fixes it, at the price of a second model call, roughly doubling latency for that request. A reply that links to a third-party help page passes the regex and is repaired by the domain validator's fix with no extra call. The code reads validation_passed explicitly, because with reask budgets exhausted the outcome can carry an invalid value rather than raising.
Notice the separate redacted argument: the PII validator is a tripwire, not the redaction step. Redact deterministically before the call and let the guard catch what slipped through.
Failure modes
- Silent pass by default. Validators without
on_failusenoop. Grep for everyuse(and confirm the action is explicit. - Overwritten validators. Calling
usetwice with the sameonvalue replaces the earlier validators for that target, per the method's docstring. Pass all validators for a target in one call. - Unread outcomes. Code that takes
validated_outputwithout checkingvalidation_passedships failed answers. - Reask amplification. Each reask is another full call. Under load, a validator with a high false-positive rate multiplies token spend and tail latency; cap
num_reasksat 1 or 2 and alert on reask rate. - Reask as an injection channel. Error messages that quote the offending output feed attacker-controlled text back into the prompt. Keep messages short and descriptive.
- Model-backed validators are fallible. PII and toxicity validators carry their own false negatives and need their own dependencies, such as NLTK data the docs say pip does not fetch. Treat them as probabilistic layers, not proofs.
- Fixes that change meaning. A fix that deletes a sentence can turn a correct answer into a misleading one. Prefer
refrainover aggressive fixes for safety-relevant content.
Operating guards and the trade-offs
Run guards in-process for the lowest latency, or as a service: guardrails create scaffolds a config and guardrails start --config=./config.py serves it, with an OpenAI-compatible route at /guards/<guard-name>/openai/v1/ that existing OpenAI SDK clients can target by changing base_url. The service form centralises policy across languages and teams; the price is a network hop and one more component to scale.
Measure four numbers per guard: block rate per validator, reask rate, added latency at p95, and escalations. Start every new validator in noop against production traffic, review what it flags, then switch to the enforcing action. Pin the framework and every validator package, because validator thresholds and models change between releases, and rerun your adversarial test set on each upgrade.
Trade-offs in one line each: compared with provider safety filters, Guardrails gives you policy you own and can test; compared with constrained decoding, it handles meaning but costs retries; compared with a scanner chain like LLM Guard, it adds repair and reask at the cost of a larger API surface.
What to do next
- Write the policy as a list: what must never enter the model, what must never leave it, and what a safe fallback looks like.
- Map each rule to a validator and an explicit
on_fail:exceptionfor inputs,fixorrefrainfor outputs,reaskonly for format. - Write custom validators for business rules and unit-test them with adversarial strings.
- Deploy new validators in
noopmode and review a week of flags before enforcing. - Check
validation_passedat every call site and route failures to a fallback or a human. - Cap
num_reasksand dashboard reask rate, block rate and added p95 latency. - Pin versions and rerun the adversarial suite on every upgrade.
Related reading: LLM Guard in depth, treating model output as untrusted, PII detection and redaction, prompt injection scanners and structured output and constrained decoding.