An LLM produces a sequence of tokens. Your code needs a value: a JSON object that matches a schema, a SQL statement in a known dialect, or one of five labels. Output validation is the work of closing that gap, and on a GPU serving stack it happens in two places. Constrained decoding works inside the decode loop. It masks the logits at every step so the model can only emit tokens that keep the output grammatical. Post-hoc validation runs after generation. It parses the result and checks what a grammar cannot express: ranges, cross-field rules, references to real records.
This article explains how each layer works, what it costs in GPU and CPU time, and where it breaks. Classifier-style safety checks are a different problem, covered in LLM guardrails on GPU.
Three layers: syntax, schema, semantics
It helps to separate three questions, because each one is answered by a different mechanism:
| Layer | Question | Mechanism | Where it runs |
|---|---|---|---|
| Syntax | Does it parse as JSON, SQL, or the regex? | Grammar-constrained decoding, or a parser after the fact | Mask on CPU, applied on GPU per step |
| Schema | Right keys, types, enums, required fields? | JSON Schema compiled to a grammar; a schema validator afterwards | Same as above, plus CPU after generation |
| Semantics | Is the value true and allowed? | Code: range checks, lookups, cross-field rules, a second model | CPU or another service, after generation |
Constrained decoding can guarantee the first two layers, at least for the subset of JSON Schema the backend supports and as long as generation is not truncated. It cannot help with the third. A model forced into {"amount": <number>} will always produce a number, but nothing in the grammar makes it the right number. Teams that turn on structured outputs and delete their validators have traded a visible failure (a parse error) for an invisible one (a plausible, wrong value).
How a token mask constrains decoding
At each decode step the model produces a logit vector over the whole vocabulary, about 128K entries for Llama 3-family tokenisers. The grammar engine tracks where each sequence is in the grammar and computes the set of token ids that are legal next. Every other logit is set to −∞ before sampling, so the softmax gives those tokens zero probability. After sampling, the engine consumes the chosen token and advances its state. Temperature, top-p and the rest of the sampler work as normal on the surviving tokens.
The masking step itself is cheap. In PyTorch it is a gather and a fill:
import torch
def apply_token_bitmask(logits: torch.Tensor, bitmask: torch.Tensor) -> None:
"""logits: [B, V] float; bitmask: [B, ceil(V/32)] int32, bit=1 means allowed."""
B, V = logits.shape
ids = torch.arange(V, device=logits.device)
words = bitmask[:, ids // 32] # [B, V] int32
allowed = (words >> (ids % 32)) & 1 # [B, V] 0/1
logits.masked_fill_(allowed == 0, float("-inf"))
def constrained_step(model_logits, grammar_states, sampler):
mask = torch.stack([g.next_bitmask() for g in grammar_states]).to(model_logits.device)
apply_token_bitmask(model_logits, mask)
next_ids = sampler(model_logits) # [B]
for g, t in zip(grammar_states, next_ids.tolist()):
g.accept(t) # advance the automaton
return next_idsThe bitmask packs one bit per token. For a 128K vocabulary that is about 4,000 int32 words, or 16 KB per sequence per step. Even at batch 256 the host-to-device copy is around 4 MB per step, which is small next to the memory traffic of the forward pass (see the decode compute ledger). Production engines use a fused kernel instead of the gather above, but the semantics are the same.
Where the cost goes: computing the mask
The expensive part is computing the mask, not applying it. Deciding which of 128K tokens are legal means checking each token's byte string against the automaton's current state. Done naively, that costs milliseconds per sequence per step, which is as long as the forward pass itself. Three engineering ideas make it workable:
- Overlap with the forward pass. The mask for step t + 1 depends only on the tokens up to t, so the CPU can compute it while the GPU runs step t + 1's forward pass. As long as mask computation finishes first, its cost is hidden. When it does not finish first, every sequence in the batch waits. One slow grammar can stall a whole continuous batch.
- Precomputation. Grammar engines such as XGrammar compile the grammar once per schema and cache, for each automaton state, which tokens can be decided without looking at the parser stack. Only the remaining tokens are checked at run time. So the first request with a new schema pays a compile cost, and later requests reuse the cache. If every request carries a slightly different schema, you pay it every time.
- Jump-forward decoding. When the grammar allows exactly one continuation, such as the closing
"}after the last required field or a fixed key name, the engine can append those tokens without running the model one step per token. SGLang's compressed finite-state machine work popularised this. It saves decode steps on schemas with long fixed keys, but the forced tokens must be retokenised carefully (next section).
The practical consequence is that structured output throughput depends on the grammar as well as the model. Deeply nested schemas, large enums and unbounded regex repetitions all raise mask cost. Measure tokens per second with and without constraints on your real schemas at your real batch size before you assume the overhead is free. See continuous batching for why one stalled step costs every request in the batch.
Structured outputs in vLLM
In current vLLM (the guided_* fields were removed in v0.12.0), constraints go in a structured_outputs object. On the OpenAI-compatible server it is passed through extra_body, or you use the standard response_format with a JSON schema. The backend (xgrammar, guidance, or the default auto) is chosen on the server. Check your version's documentation, because these names have changed before:
from openai import OpenAI
from pydantic import BaseModel, Field
class Invoice(BaseModel):
invoice_id: str = Field(pattern=r"^INV-[0-9]{6}$")
currency: str = Field(pattern=r"^[A-Z]{3}$")
total_cents: int = Field(ge=0)
line_count: int = Field(ge=1, le=200)
client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")
resp = client.chat.completions.create(
model="served-model",
messages=[{"role": "user", "content": "Extract the invoice fields:\n" + doc}],
response_format={
"type": "json_schema",
"json_schema": {"name": "invoice", "schema": Invoice.model_json_schema()},
},
max_tokens=400,
)
# Label classification with a closed set, via the vLLM-specific extra_body:
label = client.chat.completions.create(
model="served-model",
messages=[{"role": "user", "content": ticket_text}],
extra_body={"structured_outputs": {"choice": ["billing", "outage", "account", "other"]}},
max_tokens=5,
).choices[0].message.contentOffline, the same options are a StructuredOutputsParams inside SamplingParams(structured_outputs=...). Two details matter in production. First, not every JSON Schema keyword is enforced by every backend. String formats, minItems/maxItems on arrays and some pattern syntax vary, so treat the post-hoc schema check as mandatory, not belt and braces. Second, regex syntax depends on the backend: the vLLM documentation notes that xgrammar, guidance and outlines use Rust-style regex.
The arithmetic of retries
The alternative to constraining is to generate freely, validate, and retry on failure. If one attempt passes with probability p and attempts are independent, the expected number of attempts is 1/p, and the chance of still failing after k attempts is (1 − p)ᵏ. The cost is not only tokens. Each retry adds a full generation to the latency of the request that failed, so the tail moves much more than the mean:
| Pass rate p | Mean attempts | Fail after 3 attempts | Effect on the requests that retried |
|---|---|---|---|
| 0.99 | 1.01 | 1 in 10⁶ | 1% of requests take about 2× as long |
| 0.95 | 1.05 | 1.25 in 10⁴ | 5% take 2×; p99 latency roughly doubles |
| 0.80 | 1.25 | 0.8% | p95 and p99 are 2–3×; 25% more GPU tokens |
| 0.50 | 2.00 | 12.5% | unusable without constraints or a better prompt |
Attempts are rarely independent. A document that confused the model once will usually confuse it again at the same temperature, so the real failure-after-k rate is worse than the table. That is why retries should change something: append the validator's error message, raise the reasoning budget, or route to a larger model. A useful rule is to constrain whatever a grammar can express, retry only on semantic failures, and cap retries at two.
Semantic validation after generation
Post-hoc validation is ordinary code. Order the checks from cheapest to most expensive and return machine-readable errors, so the retry prompt can quote them:
from decimal import Decimal
from pydantic import ValidationError
class OutputRejected(Exception):
def __init__(self, errors: list[str], retryable: bool):
super().__init__("; ".join(errors))
self.errors, self.retryable = errors, retryable
def validate_invoice(raw: str, source_text: str, finish_reason: str) -> Invoice:
if finish_reason == "length":
raise OutputRejected(["truncated at max_tokens"], retryable=True)
try:
inv = Invoice.model_validate_json(raw) # syntax + schema
except ValidationError as e:
raise OutputRejected([err["msg"] for err in e.errors()], retryable=True)
errors = []
if inv.invoice_id not in source_text: # grounding: copied, not invented
errors.append(f"invoice_id {inv.invoice_id} does not appear in the document")
if inv.currency not in {"USD", "EUR", "GBP", "INR"}:
errors.append(f"currency {inv.currency} not accepted")
printed = find_printed_total(source_text) # deterministic parser, may be None
if printed is not None and Decimal(inv.total_cents) != printed * 100:
errors.append("total disagrees with the printed total")
if errors:
raise OutputRejected(errors, retryable=len(errors) == 1)
return invThe grounding check (does the extracted ID literally appear in the input?) catches most invented values for very little cost. Cross-checking against a deterministic parser wherever one exists turns the model into a second opinion instead of the only source. For streaming responses, an incremental JSON parser can reject a structurally broken stream early, but semantic checks need the full object. Do not show fields to a user before the object they belong to has been validated.
Worked example: an extraction service
Here is an illustrative case; the figures are an example, not a benchmark. An invoice-extraction service runs at batch 64 and produces about 150 output tokens per request. Unconstrained, with a good prompt, 93% of outputs parse and match the schema. The 7% that fail are retried, adding about 7.5% more output tokens and doubling latency for those requests. Turning on JSON-schema constraints raises the schema pass rate to effectively 100% for requests that finish (finish_reason == 'stop'). Two new problems appear in the logs.
First, about 0.5% of requests now hit max_tokens. The schema let a free-text notes field grow without limit, and the model filled it with repeated whitespace and filler. A grammar guarantees that output can still become valid, not that it ends. The fix is a maxLength on every string, a whitespace pattern that bans runs of whitespace, and a max_tokens sized from the schema rather than a default. Second, semantic rejections stayed near 3%, mostly totals that disagreed with the printed total. The grammar had removed the parse failures and left those errors exactly where they were. That 3% is what the retry budget is actually for.
Failure modes and controls
| Failure | Symptom | Control |
|---|---|---|
| Truncation | Valid prefix, invalid object; finish_reason = length | Bound every string and array; size max_tokens from the schema; never parse a truncated output as success |
| Whitespace or repetition runaway | Requests hit max_tokens with padded output | Whitespace pattern; maxLength; repetition penalty; alert on the length-finish rate |
| Distribution distortion | Quality drops when constrained, especially on hard fields | Let the model reason in an unconstrained field first, or generate then constrain; evaluate both ways |
| Token-boundary forcing | Odd tokenisation around forced text; lower quality after jump-forward | Use engines that retokenise forced spans; keep fixed keys short |
| Unsupported schema keywords | Constraint silently not enforced | Always run the full schema validator afterwards; test each keyword you rely on |
| Grammar compile stall | Latency spike on the first request per schema | Use a small set of stable schemas; warm caches at deploy time |
| Mask computation lag | Throughput falls at high batch with complex grammars | Profile CPU time per step; simplify the grammar; give the engine dedicated cores |
| Valid but wrong | Clean JSON, wrong values | Grounding checks, deterministic cross-checks, sampled human review |
Operating validation in production
- Track four rates per schema and model version: parse failures, schema failures, semantic failures and length finishes. Each points to a different fix.
- Log the validator's errors, the finish reason and the attempt number with each request ID. The failure-after-retry population is your evaluation set.
- Version schemas and pin them per deployment. A schema change is a model-behaviour change and deserves a canary.
- Load-test with constraints on, at production batch size, and watch CPU utilisation on the serving host. Mask computation is the new bottleneck, and it does not appear in GPU metrics.
- Put validation behind the same gateway that owns retries and budgets (see LLM gateway architecture), so the retry policy is in one place.
- For tool calls, validate arguments with the same pipeline before execution. Tool-use failure modes lists what slips through when you only check syntax.
Trade-offs
Constrained decoding buys structural correctness for some CPU per step, a compile cost per schema, and occasional quality loss when the grammar pushes the model away from the tokens it would have chosen. Unconstrained generation with validation keeps the model's natural distribution but turns every failure into extra latency and GPU tokens. The usual best answer is both: constrain the shape, validate the meaning, and keep a small, bounded retry budget for what remains.
What to do next
- List every place your system parses model output, and record which of the three layers is checked at each one.
- Move each output with a fixed shape to
response_formatorstructured_outputs, and addmaxLengthandmaxItemsbounds to the schema. - Write a post-hoc validator that rejects truncated outputs, re-runs the full schema and adds at least one grounding check.
- Measure throughput and p99 with and without constraints at your production batch size, while watching host CPU.
- Cap retries at two and make each retry change something: add the error, use a bigger model, or allow more reasoning.
- Add dashboards for parse, schema, semantic and length-finish rates, broken down by schema version.