An LLM produces a sequence of tokens. Your code needs a value: a JSON object that matches a schema, a SQL statement in a known dialect, or one of five labels. Output validation is the work of closing that gap, and on a GPU serving stack it happens in two places. Constrained decoding works inside the decode loop. It masks the logits at every step so the model can only emit tokens that keep the output grammatical. Post-hoc validation runs after generation. It parses the result and checks what a grammar cannot express: ranges, cross-field rules, references to real records.

This article explains how each layer works, what it costs in GPU and CPU time, and where it breaks. Classifier-style safety checks are a different problem, covered in LLM guardrails on GPU.

Three layers: syntax, schema, semantics

It helps to separate three questions, because each one is answered by a different mechanism:

LayerQuestionMechanismWhere it runs
SyntaxDoes it parse as JSON, SQL, or the regex?Grammar-constrained decoding, or a parser after the factMask on CPU, applied on GPU per step
SchemaRight keys, types, enums, required fields?JSON Schema compiled to a grammar; a schema validator afterwardsSame as above, plus CPU after generation
SemanticsIs the value true and allowed?Code: range checks, lookups, cross-field rules, a second modelCPU or another service, after generation

Constrained decoding can guarantee the first two layers, at least for the subset of JSON Schema the backend supports and as long as generation is not truncated. It cannot help with the third. A model forced into {"amount": <number>} will always produce a number, but nothing in the grammar makes it the right number. Teams that turn on structured outputs and delete their validators have traded a visible failure (a parse error) for an invisible one (a plausible, wrong value).

How a token mask constrains decoding

At each decode step the model produces a logit vector over the whole vocabulary, about 128K entries for Llama 3-family tokenisers. The grammar engine tracks where each sequence is in the grammar and computes the set of token ids that are legal next. Every other logit is set to −∞ before sampling, so the softmax gives those tokens zero probability. After sampling, the engine consumes the chosen token and advances its state. Temperature, top-p and the rest of the sampler work as normal on the surviving tokens.

One decode step with constrained decodingGPU: forward passlogits [batch, vocab]CPU: grammar engineper-sequence state -> bitmaskruns while the GPU computesbitmask (vocab/32 int32)Apply maskdisallowed logits = -infSample tokentemperature / top-p as usualtoken idAdvance grammar stateaccept token, next allowed setAfter generation: post-hoc validatorparse, schema, business rulesOn failurerepair, re-ask, or reject (bounded)EOS or max_tokens
The constrained decode loop. The grammar engine computes the next step's allowed set on the CPU while the GPU runs the forward pass. The mask is applied just before sampling, and a post-hoc validator checks what the grammar cannot.

The masking step itself is cheap. In PyTorch it is a gather and a fill:

import torch

def apply_token_bitmask(logits: torch.Tensor, bitmask: torch.Tensor) -> None:
    """logits: [B, V] float; bitmask: [B, ceil(V/32)] int32, bit=1 means allowed."""
    B, V = logits.shape
    ids = torch.arange(V, device=logits.device)
    words = bitmask[:, ids // 32]                       # [B, V] int32
    allowed = (words >> (ids % 32)) & 1                 # [B, V] 0/1
    logits.masked_fill_(allowed == 0, float("-inf"))

def constrained_step(model_logits, grammar_states, sampler):
    mask = torch.stack([g.next_bitmask() for g in grammar_states]).to(model_logits.device)
    apply_token_bitmask(model_logits, mask)
    next_ids = sampler(model_logits)                     # [B]
    for g, t in zip(grammar_states, next_ids.tolist()):
        g.accept(t)                                      # advance the automaton
    return next_ids

The bitmask packs one bit per token. For a 128K vocabulary that is about 4,000 int32 words, or 16 KB per sequence per step. Even at batch 256 the host-to-device copy is around 4 MB per step, which is small next to the memory traffic of the forward pass (see the decode compute ledger). Production engines use a fused kernel instead of the gather above, but the semantics are the same.

Where the cost goes: computing the mask

The expensive part is computing the mask, not applying it. Deciding which of 128K tokens are legal means checking each token's byte string against the automaton's current state. Done naively, that costs milliseconds per sequence per step, which is as long as the forward pass itself. Three engineering ideas make it workable:

  • Overlap with the forward pass. The mask for step t + 1 depends only on the tokens up to t, so the CPU can compute it while the GPU runs step t + 1's forward pass. As long as mask computation finishes first, its cost is hidden. When it does not finish first, every sequence in the batch waits. One slow grammar can stall a whole continuous batch.
  • Precomputation. Grammar engines such as XGrammar compile the grammar once per schema and cache, for each automaton state, which tokens can be decided without looking at the parser stack. Only the remaining tokens are checked at run time. So the first request with a new schema pays a compile cost, and later requests reuse the cache. If every request carries a slightly different schema, you pay it every time.
  • Jump-forward decoding. When the grammar allows exactly one continuation, such as the closing "} after the last required field or a fixed key name, the engine can append those tokens without running the model one step per token. SGLang's compressed finite-state machine work popularised this. It saves decode steps on schemas with long fixed keys, but the forced tokens must be retokenised carefully (next section).

The practical consequence is that structured output throughput depends on the grammar as well as the model. Deeply nested schemas, large enums and unbounded regex repetitions all raise mask cost. Measure tokens per second with and without constraints on your real schemas at your real batch size before you assume the overhead is free. See continuous batching for why one stalled step costs every request in the batch.

Structured outputs in vLLM

In current vLLM (the guided_* fields were removed in v0.12.0), constraints go in a structured_outputs object. On the OpenAI-compatible server it is passed through extra_body, or you use the standard response_format with a JSON schema. The backend (xgrammar, guidance, or the default auto) is chosen on the server. Check your version's documentation, because these names have changed before:

from openai import OpenAI
from pydantic import BaseModel, Field

class Invoice(BaseModel):
    invoice_id: str = Field(pattern=r"^INV-[0-9]{6}$")
    currency: str = Field(pattern=r"^[A-Z]{3}$")
    total_cents: int = Field(ge=0)
    line_count: int = Field(ge=1, le=200)

client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")

resp = client.chat.completions.create(
    model="served-model",
    messages=[{"role": "user", "content": "Extract the invoice fields:\n" + doc}],
    response_format={
        "type": "json_schema",
        "json_schema": {"name": "invoice", "schema": Invoice.model_json_schema()},
    },
    max_tokens=400,
)

# Label classification with a closed set, via the vLLM-specific extra_body:
label = client.chat.completions.create(
    model="served-model",
    messages=[{"role": "user", "content": ticket_text}],
    extra_body={"structured_outputs": {"choice": ["billing", "outage", "account", "other"]}},
    max_tokens=5,
).choices[0].message.content

Offline, the same options are a StructuredOutputsParams inside SamplingParams(structured_outputs=...). Two details matter in production. First, not every JSON Schema keyword is enforced by every backend. String formats, minItems/maxItems on arrays and some pattern syntax vary, so treat the post-hoc schema check as mandatory, not belt and braces. Second, regex syntax depends on the backend: the vLLM documentation notes that xgrammar, guidance and outlines use Rust-style regex.

The arithmetic of retries

The alternative to constraining is to generate freely, validate, and retry on failure. If one attempt passes with probability p and attempts are independent, the expected number of attempts is 1/p, and the chance of still failing after k attempts is (1 − p)ᵏ. The cost is not only tokens. Each retry adds a full generation to the latency of the request that failed, so the tail moves much more than the mean:

Pass rate pMean attemptsFail after 3 attemptsEffect on the requests that retried
0.991.011 in 10⁶1% of requests take about 2× as long
0.951.051.25 in 10⁴5% take 2×; p99 latency roughly doubles
0.801.250.8%p95 and p99 are 2–3×; 25% more GPU tokens
0.502.0012.5%unusable without constraints or a better prompt

Attempts are rarely independent. A document that confused the model once will usually confuse it again at the same temperature, so the real failure-after-k rate is worse than the table. That is why retries should change something: append the validator's error message, raise the reasoning budget, or route to a larger model. A useful rule is to constrain whatever a grammar can express, retry only on semantic failures, and cap retries at two.

Semantic validation after generation

Post-hoc validation is ordinary code. Order the checks from cheapest to most expensive and return machine-readable errors, so the retry prompt can quote them:

from decimal import Decimal
from pydantic import ValidationError

class OutputRejected(Exception):
    def __init__(self, errors: list[str], retryable: bool):
        super().__init__("; ".join(errors))
        self.errors, self.retryable = errors, retryable

def validate_invoice(raw: str, source_text: str, finish_reason: str) -> Invoice:
    if finish_reason == "length":
        raise OutputRejected(["truncated at max_tokens"], retryable=True)
    try:
        inv = Invoice.model_validate_json(raw)          # syntax + schema
    except ValidationError as e:
        raise OutputRejected([err["msg"] for err in e.errors()], retryable=True)
    errors = []
    if inv.invoice_id not in source_text:              # grounding: copied, not invented
        errors.append(f"invoice_id {inv.invoice_id} does not appear in the document")
    if inv.currency not in {"USD", "EUR", "GBP", "INR"}:
        errors.append(f"currency {inv.currency} not accepted")
    printed = find_printed_total(source_text)          # deterministic parser, may be None
    if printed is not None and Decimal(inv.total_cents) != printed * 100:
        errors.append("total disagrees with the printed total")
    if errors:
        raise OutputRejected(errors, retryable=len(errors) == 1)
    return inv

The grounding check (does the extracted ID literally appear in the input?) catches most invented values for very little cost. Cross-checking against a deterministic parser wherever one exists turns the model into a second opinion instead of the only source. For streaming responses, an incremental JSON parser can reject a structurally broken stream early, but semantic checks need the full object. Do not show fields to a user before the object they belong to has been validated.

Worked example: an extraction service

Here is an illustrative case; the figures are an example, not a benchmark. An invoice-extraction service runs at batch 64 and produces about 150 output tokens per request. Unconstrained, with a good prompt, 93% of outputs parse and match the schema. The 7% that fail are retried, adding about 7.5% more output tokens and doubling latency for those requests. Turning on JSON-schema constraints raises the schema pass rate to effectively 100% for requests that finish (finish_reason == 'stop'). Two new problems appear in the logs.

First, about 0.5% of requests now hit max_tokens. The schema let a free-text notes field grow without limit, and the model filled it with repeated whitespace and filler. A grammar guarantees that output can still become valid, not that it ends. The fix is a maxLength on every string, a whitespace pattern that bans runs of whitespace, and a max_tokens sized from the schema rather than a default. Second, semantic rejections stayed near 3%, mostly totals that disagreed with the printed total. The grammar had removed the parse failures and left those errors exactly where they were. That 3% is what the retry budget is actually for.

Failure modes and controls

FailureSymptomControl
TruncationValid prefix, invalid object; finish_reason = lengthBound every string and array; size max_tokens from the schema; never parse a truncated output as success
Whitespace or repetition runawayRequests hit max_tokens with padded outputWhitespace pattern; maxLength; repetition penalty; alert on the length-finish rate
Distribution distortionQuality drops when constrained, especially on hard fieldsLet the model reason in an unconstrained field first, or generate then constrain; evaluate both ways
Token-boundary forcingOdd tokenisation around forced text; lower quality after jump-forwardUse engines that retokenise forced spans; keep fixed keys short
Unsupported schema keywordsConstraint silently not enforcedAlways run the full schema validator afterwards; test each keyword you rely on
Grammar compile stallLatency spike on the first request per schemaUse a small set of stable schemas; warm caches at deploy time
Mask computation lagThroughput falls at high batch with complex grammarsProfile CPU time per step; simplify the grammar; give the engine dedicated cores
Valid but wrongClean JSON, wrong valuesGrounding checks, deterministic cross-checks, sampled human review

Operating validation in production

  • Track four rates per schema and model version: parse failures, schema failures, semantic failures and length finishes. Each points to a different fix.
  • Log the validator's errors, the finish reason and the attempt number with each request ID. The failure-after-retry population is your evaluation set.
  • Version schemas and pin them per deployment. A schema change is a model-behaviour change and deserves a canary.
  • Load-test with constraints on, at production batch size, and watch CPU utilisation on the serving host. Mask computation is the new bottleneck, and it does not appear in GPU metrics.
  • Put validation behind the same gateway that owns retries and budgets (see LLM gateway architecture), so the retry policy is in one place.
  • For tool calls, validate arguments with the same pipeline before execution. Tool-use failure modes lists what slips through when you only check syntax.

Trade-offs

Constrained decoding buys structural correctness for some CPU per step, a compile cost per schema, and occasional quality loss when the grammar pushes the model away from the tokens it would have chosen. Unconstrained generation with validation keeps the model's natural distribution but turns every failure into extra latency and GPU tokens. The usual best answer is both: constrain the shape, validate the meaning, and keep a small, bounded retry budget for what remains.

What to do next

  1. List every place your system parses model output, and record which of the three layers is checked at each one.
  2. Move each output with a fixed shape to response_format or structured_outputs, and add maxLength and maxItems bounds to the schema.
  3. Write a post-hoc validator that rejects truncated outputs, re-runs the full schema and adds at least one grounding check.
  4. Measure throughput and p99 with and without constraints at your production batch size, while watching host CPU.
  5. Cap retries at two and make each retry change something: add the error, use a bigger model, or allow more reasoning.
  6. Add dashboards for parse, schema, semantic and length-finish rates, broken down by schema version.
Key takeaway: Validate LLM output in layers. Constrained decoding masks logits at every step and guarantees syntax and schema for outputs that finish, at the cost of CPU time that must hide behind the forward pass. Post-hoc code checks truncation, the full schema and the meaning of the values. Retries are for semantic failures only, bounded, and each one should change something.