Asking a language model for JSON looks like a formatting problem. It is a probability problem. The model defines a distribution over strings; your schema picks out a tiny valid subset; and every technique for getting structured output (retrying, JSON mode, grammar-constrained decoding, a classifier head) is a different way of moving probability mass onto that subset. They differ in what distribution you end up sampling from, what they cost, and whether a type error is merely unlikely or impossible.

This page works through the math. The engineering how-to for grammars and libraries is in guided decoding and structured output; here we compare parallel rejection sampling, per-step constrained decoding and non-autoregressive typed heads, show exactly where constrained decoding distorts the model, and build a typed decision model in PyTorch that cannot emit a malformed object.

The target: the model's distribution restricted to valid outputs

Let p(y | x) be the model's distribution over output strings and V the set of strings that parse and satisfy the schema. What we usually want is the model's own belief restricted to valid outputs:

p_V(y | x) = p(y | x) * 1[y in V] / Z(x),        Z(x) = sum over y in V of p(y | x)

Z(x) is the probability that an unconstrained sample is valid. Call it q. Every method below either samples from p_V exactly, approximates it, or replaces the string distribution with something else entirely. Keeping p_V as the reference point is what lets us say one method is biased and another is not.

Three routes from a model to a typed objectHidden state hafter prefillFree sampling, N in parallelvalidate, keep first validMasked autoregressive decodeautomaton state picks a token maskTyped heads, one passsoftmax per fieldExact conditionalcost 1 / P(valid) samplesAlways valid, skewedlocal renormalisationValid by constructionfields independent unless chainedOnly the bottom route has no parser: its codomain is the schema itself
The three routes share a prefill. They differ in what happens after: sample and check, mask while sampling, or skip token generation and read typed fields straight off the hidden state.

Rejection and parallel sampling: exact, priced in samples

Rejection sampling draws from p, discards invalid outputs and repeats. Each accepted sample is distributed exactly as p_V, because the acceptance test does not look at anything except validity. The price is the number of attempts, geometric with mean 1 / q.

Parallel sampling turns that cost from latency into throughput. Draw N samples in one batch, sharing the prompt's KV cache, and return the first valid one in index order. Since the samples are independent, the first valid one is still an exact draw from p_V. The probability that none is valid is (1 - q) to the power N:

q = P(valid)N = 1N = 3N = 8
0.900.100 fail0.001 fail1e-8 fail
0.500.500 fail0.125 fail0.004 fail
0.050.950 fail0.857 fail0.663 fail

The table shows where rejection works and where it does not. A strong model asked for a small schema has q near 0.9 or above, and three parallel samples make failure negligible at roughly the latency of one. A small model asked for a deeply nested schema might have q of 0.05, and no reasonable N saves it. Validation also has to be complete: if your validator checks syntax but not enum membership, "valid" means less than you think.

Constrained decoding: local renormalisation, global skew

Constrained decoding compiles the grammar into an automaton, tracks its state while generating, and at every step sets the logits of tokens that would leave the grammar to minus infinity before the softmax. With mask m_t over the vocabulary:

p~(y_t | y_<t, x) = p(y_t | y_<t, x) * m_t(y_t) / Z_t
p~(y | x)         = product over t of p~(y_t | y_<t, x)

Every output is valid, at the cost of one mask per step. But p~ is not p_V. The global conditional divides once by Z(x); constrained decoding divides at each step by a local Z_t, which only knows whether the next token is allowed, not how much valid probability lies beyond it. Prefixes that the model strongly prefers but that rarely lead to valid completions get their full prefix mass anyway, and that mass is then squeezed into whatever valid continuation remains.

A two-step example makes it concrete. The first token is A with probability 0.9 or B with 0.1. After A, the only valid next token x has probability 0.01. After B, x has probability 1.0.

true mass on valid strings:   P(A x) = 0.9 * 0.01 = 0.009     P(B x) = 0.1 * 1.0 = 0.1
exact conditional p_V:        A x -> 0.009 / 0.109 = 0.083    B x -> 0.917
constrained decoding p~:      step 1 allows both, keeps 0.9 / 0.1
                              step 2 masks to x, renormalises 0.01 -> 1.0
                              A x -> 0.900                    B x -> 0.100

The model, conditioned on producing something valid, prefers B x eleven to one. Constrained decoding returns A x nine times in ten. This is the distortion analysed by Park et al. in Grammar-Aligned Decoding (NeurIPS 2024), whose ASAp algorithm reduces it by learning approximate expected future validity from earlier samples. In JSON the effect shows up when a model wants to write a field the schema forbids, or a free-text explanation before the answer: the mask lets it start down a path and then forces an unnatural completion. Rejection sampling, by contrast, returns B x every time it succeeds.

Masks also interact with tokenisation. One vocabulary token can contain a closing quote, a comma and a space, so the masker must run the automaton over every candidate token's characters. Compilers precompute what they can: Outlines (Willard and Louf, 2023) maps each finite-automaton state to an allowed-token set, and XGrammar handles recursive grammars with a pushdown automaton, precomputing the tokens whose validity does not depend on the stack and checking the rest at run time.

JSON mode and strict schemas on the same map

Provider features sit at different points on this map. OpenAI's JSON mode guarantees the output parses as JSON but not that it matches your schema: it constrains syntax, so q for the schema is higher but not 1. OpenAI's Structured Outputs with a strict JSON schema guarantees schema conformance through constrained decoding, and in return restricts the schema language (for example, every property must be listed as required and additional properties must be disallowed). Both inherit the distortion above, and neither checks meaning: a schema-valid refund amount can still be wrong. Some studies, such as Tam et al. (2024), also report that strict format constraints can lower accuracy on reasoning tasks, which is consistent with the mass-squeezing argument: the constraint removes the path the model would have taken.

MethodValid outputSamples from p_VCost
Free sampling + validateProbability qYes, when accepted1 / q samples
N parallel + first valid1 - (1 - q)^NYesN samples, one latency
JSON modeSyntax always, schema with higher qNoMask per step
Strict schema decodingAlwaysNo, locally renormalisedMask per step + compile
Typed headsAlways, by constructionDifferent model entirelyOne forward pass

Non-autoregressive outputs and the multimodality problem

Autoregressive decoding is slow because each token waits for the previous one. Non-autoregressive models, introduced for translation by Gu et al. (2018), predict all positions at once and factor the output as a product of independent per-position distributions given x. That is fast, and it has a well-known failure: multimodality. If the true answer has two coherent modes, a factorised model can mix them.

Structured outputs hit the same wall. Suppose a payment record has country and currency, and the truth is US with USD or DE with EUR, half each. The marginals are 0.5 and 0.5 for each field, so a factorised model assigns 0.25 to US with EUR. Sampling fields independently gives an incoherent record half the time. The gap between the joint and the product of marginals is the total correlation, KL(p(a, b) || p(a) p(b)), here one bit. Whenever fields are strongly correlated given the input, independent fields lose.

The typed decision model

A typed decision model takes the fixed-shape case seriously. When the output is a record with a known schema of enums, booleans and bounded numbers, do not generate characters at all. Run the encoder or language model once, take a pooled hidden state, and attach one head per field: a softmax over the K values of each enum, a sigmoid for each boolean, and a softmax over bins for a bounded amount. The output space is the Cartesian product of the field types, so every argmax and every sample is an inhabitant of the type. There is no parser, so there is nothing to fail. That is the sense in which it cannot emit a type error: the function's codomain is the schema.

import torch, torch.nn as nn, torch.nn.functional as F

SCHEMA = {
    "intent":   ("enum", ["refund", "billing", "bug", "account", "other"]),
    "priority": ("enum", ["low", "normal", "high", "urgent"]),
    "escalate": ("bool", None),
    "amount":   ("binned", [0, 10, 25, 50, 100, 250, 500, 1000, 5000]),  # bin edges
}

class TypedDecisionHead(nn.Module):
    def __init__(self, d_model, schema=SCHEMA):
        super().__init__()
        self.schema = schema
        self.heads = nn.ModuleDict()
        for name, (kind, spec) in schema.items():
            out = {"enum": len(spec or []), "bool": 1, "binned": len(spec or []) - 1}[kind]
            self.heads[name] = nn.Linear(d_model, out)
        # chain: priority and escalate also see the predicted intent
        self.intent_emb = nn.Embedding(len(schema["intent"][1]), d_model)

    def forward(self, h, intent=None):
        logits = {"intent": self.heads["intent"](h)}
        if intent is None:                                  # inference: use own prediction
            intent = logits["intent"].argmax(-1)
        h2 = h + self.intent_emb(intent)                    # teacher forcing in training
        for name in ("priority", "escalate", "amount"):
            logits[name] = self.heads[name](h2)
        return logits

    def loss(self, logits, target):
        total = 0.0
        for name, (kind, _) in self.schema.items():
            if kind == "bool":
                total = total + F.binary_cross_entropy_with_logits(
                    logits[name].squeeze(-1), target[name].float())
            else:
                total = total + F.cross_entropy(logits[name], target[name])
        return total

    @torch.no_grad()
    def decode(self, logits):
        """One example: logits from h of shape (1, d_model) or (d_model,)."""
        logits = {k: v.reshape(v.shape[-1]) for k, v in logits.items()}
        rec, conf = {}, {}
        for name, (kind, spec) in self.schema.items():
            if kind == "bool":
                pr = torch.sigmoid(logits[name].squeeze(-1))
                rec[name], conf[name] = bool(pr > 0.5), float(torch.maximum(pr, 1 - pr))
            else:
                pr = logits[name].softmax(-1)
                i = int(pr.argmax())
                rec[name] = spec[i] if kind == "enum" else (spec[i], spec[i + 1])
                conf[name] = float(pr[i])
        return rec, conf            # rec always matches SCHEMA; conf feeds the policy layer

Two design points matter. First, the chained intent embedding is a field-level autoregression: a few sequential head evaluations rather than dozens of tokens, enough to recover the correlations that a fully factorised head would lose. Second, for a small cross-product, say intent times priority at 20 combinations, a single joint softmax with impossible combinations masked out is better still. Masking a joint softmax renormalises once over complete records, so unlike token masking it is exactly the conditional of that head's distribution.

The per-field probabilities are another advantage. They come from a softmax trained with cross-entropy, can be calibrated with temperature scaling, and plug straight into a cost-based decision policy. A number a model writes inside its JSON is not a probability of anything. How logits become probabilities is covered in the output projection and logits.

Worked example: ticket triage

Take support-ticket triage with the schema above. A generated JSON object for it runs to about 40 output tokens. At an illustrative 25 ms per decode step that is about one second after prefill, plus the risk of an invalid enum value if you sample freely. The typed head adds a few matrix multiplies to the prefill: the answer arrives with the prompt's forward pass, with a confidence per field.

What you give up is expressiveness. The head cannot answer with a new intent, explain itself, or fill a free-text field. A common split is to use typed heads for the decision fields that drive automation and let a generative call produce the human-readable explanation afterwards, conditioned on the decided values. If you must generate the whole object, prefer N parallel samples with full validation when q is high, since they sample the model's real conditional, and constrained decoding when q is low or latency rules out retries. Sampling math covers how temperature and top-p change q itself.

Failure modes

  • Valid but wrong. Every method here guarantees shape, not truth. Measure field accuracy, not parse rate.
  • Mass squeezing. Constrained decoding on a prompt where the model wants to say something the schema forbids yields confident nonsense. Watch the mean per-token log-probability of constrained outputs; a drop flags the case.
  • Incomplete validators. Rejection sampling is only as exact as the validity test. Validate against the same schema the consumer uses.
  • Factorised incoherence. Independent typed heads produce combinations that never occur together. Chain correlated fields or use a masked joint head.
  • Schema drift. Adding an enum value to a typed head means retraining or at least fine-tuning that head; with generative output it only means changing the prompt. Version the schema with the weights.
  • Grammar compile cost. Large or frequently changing schemas can make the first constrained request slow. Cache compiled grammars by schema hash.

What to do next

  1. Measure q, the unconstrained valid rate, for your schema and model on 500 real prompts.
  2. If q is above about 0.8, try 3 parallel samples with full schema validation before reaching for constrained decoding.
  3. If you use JSON mode, add schema validation; it guarantees syntax only.
  4. Compare constrained and rejection outputs on the same prompts and inspect cases where they disagree.
  5. For fixed decision fields, prototype a typed head on a pooled hidden state and compare field accuracy and latency.
  6. Chain or jointly model correlated fields; check incoherent-combination rates.
  7. Calibrate per-field probabilities and route them to a decision policy instead of parsing self-reported confidence.
Key takeaway: Structured output is a choice of distribution. Parallel rejection sampling returns the model's exact conditional on validity at a cost of 1 / q samples; constrained decoding and strict schemas always parse but renormalise locally and can skew the answer; typed heads skip token generation and make the schema the codomain, so a malformed object is impossible and only wrong values remain.