Every major model API can now constrain a response to a JSON Schema. That feature removes a whole class of bugs: no more unbalanced braces, no stray prose before the JSON, no missing required keys. It also creates a new and quieter class of bugs, because a response that is perfectly valid against the schema can still be wrong. The total can disagree with the line items. A date can be invented. A field can be confidently filled when the document never mentioned it.
This article treats structured output as a contract between a model and its consumer and follows it through its life: one schema source, compiled per provider, checks before and after parsing, retry policy and versioning. The mechanism underneath, constrained decoding, is covered in structured output and constrained decoding; here the focus is the engineering around it, with a worked invoice-extraction example in Python.
Three levels of guarantee
It helps to separate what a structured response can promise. Syntax: the text parses as JSON. Shape: the parsed value matches the schema, so required keys exist, types are right and enums hold. Meaning: the values are correct for this input. Prompt-only JSON gives you none of these reliably. A provider's JSON mode gives you syntax. Strict schema-constrained output gives you syntax and most of shape, within the subset of JSON Schema the provider can enforce. Nothing at the decoding layer gives you meaning.
So there are always two halves: the request side, which makes the model emit the right shape, and the acceptance side, which decides whether to trust it. Deleting validation code after adopting strict mode throws away the second half.
The architecture
The schema lives once, as code. A compiler emits the strict schema for the API plus the constraints the decoder cannot enforce. After the call, a gate checks why generation stopped, the full schema validates the object, and semantic checks compare it with the source. Failures go to a bounded retry or a review queue, never silently to the consumer.
One source of truth
Write the schema in your application language, not as a JSON file that drifts from the types your code uses. In Python that is usually a Pydantic model; in TypeScript, Zod. The model below is the running example: an invoice extractor.
from datetime import date
from decimal import Decimal
from enum import Enum
from typing import Literal, Optional
from pydantic import BaseModel, ConfigDict, Field
class Currency(str, Enum):
USD = "USD"; EUR = "EUR"; GBP = "GBP"; INR = "INR"; OTHER = "OTHER"
class LineItem(BaseModel):
model_config = ConfigDict(extra="forbid")
description: str
quantity: Decimal = Field(gt=0)
unit_price: Decimal = Field(ge=0)
amount: Decimal = Field(ge=0)
class Invoice(BaseModel):
model_config = ConfigDict(extra="forbid")
schema_version: Literal["invoice.v2"]
vendor_name: str
invoice_number: Optional[str] = Field(description="null if not printed on the document")
invoice_date: Optional[date]
currency: Currency
line_items: list[LineItem]
total: Decimal = Field(ge=0)
total_evidence: str = Field(description="exact text from the document showing the total")Notice four deliberate choices. Every object forbids extra keys. Optional facts are typed as nullable, with the schema description telling the model to use null rather than guess. The currency enum has an OTHER escape hatch so the model is never forced to pick a wrong value. And total_evidence asks for the exact text that supports the most important number, which makes the semantic check below possible. The constant schema_version field travels with every output and makes stored results self-describing.
Compiling to the strict subset
Strict modes accept a subset of JSON Schema, and the subsets differ. Anthropic's documentation for Claude's structured outputs lists support for the basic types, enum, const, anyOf, allOf, $ref with definitions and common string formats such as date and email; it requires additionalProperties: false on objects and does not support recursive schemas, numeric constraints such as minimum, or string length constraints. OpenAI's strict mode requires additionalProperties: false everywhere and every property listed in required, with optional values expressed as a union with null. Both providers change these lists over time, so re-read the current documentation when you upgrade.
The portable strategy is to compile to the intersection: forbid extra properties, mark every key required and model absence as null, and strip constraints the decoder cannot enforce while recording them so they are checked after the call.
UNSUPPORTED = {"minimum", "maximum", "exclusiveMinimum", "exclusiveMaximum",
"multipleOf", "minLength", "maxLength", "pattern", "minItems", "maxItems"}
def to_strict(node, path="$", stripped=None):
"""Return (strict_schema, stripped_rules). Conservative: the intersection of strict modes."""
stripped = [] if stripped is None else stripped
if isinstance(node, list):
return [to_strict(n, path, stripped)[0] for n in node], stripped
if not isinstance(node, dict):
return node, stripped
out = {}
for key, value in node.items():
if key in UNSUPPORTED:
stripped.append((path, key, value)) # enforce after the call
continue
if key == "properties":
out[key] = {k: to_strict(v, f"{path}.{k}", stripped)[0] for k, v in value.items()}
else:
out[key] = to_strict(value, path, stripped)[0]
if out.get("type") == "object" or "properties" in out:
out["additionalProperties"] = False
out["required"] = list(out.get("properties", {})) # every key present; absence = null
out.pop("default", None) # defaults have no meaning to a decoder
return out, stripped
strict_schema, stripped = to_strict(Invoice.model_json_schema())Anthropic's Python and TypeScript SDKs already strip and client-side validate unsupported constraints; an explicit compiler still makes the stripped list visible, works for every provider and gives CI something to pin.
Making the call
On Claude the schema goes in output_config.format with type json_schema; the older top-level output_format parameter is deprecated. The SDK's client.messages.parse() helper wraps the same request and returns a validated object; the explicit form is shown here so every step is visible.
import json
import anthropic
client = anthropic.Anthropic()
def extract_invoice(document_text: str) -> dict:
resp = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
system="Extract the invoice. Use null for any field not present. Never estimate amounts.",
messages=[{"role": "user", "content": document_text}],
output_config={"format": {"type": "json_schema", "schema": strict_schema}},
)
if resp.stop_reason == "refusal":
raise ExtractionRejected("model declined; route to human review")
if resp.stop_reason == "max_tokens":
raise ExtractionTruncated("output cut off; raise max_tokens or split the document")
return json.loads(resp.content[-1].text)
# The same contract on OpenAI's Chat Completions API:
# response_format={"type": "json_schema",
# "json_schema": {"name": "invoice", "strict": True, "schema": strict_schema}}Several habits from older JSON prompting break on current models. Assistant-message prefill (starting the reply with an open brace) is rejected on recent Claude models; use the format parameter instead. Sampling parameters such as temperature are rejected on the newest Claude models, so do not add them 'for determinism'. Forcing a specific tool just to obtain JSON is also rejected on some current models; structured output is the supported route. Structured output cannot be combined with citations on Claude, which returns a 400. And a new schema pays a one-time compilation cost (Anthropic caches compiled schemas for 24 hours), so expect a slow first call after a schema change.
Gate on why generation stopped
A schema guarantee only holds for a response that finished normally. Two stop reasons void it. A refusal may produce output that does not match the schema at all; treat it as a routing decision, not a parse error, and never retry it in a tight loop. max_tokens means the JSON was cut off mid-value; the fix is a larger limit or a smaller task (split long documents, or extract line items in pages), not a tolerant parser that closes the brackets for you and silently drops the tail.
Check the stop reason before you call json.loads. It is a two-line guard and it converts the most confusing class of production errors into clearly labelled ones. If you need to salvage truncated or malformed text from models without strict decoding, the ladder in parsing model output shows how to do it with bounded effort.
Validate twice: shape, then meaning
The first validation reapplies the full schema, including every constraint the compiler stripped. A quantity of zero or a negative price is valid JSON and valid against the strict schema, but not against Field(gt=0). The second validation checks meaning against the input.
from decimal import Decimal
from pydantic import ValidationError
def accept(raw: dict, source_text: str) -> Invoice:
inv = Invoice.model_validate(raw) # full schema, including the stripped constraints
problems = []
line_sum = sum((li.amount for li in inv.line_items), Decimal("0"))
for li in inv.line_items:
if abs(li.quantity * li.unit_price - li.amount) > Decimal("0.01"):
problems.append(f"line '{li.description}': qty x price != amount")
if inv.line_items and abs(line_sum - inv.total) > Decimal("0.01"):
problems.append(f"line items sum to {line_sum}, total says {inv.total}")
if inv.total_evidence not in source_text:
problems.append("total_evidence is not a substring of the document")
if problems:
raise SemanticMismatch(problems) # do not auto-retry: send to review
return invWorked example. A supplier invoice lists three lines: 10 licences at 100.00 (1,000.00), a support plan at 180.00 and a setup fee at 25.00. The printed total is 1,205.00. The model returns valid JSON with all three lines correct but a total of 1,250.00, a digit transposition. Strict decoding accepts it; Pydantic accepts it, since 1,250.00 is a non-negative decimal. The semantic check computes a line sum of 1,205.00 and fails. It also finds that the total_evidence string the model returned, 'Total due: 1,250.00', does not occur in the document. Both signals point the same way, and the invoice goes to review with the reasons attached.
Note what the code does not do: ask the model to fix the total. A retry prompt quoting 1,205.00 invites the model to copy your number back, passing the check without improving trust. Retry automatically only for shape errors (a missing key on a non-strict path, an unsupported value), at most once or twice, with the validator's message included. Route meaning errors to a person or a second, independent extraction.
Designing fields the model can fill honestly
- Prefer null to guessing. Every fact that may be absent is nullable, and the description says when to use null. A required non-null field forces invention.
- Give enums an exit.
OTHERorUNKNOWNvalues, plus an optional free-text field, beat a closed list that forces the nearest wrong label. - Ask for evidence on high-stakes fields. A quoted span is cheap to generate and cheap to verify with a substring check.
- Keep the schema shallow and named clearly. Field names and descriptions are part of the prompt.
invoice_datewith a description beatsd1. - Compute what you can. Extract printed values and do arithmetic in code.
If the output is a tool call rather than a final answer, the same rules apply to tool input schemas; see function calling and tool use for strict tool definitions.
Versioning the contract
Schemas change: a new field, a renamed enum value, a split address. Treat each change like an API change. Additive, nullable fields are backward compatible for consumers that ignore unknown keys. Renames, removals, type changes and new enum values that consumers switch on are breaking: bump schema_version, write an explicit migration for stored data, and deploy consumers before producers.
Keep a folder of real past outputs and replay them in CI against the current models, and pin the compiled strict schema so an innocent refactor of the Pydantic model cannot change what the API receives without someone noticing.
# tests/test_invoice_contract.py: replay stored outputs against the current schema in CI.
import json, pathlib, pytest
SAMPLES = sorted(pathlib.Path("contracts/invoice/samples").glob("*.json"))
@pytest.mark.parametrize("path", SAMPLES, ids=lambda p: p.name)
def test_old_outputs_still_parse(path):
raw = json.loads(path.read_text())
if raw.get("schema_version") == "invoice.v1":
raw = upgrade_v1_to_v2(raw) # explicit, tested migration
Invoice.model_validate(raw)
def test_strict_schema_is_stable():
current = json.dumps(to_strict(Invoice.model_json_schema())[0], sort_keys=True)
pinned = pathlib.Path("contracts/invoice/strict_schema.json").read_text()
assert current == pinned, "schema changed: bump schema_version and regenerate"Pair this with a small labelled evaluation set. Schema validity rates are close to 100 percent with strict decoding and therefore tell you little; field-level accuracy against labelled documents is the metric that moves when you change prompts, schemas or models.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| 400 on the request | Schema uses an unsupported keyword or misses additionalProperties false | Compile to the strict subset; pin the output in CI |
| Valid JSON, invented values | Required non-null fields with no honest option | Nullable fields, enum exits, evidence spans |
| Unparseable tail | max_tokens stop reason | Gate on stop reason; raise limit or split input |
| Off-schema output | Refusal stop reason | Route to review; do not loop |
| Slow first call after deploy | New schema compilation | Warm with a request at deploy time |
| Consumers break after a release | Breaking schema change without a version bump | schema_version, migrations, replay tests |
| Retry 'fixes' semantic errors | Feedback leaked the expected answer | Retry shape errors only; review meaning errors |
Trade-offs
Strict structured output costs a little flexibility: you give up schema features the decoder cannot enforce, and very large or deeply nested schemas make compilation slower and responses harder for the model to fill well. Plain JSON instructions with validation and repair still work on models or endpoints without strict mode, at the price of a parse failure rate you must measure. Tool use is the right fit when the model must choose between several actions; a single response format is simpler when there is exactly one expected shape. When output is read by people rather than code, structured output may be the wrong tool entirely; see output formatting.
What to do next
- Move every JSON contract your prompts produce into a typed model (Pydantic or Zod) with
extra='forbid'and aschema_versionconstant. - Write or adopt a compiler to the strict subset and pin its output in CI; list the stripped constraints explicitly.
- Switch requests to the provider's strict format parameter, remove prefill and sampling parameters that current models reject, and add a stop-reason gate before parsing.
- Re-validate every response against the full schema, then add at least one semantic check per high-value field (sums, substrings, cross-references).
- Set the retry policy: shape errors at most twice with the validator message, meaning errors to review.
- Build a labelled set of 50 to 200 real inputs and track field-level accuracy whenever the prompt, schema or model changes.