Every application that calls a language model eventually does the same thing: it takes a string the model produced and turns it into something the rest of the code can use, such as an object, a label, a number or a function call. That step is output parsing, and it is where many LLM features fail in production. A model that returns valid JSON 99 percent of the time sends a malformed or truncated response every few seconds at scale.
Constrained decoding and provider structured-output modes reduce syntax errors dramatically, and they are covered in structured output architecture. They do not remove the need for a parser. Responses still get cut off, refused, or filled with values that are well-formed but wrong. This article designs the application-side parsing layer: its stages, its error classes, and how each failure is routed, with working Python you can adapt.
First principles: model output is untrusted input
The model is a probabilistic text generator. Even when it follows instructions well, its output is closer to user input than to a function's return value: it can be incomplete, malformed, subtly wrong, or shaped by text an attacker placed in the context. So treat the parser as a trust boundary. On one side is arbitrary text; on the other side, only values that have been extracted, parsed, validated and normalised may pass.
Two design rules follow. The parser never guesses silently: it either returns a validated object or raises an error that says which stage failed and why. And each stage is narrow, so its failures are distinguishable. A truncated response, a refusal and a wrong enum value look similar as exceptions from a single call to a JSON parser, but they need completely different handling.
Stage 1: classify how the response ended
Before looking at the text, look at why generation stopped. Every major API reports a stop or finish reason, under different names. Normalise it to a small set in your client wrapper: completed normally, hit the token limit, called a tool, or refused. A response that hit the token limit is truncated by definition. Repairing truncated JSON by closing brackets produces an object that parses and is missing data, which is worse than an error. Route it to a bigger limit or a smaller request instead.
Stage 2: extract the payload
Unless output was constrained, the payload may sit inside prose, in a Markdown code fence, or after a preamble such as Here is the JSON. Extraction should be deterministic and ordered: try a fenced block first, then the first balanced object or array. A naive regular expression from the first brace to the last brace breaks when prose after the payload contains a brace, and counting braces breaks when a string value contains one. Scan properly, respecting strings and escapes:
import re
FENCE = re.compile(r"```(?:json)?\s*\n(.*?)```", re.S)
def balanced_spans(text):
# Yield top-level {...} or [...] spans, skipping braces inside strings.
i = 0
while i < len(text):
if text[i] not in "{[":
i += 1
continue
stack, in_str, esc = [], False, False
for j in range(i, len(text)):
c = text[j]
if in_str:
if esc: esc = False
elif c == "\\": esc = True
elif c == '"': in_str = False
elif c == '"': in_str = True
elif c in "{[": stack.append("}" if c == "{" else "]")
elif c in "}]":
if not stack or stack.pop() != c:
break
if not stack:
yield text[i:j + 1]
i = j
break
i += 1
def extract(text):
m = FENCE.search(text)
if m:
return m.group(1).strip()
return next(balanced_spans(text), None)When the payload is in XML-style tags, which many prompts use because tags are easy for models to produce reliably (see XML delimiters in prompts), extraction is a non-greedy match between the opening and closing tag, taking the last occurrence if the model restates its answer.
Stage 3: a bounded tolerant-parse ladder
Try a strict parse first. If it fails, apply a short, fixed list of repairs, each fixing one well-understood, meaning-preserving defect, and try again after each. Stop at the end of the list. Good rungs: strip a byte-order mark and surrounding whitespace, remove trailing commas before a closing bracket, and replace typographic quotes used as JSON delimiters. Bad rungs: closing unterminated strings, inventing missing brackets, evaluating the text as Python, or asking a second model to fix it silently. Those guess at meaning. Record which rung succeeded, because a rising rate on any rung is an early signal that a prompt or model version changed.
import json
TRAILING_COMMA = re.compile(r",\s*([}\]])")
LADDER = [
("strict", lambda s: s),
("strip", lambda s: s.strip().lstrip("")),
("trailing_comma", lambda s: TRAILING_COMMA.sub(r"\1", s)),
]
class ParseError(Exception):
def __init__(self, kind, detail):
super().__init__(f"{kind}: {detail}")
self.kind, self.detail = kind, detail
def parse_json(payload, metrics):
last_err, s = None, payload
for name, fix in LADDER:
s = fix(s)
try:
value = json.loads(s)
metrics.incr(f"parse.rung.{name}")
return value
except json.JSONDecodeError as e:
last_err = e
raise ParseError("syntax", f"{last_err.msg} at line {last_err.lineno} col {last_err.colno}")Note the trailing-comma rule is itself a small risk, since it can change a string value that contains a comma before a bracket. That is why it runs only after the strict parse fails, and why the fixed list stays short.
Stages 4 and 5: schema and semantic validation
A parsed value is a tree of dicts and lists, not your domain type. Validate it with a schema library such as Pydantic, which gives typed objects and field-level errors. Put everything a schema can express into the schema: required fields, enums, ranges, string formats. Put everything it cannot express into semantic checks that run after: totals that must equal the sum of line items, dates in order, identifiers that must exist in your database.
from decimal import Decimal
from pydantic import BaseModel, Field, ValidationError, model_validator
class LineItem(BaseModel):
sku: str = Field(pattern=r"^[A-Z0-9-]{4,20}$")
qty: int = Field(ge=1, le=10_000)
unit_price: Decimal = Field(ge=0)
class Invoice(BaseModel):
schema_version: int = 2
invoice_id: str
currency: str = Field(pattern=r"^[A-Z]{3}$")
items: list[LineItem] = Field(min_length=1, max_length=500)
total: Decimal
@model_validator(mode="after")
def total_matches(self):
expected = sum(i.qty * i.unit_price for i in self.items)
if abs(expected - self.total) > Decimal("0.01"):
raise ValueError(f"total {self.total} != sum of items {expected}")
return self
def validate(value):
try:
return Invoice.model_validate(value)
except ValidationError as e:
# Narrow feedback: field path and message, not the whole schema
details = "; ".join(f"{'.'.join(map(str, x['loc']))}: {x['msg']}" for x in e.errors())
raise ParseError("schema", details)Retry policy for these errors is covered in structured output architecture: cap repairs at one or two, send back only the failing field path and rule, and treat a sustained repair rate as a schema problem to fix, not a retry count to raise.
The error taxonomy, and why routing is the architecture
Raise one exception type carrying a class, then route on the class in a single place. That keeps retry logic, metrics and user-facing messages consistent across every feature that calls a model.
| Error class | How you detect it | Route |
|---|---|---|
| Truncated | Stop reason says the token limit was hit | Do not repair; raise the limit or ask for less, then retry once |
| Refused | Stop reason or a refusal field, or no payload plus refusal-like text | Surface to the caller; retrying the same prompt rarely helps |
| No payload | Extraction finds nothing | Retry once with a reminder of the output format |
| Syntax | Strict parse fails and the ladder cannot fix it | Retry with the parser error message |
| Schema | Validation error: missing field, wrong type, bad enum | Retry with the field path and the rule broken |
| Semantic | Business rule fails: totals do not add up, unknown ID | Retry with the rule, or send to human review |
| Unsafe | Size, depth or content limits exceeded | Fail closed and log; never retry blindly |
Worked example: one bad response, end to end
An invoice-extraction call returns: a sentence of preamble, then a fenced block containing an invoice with two line items, a trailing comma after the second item, and a total of 150.00 where the items add up to 105.00. The stop reason is normal.
Stage 1 passes, since the response completed. Stage 2 finds the fence and extracts its contents. Stage 3 fails the strict parse, and the trailing-comma rung succeeds; the metric records that rung. Stage 4 validates types and patterns successfully. Stage 5 fails: the total does not match. The parser raises a semantic error with the message total 150.00 != sum of items 105.00. The router sends one retry that includes the original document and that single sentence. If the retry fails the same way, the document goes to a review queue, because a model that twice misreads a total may be looking at a scanned document it cannot read, and a third attempt costs money without new information.
Beyond JSON: labels, numbers and tool arguments
Classification outputs are the most common target and the most often under-parsed. Normalise case and whitespace, map known synonyms to canonical labels through an explicit table, and treat anything else as an error, never as a default class. Always offer the model an explicit unknown label, so it is not forced to guess. Numbers need the same care: parse units explicitly, reject ambiguous formats such as 1,234 in a locale-sensitive context, and keep money in decimal types.
Tool calls are output parsing too. The arguments arrive as a JSON string or an already-decoded object depending on the API, and they go through the same stages: validate against the tool's schema, run semantic checks, and only then execute. Never execute on partially streamed arguments, and make side-effecting tools idempotent, so a retried call cannot charge a card twice.
Parsed output is still untrusted
Passing validation means the value has the right shape, not that it is safe. A string field can contain SQL, shell metacharacters, HTML or a path such as ../../etc/passwd, especially when the model read documents or web pages an attacker controls; see prompt injection defence. Use parameterised queries, never string concatenation. Escape on output. Check that referenced identifiers belong to the requesting user, because the model can name any ID it saw. Cap payload size, list lengths and nesting depth before parsing, so a runaway response cannot exhaust memory.
Schema versioning and testing
Prompts and schemas change together, so version them together: put a schema_version field in the output, deploy parser changes that accept both the old and new version before the prompt that emits the new one, and remove the old one only after traffic has moved. Additive, optional fields are safe; renames and type changes need that two-step rollout. A prompt registry that stores the schema with each prompt version makes this manageable.
Test the parser without a model. Keep a golden corpus of real bad outputs from logs: fenced and unfenced payloads, preambles, trailing commas, truncations, refusals, wrong enums, braces inside strings, oversized lists. Assert the exact error class for each. Add property-based tests that serialise valid objects in many formats and check that parsing returns them unchanged. In production, emit per-stage metrics: extraction failures, ladder rung usage, error class counts, retries and final failures per prompt version, and connect them to your evaluation pipeline, because a parse rate says nothing about whether the values are right.
If you stream structured output to a user interface, the streaming rules in the structured output article apply: render fields only once closed, and never trigger side effects from a partial object.
What to do next
- Wrap your model client so every response carries a normalised stop reason, and route truncations before parsing.
- Replace ad hoc JSON loading with one parse function that runs extract, strict parse, a short ladder, schema validation and semantic checks.
- Define one error type with the classes in the table, and route retries, review queues and user messages in one place.
- Add an explicit unknown option to every classification and extraction schema.
- Put size and depth limits in front of the parser, and parameterise every query built from model output.
- Add a schema_version field and a two-step rollout rule for breaking schema changes.
- Build a golden corpus of at least twenty real malformed outputs from logs and assert the error class for each in CI.
- Dashboard ladder rung usage and error classes per prompt version, and alert when any rate moves.