Most production uses of a small language model are not chat. They are extraction and classification jobs: an email into a ticket, a PDF into an invoice record. The output goes into another program, so it has to be structured and right. A 3-billion-parameter model often produces clean JSON, until it meets a long document, an unusual layout or a missing field.
This article treats structured output as a pipeline problem rather than a prompting trick. It explains how small models fail, the three levers you have (prompting, constrained decoding and fine-tuning), how to design schemas that a small model can fill, and how to wrap the model in validation, semantic checks, repair and escalation. A worked invoice-extraction example with Ollama and pydantic runs through the whole article. The token-level mechanics of grammar-constrained generation are covered in guided decoding for SLMs; here we care about what those mechanics guarantee and what they do not.
How small models fail at structure
It helps to separate three layers of failure, because each needs a different fix.
Syntax failures are output that does not parse: a trailing comma, an unclosed brace, a preamble such as "Here is the JSON:", markdown fences, or output cut off at the token limit. Small models produce these more often than large ones, particularly after long inputs, when the instruction sits thousands of tokens back.
Schema failures parse but do not match the contract: a key named invoiceNo instead of invoice_number, a string where a number was required, an enum value paraphrased as "US dollars", or an array of 400 near-identical line items from a repetition loop.
Semantic failures are valid, well-typed and wrong. The model reads the due date as the invoice date, copies a subtotal as the total, or invents an invoice number when the document has none. Nothing downstream notices. Small models invent plausible values for missing fields more readily than large models, and are weaker at arithmetic and at picking the right one of several similar numbers.
Constrained decoding eliminates the first layer and most of the second. Nothing in the decoder can fix the third. The rest of the pipeline exists for the third layer.
Three levers: prompt, constrain, fine-tune
You can push a small model toward correct structure in three ways, and in practice you combine them.
| Lever | What it fixes | Cost | Limits |
|---|---|---|---|
| Prompting: schema in the prompt, one or two examples, explicit null rule | Most semantic confusions about what each field means | Input tokens on every call | Does not guarantee syntax; small models forget instructions over long inputs |
| Constrained decoding: a grammar or JSON schema masks invalid tokens | All syntax failures, most schema failures | Schema compile time, some per-token overhead, backend support varies | Guarantees shape, not truth; can push the model into poor values |
| Fine-tuning: LoRA on input and gold-JSON pairs | Field semantics, null behaviour, domain layouts | Labelled data, training runs, a model to maintain | Needs retraining when the schema changes |
The usual order is prompt plus constrain first, measure, and fine-tune only when field-level accuracy stalls below target. Fine-tuning is the lever that moves semantic accuracy, which is why it matters more for small models than large ones.
The pipeline
The architecture below is what a production extraction service built on a small model tends to look like. The model is one box among eight.
Each record leaves tagged with its path. The share that pass first time, need a repair, or escalate is your most useful operational number: it moves before accuracy dashboards do, and it flags new document templates in the input stream.
Designing schemas a small model can fill
Schema design is the cheapest accuracy improvement available, and it matters more for small models, which have less capacity to cope with awkward contracts.
- Keep it flat and short. Every level of nesting and every extra field is more to track. If downstream needs nesting, extract flat and reshape in code.
- Use descriptive key names. The model reads the keys:
invoice_datecarries meaning,d1carries none. - Make absence explicit. Declare fields nullable and say that null is the answer when the text lacks the value. A required string forces the model to write something, so it invents.
- Use enums for closed sets. An enum turns a paraphrase problem into a choice among a few tokens.
- Put evidence first. Generation is left to right, so a field can only condition on the fields before it. A short
evidencefield, where the model copies the lines it relies on, gives later fields something concrete to condition on. - Bound everything. Set
maxLengthon strings and a maximum item count on arrays, so a repetition loop ends as a short record you can catch rather than a 4,000-token runaway. - Avoid exotic schema features. Complex
oneOfunions, regex patterns and recursion compile into large grammars, and some backends enforce only a subset of JSON Schema. Check which keywords yours enforces.
Here is the schema for the running example, an invoice extractor, written with pydantic so the same class produces the JSON schema, validates the output and gives you typed objects.
from datetime import date
from decimal import Decimal
from enum import Enum
from typing import Annotated
from pydantic import BaseModel, Field, WithJsonSchema
# plain number in the schema; pydantic's default for Decimal is an anyOf with a regex
Money = Annotated[Decimal, WithJsonSchema({"type": "number"})]
class Currency(str, Enum):
USD = "USD"
EUR = "EUR"
GBP = "GBP"
class LineItem(BaseModel):
description: str = Field(max_length=120)
quantity: Money
unit_price: Money
amount: Money
class Invoice(BaseModel):
# evidence first: the model copies the text it relies on before committing to values
evidence: str = Field(max_length=400, description="verbatim lines holding number, date, total")
invoice_number: str | None = Field(max_length=40)
invoice_date: date | None
vendor_name: str | None = Field(max_length=120)
currency: Currency | None
line_items: list[LineItem] = Field(max_length=30)
total: Money | None
Worked example: invoice extraction with Ollama
Ollama's chat call takes a format argument that accepts a JSON schema, which it enforces during generation, and the reply is parsed back with pydantic's model_validate_json. The function below runs a first pass, validates, applies semantic checks, and on failure gives the model exactly one repair attempt with the error list before handing the document to an escalation path.
import json
from ollama import chat
from pydantic import ValidationError
SCHEMA = Invoice.model_json_schema()
SYSTEM = (
"Extract the invoice as JSON matching the schema. Copy values exactly as printed. "
"If a field is not in the text, use null. Never guess a value."
)
def extract(doc_text: str, model: str = "qwen2.5:3b") -> tuple[Invoice | None, list[str]]:
messages = [
{"role": "system", "content": SYSTEM},
{"role": "user", "content": doc_text},
]
for attempt in range(2): # first pass + one repair
reply = chat(model=model, messages=messages, format=SCHEMA,
options={"temperature": 0, "num_predict": 1024})
raw = reply.message.content
try:
inv = Invoice.model_validate_json(raw)
except ValidationError as e:
errors = [f"{'.'.join(map(str, x['loc']))}: {x['msg']}" for x in e.errors()]
else:
errors = semantic_errors(inv, doc_text)
if not errors:
return inv, []
messages += [
{"role": "assistant", "content": raw},
{"role": "user", "content": "Fix these problems and return the full JSON again:\n"
+ "\n".join(errors)},
]
return None, errors # caller escalates
def semantic_errors(inv: Invoice, doc_text: str) -> list[str]:
errs = []
for i, li in enumerate(inv.line_items):
if abs(li.quantity * li.unit_price - li.amount) > Decimal("0.01"):
errs.append(f"line_items.{i}: quantity x unit_price != amount")
if inv.total is not None and inv.line_items:
if abs(sum(li.amount for li in inv.line_items) - inv.total) > Decimal("0.05"):
errs.append("total does not equal the sum of line item amounts (check tax lines)")
if inv.invoice_number and inv.invoice_number not in doc_text:
errs.append("invoice_number does not appear in the document text")
return errsTemperature is zero, because extraction wants the most likely reading. num_predict caps output so a runaway cannot stall a worker. The repair message quotes specific errors, because a small model repairs far better with a pointer to the broken field. There is exactly one repair: a second attempt rarely succeeds where the first failed, and it doubles tail latency. The model tag is only an example.
The semantic checks are where domain knowledge enters. Line-item arithmetic catches misread quantities. The sum check catches a subtotal copied as the total, while hinting at tax lines. The substring check cheaply catches the most damaging failure, an invented identifier: an invoice number absent from the source text did not come from it.
What constrained decoding guarantees, and what it costs
A constrained decoder compiles the schema into a grammar and, at each step, masks the logits of every token that would make the output invalid. The output is therefore always parseable and schema-shaped, provided the generation is not cut off by the token limit first. That is a strong guarantee and worth having. It is also narrower than it looks.
First, it constrains shape, not content. If the model's probability mass wanted to write prose, the mask forces it down a path it did not prefer, and the values it produces there can be worse than an unconstrained answer would have been: empty strings, a placeholder such as N/A in a string field, or a number pulled from the wrong line. This is why the null rule matters so much. A nullable field gives the model a legitimate way out; a required string forces a guess.
Second, it has costs. Large schemas take time to compile, so cache compiled grammars per schema version. Masking adds per-token overhead that varies by backend; engines such as XGrammar and llguidance keep it small, but measure it at your batch sizes. Some models fill grammar-permitted whitespace with long runs of newlines, which a token cap contains.
Third, backends differ. llama.cpp, Ollama, vLLM, SGLang and Outlines all offer schema-constrained generation, with different JSON Schema coverage and request syntax that has changed between releases. Pin versions and keep a test that sends your real schema and asserts enforcement, because some stacks silently ignore unsupported keywords. The masking internals are covered in structured output architecture.
Fine-tuning for structure
When prompting and constraining leave semantic accuracy short of the target, a LoRA fine-tune on your own documents is usually the next step, and on a narrow extraction task a tuned small model can match or beat a much larger general model. The data format matters more than the volume.
- Serialize targets exactly as serving produces them: same key order, whitespace and null handling.
- Include documents with genuinely missing fields and null targets, or the model learns to always invent a value.
- Mask the loss on the prompt so the model trains only on the JSON.
- Keep the constraint on after tuning; it costs little and removes rare syntax slips.
- Version the adapter with the schema: an adapter trained on schema v3 serving schema v4 is a silent regression.
The mechanics of adapter training are in LoRA for small models. If the output is a call to a tool rather than a record, the same principles apply, with the extra step of choosing the tool, covered in SLM function calling.
Evaluating field by field
Once decoding is constrained, parse rate and schema-valid rate sit near 100 percent and say little. The numbers that matter are per field. Build a labelled set of a few hundred real documents, including ones with missing fields, and score each field, treating null as a real answer.
def norm(v):
return None if v is None else str(v).strip().casefold()
def score_field(golds, preds, f):
"""Null is a real answer: a value where gold is null counts as a hallucination."""
right = halluc = absent = 0
for g, p in zip(golds, preds):
gv, pv = norm(g.get(f)), norm(p.get(f)) if p else None
right += gv == pv
if gv is None:
absent += 1
halluc += pv is not None
return {"accuracy": right / len(golds),
"hallucination_rate": halluc / absent if absent else 0.0}Normalize before comparing, as norm does minimally, or you will measure formatting instead of reading. Track the hallucination rate separately, the share of documents where the gold value is null and the model produced something, because it is the failure users trust least. Report the pipeline path mix alongside accuracy, and re-run the set on every change of model, prompt, schema or backend version. The broader evaluation setup, including regression gates, is in SLM evaluation architecture.
Failure modes in production
| Symptom | Likely cause | Fix |
|---|---|---|
| Values present for fields absent from the document | Required, non-nullable fields; no null examples in the prompt or training data | Make fields nullable, state the null rule, add null cases to examples and training data |
| Output truncated mid-object | Token cap below the longest legitimate record, or a repetition loop | Bound arrays and strings in the schema, size the cap from real maxima, log truncations |
| Accuracy drops on one customer only | New document layout outside the prompt examples or training set | Watch the repair and escalation rate per source, add examples, retrain |
| Constraint appears ignored | Backend does not support a schema keyword, or the request field name changed between versions | Contract test that sends the real schema and asserts enforcement |
What to do next
- Collect 200 to 500 real inputs, including ones with missing fields, and label the expected JSON with nulls where values are absent.
- Write the schema as a pydantic model: flat, descriptive keys, enums for closed sets, nullable fields, bounds on strings and arrays, an evidence field first.
- Turn on schema-constrained decoding in your serving stack, and add a contract test that proves the constraint is enforced.
- Wrap the model in parse, schema validation, semantic checks, one repair with specific errors, and an escalation path.
- Score per field with null as a real answer, and track hallucination rate and the first-pass, repaired and escalated mix.
- If a field stays below target, fine-tune a LoRA adapter on your labelled data in the exact serving serialization, and version it with the schema.