A language model's output is only useful if whatever consumes it can use it. A person reading a chat window wants readable prose with a little structure; a web page that renders it needs Markdown it can sanitise; a program needs fields it can parse; an SMS gateway needs plain text with no asterisks. Output formatting is the discipline of choosing the right shape for each consumer, specifying it so the model produces it consistently, and checking that it did.

This page covers that discipline end to end: a decision table for choosing formats, how to write a format specification the model follows, where to put reasoning relative to the answer, using tags to extract parts of mixed output, length control, rendering safely, streaming, and measuring compliance, with a worked support-ticket example. Two neighbouring topics have their own pages: structured output architecture covers schemas and constrained decoding, and output parsing covers turning text into trusted, typed data. This page is about the choice and specification that come before both.

Advertisement

Format is part of the interface

Treat model output like any other interface between components: it has a producer, one or more consumers and a contract. The failures are the usual interface failures. The producer emits something the consumer does not accept, such as a JSON object preceded by 'Sure, here is the JSON'. Or the contract is ambiguous, for example 'use bullet points' without saying how many.

The model produces text left to right, one token at a time, conditioned on everything before it. That gives formatting three levers. The instructions and examples in the prompt set the distribution the model samples from. The decoding controls offered by your API, such as stop sequences, token limits and, where available, schema-constrained generation, bound what can come out. The post-processing in your code extracts, validates and repairs. Reliable systems use all three, and the stricter the consumer, the more weight shifts from instructions to decoding and validation.

Format specinstructions + examplesModelgenerates tokensRaw text+ stop reasonExtractortags, JSON, fencesValidator / lintschema, rules, lengthRendererMarkdown to sanitised HTMLProgramtyped recordRetry, repair or fallbackcount every failurefor peoplefor codefailConsumers decide the format; each arrow into them is a place to measure compliance.
The output-formatting pipeline. The format spec shapes generation, the extractor and validator enforce the contract, and each consumer gets the shape it needs; failures are routed to retry, repair or fallback and counted.

Choose the format from the consumer

ConsumerFormatWhyEnforce with
Person in a chat UILight MarkdownReadable, and renders wellSanitising renderer and a lint for banned constructs
Person through a plain-text channelPlain prose, no markupAsterisks and hashes show up literallyLint for Markdown syntax
Program with fixed fieldsJSON matching a schemaTyped, and validation is standardConstrained decoding where available, then schema validation
Program extracting part of mixed outputXML-style tagsCheap to extract, tolerant of prose around themTag extractor and presence checks
Bulk recordsJSON Lines, one object per lineStreams well; one bad row is not fatalParse per line and count rejects
Closed-set decisionA single label on its own lineTrivial to matchExact match against the allowed set
CodeOne fenced block with a language tagUnambiguous boundariesExtract the fence, then compile or test

Two choices deserve a warning. CSV looks simple, but models are unreliable at quoting fields that contain commas, quotes or newlines, so prefer JSON Lines unless a downstream tool demands CSV, and if it does, validate the column count of every row. Heavy Markdown, such as nested lists, tables and headings, is rarely what a reader wants from a short answer. Many models format heavily by default, so ask for prose explicitly when prose is what you want.

Advertisement

Writing a format specification

A good format spec reads like an interface definition, not a wish. It names every part and its order, says what each part may contain, gives limits in units the model can follow, covers edge cases (empty results, refusals, other languages) and states that nothing else should be emitted. It works best when three things agree: the instructions, the examples and the style of the prompt itself.

  • Examples beat instructions. Models imitate the shape of few-shot examples more strongly than they follow prose rules, so an example that breaks the spec, with a heading where headings are banned, teaches the model to break it. Generate examples from the spec and lint them with the same code that lints outputs.
  • Say what to do, not only what to avoid. 'Write flowing paragraphs' is followed more reliably than 'do not use bullet points', which puts the unwanted pattern in front of the model.
  • The prompt's own style leaks. A prompt written in dense Markdown tends to get Markdown back; a prompt in plain paragraphs nudges toward prose.
  • Specify edge cases. Say what to emit when there is nothing to report, for example an empty list or a fixed sentence, so the model does not invent a third shape.
  • Delimit inputs as well as outputs. Wrapping the input document in its own tags stops the model confusing input structure with the required output structure; prompt delimiters covers the options.

Reasoning first, answer second

Because generation runs left to right, whatever the model writes first conditions what comes after. If the format puts the answer first and the justification second, the model commits before it has worked anything out, and the justification becomes a rationalisation. For tasks that benefit from working, put a reasoning section before the answer and extract only the answer for the consumer. The same applies inside JSON: field order in the schema is generation order, so a reasoning field should come before verdict, not after it.

This costs tokens and latency, and the reasoning section is also text you must not leak to users by accident, which is why the extractor matters. Some reasoning-model APIs think internally before answering, which makes a visible reasoning section less necessary; check your provider's guidance rather than assuming either way. For simple lookups and formatting-only tasks, skip the reasoning section entirely.

Worked example: triaging support tickets

A support tool drafts customer replies. Two consumers read the output: the agent-facing UI, which shows the model's notes, and the customer-facing email, which renders a restricted Markdown subset to HTML. One JSON object with a Markdown string inside would work, but Markdown inside a JSON string means escaped newlines and quotes, which models get wrong more often than plain tags, and it hides the reply from anyone reading logs. So the design uses two tagged sections: analysis first, as the reasoning, then the reply.

FORMAT_SPEC = """
Respond in exactly two parts, in this order, and nothing else.

<analysis>
Your working notes about the ticket. Never shown to the customer. Plain text.
</analysis>

<reply>
The message to the customer. Markdown limited to paragraphs, bullet lists,
**bold** and inline code. No headings, tables or images. Links only to
https://help.example.com/. Three short paragraphs at most.
</reply>

Write the reply in the language the customer used.
If you cannot help, still produce both parts and say so politely in the reply.
"""

The extractor refuses truncated outputs, requires both tags exactly once, and lints the reply against the restrictions in the spec.

import re
from dataclasses import dataclass

TAG = re.compile(r"<(analysis|reply)>\s*(.*?)\s*</\1>", re.S)

class FormatError(ValueError):
    pass

@dataclass
class Parsed:
    analysis: str
    reply: str

LENGTH_STOPS = {"max_tokens", "length"}   # the value differs by provider; check yours

def extract(text: str, stop_reason: str) -> Parsed:
    # Assumes no stop sequence is set, so both closing tags must be present.
    if stop_reason in LENGTH_STOPS:
        raise FormatError("truncated")               # never ship a cut-off answer
    found = {}
    for name, body in TAG.findall(text):
        if name in found:
            raise FormatError(f"duplicate {name}")
        found[name] = body
    missing = {"analysis", "reply"} - found.keys()
    if missing:
        raise FormatError(f"missing {sorted(missing)}")
    return Parsed(found["analysis"], found["reply"])

FORBIDDEN = [
    (re.compile(r"^#{1,6} ", re.M), "heading"),
    (re.compile(r"!\[[^\]]*\]\("), "image"),
    (re.compile(r"\]\((?!https://help\.example\.com/)"), "off-domain link"),
    (re.compile(r"^\|.*\|\s*$", re.M), "table"),
]

def lint_reply(md: str, max_paragraphs: int = 3) -> list[str]:
    problems = [label for rx, label in FORBIDDEN if rx.search(md)]
    if len([b for b in md.split("\n\n") if b.strip()]) > max_paragraphs:
        problems.append("too long")
    return problems

In testing, three failure types showed up. The model sometimes added a sign-off after the closing reply tag, which is harmless because the extractor ignores text outside tags. It occasionally used a heading in long replies; adding the rule to the spec and fixing one bad example cured most of it, and the lint catches the rest and triggers one retry with the lint message appended. And a customer ticket that quoted a previous email containing the literal text of a closing tag broke extraction once. Wrapping the ticket in its own tags and escaping angle brackets in the input fixed that.

Length, truncation and stop sequences

Models estimate length poorly in words and well in structure. 'At most three bullets of one sentence each' is followed far more consistently than '50 words'. Treat any numeric word limit as a soft target and enforce the real limit in code. The token limit you pass to the API is a different thing: it does not make the model write shorter, it cuts the output off, usually mid-sentence and, for JSON, mid-object. Always check the stop reason your API returns and treat a length stop as a failure, never as a short answer.

Stop sequences end generation when a given string appears, which is useful for ending output at a closing tag and preventing trailing chatter. In most APIs the stop string itself is not included in the returned text, so the extractor must accept an answer whose closing tag is missing when the stop reason says a stop sequence fired. Some APIs also let you pre-fill the start of the assistant's turn, for example with an opening tag or brace, which removes preambles. Support for this varies by provider and model, so confirm it for your model before designing around it.

Rendering and streaming safely

Rendering model Markdown as HTML is rendering untrusted input. Convert Markdown to HTML with raw HTML disabled, then pass it through an allowlist sanitiser. Be strict about images and links: if any part of the prompt can be influenced by an attacker, such as a retrieved web page or a customer email, a Markdown image whose URL carries data in its query string can leak that data the moment the page loads. Allow links only to domains you control, and do not render remote images from model output at all; see prompt injection defence for the wider threat model.

Streaming changes the choice of format. Markdown and plain prose can be shown token by token, although a half-finished table or code fence will flicker as it renders. A JSON object is not parseable until it closes, so either stream a human-readable part first and the structured part last, use JSON Lines so each completed line is usable, or use an incremental parser. Tagged sections stream well: show the reply section as it arrives and hide everything until its opening tag has appeared.

Measuring format compliance

Formatting quietly regresses when prompts, models or input distributions change, so measure it like any other quality metric. Run a fixed evaluation set through the pipeline on every prompt or model change and track the parse rate, the clean rate (parsed and passing lint), the truncation rate, and the length distribution, not just its mean.

def format_report(samples):
    """samples: list of (raw_text, stop_reason) from a fixed eval set."""
    counts = {"parsed": 0, "clean": 0, "truncated": 0}
    lengths = []
    for text, stop_reason in samples:
        try:
            parsed = extract(text, stop_reason)
        except FormatError as err:
            counts["truncated"] += str(err) == "truncated"
            continue
        counts["parsed"] += 1
        counts["clean"] += not lint_reply(parsed.reply)
        lengths.append(len(parsed.reply.split()))
    lengths.sort()
    rates = {k: v / len(samples) for k, v in counts.items()}
    rates["p95_words"] = lengths[int(0.95 * (len(lengths) - 1))] if lengths else None
    return rates

Format metrics say nothing about whether the content is right, so pair them with task metrics; prompt evaluation covers building those sets. In production, log every extraction or lint failure with the raw output, and alert on the rate rather than on individual failures.

Failure modes

SymptomCauseFix
Preamble before the payloadChat-tuned default behaviour.Tags plus extraction, a stop sequence, or pre-fill where supported.
Output cut off mid-structureToken limit reached.Check the stop reason, raise the limit, shorten the requested output.
Heavy Markdown where prose was wantedModel default, Markdown-heavy prompt or examples.Ask for prose positively, rewrite the prompt in prose, fix examples.
Broken JSON escaping in long text fieldsNewlines and quotes inside strings.Move long text into a tagged section or use constrained decoding.
Answer contradicts its own justificationAnswer generated before reasoning.Reorder so reasoning comes first.
Format drifts after a model upgradeDifferent defaults.Re-run the format eval set before switching.

Trade-offs

  • Strictness against quality: the tightest formats, such as a bare label or a rigid schema, are the easiest to parse but leave no room for the model to say it is unsure. Add an explicit 'unknown' option or a notes field.
  • Reasoning sections improve hard answers and cost tokens, latency and leak risk.
  • Retries fix most format failures cheaply, but a retry rate above a few percent means the specification or the examples are wrong.

What to do next

  1. List every consumer of each prompt's output and pick its format from the decision table.
  2. Rewrite each format spec to name parts, order, allowed content, limits in structural units and edge-case behaviour.
  3. Lint your few-shot examples with the same code that lints outputs.
  4. Put reasoning before answers, in text and in schema field order, where the task needs working.
  5. Check the stop reason on every call and treat length stops as failures.
  6. Render Markdown through a sanitiser with images off and links restricted to your domains.
  7. Build a format eval set and track parse, clean and truncation rates on every change.
Key takeaway: Output formatting is interface design: the consumer decides the shape, the prompt specifies it, decoding controls bound it, and code extracts and checks it. Choose light Markdown for people, plain prose for plain channels, schemas for programs and tags for mixed output. Write specs that name parts, order and structural limits, keep examples consistent with them, and put reasoning before answers. Treat truncation as failure, sanitise anything you render, and measure parse, clean and truncation rates on every change, because format quality regresses silently.