Streaming changes what a prompt is for. When a response arrives all at once, only its final content matters, and the order of its parts is a matter of style. When it is streamed, the user starts reading the first sentence while the model is still generating the tenth, so order becomes latency. A polite three-sentence preamble that costs nothing in a batch API costs the user two seconds of staring at filler.

This article treats streaming as a design constraint on both the prompt and the client. It covers which metric to optimise, how to write prompts that put useful content first, which output formats survive being rendered half-finished, how to route a single stream into several parts of a UI, how to moderate output that is already on screen, and how cancellation should work. The examples are provider-neutral and use plain Server-Sent Events.

Advertisement

The metric that matters: time to useful token

Most dashboards report time to first token: the delay between sending a request and receiving the first byte of output. It measures queueing and prompt processing, which infrastructure controls. It does not measure what the user experiences, because the first tokens are often filler such as "Great question! Let me help you with that."

A better metric is time to useful token: when the first token that carries the answer arrives. It equals time to first token plus the time spent generating preamble. At a decode speed of 50 tokens per second, a 40-token preamble adds 0.8 seconds, which is often more than the whole time to first token of a well-run service. Infrastructure can shorten the queue and the prefill; only the prompt can delete the preamble.

A streamed answer passes through four stages; prompt design decides what the user sees firstModeltokens in prompt orderchunksServer relaySSE, moderation windoweventsClient parsermarkers, NDJSON, fencessafe textRendererincremental DOMUser: Stopcancelabort propagates upstreamTimeline of one answerqueue + prefillfirst tokenTTFTpreamble tokensuser waitsfirst usefulTTUTrest of answerInfrastructure shortens the grey and yellow boxes. Only the prompt can remove the red one.Answer-first prompting moves the green box left without making the model any faster.
Where streamed latency comes from. Prompt design controls the red preamble box and the order of everything after it; the client controls what is safe to show and how cancellation propagates.

Measure it by tagging the first token of each answer section in your client, or by logging the full stream and labelling, offline, the character offset at which the answer begins. Report the median and the 95th percentile, next to time to first token, so you can see whether a slow experience is an infrastructure problem or a prompt problem.

Prompting for order: answer first

Models tend to frame before they answer, because a lot of their training text does. Streaming prompts must override that explicitly and specifically. Saying "be concise" is not enough; say what comes first, what comes next and where caveats go.

SYSTEM
You are the support assistant for Acme Billing. Your answers are streamed to
the user token by token, so order matters:

1. Start with the answer itself in one or two sentences. Do not greet, do not
   restate the question, do not say what you are about to do.
2. Then give the steps, as a numbered list, one step per line.
3. Put caveats after the steps, never before them.
4. If you need a table, put it last.
5. When you cite help articles, end with a line containing exactly
   <<SOURCES>> followed by one article id per line.

Each rule maps to a streaming effect. The answer-first rule moves the useful token forward. One step per line gives the renderer natural points to update the display. Caveats after steps means a user who stops reading early still has the instructions. Tables last, because a table cannot be shown sensibly until its header and several rows have arrived. The explicit sources marker lets the client route citations to a separate panel, as shown below.

Answer-first does have a cost. Chain-of-thought style reasoning before the answer can improve accuracy on multi-step problems, and an answer-first prompt removes that space. For reasoning-heavy tasks, use a model with a separate reasoning phase, which is not shown as answer text, or accept the delay and show a progress state rather than filler. Do not ask the model to reason aloud in the visible answer and then summarise; users read the reasoning, and it may contradict the final answer.

Advertisement

Formats that render safely while incomplete

Markdown is the default output format for chat interfaces, and it has three constructs that break when rendered half-finished. An open code fence turns everything after it into code until it closes. A table row without its closing pipe renders as a broken line. A link whose URL has not finished renders as raw brackets. Users notice the flicker on every chunk.

The fix belongs in the client: render a repaired copy of the text received so far, and keep the raw text unchanged for the next chunk.

def renderable(markdown_so_far: str) -> str:
    """Make a partial markdown answer safe to render on every chunk."""
    text = markdown_so_far
    # 1. An odd number of ``` fences means a code block is still open: close it.
    if text.count("```") % 2 == 1:
        text += "\n```"
    # 2. Hide a half-written table row; it would render as a broken pipe line.
    lines = text.split("\n")
    if lines and lines[-1].lstrip().startswith("|") and not lines[-1].rstrip().endswith("|"):
        lines = lines[:-1]
    # 3. Hide a half-written link: "[label](http" renders as raw text.
    last = lines[-1] if lines else ""
    if last.rfind("[") > last.rfind(")"):
        lines[-1] = last[: last.rfind("[")]
    return "\n".join(lines)

Two rules on the prompt side make this simpler. Ask for code blocks to be preceded by a sentence, so the fence never arrives as the very first token, and avoid nested lists deeper than two levels, which renderers re-indent unpredictably as they grow. If your UI renders HTML from markdown, sanitise after repair on every update, not once at the end, because the partial answer is shown to the user too.

Structured output that streams

A single JSON object cannot be parsed until its final brace arrives, so a prompt that asks for a JSON array of ten results makes the user wait for all ten. Two approaches fix this.

The first is to change the format: ask for newline-delimited JSON, one complete object per line. Each line is parseable as soon as its newline arrives, so the client can show the first result while the model writes the second. The prompt must be explicit about it, as in "output one JSON object per line, no surrounding array, no prose", and the parser must tolerate a bad line rather than abort the stream.

import json

def ndjson_items(chunks):
    """Yield one parsed object per completed line as the stream arrives.
    Prompt: 'Output one JSON object per line, no array, no prose.'"""
    buf = ""
    for chunk in chunks:
        buf += chunk
        while "\n" in buf:
            line, buf = buf.split("\n", 1)
            line = line.strip()
            if not line:
                continue
            try:
                yield json.loads(line)
            except json.JSONDecodeError:
                yield {"_error": "bad_line", "raw": line[:200]}   # surface, do not crash
    if buf.strip():                                    # stream ended without a final newline
        try:
            yield json.loads(buf)
        except json.JSONDecodeError:
            yield {"_error": "truncated_line", "raw": buf[:200]}

The second is a tolerant partial-JSON parser that closes open strings, arrays and objects to produce a best-effort view of the incomplete document. It suits a single large object, such as a form being filled, but values change as they arrive, so render only fields that are complete. In both cases, validate every finished object against the schema; streaming does not relax correctness. The structured JSON output article covers schemas and constrained decoding.

Markers: one stream, several panes

Interfaces often show more than one thing: the answer, a sources panel, suggested follow-up questions. With one request you can ask the model to separate them with sentinel markers, such as a line containing only a sources marker. The hazard is that a chunk boundary can split a marker in half, so a naive search shows the user part of a marker and then fails to switch panes.

The router below handles this by holding back only the few characters that could still become a marker. The added delay is at most the length of the longest marker minus one character, which is invisible.

class MarkerRouter:
    """Split a token stream into panes at sentinel markers, even when a marker
    is split across chunks. Holds back at most len(longest marker) - 1 chars."""

    def __init__(self, markers):
        self.markers = markers              # e.g. {"<<SOURCES>>": "sources"}
        self.pane = "answer"
        self.buf = ""
        self.hold = max(len(m) for m in markers) - 1

    def feed(self, chunk):
        self.buf += chunk
        out = []
        while True:
            hits = [(self.buf.find(m), m) for m in self.markers if m in self.buf]
            if not hits:
                break
            i, m = min(hits)
            if i:
                out.append((self.pane, self.buf[:i]))
            self.pane = self.markers[m]
            self.buf = self.buf[i + len(m):]
        safe = len(self.buf) - self.hold
        if safe > 0:                        # emit all but a possible marker prefix
            out.append((self.pane, self.buf[:safe]))
            self.buf = self.buf[safe:]
        return out

    def close(self):
        rest, self.buf = self.buf, ""
        return [(self.pane, rest)] if rest else []

Choose markers that never occur in normal text, keep them on their own line, and treat a missing marker as normal: if the model forgets to emit sources, the answer pane simply holds everything. Do not ask the model to emit HTML or JSON wrappers around a streaming answer for this purpose; markers are cheaper and fail more gracefully.

A client that ties it together

The client reads Server-Sent Events from your own relay, feeds each data line through the router and renders each pane with the repair function. A cancellation event, set when the user presses Stop, breaks the loop; closing the response closes the connection, and your relay must then cancel the upstream model request so generation, and billing, actually stop.

import asyncio, httpx, json

async def stream_answer(url, payload, on_text, cancel: asyncio.Event):
    router = MarkerRouter({"<<SOURCES>>": "sources"})
    async with httpx.AsyncClient(timeout=httpx.Timeout(10.0, read=60.0)) as client:
        async with client.stream("POST", url, json=payload) as resp:
            resp.raise_for_status()
            async for line in resp.aiter_lines():           # Server-Sent Events
                if cancel.is_set():
                    break                                   # closing the response aborts upstream
                if not line.startswith("data:"):
                    continue
                data = line[5:].strip()
                if data == "[DONE]":
                    break
                chunk = json.loads(data)["text"]           # relay JSON-encodes, so newlines survive
                for pane, text in router.feed(chunk):
                    on_text(pane, text)
    for pane, text in router.close():
        on_text(pane, text)

Keep the relay between browser and model, rather than streaming directly from a provider to the browser. The relay holds credentials, applies moderation, logs the stream for evaluation and enforces backpressure if a slow client cannot keep up. The Server-Sent Events introduction covers reconnection and the event format.

Moderating output that is already visible

Batch responses can be checked before anyone sees them. A stream cannot, unless you delay it. The practical pattern is a sliding buffer at the relay: hold the most recent few sentences, run a fast classifier or rule set on the buffer, and release text only once it has passed. The buffer size is a direct trade between safety and latency.

StrategyUser-visible delayRisk
No buffer, check after completionNoneHarmful text shown, then retracted
Sentence buffer, fast classifierAbout one sentenceProblems spanning sentences slip through
Paragraph bufferSeveral secondsFeels sluggish; most of streaming's benefit lost
Full response, then fake streamWhole generation timeSafe, but is not streaming

Whatever you choose, design the retraction path: if a check fails after text was shown, replace the message with a clear notice rather than silently deleting it, and log the event. Prompt-side, instruct the model to state refusals in the first sentence, so a refusal never follows half an answer. The guardrails article covers what to check.

Reasoning phases and tool calls

Two kinds of output are not answer text. Reasoning models may produce a long thinking phase first; tool-using agents pause while a tool runs. Both create silence after the first token, and silence reads as failure. Show a status line instead: "Checking your invoice history" is better than a spinner, and much better than raw reasoning text, which is often long, tentative and occasionally wrong.

Status lines can come from the client, mapped from tool names, which is reliable and translatable, or from the model, by asking it to write one short sentence before each tool call. The first is safer. Either way, the prompt should still put the final answer first once the model starts writing it.

Failure modes

SymptomCauseFix
Fast first token, slow answerPreamble and restated questionAnswer-first rules; measure time to useful token
Whole page renders as codeUnclosed code fence mid-streamRepair before every render
Results appear all at onceOne JSON array requestedNDJSON, one object per line
Marker fragments visibleMarker split across chunksHold back marker length minus one
Stop pressed, billing continuesUpstream request not cancelledPropagate abort through the relay
Answer contradicts visible reasoningReasoning written into the answerSeparate reasoning phase or status lines

What to do next

  1. Log full streams for a sample of real traffic and measure time to useful token next to time to first token.
  2. Rewrite the system prompt with explicit ordering rules: answer first, steps next, caveats and tables last.
  3. Add a markdown repair step before every render, and test it on answers cut at every chunk position.
  4. Switch any list-shaped structured output to NDJSON and validate each object as it completes.
  5. Put a relay between client and model that handles moderation buffering, cancellation and logging.
  6. Read the backpressure and circuit breakers article before scaling the relay.
Key takeaway: Once output is streamed, order is latency: users read the first sentence while the rest is generated, so the prompt must put the answer first and the preamble nowhere. Measure time to useful token rather than time to first token, choose formats that render safely when incomplete, stream structured data as one object per line, route panes with markers that survive chunk boundaries, buffer moderation deliberately, and make Stop actually stop the upstream request.