A language model produces its answer one token at a time. Each decode step runs the whole network once and emits one token (or a few, with speculative decoding), so a 600-token answer is 600 sequential steps. Without streaming, the user stares at a spinner until the last step finishes. With streaming, the first words appear after a few hundred milliseconds and the rest arrive at reading speed. Total generation time is identical; perceived latency is not.
Streaming looks like a checkbox (stream=true) but it is really an architecture: a chain of components between the GPU and the screen, every one of which must pass bytes through as they arrive, survive a connection that stays open for a minute, and react when the user goes away. This article walks that chain hop by hop, with code for the parts that are easy to get wrong, then covers failure modes, metrics and trade-offs.
Why stream: two latencies, not one
Two numbers describe a streamed response. Time to first token (TTFT) is how long the user waits before anything appears: queueing, the prefill pass over the prompt, the first decode step, and the first-byte delay of every network hop. Inter-token latency (also called time per output token) is the gap between subsequent tokens and is set mostly by the decode step and the batch the engine is running. A non-streamed response is judged only on total time, which is roughly TTFT plus output tokens times inter-token latency.
Worked example: a chat answer of 600 tokens, TTFT of 400 ms, 25 ms per token. Total time is 0.4 + 600 x 0.025 = 15.4 seconds. Unstreamed, the user waits 15.4 seconds for anything. Streamed, they start reading at 0.4 seconds, and at 40 tokens per second (roughly 30 words per second) the text arrives faster than anyone reads it.
The path from decode step to pixels
The engine emits token IDs. A detokenizer turns them into text deltas. The model API frames the deltas as server-sent events. Your backend usually relays them, adding authentication, logging, filtering or tool execution. A reverse proxy, load balancer or CDN sits in front of your backend. The browser reads the body incrementally and a renderer paints. Every hop is a buffer, and a buffer that waits to fill before flushing turns your stream into one late response, silently.
Server-sent events and the wire formats
Most LLM APIs stream with server-sent events (SSE), a plain-text format over an ordinary HTTP response with Content-Type: text/event-stream. An event is one or more lines of the form field: value, terminated by a blank line. The fields are data (the payload; several data lines join with newlines), event (a type name), id (for resumption) and retry. A line starting with a colon is a comment, which servers use as a keep-alive. Chunk boundaries on the wire mean nothing: one TCP read can hold half an event or five events, so parsers must buffer until they see the blank line.
Two wire shapes dominate. OpenAI-style chat completions send unnamed events whose data is a JSON chunk with choices[0].delta.content, and end the stream with the literal data: [DONE]; token usage arrives in a final chunk only if you request it with stream_options: {"include_usage": true}. Anthropic's Messages API sends named events in a fixed order: message_start, then for each content block content_block_start, a run of content_block_delta and content_block_stop, then message_delta (stop reason and output usage) and message_stop, with ping and error events possible in between. Whatever the provider, write the parser once, against the format, not against one provider's typical chunk sizes:
def sse_events(byte_chunks):
"""Yield (event, data) pairs from an iterable of raw byte chunks."""
buf = ""
decoder = codecs.getincrementaldecoder("utf-8")()
for chunk in byte_chunks:
buf += decoder.decode(chunk) # never split a multi-byte char
buf = buf.replace("\r\n", "\n")
while "\n\n" in buf:
raw, buf = buf.split("\n\n", 1)
event, data = "message", []
for line in raw.split("\n"):
if not line or line.startswith(":"):
continue # keep-alive comment
field, _, value = line.partition(":")
value = value[1:] if value.startswith(" ") else value
if field == "event":
event = value
elif field == "data":
data.append(value)
if data:
yield event, "\n".join(data)Note the incremental UTF-8 decoder. Decoding each network chunk on its own corrupts any character whose bytes straddle a chunk boundary, which for CJK text and emoji happens constantly.
Incremental detokenization and stop sequences
If you run your own engine, the detokenizer is yours to get right, and it has two traps. The first is that tokens are not characters. Byte-level BPE vocabularies contain tokens that are fragments of a UTF-8 sequence, so decoding token by token produces replacement characters halfway through an emoji. The fix is to decode the growing sequence and emit only the new, complete text: keep an offset into the decoded string and hold back output while the tail ends in an incomplete character.
The second trap is stop sequences. If the stop string is \nUser: and you stream each delta immediately, the client sees \nUs before the engine discovers the match and stops. The engine has to hold back any suffix that could be the start of a stop sequence, and release it only once it is proven harmless:
def stream_with_stops(deltas, stops):
pending = ""
for delta in deltas:
pending += delta
hit = min((pending.find(s) for s in stops if s in pending), default=-1)
if hit >= 0:
yield pending[:hit] # emit up to the stop, then end
return
# keep the longest suffix that is a prefix of some stop string
keep = 0
for s in stops:
for k in range(min(len(s) - 1, len(pending)), 0, -1):
if s.startswith(pending[-k:]):
keep = max(keep, k)
break
if len(pending) > keep:
yield pending[:len(pending) - keep]
pending = pending[len(pending) - keep:]
yield pending # stream ended without a stopHold-back is bounded by the longest stop string minus one character, so it costs a few tokens of latency at most. Any wrapper that adds its own stop logic must repeat it.
The backend relay: forward, keep alive, cancel
Most applications do not hand the provider stream straight to the browser. A backend relay keeps the API key server-side, enforces quotas, logs the full response, runs tools and may rewrite events into its own schema. A relay has four jobs: forward each event as soon as it arrives, send keep-alives when the model is quiet (long tool calls, long prefill), notice when the client disconnects, and propagate that cancellation upstream. A Starlette/FastAPI sketch:
from starlette.responses import StreamingResponse
async def chat(request):
body = await request.json()
upstream = await provider.open_stream(body) # async iterator of (event, data)
async def relay():
text = []
try:
async for event, data in with_heartbeat(upstream, every=15):
if await request.is_disconnected():
break # user left: stop paying
if event == "heartbeat":
yield ": keep-alive\n\n"
continue
text.append(extract_text(event, data))
yield f"event: {event}\ndata: {data}\n\n"
except ProviderError as e:
yield f"event: error\ndata: {json.dumps({'message': str(e)})}\n\n"
finally:
await upstream.aclose() # cancels generation upstream
await audit_log.write(request, "".join(text))
return StreamingResponse(relay(), media_type="text/event-stream", headers={
"Cache-Control": "no-cache",
"X-Accel-Buffering": "no", # nginx: do not buffer this response
})The finally block is the important part. Closing the upstream iterator closes the upstream HTTP connection, and a well-behaved inference server treats that as an abort, removing the sequence from its running batch and freeing its KV-cache blocks. Skip it and every abandoned tab keeps a GPU slot busy generating text nobody will read. In continuous-batching servers that is not just wasted money; it is throughput stolen from live users.
Proxies, load balancers and other stream killers
Intermediaries are where streams die quietly. The usual suspects:
| Hop | What goes wrong | Fix |
|---|---|---|
| nginx reverse proxy | proxy_buffering on by default collects the response before sending | proxy_buffering off; for the route, or the X-Accel-Buffering: no header |
| Compression middleware | gzip waits to fill a block, so tokens arrive in bursts | exclude text/event-stream from compression |
| Load balancer idle timeout | connection cut during a long silent tool call or prefill | comment keep-alives every 10-20 s; raise the idle timeout |
| Serverless platforms | some buffer whole responses or cap response duration | check the platform's streaming support before designing on it |
Test streaming end to end through the real path, not against localhost. A simple check: curl -N against the public URL and watch whether tokens trickle or arrive at once.
The client: reading and rendering a stream
The browser's EventSource only issues GET requests without a body and cannot set an Authorization header, so LLM front ends usually POST with fetch and read the body stream themselves, feeding bytes through the same parser logic. Two client-side details matter. Use TextDecoder with {stream: true} for the same multi-byte reason as the server. And do not re-render on every token: re-parsing Markdown and re-highlighting code 40 times a second makes long answers stutter. Append deltas to a buffer and paint once per animation frame:
const res = await fetch("/api/chat", { method: "POST", body, signal: controller.signal });
const reader = res.body.getReader();
const decoder = new TextDecoder();
let pending = "", text = "", scheduled = false;
for (;;) {
const { value, done } = await reader.read();
if (done) break;
pending += decoder.decode(value, { stream: true });
let i;
while ((i = pending.indexOf("\n\n")) >= 0) {
const evt = parseEvent(pending.slice(0, i)); // same rules as the server parser
pending = pending.slice(i + 2);
if (evt.type === "error") throw new Error(evt.data);
text += deltaText(evt);
}
if (!scheduled) {
scheduled = true;
requestAnimationFrame(() => { scheduled = false; render(text); });
}
}
// "Stop generating" button: controller.abort() closes the socket, which the relay sees.
Streaming tool calls and structured output
Structured output and tool calls stream too, as fragments of JSON. Anthropic sends tool arguments as input_json_delta events carrying partial_json strings; OpenAI-style chunks carry tool_calls[i].function.arguments fragments indexed by call. The fragments are not valid JSON individually and may split a string or number anywhere. The rule is simple: accumulate per block or per call index, parse once the block is closed, and only then act. Executing a tool on partially streamed arguments is how an agent deletes the wrong file. A tolerant partial-JSON parser is fine for showing a form filling in on screen; it is never fine for side effects.
Failure modes unique to streaming
The failure modes specific to streaming:
- Errors after the 200. The status line and headers left before generation finished, so an overload or a safety stop mid-answer cannot become a 5xx. It arrives as an error event or simply as a truncated stream. Clients must treat a stream that ends without the terminal event (
[DONE]ormessage_stop) as failed, not as complete. - Retries duplicate text. Generation is not resumable in general: a retry samples a new answer. Either discard what was shown and restart visibly, or have the relay persist the text it has sent under a response ID and serve reconnects from that store (the SSE
Last-Event-IDheader exists for exactly this replay). - Lost usage accounting. Token usage usually arrives at the end. A cut stream has none, so meter from the relay's own count of forwarded tokens, or from the provider's billing records, not only from the final event.
- Connection pressure. A streamed request holds a socket and a worker for its whole duration. Thread-per-request servers run out of threads; use async I/O in the relay and size connection limits for concurrent streams, not requests per second.
- Zombie generations. Missing cancellation propagation means abandoned requests keep consuming GPU time. Measure it: compare tokens generated upstream to tokens delivered downstream.
Metrics that tell you the stream is healthy
Instrument the relay, since it sees both sides. Record TTFT (request received to first content byte sent), inter-token latency percentiles, time to last token, output tokens, and the end state of every stream: completed, client-cancelled, upstream error, or timed out. A rising client-cancel rate is often the first sign that answers are slow or bad. Alert on p95 TTFT and on streams with no terminal event.
Transport trade-offs
| Transport | Strengths | Weaknesses |
|---|---|---|
| SSE over HTTP/1.1 or HTTP/2 | plain HTTP, works through most proxies, trivial to debug with curl, provider-native | one direction; client cancel means closing the request |
| WebSocket | bidirectional: interrupt, steer, send audio while receiving | stateful connections, sticky load balancing, your own framing and reconnect logic |
| gRPC server streaming | typed messages, flow control, good service to service | browsers need a gateway; overkill for one-way text |
| No streaming | simplest; whole-response validation before display | worst perceived latency; fine for batch and back-end jobs |
Default to SSE for user-facing text. Choose WebSockets when the client must talk while the model is talking, as in voice agents. Turn streaming off for batch jobs, evaluations and any pipeline where a machine consumes the whole answer; there it only adds failure modes.
What to do next
Related reading on this site: continuous batching and vLLM batching and the KV cache explain the engine that sets inter-token latency; speculative decoding explains how several tokens can arrive per step; LLM tool use and structured output cover what the streamed JSON means.
- Run
curl -Nagainst your public endpoint and confirm tokens trickle; fix any buffering hop before anything else. - Replace per-chunk UTF-8 decoding with incremental decoders on both server and client.
- Add a
finallyin the relay that closes the upstream stream, and verify on the inference server that a closed tab frees its slot. - Send comment keep-alives every 10-20 seconds and check load-balancer idle timeouts against your longest tool call.
- Treat a stream without a terminal event as a failure in the client, with a visible retry.
- Accumulate tool-call JSON per block and execute only after the block closes.
- Log TTFT, inter-token latency, end state and generated-versus-delivered tokens per stream, and alert on p95 TTFT.