Streaming does not make a language model faster. A 400-token answer takes the same GPU time whether the client receives it in one response or in 400 small events. What streaming changes is when the user sees the first useful output, and that changes everything about how a serving stack has to be built: the HTTP response stays open for seconds or minutes, every proxy between the GPU and the browser must forward bytes immediately, a closed browser tab has to stop GPU work, and an error that happens after the 200 status line has already been sent can no longer be reported with a status code.
This article follows a token from the engine's decode step to the reader's screen. It covers the two latency numbers that streaming exposes (time to first token and inter-token latency), the Server-Sent Events wire format most LLM APIs use, incremental detokenisation, cancellation, backpressure, proxies and timeouts, and how to measure the result from the client side. It is not about StreamingLLM, the attention-sink technique for long contexts; that is covered in the StreamingLLM article.
TTFT and inter-token latency
Two numbers describe a streamed response. Time to first token (TTFT) is the gap between the client sending the request and receiving the first content token. It is the sum of queue wait, prefill of the whole prompt, one decode step, detokenisation and the network. Inter-token latency (ITL), sometimes reported as time per output token, is the gap between consecutive tokens. With continuous batching every running sequence advances one token per engine step, so ITL is essentially the duration of one batched decode step plus whatever delays the flush.
The two are driven by different resources. Prefill is compute bound and grows with prompt length; decode is memory-bandwidth bound and grows with batch size and total KV cache read per step. That is why a single knob rarely fixes both, and why the serving literature treats them as separate SLOs; the serving SLO article shows how to set percentile targets for each.
The delivery path
A typical open-source stack splits into an asynchronous API server and an engine. The engine runs a loop: admit waiting requests that fit in KV cache memory, run a step that mixes prefill chunks for new requests with one decode token for every running request, sample, and push each new token id into the output queue that belongs to its request. The API server holds one coroutine per open HTTP connection; each coroutine awaits its queue, turns token ids into text, frames the text as an event and writes it to the socket.
If the engine wrote to sockets directly, one slow client could stall a step serving two hundred other streams. The queue is the isolation boundary.
The wire format: Server-Sent Events
Most LLM HTTP APIs stream with Server-Sent Events: a normal HTTP response with Content-Type: text/event-stream, whose body is a sequence of events. Each event is one or more field: value lines terminated by a blank line. Lines starting with a colon are comments, which clients ignore and which make good heartbeats. The OpenAI-style convention that many servers copy sends one JSON chunk per data: line and finishes with a literal data: [DONE] sentinel. The protocol itself is described in the SSE introduction; the LLM-specific parts look like this:
HTTP/1.1 200 OK
Content-Type: text/event-stream
Cache-Control: no-cache
X-Accel-Buffering: no
data: {"id":"r1","choices":[{"index":0,"delta":{"role":"assistant"}}]}
data: {"id":"r1","choices":[{"index":0,"delta":{"content":"Stream"}}]}
data: {"id":"r1","choices":[{"index":0,"delta":{"content":"ing works"}}]}
: keepalive
data: {"id":"r1","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"id":"r1","choices":[],"usage":{"prompt_tokens":812,"completion_tokens":3}}
data: [DONE]Three details trip people up. First, the status line is committed before generation starts, so a failure at token 200 cannot become a 500; it must be sent as an error event in-band and the client must look for it. Second, usage accounting arrives at the end, so a client that disconnects early never receives it; billing and quota must be computed server-side from what was actually generated, not from what the client saw. Third, a missing terminator is a signal: if the stream ends without a finish reason or [DONE], the client should treat the answer as truncated, not complete.
Incremental detokenisation and stop strings
Tokens are not characters. A tokenizer may split one emoji or one CJK character across several byte-level tokens, and many tokenizers prepend or merge spaces in ways that depend on neighbouring tokens. Decoding each token in isolation therefore produces broken characters and wrong spacing. The fix is incremental detokenisation: decode a short window of recent tokens, compare it with what was already emitted, and send only the new suffix once it no longer ends in an incomplete character.
Stop strings add a second constraint. If the stop sequence is \n\nUser: and the model has produced \n\nUs, the server must not emit those characters yet, because the next token may complete the stop string and they would then have to be retracted, which a stream cannot do. The server holds back the longest suffix that is a prefix of any stop string. A compact version of both ideas:
class IncrementalDetok:
"""Emit only stable text: complete characters, never a possible stop-string prefix."""
def __init__(self, tokenizer, stops=()):
self.tok, self.stops = tokenizer, list(stops)
self.ids, self.text_sent = [], ""
def _holdback(self, text):
keep = 0
for s in self.stops:
for k in range(1, len(s)):
if text.endswith(s[:k]):
keep = max(keep, k)
return keep
def push(self, token_id):
"""Return (new_text, stopped)."""
self.ids.append(token_id)
full = self.tok.decode(self.ids)
if full.endswith("\ufffd"): # incomplete UTF-8 sequence: wait
return "", False
for s in self.stops:
i = full.find(s, max(0, len(self.text_sent) - len(s)))
if i != -1:
out = full[len(self.text_sent):i]
self.text_sent = full[:i]
return out, True
stable = len(full) - self._holdback(full)
out = full[len(self.text_sent):stable]
self.text_sent = full[:stable]
return out, FalseRe-decoding the whole id list is quadratic; production servers decode a short sliding window instead. Keep the simple version as a test oracle for the optimised one.
The streaming endpoint: framing, heartbeats and cancellation
The endpoint coroutine has four jobs: read the queue, frame events, notice disconnects, and keep idle connections alive. The sketch below uses an ASGI framework; engine is any object with an async generate iterator and an abort method, which is the shape of the common open-source engines.
import asyncio, json
HEARTBEAT_S = 15
async def sse_stream(request, engine, req_id, prompt, params, detok):
gen = engine.generate(prompt, params, req_id).__aiter__()
nxt = asyncio.ensure_future(gen.__anext__()) # never cancel it: that would end gen
sent_tokens = 0
try:
while True:
done, _ = await asyncio.wait({nxt}, timeout=HEARTBEAT_S)
if not done:
yield ": keepalive\n\n" # comment line: keeps proxies open
continue
try:
out = nxt.result()
except StopAsyncIteration:
break
nxt = asyncio.ensure_future(gen.__anext__())
if await request.is_disconnected():
raise asyncio.CancelledError
text, stopped = detok.push(out.token_id)
sent_tokens += 1
if text:
chunk = {"id": req_id, "choices": [{"index": 0, "delta": {"content": text}}]}
yield "data: " + json.dumps(chunk) + "\n\n"
if stopped or out.finished:
break
yield "data: " + json.dumps({"id": req_id, "choices": [],
"usage": {"completion_tokens": sent_tokens}}) + "\n\n"
yield "data: [DONE]\n\n"
except asyncio.CancelledError:
await engine.abort(req_id) # frees the sequence and its KV blocks
raise
except Exception as exc: # status 200 already sent: report in-band
yield "data: " + json.dumps({"error": {"message": str(exc)}}) + "\n\n"
await engine.abort(req_id)
finally:
nxt.cancel()
record_usage(req_id, sent_tokens) # bill what was generated, not what was readThe abort path is the one that saves money. Without it, a user who closes the tab after two seconds leaves a sequence decoding to its maximum length, holding KV cache blocks that could have admitted another request. In load tests with realistic abandonment, missing cancellation shows up as lower admitted concurrency and longer queue waits, which in turn raises TTFT for everyone else. Continuous batching makes the freed slot reusable on the very next step, so cancellation is worth implementing carefully.
Worked example: a latency budget and a long prompt
A worked budget makes the trade-offs concrete. The figures below are illustrative arithmetic for one replica, not a benchmark of any particular GPU or engine.
| Component | Assumption | Contribution |
|---|---|---|
| Queue wait | replica at 85% of admitted capacity | 40 ms at p50, 400 ms at p99 |
| Prefill | 2,000-token prompt at 10,000 prefill tokens/s | 200 ms |
| First decode step | batch of 64 sequences | 25 ms |
| Detokenise, frame, network | same region | 15 ms |
| TTFT | sum | about 280 ms at p50 |
| ITL | one 25 ms step per token | 40 tokens/s per stream |
| Total for 400 tokens | 280 ms + 399 x 25 ms | about 10.3 s |
Now a new request arrives with a 16,000-token prompt. Prefilled in one piece at the same rate it would occupy 1.6 seconds of GPU time, and every one of the 64 running streams would freeze for that long because they share the step. Readers see a stall mid-sentence. Chunked prefill caps the prefill tokens per step (for example 512), so the long prompt is spread over about 31 steps that each also advance the running streams; each of those steps costs about 51 ms of prefill plus the 25 ms decode, so ITL roughly triples for a few seconds instead of freezing for 1.6. The cost is a longer TTFT for the big request. Separating prefill and decode onto different GPUs removes the interference entirely at the cost of moving KV cache between them, as described in the disaggregated serving article.
Backpressure and slow clients
The engine produces tokens at GPU speed; a phone on a weak network may read them more slowly. If the per-request queue is unbounded, a stuck client accumulates memory on the API server indefinitely. If the engine blocks on a full queue, one slow client stalls the batch. Neither is acceptable, so the usual design is a bounded queue with a policy:
- Coalesce. When the socket is not writable, append new text to a pending buffer and send one larger event when it is. Readers do not notice 50 ms of batching; they do notice a frozen stream.
- Deadline. If nothing has been written for a configured interval (for example 30 seconds) while tokens are pending, treat the client as gone and abort the request.
- Never block the engine. The engine's put into the queue must be non-blocking. Overflow is handled on the API side by coalescing, not by back-pressuring the GPU loop.
Proxies, load balancers and retries
Most streaming bugs in production are not in the model server. They are in a component between it and the user that buffers, compresses or times out:
| Hop | Symptom | Fix |
|---|---|---|
| nginx reverse proxy | whole answer arrives at once | proxy_buffering off; or response header X-Accel-Buffering: no |
| Compression middleware | chunks held until a compression block fills | disable gzip for text/event-stream |
| Cloud load balancer | streams cut at a fixed idle time | raise the idle timeout above the longest silence; send heartbeat comments |
| HTTP/1.1 browsers | few concurrent streams per origin | serve over HTTP/2 so streams multiplex on one connection |
| Service mesh sidecar | retries a half-sent stream | disable retries for streaming routes; retry only before the first byte |
Retrying before the first byte is safe; retrying after part of an answer was shown appends a second, different answer.
Measuring what users experience
Server-side engine metrics stop at the queue. To know what users experience, measure from a client in the same network position as your users. A minimal probe:
import time, json, httpx
def probe(url, payload):
t0 = time.perf_counter()
stamps = []
with httpx.stream("POST", url, json={**payload, "stream": True}, timeout=120) as r:
for line in r.iter_lines():
if not line.startswith("data: ") or line == "data: [DONE]":
continue
evt = json.loads(line[6:])
if evt.get("choices") and evt["choices"][0]["delta"].get("content"):
stamps.append(time.perf_counter())
ttft = stamps[0] - t0
gaps = sorted(b - a for a, b in zip(stamps, stamps[1:]))
return {"ttft_ms": 1000 * ttft,
"itl_p50_ms": 1000 * gaps[len(gaps) // 2],
"itl_p99_ms": 1000 * gaps[int(len(gaps) * 0.99)],
"events": len(stamps)}Two warnings. Events are not tokens once coalescing is on, so report gaps between events and count tokens from the usage block. And run many probes in parallel, because ITL depends on what else shares the batch. If the probe's TTFT is far above the engine's own TTFT histogram, the difference is queueing or buffering outside the engine, which is exactly the part this article is about.
Failure modes and trade-offs
| Failure | What users see | Root cause and fix |
|---|---|---|
| Buffering proxy | spinner, then full answer | turn buffering off per route; test through the real ingress path |
| No cancellation | nothing; costs and queue waits rise | abort on disconnect, plus a write deadline |
| Prefill interference | mid-sentence stalls when long prompts arrive | chunked prefill or disaggregated prefill |
| Broken characters | replacement glyphs in non-English text | incremental detokenisation with UTF-8 completion check |
| Leaked stop string | partial stop sequence visible | hold back stop-string prefixes |
| Mid-stream error hidden | answer simply ends | in-band error event; client checks finish reason |
| Idle timeout | long reasoning answers cut off | heartbeat comments and longer LB idle timeout |
The trade-offs are mostly between smoothness and efficiency. Larger batches raise throughput but lengthen each step and therefore ITL. Coalescing saves CPU but makes text arrive in bursts. Disaggregation protects ITL but adds a KV transfer to TTFT and more machines to operate.
What to do next
- Run the probe through your real ingress (CDN, load balancer, proxy) and directly against the engine; any TTFT gap between the two is buffering or queueing to fix first.
- Verify cancellation: start a long generation, close the client after one second, and confirm the engine's running-sequence count drops on the next step.
- Add heartbeat comments and set load balancer idle timeouts above your longest expected silence, including long reasoning phases before the first visible token.
- Test detokenisation on emoji, CJK and right-to-left text, and with stop strings that span token boundaries.
- Enable chunked prefill, replay a trace that mixes short and very long prompts, and compare ITL p99 for the short streams before and after.
- Make clients treat a stream without a finish reason or terminator as truncated, and surface in-band error events to users.