Streaming does not make a language model faster. A 400-token answer takes the same GPU time whether the client receives it in one response or in 400 small events. What streaming changes is when the user sees the first useful output, and that changes everything about how a serving stack has to be built: the HTTP response stays open for seconds or minutes, every proxy between the GPU and the browser must forward bytes immediately, a closed browser tab has to stop GPU work, and an error that happens after the 200 status line has already been sent can no longer be reported with a status code.

This article follows a token from the engine's decode step to the reader's screen. It covers the two latency numbers that streaming exposes (time to first token and inter-token latency), the Server-Sent Events wire format most LLM APIs use, incremental detokenisation, cancellation, backpressure, proxies and timeouts, and how to measure the result from the client side. It is not about StreamingLLM, the attention-sink technique for long contexts; that is covered in the StreamingLLM article.

TTFT and inter-token latency

Two numbers describe a streamed response. Time to first token (TTFT) is the gap between the client sending the request and receiving the first content token. It is the sum of queue wait, prefill of the whole prompt, one decode step, detokenisation and the network. Inter-token latency (ITL), sometimes reported as time per output token, is the gap between consecutive tokens. With continuous batching every running sequence advances one token per engine step, so ITL is essentially the duration of one batched decode step plus whatever delays the flush.

The two are driven by different resources. Prefill is compute bound and grows with prompt length; decode is memory-bandwidth bound and grows with batch size and total KV cache read per step. That is why a single knob rarely fixes both, and why the serving literature treats them as separate SLOs; the serving SLO article shows how to set percentile targets for each.

The delivery path

A typical open-source stack splits into an asynchronous API server and an engine. The engine runs a loop: admit waiting requests that fit in KV cache memory, run a step that mixes prefill chunks for new requests with one decode token for every running request, sample, and push each new token id into the output queue that belongs to its request. The API server holds one coroutine per open HTTP connection; each coroutine awaits its queue, turns token ids into text, frames the text as an event and writes it to the socket.

Token streaming: from the GPU step loop to the reader's screenClientSSE readerLB / proxybuffering OFFAPI serverHTTP + detokenizerPer-request queueone per streamtoken idsEngine scheduler loopadmit, prefill chunk, decode stepfan outGPUone token per running seqDisconnect detectedabort(request_id)free KV blockssocket closedTTFT = queue wait + prefill + first decode step + network. ITL = one batched decode step + flush.Every hop between the GPU and the client must flush per event, or streaming silently turns back into batch.
The delivery path. The engine never talks to sockets; it fills per-request queues. The API server owns framing, detokenisation and disconnect handling.

If the engine wrote to sockets directly, one slow client could stall a step serving two hundred other streams. The queue is the isolation boundary.

The wire format: Server-Sent Events

Most LLM HTTP APIs stream with Server-Sent Events: a normal HTTP response with Content-Type: text/event-stream, whose body is a sequence of events. Each event is one or more field: value lines terminated by a blank line. Lines starting with a colon are comments, which clients ignore and which make good heartbeats. The OpenAI-style convention that many servers copy sends one JSON chunk per data: line and finishes with a literal data: [DONE] sentinel. The protocol itself is described in the SSE introduction; the LLM-specific parts look like this:

HTTP/1.1 200 OK
Content-Type: text/event-stream
Cache-Control: no-cache
X-Accel-Buffering: no

data: {"id":"r1","choices":[{"index":0,"delta":{"role":"assistant"}}]}

data: {"id":"r1","choices":[{"index":0,"delta":{"content":"Stream"}}]}

data: {"id":"r1","choices":[{"index":0,"delta":{"content":"ing works"}}]}

: keepalive

data: {"id":"r1","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: {"id":"r1","choices":[],"usage":{"prompt_tokens":812,"completion_tokens":3}}

data: [DONE]

Three details trip people up. First, the status line is committed before generation starts, so a failure at token 200 cannot become a 500; it must be sent as an error event in-band and the client must look for it. Second, usage accounting arrives at the end, so a client that disconnects early never receives it; billing and quota must be computed server-side from what was actually generated, not from what the client saw. Third, a missing terminator is a signal: if the stream ends without a finish reason or [DONE], the client should treat the answer as truncated, not complete.

Incremental detokenisation and stop strings

Tokens are not characters. A tokenizer may split one emoji or one CJK character across several byte-level tokens, and many tokenizers prepend or merge spaces in ways that depend on neighbouring tokens. Decoding each token in isolation therefore produces broken characters and wrong spacing. The fix is incremental detokenisation: decode a short window of recent tokens, compare it with what was already emitted, and send only the new suffix once it no longer ends in an incomplete character.

Stop strings add a second constraint. If the stop sequence is \n\nUser: and the model has produced \n\nUs, the server must not emit those characters yet, because the next token may complete the stop string and they would then have to be retracted, which a stream cannot do. The server holds back the longest suffix that is a prefix of any stop string. A compact version of both ideas:

class IncrementalDetok:
    """Emit only stable text: complete characters, never a possible stop-string prefix."""

    def __init__(self, tokenizer, stops=()):
        self.tok, self.stops = tokenizer, list(stops)
        self.ids, self.text_sent = [], ""

    def _holdback(self, text):
        keep = 0
        for s in self.stops:
            for k in range(1, len(s)):
                if text.endswith(s[:k]):
                    keep = max(keep, k)
        return keep

    def push(self, token_id):
        """Return (new_text, stopped)."""
        self.ids.append(token_id)
        full = self.tok.decode(self.ids)
        if full.endswith("\ufffd"):           # incomplete UTF-8 sequence: wait
            return "", False
        for s in self.stops:
            i = full.find(s, max(0, len(self.text_sent) - len(s)))
            if i != -1:
                out = full[len(self.text_sent):i]
                self.text_sent = full[:i]
                return out, True
        stable = len(full) - self._holdback(full)
        out = full[len(self.text_sent):stable]
        self.text_sent = full[:stable]
        return out, False

Re-decoding the whole id list is quadratic; production servers decode a short sliding window instead. Keep the simple version as a test oracle for the optimised one.

The streaming endpoint: framing, heartbeats and cancellation

The endpoint coroutine has four jobs: read the queue, frame events, notice disconnects, and keep idle connections alive. The sketch below uses an ASGI framework; engine is any object with an async generate iterator and an abort method, which is the shape of the common open-source engines.

import asyncio, json

HEARTBEAT_S = 15

async def sse_stream(request, engine, req_id, prompt, params, detok):
    gen = engine.generate(prompt, params, req_id).__aiter__()
    nxt = asyncio.ensure_future(gen.__anext__())   # never cancel it: that would end gen
    sent_tokens = 0
    try:
        while True:
            done, _ = await asyncio.wait({nxt}, timeout=HEARTBEAT_S)
            if not done:
                yield ": keepalive\n\n"           # comment line: keeps proxies open
                continue
            try:
                out = nxt.result()
            except StopAsyncIteration:
                break
            nxt = asyncio.ensure_future(gen.__anext__())
            if await request.is_disconnected():
                raise asyncio.CancelledError
            text, stopped = detok.push(out.token_id)
            sent_tokens += 1
            if text:
                chunk = {"id": req_id, "choices": [{"index": 0, "delta": {"content": text}}]}
                yield "data: " + json.dumps(chunk) + "\n\n"
            if stopped or out.finished:
                break
        yield "data: " + json.dumps({"id": req_id, "choices": [],
                                     "usage": {"completion_tokens": sent_tokens}}) + "\n\n"
        yield "data: [DONE]\n\n"
    except asyncio.CancelledError:
        await engine.abort(req_id)                 # frees the sequence and its KV blocks
        raise
    except Exception as exc:                       # status 200 already sent: report in-band
        yield "data: " + json.dumps({"error": {"message": str(exc)}}) + "\n\n"
        await engine.abort(req_id)
    finally:
        nxt.cancel()
        record_usage(req_id, sent_tokens)          # bill what was generated, not what was read

The abort path is the one that saves money. Without it, a user who closes the tab after two seconds leaves a sequence decoding to its maximum length, holding KV cache blocks that could have admitted another request. In load tests with realistic abandonment, missing cancellation shows up as lower admitted concurrency and longer queue waits, which in turn raises TTFT for everyone else. Continuous batching makes the freed slot reusable on the very next step, so cancellation is worth implementing carefully.

Worked example: a latency budget and a long prompt

A worked budget makes the trade-offs concrete. The figures below are illustrative arithmetic for one replica, not a benchmark of any particular GPU or engine.

ComponentAssumptionContribution
Queue waitreplica at 85% of admitted capacity40 ms at p50, 400 ms at p99
Prefill2,000-token prompt at 10,000 prefill tokens/s200 ms
First decode stepbatch of 64 sequences25 ms
Detokenise, frame, networksame region15 ms
TTFTsumabout 280 ms at p50
ITLone 25 ms step per token40 tokens/s per stream
Total for 400 tokens280 ms + 399 x 25 msabout 10.3 s

Now a new request arrives with a 16,000-token prompt. Prefilled in one piece at the same rate it would occupy 1.6 seconds of GPU time, and every one of the 64 running streams would freeze for that long because they share the step. Readers see a stall mid-sentence. Chunked prefill caps the prefill tokens per step (for example 512), so the long prompt is spread over about 31 steps that each also advance the running streams; each of those steps costs about 51 ms of prefill plus the 25 ms decode, so ITL roughly triples for a few seconds instead of freezing for 1.6. The cost is a longer TTFT for the big request. Separating prefill and decode onto different GPUs removes the interference entirely at the cost of moving KV cache between them, as described in the disaggregated serving article.

Backpressure and slow clients

The engine produces tokens at GPU speed; a phone on a weak network may read them more slowly. If the per-request queue is unbounded, a stuck client accumulates memory on the API server indefinitely. If the engine blocks on a full queue, one slow client stalls the batch. Neither is acceptable, so the usual design is a bounded queue with a policy:

  • Coalesce. When the socket is not writable, append new text to a pending buffer and send one larger event when it is. Readers do not notice 50 ms of batching; they do notice a frozen stream.
  • Deadline. If nothing has been written for a configured interval (for example 30 seconds) while tokens are pending, treat the client as gone and abort the request.
  • Never block the engine. The engine's put into the queue must be non-blocking. Overflow is handled on the API side by coalescing, not by back-pressuring the GPU loop.

Proxies, load balancers and retries

Most streaming bugs in production are not in the model server. They are in a component between it and the user that buffers, compresses or times out:

HopSymptomFix
nginx reverse proxywhole answer arrives at onceproxy_buffering off; or response header X-Accel-Buffering: no
Compression middlewarechunks held until a compression block fillsdisable gzip for text/event-stream
Cloud load balancerstreams cut at a fixed idle timeraise the idle timeout above the longest silence; send heartbeat comments
HTTP/1.1 browsersfew concurrent streams per originserve over HTTP/2 so streams multiplex on one connection
Service mesh sidecarretries a half-sent streamdisable retries for streaming routes; retry only before the first byte

Retrying before the first byte is safe; retrying after part of an answer was shown appends a second, different answer.

Measuring what users experience

Server-side engine metrics stop at the queue. To know what users experience, measure from a client in the same network position as your users. A minimal probe:

import time, json, httpx

def probe(url, payload):
    t0 = time.perf_counter()
    stamps = []
    with httpx.stream("POST", url, json={**payload, "stream": True}, timeout=120) as r:
        for line in r.iter_lines():
            if not line.startswith("data: ") or line == "data: [DONE]":
                continue
            evt = json.loads(line[6:])
            if evt.get("choices") and evt["choices"][0]["delta"].get("content"):
                stamps.append(time.perf_counter())
    ttft = stamps[0] - t0
    gaps = sorted(b - a for a, b in zip(stamps, stamps[1:]))
    return {"ttft_ms": 1000 * ttft,
            "itl_p50_ms": 1000 * gaps[len(gaps) // 2],
            "itl_p99_ms": 1000 * gaps[int(len(gaps) * 0.99)],
            "events": len(stamps)}

Two warnings. Events are not tokens once coalescing is on, so report gaps between events and count tokens from the usage block. And run many probes in parallel, because ITL depends on what else shares the batch. If the probe's TTFT is far above the engine's own TTFT histogram, the difference is queueing or buffering outside the engine, which is exactly the part this article is about.

Failure modes and trade-offs

FailureWhat users seeRoot cause and fix
Buffering proxyspinner, then full answerturn buffering off per route; test through the real ingress path
No cancellationnothing; costs and queue waits riseabort on disconnect, plus a write deadline
Prefill interferencemid-sentence stalls when long prompts arrivechunked prefill or disaggregated prefill
Broken charactersreplacement glyphs in non-English textincremental detokenisation with UTF-8 completion check
Leaked stop stringpartial stop sequence visiblehold back stop-string prefixes
Mid-stream error hiddenanswer simply endsin-band error event; client checks finish reason
Idle timeoutlong reasoning answers cut offheartbeat comments and longer LB idle timeout

The trade-offs are mostly between smoothness and efficiency. Larger batches raise throughput but lengthen each step and therefore ITL. Coalescing saves CPU but makes text arrive in bursts. Disaggregation protects ITL but adds a KV transfer to TTFT and more machines to operate.

What to do next

  1. Run the probe through your real ingress (CDN, load balancer, proxy) and directly against the engine; any TTFT gap between the two is buffering or queueing to fix first.
  2. Verify cancellation: start a long generation, close the client after one second, and confirm the engine's running-sequence count drops on the next step.
  3. Add heartbeat comments and set load balancer idle timeouts above your longest expected silence, including long reasoning phases before the first visible token.
  4. Test detokenisation on emoji, CJK and right-to-left text, and with stop strings that span token boundaries.
  5. Enable chunked prefill, replay a trace that mixes short and very long prompts, and compare ITL p99 for the short streams before and after.
  6. Make clients treat a stream without a finish reason or terminator as truncated, and surface in-band error events to users.
Key takeaway: Streaming moves no work off the GPU; it changes when output becomes visible and turns every response into a long-lived connection. Treat TTFT and inter-token latency as separate budgets, keep the engine loop isolated from sockets with per-request queues, detokenise incrementally, abort on disconnect so KV cache is freed, and make every hop to the client flush per event. Measure from the client, because that is the only place the whole path is visible.