A language model produces its answer one token at a time, and a ten-second answer that appears all at once feels broken while the same answer streamed from the first quarter-second feels fast. Almost every chat product therefore streams, and almost all of them use Server-Sent Events: model providers stream to your server over SSE, and your server streams to the browser the same way. The format is simple, which is exactly why teams underestimate the pattern. The hard parts are around it: a request that must be a POST, an error that arrives after the 200 status, a proxy that buffers the whole response, a closed tab that keeps spending tokens, and a phone that changes networks halfway through an answer.
This article builds the pattern end to end, with real server and client code. It assumes you know what SSE is; if not, start with Server-Sent Events: one-way server-to-client streaming.
Why SSE fits token streaming, and why not EventSource
Token streaming is one long request with a one-way flow of small text messages: SSE's shape, a normal HTTP response with content type text/event-stream that stays open while the server writes events. It passes through ordinary HTTP infrastructure, authenticates like any request, and needs no protocol upgrade. A WebSocket adds a second, bidirectional channel you do not need for one answer, though it can make sense for a session with many interleaved messages; the trade-offs are in SSE vs WebSocket.
The browser's built-in EventSource is usually the wrong client for this job. It only issues GET requests, cannot send a request body, and cannot set custom headers such as an Authorization bearer token. A chat request carries a conversation, model options and tool definitions, which belong in a POST body. So LLM front ends call fetch with a POST, read the response body as a stream, and parse the SSE format themselves. The wire format is the same; you simply take over the parser, and with it the reconnect behaviour that EventSource would otherwise provide.
The shape of the system
The relay holds the API key, enforces quotas, records usage and translates each provider's stream into one schema.
The wire format, as a parser sees it
An SSE stream is UTF-8 text made of events separated by a blank line. Each line is a field: data: carries payload (several data lines in one event are joined with newlines), event: names the event type, id: sets the last event ID, retry: suggests a reconnect delay, and a line starting with a colon is a comment that clients ignore, which makes it the standard heartbeat. Lines may end in CRLF, LF or CR.
The trap for hand-written parsers: the network delivers bytes, not lines. A chunk can end mid-line or mid-character, so decode with a streaming decoder, buffer the unfinished line, and dispatch only on the blank line.
Define your own event schema
Providers each stream differently. OpenAI's Chat Completions API sends unnamed data events whose JSON carries a choices[0].delta and ends with a literal data: [DONE]. Its Responses API sends typed events such as response.created, response.output_text.delta and response.completed, each with a sequence_number. Anthropic's Messages API sends named events: message_start, content_block_start, content_block_delta, content_block_stop, message_delta and message_stop, plus ping and error events. Formats evolve, so check the provider reference when you write the adapter, and do not let any of them leak into your front end.
Instead, define a small schema of your own and keep it stable:
id: g7f3:0
event: meta
data: {"generation_id":"g7f3","model":"large-2026-09"}
id: g7f3:1
event: delta
data: {"text":"Logs flow "}
id: g7f3:2
event: delta
data: {"text":"through three stages."}
: keep-alive
id: g7f3:3
event: usage
data: {"input_tokens":812,"output_tokens":9}
id: g7f3:4
event: done
data: {"finish_reason":"stop"}Five event types cover most products: meta once at the start, delta for text, optional tool events for tool calls, usage near the end, and exactly one terminal event, either done or error. The terminal event matters: a stream that simply stops is indistinguishable from a dropped connection, so the client must treat a stream that ends without done or error as interrupted. The IDs combine the generation and a sequence number, which is what makes resume possible later.
The relay server
Here is a relay in Python with Starlette (FastAPI exposes the same objects). provider_stream stands for your adapter, an async iterator that turns one provider's stream into (type, payload) pairs.
import asyncio, json, uuid
from starlette.requests import Request
from starlette.responses import StreamingResponse
HEARTBEAT_S = 15
_END = object()
def sse(event_type, data, eid=None):
head = f"id: {eid}\n" if eid is not None else ""
return f"{head}event: {event_type}\ndata: {json.dumps(data)}\n\n"
async def chat(request: Request):
body = await request.json()
gen_id = uuid.uuid4().hex[:8]
async def events():
queue = asyncio.Queue()
upstream = provider_stream(body["messages"], model=body.get("model"))
async def pump(): # reads the provider; never cancelled mid-read by a timeout
try:
async for item in upstream:
await queue.put(item)
await queue.put(_END)
except Exception as exc:
await queue.put(exc)
task = asyncio.create_task(pump())
seq = 0
yield sse("meta", {"generation_id": gen_id}, f"{gen_id}:{seq}")
try:
while True:
try:
item = await asyncio.wait_for(queue.get(), HEARTBEAT_S)
except asyncio.TimeoutError:
yield ": keep-alive\n\n" # long tool call: keep proxies and the client alive
continue
if item is _END:
seq += 1
yield sse("done", {"finish_reason": "stop"}, f"{gen_id}:{seq}")
return
if isinstance(item, Exception): # after the first byte the status is already 200
log_exception(gen_id, item)
yield sse("error", {"message": "generation failed", "retryable": True})
return
if await request.is_disconnected():
return # finally cancels upstream
kind, payload = item
seq += 1
yield sse(kind, payload, f"{gen_id}:{seq}")
finally:
task.cancel() # stop the pump and let it unwind first,
await asyncio.gather(task, return_exceptions=True)
await upstream.aclose() # then close the provider stream: billing stops
return StreamingResponse(events(), media_type="text/event-stream", headers={
"Cache-Control": "no-cache",
"X-Accel-Buffering": "no", # nginx: do not buffer this response
})Three details carry the weight. The heartbeat keeps intermediaries from closing an idle connection during a tool call. The finally block cancels the pump, waits for it to unwind (closing a generator that is still mid-read raises), then closes the upstream stream on every exit path, which stops generation and billing when the user leaves. And errors after the first byte become an in-band error event, because the status line was sent long ago. The pump task and queue exist so that a heartbeat timeout never cancels a read in the middle of the provider stream; the only thing that waits with a timeout is the queue.
The browser client
The client posts the conversation, reads the body as a stream of text, parses events and uses an AbortController for the Stop button.
async function streamChat(messages, onEvent, signal) {
const res = await fetch("/api/chat", {
method: "POST",
headers: { "Content-Type": "application/json", "Authorization": `Bearer ${token}` },
body: JSON.stringify({ messages }),
signal,
});
if (!res.ok) throw new Error(`HTTP ${res.status}`); // errors before streaming starts
const reader = res.body.pipeThrough(new TextDecoderStream()).getReader();
let buf = "", ev = { type: "message", data: [], id: null }, terminal = false;
for (;;) {
const { value, done } = await reader.read();
if (done) break;
buf += value;
const lines = buf.split(/\r\n|\r|\n/);
buf = lines.pop(); // keep the unfinished line
for (const line of lines) {
if (line === "") { // blank line: dispatch
if (ev.data.length) {
const msg = { type: ev.type, id: ev.id, data: JSON.parse(ev.data.join("\n")) };
if (msg.type === "done" || msg.type === "error") terminal = true;
onEvent(msg);
}
ev = { type: "message", data: [], id: ev.id };
} else if (line.startsWith(":")) { // comment / heartbeat
} else {
const i = line.indexOf(":");
const field = i < 0 ? line : line.slice(0, i);
const val = i < 0 ? "" : line.slice(i + 1).replace(/^ /, "");
if (field === "data") ev.data.push(val);
else if (field === "event") ev.type = val;
else if (field === "id") ev.id = val;
}
}
}
if (!terminal) throw new Error("stream interrupted"); // caller may resume from last id
}
const ctl = new AbortController();
stopButton.onclick = () => ctl.abort();One simplification: a CRLF split across two chunks reads as two line endings. Servers that write only LF, like the relay above, never trigger it.
Cancellation is a cost control, not a nicety
Generation is billed per output token, and the provider keeps generating until its own connection closes. So the abort must travel the whole chain: Stop button, AbortController, the browser closes the connection, the proxy closes its upstream connection, your relay notices, and your relay closes the provider stream. Any link that fails to propagate turns abandoned tabs into spend.
Servers often notice a vanished client only on their next write, another reason for heartbeats. Test it: close a tab mid-answer and confirm in usage logs that output tokens stopped within seconds.
Buffering: the first-token killer
The most common production bug is a stream that works locally and arrives all at once in production. Something between server and browser is buffering: a reverse proxy collecting the response before forwarding it, response compression waiting for a full block, or a framework that buffers until the handler returns. The symptom is that time to first token equals total generation time.
- nginx: set
proxy_buffering offfor the route, or sendX-Accel-Buffering: nofrom the app. - Compression: exclude
text/event-streamfrom gzip or brotli middleware, or make sure it flushes per event. - Load balancers: raise the idle timeout above your longest silent gap, and keep heartbeats well inside it; many managed balancers default to about a minute.
Measure time to first delta at the client; the server can flush perfectly while a proxy holds everything.
Resumable generations
By default the generation is tied to the connection: if a mobile client changes networks mid-answer, the stream dies and the user must regenerate, paying again. For short chat answers that is usually acceptable. For long agent runs, reports or code generation that take minutes, decouple generation from delivery.
A worker runs the generation and appends each event, with its ID, to a log keyed by generation ID, for example a Redis stream or a database table with a time-to-live. The SSE endpoint reads from that log rather than from the provider. A reconnecting client sends the last ID it saw, and the endpoint replays everything after it and then follows live events. The replay-then-follow race, cursor design and status codes are covered in SSE reconnection and Last-Event-ID.
The cost: you store every token, and a disconnect no longer stops spending, so set a timeout after which a generation with no connected client is cancelled.
Worked example: one answer on a moving phone
Follow a single 900-token answer. At 0 ms the client POSTs. The relay authenticates, writes meta immediately, and the browser renders a typing indicator. The provider's first delta arrives a few hundred milliseconds later and is forwarded at once. The client appends deltas to a buffer and re-renders at most once per animation frame, because re-rendering Markdown on every delta is the main client-side cost.
At token 400 the phone switches from Wi-Fi to cellular. The fetch fails without a terminal event, so the client knows the answer is incomplete. With a generation log, it reconnects with the last ID, g7f3:212, receives the events generated while it was offline in one burst, and continues live; the user sees a pause, not a restart. Without a log, the client shows the partial answer with a Retry control, and the relay's finally block has already cancelled the upstream request.
Failure modes
- Everything arrives at the end. A buffering proxy or compression layer. Fix the route config and measure at the client.
- Streams cut at a fixed duration. An idle or maximum-duration timeout at a load balancer or gateway, usually during a long tool call. Heartbeats and higher limits.
- Silent truncation. The client treats end of body as success. Require a terminal event.
- Abandoned tabs keep spending. Cancellation does not reach the provider. Close upstream in a finally path and test it.
- Tabs starving each other. HTTP/1.1 caps connections per host; serve over HTTP/2 or HTTP/3.
- Broken rendering mid-stream. Half a Markdown code fence or table renders oddly until it closes. Use a renderer that tolerates incomplete input, and never insert model output as raw HTML; sanitise it like any untrusted input.
- Double generation on retry. A client retries the POST after a drop and starts a second paid generation. Send an idempotency key with the request and reattach to the existing generation.
Trade-offs
| Choice | When it wins | What it costs |
|---|---|---|
| SSE over fetch | One-way answer streams through ordinary HTTP | You own parsing and reconnect |
| EventSource | GET-able streams with cookie auth, free reconnect | No body, no custom headers |
| WebSocket | Many interleaved messages each way in one session | Separate scaling, auth and proxy story |
| Generation log | Long generations, mobile users | Token storage, worker lifecycle, cancellation policy |
For the scaling side, including fan-out and connection counts, see scaling SSE horizontally; for slow consumers, backpressure in streaming.
What to do next
- Write down your event schema with a single terminal event, and make the client treat any stream without it as interrupted.
- Put a relay in front of every provider and write one adapter per provider stream format.
- Add a comment heartbeat every 10 to 20 seconds and confirm your load balancer's idle timeout is comfortably longer.
- Disable buffering and compression on the streaming route, then measure time to first delta from a real browser in production.
- Close the upstream request on every exit path and prove it: close a tab mid-answer and check that output tokens stop.
- Add idempotency keys so a retried POST cannot start a second paid generation.
- If generations run for minutes or users are mobile, add a generation log with ID-based resume and a no-client cancellation timeout.