Server-Sent Events have a reputation for being forgiving. The browser reconnects on its own, the wire format is plain text, and a broken connection usually heals in a few seconds. That reputation hides a real design problem: once the server has sent a 200 and the first bytes, it has no HTTP-level way to report anything else. A query that times out, a model that refuses, a permission revoked mid-stream or a malformed event all have to be expressed inside the stream, or they will be expressed as a dropped connection that the browser quietly retries forever.
This article classifies failures by layer, designs an in-band error event that does not collide with the browser's own error signal, handles failures after the response is committed, builds a client dispatcher that survives bad events, and shows how to stop a stream on purpose.
Four layers, four kinds of failure
Every SSE failure comes from one of four layers, each with its own signal and correct response. Mixing them up causes most bad error handling, such as retrying a permission error forever.
| Layer | Example | What the client sees | Right response |
|---|---|---|---|
| Connection setup | 401, 403, 503, wrong Content-Type | EventSource: one error event, readyState CLOSED | Classify by status (fetch) or by state (EventSource); refresh auth or back off |
| Transport after setup | Proxy idle timeout, Wi-Fi switch, server restart | error event, readyState CONNECTING, automatic retry | Usually nothing; show degraded only after several attempts |
| Application, mid-stream | Query timeout, upstream 429, revoked access | Whatever the server chooses to send | A typed problem event, then continue or end deliberately |
| Client processing | Invalid JSON, unknown event version, handler bug | An exception inside your listener | Isolate per event, count, never kill the stream |
Connection setup and status codes are covered in SSE reconnection and Last-Event-ID, and how readyState behaves across hidden tabs, sleep and token expiry is covered in the SSE client lifecycle. The rest of this article concentrates on the two layers those pages leave open: application errors after the stream has started, and errors in your own client code.
What the specification lets you do
A few rules from the WHATWG HTML specification decide what an error design can and cannot do. They are worth knowing exactly, because each one has caused a production incident somewhere.
- Setup failures are fatal. If the response status is not 200 or the
Content-Typeis nottext/event-stream, the browser fails the connection and does not retry. The spec names204 No Contentas the way to tell a client to stop reconnecting. - Everything after setup reconnects. A body that ends normally and a network error both lead the browser to reestablish the connection after the reconnection time. A server that closes the stream to signal "finished" has in fact asked for a reconnect.
retrymust be digits. The field is honoured only if its value is ASCII digits;retry: 5sis silently ignored.- An unterminated event is dropped. If the body ends before the blank line that terminates an event, that event is discarded. An error message written just before a crash may never be dispatched.
- An event without data is not dispatched. If the data buffer is empty when the blank line arrives, nothing fires.
event: donefollowed by a blank line does nothing at all. - The
eventfield sets the DOM event type. A server lineevent: errordispatches an event whose type iserror.
The last rule matters most. Many tutorials suggest sending a custom event: error, but in the browser it lands in the same onerror handler as real connection failures, and every helper that assumes error means "connection trouble" will misbehave. Pick a name the platform does not use, such as problem.
In-band problem events
An in-band error is a normal SSE event with its own type and a JSON body. Borrowing field names from RFC 9457 problem details keeps it familiar to anyone who has read your HTTP error responses, and adding a retry hint and a fatal flag lets the client act without guessing.
event: problem
id: 18342
data: {"type":"https://api.example.com/problems/upstream-timeout",
data: "title":"Pricing service timed out","status":504,
data: "fatal":false,"retryAfterMs":2000,"scope":"stream",
data: "traceId":"4bf92f3577b34da6"}Multiple data: lines are joined with newlines, so the JSON above parses as one object. Four fields carry the behaviour:
fatalsays whether the stream will continue. Non-fatal problems are informational: a partial result, a skipped item, a degraded dependency.scopesays what failed: the wholestream, oneitem(with its ID), or thesession(authorization).retryAfterMslets the server pace the client during an incident instead of relying on the browser's default reconnection time.traceIdties the event to server logs. Never put stack traces or internal hostnames in the body; the stream is readable by the user and by any extension in their browser.
Give problem events an id like any other event when they are part of the replayable history, so a client that reconnects learns about a skipped item it missed. Leave the id off for transient notices, such as "pricing is slow", that would be misleading on replay.
Errors after the 200 is committed
The hardest errors happen after the response is committed. The status line has gone, so you cannot switch to a 500, and the decision is between three exits: report and continue, report and finish, or abort. The handler below makes that decision explicitly and makes sure the problem event is complete and flushed before anything else happens.
// Node.js, framework-free. send() writes one complete event, blank line included.
function send(res, { event, id, data }) {
let out = '';
if (event) out += `event: ${event}\n`;
if (id !== undefined) out += `id: ${id}\n`;
for (const line of JSON.stringify(data).split('\n')) out += `data: ${line}\n`;
res.write(out + '\n');
}
async function streamReport(req, res) {
res.writeHead(200, { 'Content-Type': 'text/event-stream',
'Cache-Control': 'no-store', 'X-Accel-Buffering': 'no' });
res.write('retry: 3000\n\n'); // digits only
try {
for await (const row of producer(req)) {
try {
send(res, { event: 'row', id: row.seq, data: row });
} catch (err) { // one bad row: report, keep going
send(res, { event: 'problem', id: row.seq, data: problem(err, 'item', false) });
}
}
send(res, { event: 'end', data: { ok: true } }); // data is required
} catch (err) { // producer died: report, then finish
send(res, { event: 'problem', data: problem(err, 'stream', isFatal(err)) });
send(res, { event: 'end', data: { ok: false } });
} finally {
res.end(); // a clean end, never a socket reset
}
}Three details make this robust. The terminal event carries data, because an event with an empty data buffer is never dispatched. Each event is written as one string ending in a blank line, so a crash can lose a whole event but never leave half of one to be silently discarded. And the stream ends with res.end() after the terminal event, so the client has already decided whether to reconnect before the browser's reconnect logic sees the closed body.
Abort instead of ending only when you cannot trust what you have already sent, for example when a downstream consistency check fails. Then destroy the socket and let the client resume from its last good Last-Event-ID. Keep long-lived streams alive with comment lines, as in heartbeat and keepalive strategies, so proxies do not turn quiet periods into transport errors.
A client dispatcher that survives bad events
On the client, the goal is that one bad event never takes the stream down, and that every error reaches the UI as one of three states: live, degraded (retrying or partial), or stopped (needs the user or a new session). A thrown exception inside a listener is reported to the console but does not close the EventSource, so a crashing handler fails silently: the stream stays open and the UI stops updating. Wrap every handler.
const es = new EventSource('/reports/42/stream');
const stats = { badEvents: 0, transportErrors: 0 };
function on(type, handler) {
es.addEventListener(type, (ev) => {
let msg;
try { msg = JSON.parse(ev.data); }
catch { stats.badEvents++; report('parse', type, ev.lastEventId); return; }
try { handler(msg, ev); }
catch (err) { stats.badEvents++; report('handler', type, ev.lastEventId, err); }
});
}
on('row', (row) => table.upsert(row));
on('problem', (p) => {
if (p.scope === 'item') return table.markSkipped(p);
if (p.scope === 'session') { es.close(); return ui.set('stopped', 'Sign in again'); }
ui.set(p.fatal ? 'stopped' : 'degraded', p.title);
if (p.fatal) es.close();
});
on('end', (e) => { es.close(); ui.set(e.ok ? 'done' : 'stopped'); });
es.onopen = () => { stats.transportErrors = 0; ui.set('live'); };
es.onerror = () => { // transport only: no server event is named 'error'
if (es.readyState === EventSource.CLOSED) return ui.set('stopped', 'Connection refused');
if (++stats.transportErrors >= 3) ui.set('degraded', 'Reconnecting...');
};Notice the threshold on transport errors. A single reconnect is normal behaviour on mobile networks and behind proxies with idle timeouts, and flashing a warning for it trains users to ignore warnings. Count attempts, reset on open, and only escalate when the pattern persists.
Stopping a stream on purpose
Because a closed body means "reconnect", stopping a stream on purpose takes two cooperating parts. The server sends a terminal event (end, or a problem with fatal: true) and the client calls es.close() when it sees it. That covers the normal case. For the client that missed the terminal event, perhaps because its connection dropped at the same moment, the server must also answer the next reconnect with 204 No Content, which the specification defines as the stop signal. Keep a short-lived record of finished stream IDs so the reconnect can be recognised.
During incidents, send retry: 30000 (randomised per connection, since the browser applies it exactly) before ending streams you are shedding, so the herd returns over a wider window.
fetch-based clients: classify, then retry
If you read the stream with fetch and a ReadableStream parser, as most LLM front ends do (see SSE for LLM streaming), you get status codes back and lose automatic reconnection. That trade lets you classify properly:
function classify(err, res) {
if (err?.name === 'AbortError') return 'cancelled'; // user navigated or pressed stop
if (res && (res.status === 401 || res.status === 403)) return 'auth';
if (res && (res.status === 204 || res.status === 404 || res.status === 410)) return 'stop';
if (res && res.status === 429) return 'throttled'; // honour Retry-After
if (res && res.status >= 400 && res.status < 500) return 'fatal';
return 'transient'; // 5xx, network, mid-body drop
}
function backoff(attempt, baseMs = 500, capMs = 30000) {
return Math.random() * Math.min(capMs, baseMs * 2 ** attempt); // full jitter
}Only transient and throttled retry. auth refreshes the credential once and then stops. Everything else ends the stream and tells the user. The same backoff shape is discussed for WebSockets in WebSocket reconnection strategies.
Worked example: a failover during a 50,000-row export
Consider a reporting page that streams 50,000 rows of a quarterly export. At row 31,200 the database replica serving the query is failed over and the cursor dies with a connection error. Rows 18,004 and 18,005 earlier had a currency code the serializer did not recognise.
With naive handling, the serializer exception at row 18,004 crashed the handler, the browser reconnected with Last-Event-ID: 18003, hit the same row, and looped. With the patterns above:
- Rows 18,004 and 18,005 each produce an item-scoped
problemevent with their IDs. The client marks two rows as skipped; the stream continues. - At 31,200 the producer throws. The handler sends a stream-scoped problem with
fatal: false,retryAfterMs: 5000, thenendwithok: false, then ends the body cleanly. - The client shows "degraded: replica failover, resuming" and opens a new stream after five seconds with the last row ID as a query parameter, because a fresh
EventSourcedoes not sendLast-Event-ID. - Telemetry records two item problems and one stream problem with the same trace ID as the server log, so the on-call engineer sees a failover, not a mystery.
Failure modes
| Failure mode | Symptom | Fix |
|---|---|---|
Server sends event: error | Connection handler fires with data; UI says "offline" | Rename to problem or similar |
| Terminal event with no data line | Client never closes; endless reconnects | Always include a data: line |
| Close body to mean "done" | Finished jobs restart every few seconds | Terminal event plus close(), and 204 on reconnect |
| Unhandled listener exception | Stream open, UI frozen, no error shown | Wrap handlers; count and report bad events |
| Poison event replayed on resume | Same crash after every reconnect | Item-scoped problem, advance past the ID |
| Error body leaks internals | Stack traces visible in DevTools | Problem type and trace ID only |
retry: 5s | Ignored; default delay used | Digits only, in milliseconds |
Trade-offs
In-band errors add a schema you must version, and old clients ignore unknown event types, so ship the dispatcher before the server emits problems. Ending and resuming on every producer failure costs a reconnect and a replay: cheap for rows, expensive for LLM generations, where a non-fatal problem on an open stream is better. Moving from EventSource to fetch buys status codes but makes you own reconnection, jitter and resume.
What to do next
- Grep your server code for
event: errorand rename it; update clients in the same release. - Define a
problemevent schema withfatal,scope,retryAfterMsand a trace ID, and document it next to your HTTP error format. - Make every server write one complete event per call, and give terminal events a
data:line. - Add a finished-stream record and answer reconnects to finished streams with 204.
- Wrap every client listener, count parse and handler failures, and send the counts to telemetry.
- Map every error to live, degraded or stopped, and only show degraded after repeated transport errors.
- Run a game day: kill the producer mid-stream, inject a malformed event, and confirm the stream survives the first and stops cleanly when asked.