Server-Sent Events have a reputation for being forgiving. The browser reconnects on its own, the wire format is plain text, and a broken connection usually heals in a few seconds. That reputation hides a real design problem: once the server has sent a 200 and the first bytes, it has no HTTP-level way to report anything else. A query that times out, a model that refuses, a permission revoked mid-stream or a malformed event all have to be expressed inside the stream, or they will be expressed as a dropped connection that the browser quietly retries forever.

This article classifies failures by layer, designs an in-band error event that does not collide with the browser's own error signal, handles failures after the response is committed, builds a client dispatcher that survives bad events, and shows how to stop a stream on purpose.

Advertisement

Four layers, four kinds of failure

Every SSE failure comes from one of four layers, each with its own signal and correct response. Mixing them up causes most bad error handling, such as retrying a permission error forever.

LayerExampleWhat the client seesRight response
Connection setup401, 403, 503, wrong Content-TypeEventSource: one error event, readyState CLOSEDClassify by status (fetch) or by state (EventSource); refresh auth or back off
Transport after setupProxy idle timeout, Wi-Fi switch, server restarterror event, readyState CONNECTING, automatic retryUsually nothing; show degraded only after several attempts
Application, mid-streamQuery timeout, upstream 429, revoked accessWhatever the server chooses to sendA typed problem event, then continue or end deliberately
Client processingInvalid JSON, unknown event version, handler bugAn exception inside your listenerIsolate per event, count, never kill the stream

Connection setup and status codes are covered in SSE reconnection and Last-Event-ID, and how readyState behaves across hidden tabs, sleep and token expiry is covered in the SSE client lifecycle. The rest of this article concentrates on the two layers those pages leave open: application errors after the stream has started, and errors in your own client code.

ProducerDB query, model, queueSSE handlerstatus already 200Proxy / CDNbuffers, idle timeoutsBrowser transportEventSource or fetchDispatcherparse, validate, routeUI statelive, degraded, stoppedError telemetrycounts by classthrowsproblem eventbytes or dropeventstate changeclassifiedonerrorEach layer fails differently; each error should reach the UI as one of three states
Where SSE errors originate and where they should be handled: producer failures become typed events, transport failures surface through onerror, and every path ends in one of three UI states.

What the specification lets you do

A few rules from the WHATWG HTML specification decide what an error design can and cannot do. They are worth knowing exactly, because each one has caused a production incident somewhere.

  • Setup failures are fatal. If the response status is not 200 or the Content-Type is not text/event-stream, the browser fails the connection and does not retry. The spec names 204 No Content as the way to tell a client to stop reconnecting.
  • Everything after setup reconnects. A body that ends normally and a network error both lead the browser to reestablish the connection after the reconnection time. A server that closes the stream to signal "finished" has in fact asked for a reconnect.
  • retry must be digits. The field is honoured only if its value is ASCII digits; retry: 5s is silently ignored.
  • An unterminated event is dropped. If the body ends before the blank line that terminates an event, that event is discarded. An error message written just before a crash may never be dispatched.
  • An event without data is not dispatched. If the data buffer is empty when the blank line arrives, nothing fires. event: done followed by a blank line does nothing at all.
  • The event field sets the DOM event type. A server line event: error dispatches an event whose type is error.

The last rule matters most. Many tutorials suggest sending a custom event: error, but in the browser it lands in the same onerror handler as real connection failures, and every helper that assumes error means "connection trouble" will misbehave. Pick a name the platform does not use, such as problem.

Advertisement

In-band problem events

An in-band error is a normal SSE event with its own type and a JSON body. Borrowing field names from RFC 9457 problem details keeps it familiar to anyone who has read your HTTP error responses, and adding a retry hint and a fatal flag lets the client act without guessing.

event: problem
id: 18342
data: {"type":"https://api.example.com/problems/upstream-timeout",
data:  "title":"Pricing service timed out","status":504,
data:  "fatal":false,"retryAfterMs":2000,"scope":"stream",
data:  "traceId":"4bf92f3577b34da6"}

Multiple data: lines are joined with newlines, so the JSON above parses as one object. Four fields carry the behaviour:

  • fatal says whether the stream will continue. Non-fatal problems are informational: a partial result, a skipped item, a degraded dependency.
  • scope says what failed: the whole stream, one item (with its ID), or the session (authorization).
  • retryAfterMs lets the server pace the client during an incident instead of relying on the browser's default reconnection time.
  • traceId ties the event to server logs. Never put stack traces or internal hostnames in the body; the stream is readable by the user and by any extension in their browser.

Give problem events an id like any other event when they are part of the replayable history, so a client that reconnects learns about a skipped item it missed. Leave the id off for transient notices, such as "pricing is slow", that would be misleading on replay.

Errors after the 200 is committed

The hardest errors happen after the response is committed. The status line has gone, so you cannot switch to a 500, and the decision is between three exits: report and continue, report and finish, or abort. The handler below makes that decision explicitly and makes sure the problem event is complete and flushed before anything else happens.

// Node.js, framework-free. send() writes one complete event, blank line included.
function send(res, { event, id, data }) {
  let out = '';
  if (event) out += `event: ${event}\n`;
  if (id !== undefined) out += `id: ${id}\n`;
  for (const line of JSON.stringify(data).split('\n')) out += `data: ${line}\n`;
  res.write(out + '\n');
}

async function streamReport(req, res) {
  res.writeHead(200, { 'Content-Type': 'text/event-stream',
                       'Cache-Control': 'no-store', 'X-Accel-Buffering': 'no' });
  res.write('retry: 3000\n\n');                 // digits only
  try {
    for await (const row of producer(req)) {
      try {
        send(res, { event: 'row', id: row.seq, data: row });
      } catch (err) {                           // one bad row: report, keep going
        send(res, { event: 'problem', id: row.seq, data: problem(err, 'item', false) });
      }
    }
    send(res, { event: 'end', data: { ok: true } });     // data is required
  } catch (err) {                               // producer died: report, then finish
    send(res, { event: 'problem', data: problem(err, 'stream', isFatal(err)) });
    send(res, { event: 'end', data: { ok: false } });
  } finally {
    res.end();                                  // a clean end, never a socket reset
  }
}

Three details make this robust. The terminal event carries data, because an event with an empty data buffer is never dispatched. Each event is written as one string ending in a blank line, so a crash can lose a whole event but never leave half of one to be silently discarded. And the stream ends with res.end() after the terminal event, so the client has already decided whether to reconnect before the browser's reconnect logic sees the closed body.

Abort instead of ending only when you cannot trust what you have already sent, for example when a downstream consistency check fails. Then destroy the socket and let the client resume from its last good Last-Event-ID. Keep long-lived streams alive with comment lines, as in heartbeat and keepalive strategies, so proxies do not turn quiet periods into transport errors.

A client dispatcher that survives bad events

On the client, the goal is that one bad event never takes the stream down, and that every error reaches the UI as one of three states: live, degraded (retrying or partial), or stopped (needs the user or a new session). A thrown exception inside a listener is reported to the console but does not close the EventSource, so a crashing handler fails silently: the stream stays open and the UI stops updating. Wrap every handler.

const es = new EventSource('/reports/42/stream');
const stats = { badEvents: 0, transportErrors: 0 };

function on(type, handler) {
  es.addEventListener(type, (ev) => {
    let msg;
    try { msg = JSON.parse(ev.data); }
    catch { stats.badEvents++; report('parse', type, ev.lastEventId); return; }
    try { handler(msg, ev); }
    catch (err) { stats.badEvents++; report('handler', type, ev.lastEventId, err); }
  });
}

on('row', (row) => table.upsert(row));
on('problem', (p) => {
  if (p.scope === 'item') return table.markSkipped(p);
  if (p.scope === 'session') { es.close(); return ui.set('stopped', 'Sign in again'); }
  ui.set(p.fatal ? 'stopped' : 'degraded', p.title);
  if (p.fatal) es.close();
});
on('end', (e) => { es.close(); ui.set(e.ok ? 'done' : 'stopped'); });

es.onopen = () => { stats.transportErrors = 0; ui.set('live'); };
es.onerror = () => {                    // transport only: no server event is named 'error'
  if (es.readyState === EventSource.CLOSED) return ui.set('stopped', 'Connection refused');
  if (++stats.transportErrors >= 3) ui.set('degraded', 'Reconnecting...');
};

Notice the threshold on transport errors. A single reconnect is normal behaviour on mobile networks and behind proxies with idle timeouts, and flashing a warning for it trains users to ignore warnings. Count attempts, reset on open, and only escalate when the pattern persists.

Stopping a stream on purpose

Because a closed body means "reconnect", stopping a stream on purpose takes two cooperating parts. The server sends a terminal event (end, or a problem with fatal: true) and the client calls es.close() when it sees it. That covers the normal case. For the client that missed the terminal event, perhaps because its connection dropped at the same moment, the server must also answer the next reconnect with 204 No Content, which the specification defines as the stop signal. Keep a short-lived record of finished stream IDs so the reconnect can be recognised.

During incidents, send retry: 30000 (randomised per connection, since the browser applies it exactly) before ending streams you are shedding, so the herd returns over a wider window.

fetch-based clients: classify, then retry

If you read the stream with fetch and a ReadableStream parser, as most LLM front ends do (see SSE for LLM streaming), you get status codes back and lose automatic reconnection. That trade lets you classify properly:

function classify(err, res) {
  if (err?.name === 'AbortError') return 'cancelled';      // user navigated or pressed stop
  if (res && (res.status === 401 || res.status === 403)) return 'auth';
  if (res && (res.status === 204 || res.status === 404 || res.status === 410)) return 'stop';
  if (res && res.status === 429) return 'throttled';        // honour Retry-After
  if (res && res.status >= 400 && res.status < 500) return 'fatal';
  return 'transient';                                       // 5xx, network, mid-body drop
}

function backoff(attempt, baseMs = 500, capMs = 30000) {
  return Math.random() * Math.min(capMs, baseMs * 2 ** attempt);   // full jitter
}

Only transient and throttled retry. auth refreshes the credential once and then stops. Everything else ends the stream and tells the user. The same backoff shape is discussed for WebSockets in WebSocket reconnection strategies.

Worked example: a failover during a 50,000-row export

Consider a reporting page that streams 50,000 rows of a quarterly export. At row 31,200 the database replica serving the query is failed over and the cursor dies with a connection error. Rows 18,004 and 18,005 earlier had a currency code the serializer did not recognise.

With naive handling, the serializer exception at row 18,004 crashed the handler, the browser reconnected with Last-Event-ID: 18003, hit the same row, and looped. With the patterns above:

  1. Rows 18,004 and 18,005 each produce an item-scoped problem event with their IDs. The client marks two rows as skipped; the stream continues.
  2. At 31,200 the producer throws. The handler sends a stream-scoped problem with fatal: false, retryAfterMs: 5000, then end with ok: false, then ends the body cleanly.
  3. The client shows "degraded: replica failover, resuming" and opens a new stream after five seconds with the last row ID as a query parameter, because a fresh EventSource does not send Last-Event-ID.
  4. Telemetry records two item problems and one stream problem with the same trace ID as the server log, so the on-call engineer sees a failover, not a mystery.

Failure modes

Failure modeSymptomFix
Server sends event: errorConnection handler fires with data; UI says "offline"Rename to problem or similar
Terminal event with no data lineClient never closes; endless reconnectsAlways include a data: line
Close body to mean "done"Finished jobs restart every few secondsTerminal event plus close(), and 204 on reconnect
Unhandled listener exceptionStream open, UI frozen, no error shownWrap handlers; count and report bad events
Poison event replayed on resumeSame crash after every reconnectItem-scoped problem, advance past the ID
Error body leaks internalsStack traces visible in DevToolsProblem type and trace ID only
retry: 5sIgnored; default delay usedDigits only, in milliseconds

Trade-offs

In-band errors add a schema you must version, and old clients ignore unknown event types, so ship the dispatcher before the server emits problems. Ending and resuming on every producer failure costs a reconnect and a replay: cheap for rows, expensive for LLM generations, where a non-fatal problem on an open stream is better. Moving from EventSource to fetch buys status codes but makes you own reconnection, jitter and resume.

What to do next

  1. Grep your server code for event: error and rename it; update clients in the same release.
  2. Define a problem event schema with fatal, scope, retryAfterMs and a trace ID, and document it next to your HTTP error format.
  3. Make every server write one complete event per call, and give terminal events a data: line.
  4. Add a finished-stream record and answer reconnects to finished streams with 204.
  5. Wrap every client listener, count parse and handler failures, and send the counts to telemetry.
  6. Map every error to live, degraded or stopped, and only show degraded after repeated transport errors.
  7. Run a game day: kill the producer mid-stream, inject a malformed event, and confirm the stream survives the first and stops cleanly when asked.
Key takeaway: After the first bytes, an SSE server can only report errors inside the stream, and a closed body means reconnect, not done. Classify failures by layer, send typed problem events under a name other than error, write each event whole and give terminal events data, wrap every client handler, map errors to live, degraded or stopped, and stop streams with a terminal event, close() and a 204 for late reconnects.