Server-Sent Events look simple to operate: one HTTP response that never ends, a few lines of data: text at a time. That simplicity is exactly why they are hard to observe. Standard HTTP dashboards were built for short requests, so a healthy SSE service can show a p99 latency of forty minutes, a 0 percent error rate while every client is silently reconnecting every sixty seconds, and green status codes while a proxy holds every event in a buffer. The request is the wrong unit of measurement. The connection and the event are the right ones.

This article builds an observability model for SSE from first principles: what to measure at the pipeline, the server and the client, how to instrument a server with real code, how to log and trace long-lived streams without drowning in data, which alerts catch real failures, and a worked diagnosis of a reconnect storm caused by a proxy timeout. It assumes you know the protocol basics; if not, start with our reconnect and Last-Event-ID article, which covers the wire contract this monitoring depends on.

Advertisement

Why request metrics lie about streams

A normal HTTP request has a start, an end and a status, and the latency between them tells you about user experience. An SSE response has a start, a status that is decided in the first milliseconds, and then an open-ended body that may last hours. Three things follow.

  • Duration is not latency. The time from request to response end is a connection lifetime. Mixed into a request-latency histogram, it ruins percentiles for every other route on the service. Exclude SSE routes from request-latency service level objectives and measure their lifetimes separately.
  • Status codes describe only the start. A stream that opened with 200 and then stalled, was cut by a proxy or fell hours behind is still recorded as a success. Most SSE failures happen after the status line.
  • The client hides failure. EventSource reconnects automatically after a dropped connection, so users often see a stale screen rather than an error. Server-side counters show many short successful requests; nobody is paged.

So the model has two units. The connection has an open time, a duration, a close reason and a count of events delivered. The event has a creation time, a delivery time on each connection, and a size. Every useful SSE metric is a count, gauge or histogram over one of these two.

What to measure, and where

Producerstamps created_atBroker / logevent idsSSE serverper-connection queueProxy / LBbuffer, idleBrowserEventSourceM1 Pipelinepublish rate, consumer lagM2 Serveropen streams, queue depth, writesM3 Client RUMtime to first event, errorsbeaconDerived signalsevent lag, close reasons, reconnect rate, duration shapeA stream is healthy only if all three vantage points agree
Where to instrument a Server-Sent Events delivery path: the event pipeline (M1), the SSE server and its per-connection queues (M2), and the browser through real-user monitoring (M3). The useful signals, such as end-to-end event lag and close reasons, are derived by joining them.

Instrument three vantage points, because each sees failures the others cannot.

SignalTypeWhy it matters
Open streamsGauge, per instanceCapacity and balance; a drop to zero on one node is a broken node
Connections opened, closedCounters, closed labelled by reasonChurn; close reasons separate client, server, proxy and slow-consumer causes
Connection durationHistogram with long bucketsSpikes at fixed values reveal timeouts
Resume requestsCounter: requests carrying Last-Event-IDReconnect rate as the server sees it
Replayed events per resumeHistogramHow far behind clients fall during gaps
Events and bytes sentCounters by event typeThroughput and cost
Per-connection queue depthHistogram or max gaugeSlow consumers before they exhaust memory
Slow-consumer evictionsCounterConnections closed because they could not keep up
End-to-end event lagHistogram: write time minus created_atThe number users actually feel
Time to first eventClient histogramDetects proxy buffering the server cannot see
Client errors by readyStateClient counterFatal closes versus automatic retries

Keep label cardinality bounded. Label by route, event type, instance and close reason; never by user id, connection id or topic name if topics are per-user. Those belong in logs and traces, not in metric labels, or a hundred thousand connections become a hundred thousand time series.

Advertisement

Instrumenting the server

The sketch below instruments a Starlette SSE endpoint with the Python prometheus_client library. The structure carries over to any stack: count the open in one place, count the close in a finally block with a reason, and measure every event as it is written. subscribe stands for your broker client and Event carries an id, a type, a payload and a created_at timestamp set by the producer.

import asyncio, time
from prometheus_client import Counter, Gauge, Histogram
from starlette.responses import StreamingResponse

OPEN = Gauge("sse_open_streams", "Open SSE streams", ["route"])
OPENED = Counter("sse_streams_opened_total", "Streams opened", ["route", "resumed"])
CLOSED = Counter("sse_streams_closed_total", "Streams closed", ["route", "reason"])
DURATION = Histogram("sse_stream_duration_seconds", "Stream lifetime", ["route"],
                     buckets=[1, 5, 15, 30, 55, 60, 65, 120, 300, 900, 3600, 14400])
SENT = Counter("sse_events_sent_total", "Events written", ["route", "type"])
LAG = Histogram("sse_event_lag_seconds", "created_at to write", ["route"],
                buckets=[0.05, 0.1, 0.25, 0.5, 1, 2, 5, 15, 60])
QUEUE_MAX = 1000

async def stream(request, route="orders"):
    resumed = "true" if request.headers.get("last-event-id") else "false"
    OPENED.labels(route, resumed).inc()
    OPEN.labels(route).inc()
    started, reason = time.monotonic(), "server_end"
    queue = await subscribe(request, maxsize=QUEUE_MAX)

    async def body():
        nonlocal reason
        try:
            yield "retry: 5000\n\n"
            while True:
                try:
                    ev = await asyncio.wait_for(queue.get(), timeout=15)
                except asyncio.TimeoutError:
                    yield ": keepalive\n\n"          # comment line, ignored by clients
                    continue
                if queue.qsize() > QUEUE_MAX * 0.9:
                    reason = "slow_consumer"
                    return
                yield f"id: {ev.id}\nevent: {ev.type}\ndata: {ev.json}\n\n"
                SENT.labels(route, ev.type).inc()
                LAG.labels(route).observe(max(0.0, time.time() - ev.created_at))
        except (asyncio.CancelledError, GeneratorExit):
            reason = "client_gone"                    # disconnect cancels the generator
            raise
        except Exception:
            reason = "error"
            raise
        finally:
            await queue.close()
            OPEN.labels(route).dec()
            CLOSED.labels(route, reason).inc()
            DURATION.labels(route).observe(time.monotonic() - started)

    return StreamingResponse(body(), media_type="text/event-stream",
                             headers={"Cache-Control": "no-cache",
                                      "X-Accel-Buffering": "no"})

Details that matter: the duration buckets deliberately bracket common idle timeouts (55, 60, 65 seconds) so a timeout shows up as a spike rather than disappearing inside a wide bucket. Lag uses wall-clock time against a producer timestamp, which is only as accurate as clock synchronisation between hosts; keep hosts on NTP and treat lag under about 50 milliseconds as noise. Heartbeat comments are not counted as events, so a stream that carries only heartbeats is visible as zero event throughput with open connections. And a graceful shutdown should set reason = "shutdown" before closing streams, so deploys do not look like failures.

Client-side telemetry

The browser sees what the server cannot: buffering in proxies, mobile networks dropping idle connections, and fatal failures. EventSource exposes a readyState of CONNECTING (0), OPEN (1) or CLOSED (2), and an error event that carries no status code. The useful distinction is what the state is after an error: CONNECTING means the browser will retry; CLOSED means it gave up, for example because the response was not 200 or did not have the text/event-stream content type, and nothing will recover without your code.

const t0 = performance.now();
let firstEvent = null, errors = 0;
const es = new EventSource("/events/orders");

es.addEventListener("order", (e) => {
  if (firstEvent === null) {
    firstEvent = performance.now() - t0;
    report({ metric: "sse_time_to_first_event_ms", value: firstEvent });
  }
});
es.onerror = () => {
  errors += 1;
  const fatal = es.readyState === EventSource.CLOSED;
  report({ metric: "sse_client_error", fatal, errors });
  if (fatal) scheduleManualReconnect();   // the browser will not retry by itself
};

function report(payload) {
  navigator.sendBeacon("/rum", JSON.stringify({ ...payload, page: location.pathname }));
}

Sample this telemetry, aggregate it server-side into histograms, and compare time to first event with the server's time to first write. If the server wrote within 50 milliseconds but clients report two seconds, something in between is buffering. Our client lifecycle article covers how hidden tabs and sleep affect these numbers, which matters when interpreting them.

Logs and traces for long-lived streams

Log connections, not events. One structured line when a stream closes, carrying route, a connection id, user or tenant id, resumed flag, duration, events sent, bytes sent, maximum queue depth and close reason, is enough to answer most questions and costs one line per connection. Logging every event at a thousand events per second across a hundred thousand connections creates a logging bill larger than the service. Sample event-level debug logging per connection when you need it.

Do not wrap a stream in one span. A tracing span that lasts four hours is useless in most trace backends: it arrives only when the stream closes, if at all, and cannot be read while open. Use a short span for the connection setup (authentication, subscription, replay) and record events as their own short spans or as links. The more valuable trace is the event's path: the producer starts a trace, carries the W3C traceparent value inside the event envelope through the broker, and the SSE server creates a short "deliver" span linked to it when it writes the event. That shows exactly where an event spent its time.

Correlate with an id. Return a connection id in a response header and include it in close logs and client telemetry. When a user reports a stale screen, support can find the stream, its close reason and the server it lived on.

Worked example: the sixty-second reconnect storm

A dashboard service streams order updates to around 40,000 open tabs. The numbers here are illustrative. After an infrastructure change, users report that updates "sometimes take a minute". Request error rate is 0 percent and the request latency SLO is green, because SSE routes were excluded from it.

The connection metrics tell the story in three steps. First, sse_streams_opened_total with resumed="true" has risen from a few per second to about 650 per second: almost every tab reconnects roughly once a minute. Second, the duration histogram has a tall spike in the 60 to 65 second bucket that was not there before. Third, close reasons show client_gone for those connections: from the server's view, the other side hung up. A fixed duration, a client-side hang-up and a recent infrastructure change point to an intermediary with a 60 second idle timeout. The change had replaced a load balancer whose idle timeout was several minutes with an nginx tier using the default proxy_read_timeout of 60 seconds, which closes the upstream connection when no data arrives for that long.

Why did users see delays rather than errors? On quiet order streams, no event and no heartbeat was sent for over a minute, because the heartbeat interval had been set to 90 seconds when the old balancer allowed it. The proxy closed the connection; the browser waited its retry interval, reconnected with Last-Event-ID, and the replay histogram shows small replays, so data was not lost, only late. Each reconnect also re-ran authentication and a replay query, which explained a rise in database load nobody had connected to the stream.

The fix was to send heartbeat comments every 15 seconds, well inside every timeout on the path, and raise proxy_read_timeout on SSE locations. Afterwards the 60 second spike disappeared, median connection duration returned to tens of minutes, and resumes fell back to the rate of real network changes. The lesson for alerting: the duration histogram and the resume rate caught this; status codes and request latency never would. Our heartbeat and keepalive article explains how to choose the interval from the timeouts on the path.

Alerts and dashboards

Alert on symptoms users feel, with a small number of rules:

  • Event lag: p99 of sse_event_lag_seconds above your freshness objective, for example 5 seconds, for 10 minutes.
  • Reconnect rate: resumed opens per open stream per minute above its normal baseline, which catches proxy timeouts and flapping backends.
  • Fatal client errors: rate of fatal=true client errors above a small threshold, which catches authentication and content-type breakage.
  • Slow consumers: evictions per minute rising, or queue-depth percentiles approaching the limit, which predicts memory pressure.
  • Balance: one instance holding far more or far fewer open streams than its peers, which catches broken nodes and sticky imbalance after deploys.

Dashboards should place open streams, opens split by resumed, close reasons, duration histogram, event lag and time to first event on one screen, because diagnosis almost always needs two of them side by side. Our horizontal scaling article shows how these signals feed capacity planning and reconnect-storm shedding.

Failure modes and trade-offs

FailureSignal that catches itUsual fix
Proxy idle timeoutDuration spike at a fixed value, high resume rateHeartbeats inside the smallest timeout
Proxy bufferingClient time to first event far above server first writeDisable buffering on SSE routes
Slow consumersQueue depth and evictionsBounded queues, evict and let clients resume
Stalled producerOpen streams steady, events sent at zero, lag absentAlert on throughput per route
Auth expiry loopFatal client errors, rising opens with immediate closesRefresh tokens before expiry; distinct close reason
Deploy churn read as failureClose spike with reason shutdownLabel shutdowns; drain gradually
Metrics cardinality explosionMetric backend cost and slownessNo per-user or per-connection labels

The trade-off is depth against cost. Connection-level metrics and close logs are cheap and catch most problems. Event lag needs a timestamp in every event and synchronised clocks. Client telemetry needs sampling and a collection endpoint but is the only way to see buffering and fatal client states. Per-event tracing is the most expensive and is best sampled, or turned on for one tenant while debugging.

What to do next

  1. Remove SSE routes from request-latency SLOs and dashboards, and add a separate connection-duration histogram with buckets around your proxies' idle timeouts.
  2. Add open, opened (split by resumed) and closed (split by reason) counters, and make every close path set a reason.
  3. Stamp created_at on events at the producer and record event lag at write time.
  4. Ship a sampled client beacon for time to first event and errors split by fatal versus retrying.
  5. Write one structured log line per closed connection, and return a connection id header for support lookups.
  6. List every timeout on the path (load balancer, proxy, CDN, server) and set the heartbeat interval well below the smallest.
  7. Create the five alerts above, and test them by forcing a short proxy timeout in staging.
Key takeaway: SSE needs connection and event metrics, not request metrics: open streams, opens split by resume, closes split by reason, a duration histogram that exposes timeouts, queue depth, and end-to-end event lag. Add sampled client telemetry for time to first event and fatal errors, log one line per closed connection, and trace events rather than streams. Alert on lag, reconnect rate, fatal client errors, slow consumers and imbalance, and keep heartbeats well inside every timeout on the path.