Server-Sent Events look simple to operate: one HTTP response that never ends, a few lines of data: text at a time. That simplicity is exactly why they are hard to observe. Standard HTTP dashboards were built for short requests, so a healthy SSE service can show a p99 latency of forty minutes, a 0 percent error rate while every client is silently reconnecting every sixty seconds, and green status codes while a proxy holds every event in a buffer. The request is the wrong unit of measurement. The connection and the event are the right ones.
This article builds an observability model for SSE from first principles: what to measure at the pipeline, the server and the client, how to instrument a server with real code, how to log and trace long-lived streams without drowning in data, which alerts catch real failures, and a worked diagnosis of a reconnect storm caused by a proxy timeout. It assumes you know the protocol basics; if not, start with our reconnect and Last-Event-ID article, which covers the wire contract this monitoring depends on.
Why request metrics lie about streams
A normal HTTP request has a start, an end and a status, and the latency between them tells you about user experience. An SSE response has a start, a status that is decided in the first milliseconds, and then an open-ended body that may last hours. Three things follow.
- Duration is not latency. The time from request to response end is a connection lifetime. Mixed into a request-latency histogram, it ruins percentiles for every other route on the service. Exclude SSE routes from request-latency service level objectives and measure their lifetimes separately.
- Status codes describe only the start. A stream that opened with 200 and then stalled, was cut by a proxy or fell hours behind is still recorded as a success. Most SSE failures happen after the status line.
- The client hides failure.
EventSourcereconnects automatically after a dropped connection, so users often see a stale screen rather than an error. Server-side counters show many short successful requests; nobody is paged.
So the model has two units. The connection has an open time, a duration, a close reason and a count of events delivered. The event has a creation time, a delivery time on each connection, and a size. Every useful SSE metric is a count, gauge or histogram over one of these two.
What to measure, and where
Instrument three vantage points, because each sees failures the others cannot.
| Signal | Type | Why it matters |
|---|---|---|
| Open streams | Gauge, per instance | Capacity and balance; a drop to zero on one node is a broken node |
| Connections opened, closed | Counters, closed labelled by reason | Churn; close reasons separate client, server, proxy and slow-consumer causes |
| Connection duration | Histogram with long buckets | Spikes at fixed values reveal timeouts |
| Resume requests | Counter: requests carrying Last-Event-ID | Reconnect rate as the server sees it |
| Replayed events per resume | Histogram | How far behind clients fall during gaps |
| Events and bytes sent | Counters by event type | Throughput and cost |
| Per-connection queue depth | Histogram or max gauge | Slow consumers before they exhaust memory |
| Slow-consumer evictions | Counter | Connections closed because they could not keep up |
| End-to-end event lag | Histogram: write time minus created_at | The number users actually feel |
| Time to first event | Client histogram | Detects proxy buffering the server cannot see |
| Client errors by readyState | Client counter | Fatal closes versus automatic retries |
Keep label cardinality bounded. Label by route, event type, instance and close reason; never by user id, connection id or topic name if topics are per-user. Those belong in logs and traces, not in metric labels, or a hundred thousand connections become a hundred thousand time series.
Instrumenting the server
The sketch below instruments a Starlette SSE endpoint with the Python prometheus_client library. The structure carries over to any stack: count the open in one place, count the close in a finally block with a reason, and measure every event as it is written. subscribe stands for your broker client and Event carries an id, a type, a payload and a created_at timestamp set by the producer.
import asyncio, time
from prometheus_client import Counter, Gauge, Histogram
from starlette.responses import StreamingResponse
OPEN = Gauge("sse_open_streams", "Open SSE streams", ["route"])
OPENED = Counter("sse_streams_opened_total", "Streams opened", ["route", "resumed"])
CLOSED = Counter("sse_streams_closed_total", "Streams closed", ["route", "reason"])
DURATION = Histogram("sse_stream_duration_seconds", "Stream lifetime", ["route"],
buckets=[1, 5, 15, 30, 55, 60, 65, 120, 300, 900, 3600, 14400])
SENT = Counter("sse_events_sent_total", "Events written", ["route", "type"])
LAG = Histogram("sse_event_lag_seconds", "created_at to write", ["route"],
buckets=[0.05, 0.1, 0.25, 0.5, 1, 2, 5, 15, 60])
QUEUE_MAX = 1000
async def stream(request, route="orders"):
resumed = "true" if request.headers.get("last-event-id") else "false"
OPENED.labels(route, resumed).inc()
OPEN.labels(route).inc()
started, reason = time.monotonic(), "server_end"
queue = await subscribe(request, maxsize=QUEUE_MAX)
async def body():
nonlocal reason
try:
yield "retry: 5000\n\n"
while True:
try:
ev = await asyncio.wait_for(queue.get(), timeout=15)
except asyncio.TimeoutError:
yield ": keepalive\n\n" # comment line, ignored by clients
continue
if queue.qsize() > QUEUE_MAX * 0.9:
reason = "slow_consumer"
return
yield f"id: {ev.id}\nevent: {ev.type}\ndata: {ev.json}\n\n"
SENT.labels(route, ev.type).inc()
LAG.labels(route).observe(max(0.0, time.time() - ev.created_at))
except (asyncio.CancelledError, GeneratorExit):
reason = "client_gone" # disconnect cancels the generator
raise
except Exception:
reason = "error"
raise
finally:
await queue.close()
OPEN.labels(route).dec()
CLOSED.labels(route, reason).inc()
DURATION.labels(route).observe(time.monotonic() - started)
return StreamingResponse(body(), media_type="text/event-stream",
headers={"Cache-Control": "no-cache",
"X-Accel-Buffering": "no"})Details that matter: the duration buckets deliberately bracket common idle timeouts (55, 60, 65 seconds) so a timeout shows up as a spike rather than disappearing inside a wide bucket. Lag uses wall-clock time against a producer timestamp, which is only as accurate as clock synchronisation between hosts; keep hosts on NTP and treat lag under about 50 milliseconds as noise. Heartbeat comments are not counted as events, so a stream that carries only heartbeats is visible as zero event throughput with open connections. And a graceful shutdown should set reason = "shutdown" before closing streams, so deploys do not look like failures.
Client-side telemetry
The browser sees what the server cannot: buffering in proxies, mobile networks dropping idle connections, and fatal failures. EventSource exposes a readyState of CONNECTING (0), OPEN (1) or CLOSED (2), and an error event that carries no status code. The useful distinction is what the state is after an error: CONNECTING means the browser will retry; CLOSED means it gave up, for example because the response was not 200 or did not have the text/event-stream content type, and nothing will recover without your code.
const t0 = performance.now();
let firstEvent = null, errors = 0;
const es = new EventSource("/events/orders");
es.addEventListener("order", (e) => {
if (firstEvent === null) {
firstEvent = performance.now() - t0;
report({ metric: "sse_time_to_first_event_ms", value: firstEvent });
}
});
es.onerror = () => {
errors += 1;
const fatal = es.readyState === EventSource.CLOSED;
report({ metric: "sse_client_error", fatal, errors });
if (fatal) scheduleManualReconnect(); // the browser will not retry by itself
};
function report(payload) {
navigator.sendBeacon("/rum", JSON.stringify({ ...payload, page: location.pathname }));
}Sample this telemetry, aggregate it server-side into histograms, and compare time to first event with the server's time to first write. If the server wrote within 50 milliseconds but clients report two seconds, something in between is buffering. Our client lifecycle article covers how hidden tabs and sleep affect these numbers, which matters when interpreting them.
Logs and traces for long-lived streams
Log connections, not events. One structured line when a stream closes, carrying route, a connection id, user or tenant id, resumed flag, duration, events sent, bytes sent, maximum queue depth and close reason, is enough to answer most questions and costs one line per connection. Logging every event at a thousand events per second across a hundred thousand connections creates a logging bill larger than the service. Sample event-level debug logging per connection when you need it.
Do not wrap a stream in one span. A tracing span that lasts four hours is useless in most trace backends: it arrives only when the stream closes, if at all, and cannot be read while open. Use a short span for the connection setup (authentication, subscription, replay) and record events as their own short spans or as links. The more valuable trace is the event's path: the producer starts a trace, carries the W3C traceparent value inside the event envelope through the broker, and the SSE server creates a short "deliver" span linked to it when it writes the event. That shows exactly where an event spent its time.
Correlate with an id. Return a connection id in a response header and include it in close logs and client telemetry. When a user reports a stale screen, support can find the stream, its close reason and the server it lived on.
Worked example: the sixty-second reconnect storm
A dashboard service streams order updates to around 40,000 open tabs. The numbers here are illustrative. After an infrastructure change, users report that updates "sometimes take a minute". Request error rate is 0 percent and the request latency SLO is green, because SSE routes were excluded from it.
The connection metrics tell the story in three steps. First, sse_streams_opened_total with resumed="true" has risen from a few per second to about 650 per second: almost every tab reconnects roughly once a minute. Second, the duration histogram has a tall spike in the 60 to 65 second bucket that was not there before. Third, close reasons show client_gone for those connections: from the server's view, the other side hung up. A fixed duration, a client-side hang-up and a recent infrastructure change point to an intermediary with a 60 second idle timeout. The change had replaced a load balancer whose idle timeout was several minutes with an nginx tier using the default proxy_read_timeout of 60 seconds, which closes the upstream connection when no data arrives for that long.
Why did users see delays rather than errors? On quiet order streams, no event and no heartbeat was sent for over a minute, because the heartbeat interval had been set to 90 seconds when the old balancer allowed it. The proxy closed the connection; the browser waited its retry interval, reconnected with Last-Event-ID, and the replay histogram shows small replays, so data was not lost, only late. Each reconnect also re-ran authentication and a replay query, which explained a rise in database load nobody had connected to the stream.
The fix was to send heartbeat comments every 15 seconds, well inside every timeout on the path, and raise proxy_read_timeout on SSE locations. Afterwards the 60 second spike disappeared, median connection duration returned to tens of minutes, and resumes fell back to the rate of real network changes. The lesson for alerting: the duration histogram and the resume rate caught this; status codes and request latency never would. Our heartbeat and keepalive article explains how to choose the interval from the timeouts on the path.
Alerts and dashboards
Alert on symptoms users feel, with a small number of rules:
- Event lag: p99 of
sse_event_lag_secondsabove your freshness objective, for example 5 seconds, for 10 minutes. - Reconnect rate: resumed opens per open stream per minute above its normal baseline, which catches proxy timeouts and flapping backends.
- Fatal client errors: rate of
fatal=trueclient errors above a small threshold, which catches authentication and content-type breakage. - Slow consumers: evictions per minute rising, or queue-depth percentiles approaching the limit, which predicts memory pressure.
- Balance: one instance holding far more or far fewer open streams than its peers, which catches broken nodes and sticky imbalance after deploys.
Dashboards should place open streams, opens split by resumed, close reasons, duration histogram, event lag and time to first event on one screen, because diagnosis almost always needs two of them side by side. Our horizontal scaling article shows how these signals feed capacity planning and reconnect-storm shedding.
Failure modes and trade-offs
| Failure | Signal that catches it | Usual fix |
|---|---|---|
| Proxy idle timeout | Duration spike at a fixed value, high resume rate | Heartbeats inside the smallest timeout |
| Proxy buffering | Client time to first event far above server first write | Disable buffering on SSE routes |
| Slow consumers | Queue depth and evictions | Bounded queues, evict and let clients resume |
| Stalled producer | Open streams steady, events sent at zero, lag absent | Alert on throughput per route |
| Auth expiry loop | Fatal client errors, rising opens with immediate closes | Refresh tokens before expiry; distinct close reason |
| Deploy churn read as failure | Close spike with reason shutdown | Label shutdowns; drain gradually |
| Metrics cardinality explosion | Metric backend cost and slowness | No per-user or per-connection labels |
The trade-off is depth against cost. Connection-level metrics and close logs are cheap and catch most problems. Event lag needs a timestamp in every event and synchronised clocks. Client telemetry needs sampling and a collection endpoint but is the only way to see buffering and fatal client states. Per-event tracing is the most expensive and is best sampled, or turned on for one tenant while debugging.
What to do next
- Remove SSE routes from request-latency SLOs and dashboards, and add a separate connection-duration histogram with buckets around your proxies' idle timeouts.
- Add open, opened (split by resumed) and closed (split by reason) counters, and make every close path set a reason.
- Stamp
created_aton events at the producer and record event lag at write time. - Ship a sampled client beacon for time to first event and errors split by fatal versus retrying.
- Write one structured log line per closed connection, and return a connection id header for support lookups.
- List every timeout on the path (load balancer, proxy, CDN, server) and set the heartbeat interval well below the smallest.
- Create the five alerts above, and test them by forcing a short proxy timeout in staging.