Server-Sent Events are an ordinary HTTP response with Content-Type: text/event-stream that the server keeps open and writes to whenever it has news. It needs no upgrade handshake, so it works anywhere HTTP works. The catch is that everything between browser and server is built for responses that end: proxies buffer them, load balancers time them out and compressors batch them, and none of it shows up on localhost.
This article treats the network path as a series of hops, lists what common hops do by default (checked against vendor documentation), and shows a probe that identifies the guilty hop. Relaying LLM tokens over SSE is covered in the LLM streaming pattern and running many SSE nodes in horizontal SSE scaling; this page is about the network path in between.
Three things a middlebox can do to a stream
Any intermediary can do three things to a long-lived response, and each one produces a recognisable symptom.
- Buffer. The hop collects bytes until a size threshold or the end of the response. Events arrive late, in bursts, or all at once at close; for a stream that never closes, possibly never.
- Time out. The hop closes the stream. There are two different kinds of timer: an idle timeout, which fires when no bytes have moved for N seconds and can be beaten with heartbeats, and a total (whole-response) timeout, which fires N seconds after the request regardless of traffic and cannot be beaten with heartbeats.
- Transform. The hop compresses, re-chunks or rewrites headers. Usually harmless, but unflushed compression is buffering, and some headers are invalid under HTTP/2.
EventSource reconnects after a dropped stream and sends Last-Event-ID, so timeouts are survivable if the server can resume (see SSE reconnect and Last-Event-ID). Nothing reconnects you out of buffering: a buffered stream looks healthy at both ends and is simply late.
Buffering: where it hides
Buffering is the most common SSE bug in production, and confusing because both ends report success. Look in four places.
- Your own framework. Many frameworks buffer the response body in memory and send it when the handler returns. You need an API that writes and flushes as you go, such as a streaming response type or an explicit flush after each event. Check this first by calling the app directly with no proxy in between.
- Compression middleware. gzip and brotli work on blocks. A compressor that is not flushed after every event will hold several kilobytes of events until a block fills. Either exclude
text/event-streamfrom compression or confirm that your middleware flushes the compressor on every write. Event streams are small and repetitive, so compression gains little; excluding them is the simple fix. - Reverse proxies. nginx buffers upstream responses by default (
proxy_buffering on). Either turn it off for the SSE location, or have the app sendX-Accel-Buffering: no, which nginx honours per response. Other proxies have their own switches; look for response buffering, not request buffering. - Inspecting proxies and endpoint security software. Some corporate proxies and antivirus products that decrypt TLS scan a response before releasing it, which can mean buffering the whole response. You do not control these hops, so you can only detect them and fall back (see below).
Three response headers keep well-behaved intermediaries out of the way. Cache-Control: no-cache stops anything from serving a stored copy of the stream. no-transform, added to the same header, is the HTTP caching directive that asks intermediaries not to modify the body, which some respect for compression. And there must be no Content-Length; over HTTP/1.1 the response uses chunked transfer encoding, and over HTTP/2 and HTTP/3 the stream simply stays open.
# nginx: one location for streams, everything else keeps its defaults
location /events/ {
proxy_pass http://sse_backend;
proxy_http_version 1.1;
proxy_set_header Connection "";
proxy_buffering off; # or send X-Accel-Buffering: no from the app
proxy_cache off;
gzip off; # never batch events into compression blocks
proxy_read_timeout 3600s; # default 60s: time allowed between two reads
}
Timeouts: idle versus total
Every hop has a timer, and the stream lasts only as long as the most aggressive one allows. The table lists defaults for common hops as documented by their vendors at the time of writing. Always confirm against the version you run, because managed products change.
| Hop | Timer and default | Kind | What to do |
|---|---|---|---|
| nginx | proxy_read_timeout, 60 s between two successive reads from upstream | Idle | Raise it for the stream location; heartbeat well inside it |
| AWS Application Load Balancer | Connection idle timeout, 60 s by default, configurable from 1 to 4,000 s | Idle | Raise it, and heartbeat inside it |
| Google Cloud global external Application Load Balancer | Backend service timeout, 30 s by default (60 minutes for serverless NEGs); covers the whole response | Total | Raise it to the longest stream you want, and design for reconnection anyway |
| Envoy (route) | Route timeout, 15 s by default, from end of request until the response is complete | Total | Set it to 0 for stream routes; bound streams with idle_timeout or max_stream_duration |
| Envoy (connection manager) | stream_idle_timeout, 5 minutes by default | Idle | Heartbeat inside it |
| Cloud Run | Request timeout, 5 minutes by default, up to 60 minutes | Total | Raise it; clients must reconnect at the limit |
| Cloudflare proxy | About 100 s for the origin to respond before error 524; Enterprise plans can raise it | First response | Send headers and a first comment immediately; heartbeat regularly |
Two rows deserve emphasis. The Google Cloud load balancer counts from the first request byte to the last response byte, so by default every SSE stream through it ends at 30 seconds, however busy. Envoy's route timeout works the same way and sits under many meshes and gateways. Heartbeats cannot help; change the configuration and still treat stream end as normal. For Cloudflare, the 100-second figure is documented for the wait for the origin's response. How it treats quiet gaps after headers have been sent is not something we could confirm from Cloudflare's own documentation, so assume the worst and send heartbeats.
Ending streams is not always bad. Closing each stream server-side after a few minutes keeps the resume path exercised and rebalances load after deploys.
Heartbeats and what they cannot fix
An SSE comment is a line starting with a colon. Clients ignore it, but every hop sees bytes moving, which resets idle timers. Send one whenever the stream has been quiet for a set interval, chosen well below the shortest idle timeout on the path. With nginx at 60 s and a load balancer at 60 s, 15 to 20 seconds leaves room for scheduling jitter. Heartbeat design in general is covered in heartbeat and keepalive strategies.
import asyncio, json, time
HEARTBEAT_S = 15
MAX_STREAM_S = 15 * 60 # end deliberately; EventSource will reconnect
async def event_stream(queue):
# Yield SSE frames; the framework must flush each yielded chunk.
yield ": open\n\n" # first bytes at once: beats first-byte timers
yield "retry: 3000\n\n" # reconnect delay hint in ms
started = time.monotonic()
while time.monotonic() - started < MAX_STREAM_S:
try:
ev = await asyncio.wait_for(queue.get(), timeout=HEARTBEAT_S)
except asyncio.TimeoutError:
yield f": hb {int(time.time())}\n\n" # comment line, ignored by clients
continue
yield f"id: {ev['id']}\nevent: {ev['type']}\ndata: {json.dumps(ev['data'])}\n\n"
HEADERS = {
"Content-Type": "text/event-stream",
"Cache-Control": "no-cache, no-transform",
"X-Accel-Buffering": "no",
# no Content-Length, and no Connection header (invalid in HTTP/2)
}A heartbeat resets idle timers but cannot get through a buffer; a comment stuck in a buffer is as late as an event. If heartbeats seem to help, they are filling a size-based buffer faster, which is a symptom, not a fix.
Connection limits and HTTP/2
Over HTTP/1.1, browsers allow about six concurrent connections per origin, shared across tabs. Every open EventSource holds one, so several tabs can starve the page's normal API calls. HTTP/2 and HTTP/3 multiplex streams over one connection; what counts is the protocol between browser and edge, so an HTTP/2 CDN in front of an HTTP/1.1 origin is fine.
HTTP/2 brings its own trap. RFC 9113 treats connection-specific header fields such as Connection and Keep-Alive as malformed. Hand-written SSE handlers often set Connection: keep-alive because an old tutorial did; some HTTP/2 stacks strip it and others reject the response, so the code fails only behind certain proxies. Remove it; HTTP/1.1 connections are persistent by default anyway.
Find the hop with a probe
Do not guess which hop is responsible; measure. Expose a diagnostic endpoint that emits one event per second carrying the server's send time, then read it through each hop in turn and record when each event arrives. The pattern of arrival times identifies the cause.
# sse_probe.py: print arrival time and lag of each event from a /probe endpoint
# Usage: python sse_probe.py URL [seconds]
import sys, time, urllib.request
url, limit = sys.argv[1], float(sys.argv[2]) if len(sys.argv) > 2 else 120
req = urllib.request.Request(url, headers={
"Accept": "text/event-stream",
"Accept-Encoding": "gzip, br", # ask like a browser does, or compression bugs hide
})
t0 = time.time()
with urllib.request.urlopen(req, timeout=limit) as resp:
print("status", resp.status, "encoding", resp.headers.get("Content-Encoding"))
for raw in resp:
now = time.time()
line = raw.decode("utf-8", "replace").rstrip("\n")
if line.startswith("data: "):
sent = float(line[6:])
print(f"t={now - t0:7.2f}s lag={now - sent:6.2f}s")
if now - t0 > limit:
break
print(f"stream ended at t={time.time() - t0:.1f}s")The Accept-Encoding header matters: curl does not send it by default, so a stream that works under curl and fails in the browser often has a compression problem. Run the probe against the app directly, then the ingress, then the load balancer, then the CDN hostname. The first hop where the pattern changes is the culprit. Lag compares two clocks, so read the pattern, not the absolute number.
| Arrival pattern | Likely cause |
|---|---|
| Nothing, then everything when the stream ends | Whole-response buffering: framework, proxy or inspecting middlebox |
| Bursts of several events with growing lag | Size-based buffer or compression without flush |
| Steady, then closes at a round number of seconds | A timer: compare the number with the table above |
| Closes after a quiet period only | Idle timeout shorter than your heartbeat interval |
| Fine directly, fails only through one hostname | That hop's configuration or an HTTP/2 header rejection |
| Fine for you, broken for one customer | Their corporate proxy or endpoint security software |
Worked example: five hops, five fixes
A chat product streams model output over SSE through Cloudflare, a Google Cloud global external Application Load Balancer, an nginx reverse proxy on GKE and a Python app. In production, users see the whole answer at once, and long answers stop at 30 seconds.
Step one: the probe against the pod shows one event per second with near-zero lag, so the app flushes. Step two: through the ingress, nothing arrives until the end; proxy_buffering is on. The app starts sending X-Accel-Buffering: no; if your ingress is a controller, set its equivalent buffering and read-timeout options instead (ingress-nginx, which exposed these as annotations, was retired in March 2026). Step three: through the load balancer, the stream ends at exactly 30.0 seconds, the backend service timeout default; it is raised to 3,600 seconds for the stream backend only, and the server ends streams itself after 15 minutes. Step four: through Cloudflare, events flow, and the immediate opening comment means the 100-second response wait never applies. Step five: one enterprise customer still sees whole answers at once, and the probe from their network shows buffering, so the client falls back: if no bytes arrive within five seconds of opening, it switches to long-polling.
Five hops, five independent settings, each confirmed by the probe where it was made; from the browser alone they are indistinguishable.
Failure modes
- Fixing one hop and declaring victory. Another hop may also buffer; re-run the probe at the outermost hostname after every change.
- Heartbeats against a total timeout. Teams send heartbeats every five seconds and still lose streams at 30 or 15 seconds, because the timer counts the whole response. Read the timer's definition, not its name.
- Global timeout changes. Raising the read timeout for every route lets hung backends hold connections for an hour. Scope long timeouts to the stream routes.
- No resume path. Every stream eventually ends: deploys, scale-in, timeouts, network changes. Without event IDs and resumption, each end loses data or replays everything.
- Compression added later. A site-wide brotli rollout turns every stream bursty. Run the probe in CI against staging.
Trade-offs
Long versus short streams. Long streams mean fewer reconnects and less handshake overhead, but more exposure to every timer and to stale connections after a deploy. Short streams (minutes) keep the resume path tested and spread load, at the cost of a reconnect every few minutes, which is negligible for most products.
Shared ingress versus a dedicated path. Disabling buffering on a shared ingress is quick; a separate route or hostname for streams, with its own timeouts and no compression, is easier to reason about.
SSE versus alternatives. SSE is plain HTTP, so it passes more corporate networks than WebSockets, but inspecting proxies can still buffer it. See SSE versus WebSocket; long-polling works almost everywhere, at a cost in latency.
What to do next
- Draw your hop chain from browser to app and record each hop's idle and total timeouts and buffering behaviour, from the documentation for your version.
- Add a
/probeendpoint that emits one timestamped event per second, and run the probe through every hop in turn. - Send
Cache-Control: no-cache, no-transform,X-Accel-Buffering: noand noContent-LengthorConnectionheader; send a comment as the first bytes. - Set the heartbeat interval to a quarter or a third of the shortest idle timeout, and raise total timeouts only on stream routes.
- End streams server-side after a fixed time and make sure reconnects resume from
Last-Event-ID. - Add a client-side fallback to long-polling when no bytes arrive within a few seconds of opening.