Every long-lived connection needs some periodic traffic, but the word heartbeat hides two different jobs. One is keeping the path open: NAT devices, load balancers and proxies drop connections that stay idle too long, usually without telling either end. The other is detecting a dead peer: a crashed process, a phone that walked out of coverage or a partition can leave a connection half-open for hours. The two jobs want different intervals, live at different layers, and fail in different ways.
Why silent death happens and how ping/pong detects it is explained in heartbeats and dead-connection detection. This article is the configuration guide that follows: how to choose intervals and timeouts from the path you actually have, which mechanism each protocol offers and what its defaults are, when you need an application-level heartbeat on top, and what it all costs at a million connections.
Two jobs, two constraints
Keep-alive traffic exists to reset idle timers in middleboxes. Its interval is set by the shortest idle timeout anywhere on the path. Miss it once and the connection is gone, usually silently: the middlebox forgets the flow, and the next packet from either side is dropped or answered with a reset.
Liveness detection exists to notice a dead peer quickly enough to act: reconnect, fail over, mark a user offline, release a lock. Its constraint is the detection time the product needs, which is roughly the interval plus the timeout, or interval times the number of misses tolerated. Its cost is false positives: declare death too eagerly and a slow network causes reconnect storms.
One mechanism can do both jobs, but design for each separately and then take the tighter interval. A chat app may need presence updates within 30 seconds while its load balancer allows 350 seconds idle; the presence requirement wins. A telemetry stream may tolerate five minutes of undetected death while a carrier NAT drops idle flows much sooner; the NAT wins.
The arithmetic
Three formulas cover most decisions. First, the keep-alive interval must be below the smallest idle timeout on the path with margin for jitter and one lost packet; half the smallest timeout is a sound default. Second, worst-case detection time is about interval plus timeout when one missed pong is fatal, or interval times misses when you tolerate several. Third, the server-side load is connections divided by interval.
def plan(conns, idle_timeouts_s, detect_target_s, misses=2):
interval = min(min(idle_timeouts_s) / 2, detect_target_s / misses)
return {
"interval_s": interval,
"worst_detect_s": interval * misses,
"pings_per_s": round(conns / interval),
}
# ALB at its 60 s default idle timeout, product wants detection within 60 s
print(plan(1_000_000, [60, 350], 60))
# {'interval_s': 30.0, 'worst_detect_s': 60.0, 'pings_per_s': 33333}A third of a million pings per second sounds alarming, but each is a few bytes handled by the event loop; the cost is mostly wakeups and, on mobile, radio time. What matters is that the cost is linear in connections and inversely proportional to the interval, so halving the interval doubles it. Know your idle timeouts: on AWS, an Application Load Balancer's idle timeout defaults to 60 seconds, and a Network Load Balancer's TCP idle timeout defaults to 350 seconds and can be set between 60 and 6,000 seconds. Home routers and carrier NATs do not publish theirs and can be much shorter, so for clients on the open internet measure, or assume tens of seconds.
TCP keepalive: cheap, one hop, slow by default
TCP keepalive sends empty probes on an idle socket and resets the connection if they go unanswered. On Linux the system-wide defaults are 7,200 seconds of idle time before the first probe, 75 seconds between probes and 9 probes, so a dead peer is detected after more than two hours. Those defaults are useless for both jobs. Set the values per socket instead:
import socket
def tune_keepalive(sock, idle=30, interval=10, count=3, user_timeout_ms=40_000):
sock.setsockopt(socket.SOL_SOCKET, socket.SO_KEEPALIVE, 1)
sock.setsockopt(socket.IPPROTO_TCP, socket.TCP_KEEPIDLE, idle) # Linux
sock.setsockopt(socket.IPPROTO_TCP, socket.TCP_KEEPINTVL, interval)
sock.setsockopt(socket.IPPROTO_TCP, socket.TCP_KEEPCNT, count)
# Bound how long sent data may stay unacknowledged before the kernel gives up.
sock.setsockopt(socket.IPPROTO_TCP, socket.TCP_USER_TIMEOUT, user_timeout_ms)Keepalive probes only fire when the socket is idle. If the peer vanishes while you have unacknowledged data in flight, retransmission timers govern instead, and they can take many minutes; TCP_USER_TIMEOUT caps that. Two limits make TCP keepalive insufficient on its own. The probe is answered by the peer's kernel, so a deadlocked application with a healthy kernel looks alive. And every proxy or load balancer that terminates TCP answers probes itself, so keepalive covers exactly one hop. Use it as a cheap backstop, on server-to-server links and between proxies and backends.
WebSocket ping and pong
RFC 6455 defines Ping (opcode 0x9) and Pong (opcode 0xA) control frames. A control frame's payload is at most 125 bytes, a Pong must echo the Ping's application data, and an endpoint may send an unsolicited Pong as a one-way heartbeat. Pings reach the endpoint that terminates the WebSocket, which may be a gateway rather than your application.
The asymmetry that shapes browser designs: browsers answer server pings automatically, but the JavaScript WebSocket API has no method to send a ping or observe a pong. A server can detect dead browsers with protocol pings, but a browser that wants to detect a dead server, for example to show a reconnecting banner, must use application messages:
// browser side: application-level heartbeat over a WebSocket
function heartbeat(ws, intervalMs = 25000, timeoutMs = 10000) {
let seq = 0, timer = null, deadline = null;
const tick = () => {
clearTimeout(deadline);
ws.send(JSON.stringify({ type: "hb", seq: ++seq, t: Date.now() }));
deadline = setTimeout(() => ws.close(4000, "heartbeat timeout"), timeoutMs);
};
ws.addEventListener("message", (ev) => {
clearTimeout(deadline); // any inbound message proves liveness
const msg = JSON.parse(ev.data);
if (msg.type === "hb-ack") recordRtt(Date.now() - msg.t);
});
timer = setInterval(tick, intervalMs + Math.random() * 2000); // jitter
ws.addEventListener("close", () => { clearInterval(timer); clearTimeout(deadline); });
}Treat any inbound frame as proof of life, not only the heartbeat reply, so busy connections never time out spuriously. Reconnection after a timeout needs backoff and jitter, as covered in the WebSocket guide.
HTTP/2 PING and gRPC keepalive
HTTP/2 has a PING frame with an 8-byte opaque payload; the receiver must reply with a PING carrying the same payload and the ACK flag. It measures round-trip time and liveness of the whole connection, across all its streams. gRPC builds its keepalive on it, and the grpc.io keepalive guide documents the C-core defaults: the client's keepalive time is disabled (INT_MAX), the keepalive timeout is 20 seconds, and pings without active calls are off. The server's own keepalive time is 2 hours, and by default it permits client pings no more often than every 5 minutes.
That last default is the classic trap. A client configured to ping every 30 seconds against a server left at defaults will be told off: the guide says the server eventually sends GOAWAY with the debug data too_many_pings, and the client's calls fail. Client and server settings must be changed together:
import grpc
client_opts = [
("grpc.keepalive_time_ms", 30_000), # ping after 30 s of inactivity
("grpc.keepalive_timeout_ms", 10_000), # wait 10 s for the ack
("grpc.keepalive_permit_without_calls", 1), # ping even with no active calls
("grpc.http2.max_pings_without_data", 0), # lift the default cap of 2 pings
]
server_opts = [
("grpc.http2.min_ping_interval_without_data_ms", 20_000), # allow the client's rate
("grpc.keepalive_permit_without_calls", 1),
]
channel = grpc.insecure_channel("svc:50051", options=client_opts)
server = grpc.server(executor, options=server_opts)The string keys above are the C-core channel argument names that Python passes through. Note the third client option: by default C-core lets a client send only two pings without any data or header frame in between, so an idle channel stops pinging unless that cap is lifted. On the server, grpc.http2.max_ping_strikes (default 2) is how many over-frequent pings it tolerates before the GOAWAY. Defaults and option names differ between language implementations, so check your runtime's documentation rather than copying values between them. Streaming RPC patterns are covered in bidirectional gRPC.
SSE comments and MQTT keep alive
Server-Sent Events has no ping frame, but any line beginning with a colon is a comment the client ignores. Sending a comment every 15 to 30 seconds keeps proxies from closing the response as idle. Many reverse proxies time out an upstream read after about a minute of silence by default, and nginx's proxy_read_timeout is 60 seconds. The browser's EventSource reconnects on its own when the stream drops, so detection on the client is built in; the server sees a dead client only when a write fails.
async def sse_stream(queue, send):
while True:
try:
event = await asyncio.wait_for(queue.get(), timeout=20)
await send(f"id: {event.id}\ndata: {event.json}\n\n")
except asyncio.TimeoutError:
await send(": keep-alive\n\n") # comment line; resets proxy idle timersMQTT builds keep alive into the session. The client declares a keep-alive period in CONNECT; if it has nothing else to send it sends PINGREQ and expects PINGRESP. The broker disconnects a client from which it receives no control packet within one and a half times the keep-alive period, and then publishes the client's will message if one was set, which makes keep alive the basis of presence in many IoT systems. Choose the period against the cellular NAT on the device side, not the broker; see MQTT for bidirectional messaging.
When you need an application heartbeat
Protocol-level pings stop at the first component that terminates the protocol. Add an application heartbeat when you need any of these: liveness of the application itself rather than its kernel or gateway; detection from a browser; presence semantics such as online, away and offline; or RTT and clock-offset measurements carried end to end. Give the message a sequence number and a send timestamp so the receiver can compute RTT and spot gaps, and let any application message count as a heartbeat so idle-only traffic stays small.
On mobile, heartbeats cost battery because each one can wake the cellular radio, which then stays in a high-power state for several seconds. Coalesce heartbeats with real traffic, lengthen the interval when the app is backgrounded, and for long background periods rely on the platform's push service rather than holding a socket open.
Worked example: a chat service
A chat service holds one million WebSocket connections behind an Application Load Balancer at its 60-second default, and the product wants presence accurate to about a minute. The plan above gives a 30-second interval and two tolerated misses: about 33,000 heartbeats per second across the fleet, roughly 1,700 per second on each of 20 gateway nodes, which is trivial for an event loop. The gateway sends protocol pings to clients; browsers send application heartbeats so they can show a reconnecting banner within about 40 seconds of the server disappearing.
Gateways keep TCP keepalive at 30, 10 and 3 seconds on their backend links. The first load test exposes the real constraint: when one gateway restarts, 50,000 clients detect it within seconds and reconnect at once. Jittered reconnect backoff, plus spreading heartbeat phases with the random offset in the code above, flattens the spike.
Failure modes
| Failure | Symptom | Prevention |
|---|---|---|
| Interval above an idle timeout | Connections drop at a suspiciously regular age | Measure the path; use half the smallest timeout |
| Relying on Linux keepalive defaults | Dead peers held for over two hours | Set per-socket idle, interval, count and TCP_USER_TIMEOUT |
| gRPC client pings too often | GOAWAY with too_many_pings, failed calls | Raise the server's permitted ping rate together with the client's |
| Proxy answers the ping | App deadlocked but connection looks healthy | Application heartbeat end to end |
| Only heartbeats count | Busy connections time out while streaming data | Any inbound message resets the deadline |
| Synchronised heartbeats and reconnects | Load spikes and reconnect storms | Jitter intervals and backoff |
| Aggressive mobile interval | Battery drain complaints | Coalesce, adapt when backgrounded, use push |
Trade-offs
| Decision | Option A | Option B |
|---|---|---|
| Interval | Short: fast detection, more load and battery | Long: cheap, slow detection and risk of idle drops |
| Layer | Protocol ping: free in most stacks, stops at the first proxy | Application heartbeat: end to end, costs bytes and code |
| Who drives | Server-driven: one policy for the fleet, no client changes | Client-driven: client detects dead servers, needs client releases to tune |
What to do next
- List every hop between client and application and record each idle timeout; measure the ones you cannot look up.
- Write down the detection time the product needs, and compute interval, misses and fleet-wide ping rate.
- Set TCP keepalive and TCP_USER_TIMEOUT per socket on server-side links instead of relying on system defaults.
- Configure the protocol's own mechanism: WebSocket pings, gRPC keepalive on both client and server, SSE comments, MQTT keep alive.
- Add an application heartbeat with sequence and timestamp wherever you need end-to-end or browser-side detection.
- Jitter intervals and reconnect backoff, then load-test a gateway restart.
- Alert on connection age distributions: a spike at one age means a timeout you missed.