Every long-lived connection needs some periodic traffic, but the word heartbeat hides two different jobs. One is keeping the path open: NAT devices, load balancers and proxies drop connections that stay idle too long, usually without telling either end. The other is detecting a dead peer: a crashed process, a phone that walked out of coverage or a partition can leave a connection half-open for hours. The two jobs want different intervals, live at different layers, and fail in different ways.

Why silent death happens and how ping/pong detects it is explained in heartbeats and dead-connection detection. This article is the configuration guide that follows: how to choose intervals and timeouts from the path you actually have, which mechanism each protocol offers and what its defaults are, when you need an application-level heartbeat on top, and what it all costs at a million connections.

Advertisement

Two jobs, two constraints

Keep-alive traffic exists to reset idle timers in middleboxes. Its interval is set by the shortest idle timeout anywhere on the path. Miss it once and the connection is gone, usually silently: the middlebox forgets the flow, and the next packet from either side is dropped or answered with a reset.

Liveness detection exists to notice a dead peer quickly enough to act: reconnect, fail over, mark a user offline, release a lock. Its constraint is the detection time the product needs, which is roughly the interval plus the timeout, or interval times the number of misses tolerated. Its cost is false positives: declare death too eagerly and a slow network causes reconnect storms.

One mechanism can do both jobs, but design for each separately and then take the tighter interval. A chat app may need presence updates within 30 seconds while its load balancer allows 350 seconds idle; the presence requirement wins. A telemetry stream may tolerate five minutes of undetected death while a carrier NAT drops idle flows much sooner; the NAT wins.

Every hop has its own idle timer; every layer has its own heartbeatClientapp, socketHome or carrier NATidle timer: unknownLoad balanceridle timer: e.g. 60 sProxyread timeoutServerappApplication heartbeat: end to end, proves the app is alive, carries sequence and timestampsWebSocket ping/pong, HTTP/2 PING, MQTT PINGREQ: one connection, client to terminating proxyTCP keepalive: one hopTCP keepalive: one hopTCP keepalive: one hopLower layers are cheaper but cover less of the path. Only an application heartbeat crosses every proxy.The interval must beat the shortest idle timer anywhere on the path, not just the one you configured.
Each hop has its own idle timer, and each layer's heartbeat covers a different span of the path. TCP keepalive covers one hop; protocol pings reach the terminating proxy; only an application heartbeat reaches the application.

The arithmetic

Three formulas cover most decisions. First, the keep-alive interval must be below the smallest idle timeout on the path with margin for jitter and one lost packet; half the smallest timeout is a sound default. Second, worst-case detection time is about interval plus timeout when one missed pong is fatal, or interval times misses when you tolerate several. Third, the server-side load is connections divided by interval.

def plan(conns, idle_timeouts_s, detect_target_s, misses=2):
    interval = min(min(idle_timeouts_s) / 2, detect_target_s / misses)
    return {
        "interval_s": interval,
        "worst_detect_s": interval * misses,
        "pings_per_s": round(conns / interval),
    }

# ALB at its 60 s default idle timeout, product wants detection within 60 s
print(plan(1_000_000, [60, 350], 60))
# {'interval_s': 30.0, 'worst_detect_s': 60.0, 'pings_per_s': 33333}

A third of a million pings per second sounds alarming, but each is a few bytes handled by the event loop; the cost is mostly wakeups and, on mobile, radio time. What matters is that the cost is linear in connections and inversely proportional to the interval, so halving the interval doubles it. Know your idle timeouts: on AWS, an Application Load Balancer's idle timeout defaults to 60 seconds, and a Network Load Balancer's TCP idle timeout defaults to 350 seconds and can be set between 60 and 6,000 seconds. Home routers and carrier NATs do not publish theirs and can be much shorter, so for clients on the open internet measure, or assume tens of seconds.

Advertisement

TCP keepalive: cheap, one hop, slow by default

TCP keepalive sends empty probes on an idle socket and resets the connection if they go unanswered. On Linux the system-wide defaults are 7,200 seconds of idle time before the first probe, 75 seconds between probes and 9 probes, so a dead peer is detected after more than two hours. Those defaults are useless for both jobs. Set the values per socket instead:

import socket

def tune_keepalive(sock, idle=30, interval=10, count=3, user_timeout_ms=40_000):
    sock.setsockopt(socket.SOL_SOCKET, socket.SO_KEEPALIVE, 1)
    sock.setsockopt(socket.IPPROTO_TCP, socket.TCP_KEEPIDLE, idle)       # Linux
    sock.setsockopt(socket.IPPROTO_TCP, socket.TCP_KEEPINTVL, interval)
    sock.setsockopt(socket.IPPROTO_TCP, socket.TCP_KEEPCNT, count)
    # Bound how long sent data may stay unacknowledged before the kernel gives up.
    sock.setsockopt(socket.IPPROTO_TCP, socket.TCP_USER_TIMEOUT, user_timeout_ms)

Keepalive probes only fire when the socket is idle. If the peer vanishes while you have unacknowledged data in flight, retransmission timers govern instead, and they can take many minutes; TCP_USER_TIMEOUT caps that. Two limits make TCP keepalive insufficient on its own. The probe is answered by the peer's kernel, so a deadlocked application with a healthy kernel looks alive. And every proxy or load balancer that terminates TCP answers probes itself, so keepalive covers exactly one hop. Use it as a cheap backstop, on server-to-server links and between proxies and backends.

WebSocket ping and pong

RFC 6455 defines Ping (opcode 0x9) and Pong (opcode 0xA) control frames. A control frame's payload is at most 125 bytes, a Pong must echo the Ping's application data, and an endpoint may send an unsolicited Pong as a one-way heartbeat. Pings reach the endpoint that terminates the WebSocket, which may be a gateway rather than your application.

The asymmetry that shapes browser designs: browsers answer server pings automatically, but the JavaScript WebSocket API has no method to send a ping or observe a pong. A server can detect dead browsers with protocol pings, but a browser that wants to detect a dead server, for example to show a reconnecting banner, must use application messages:

// browser side: application-level heartbeat over a WebSocket
function heartbeat(ws, intervalMs = 25000, timeoutMs = 10000) {
  let seq = 0, timer = null, deadline = null;
  const tick = () => {
    clearTimeout(deadline);
    ws.send(JSON.stringify({ type: "hb", seq: ++seq, t: Date.now() }));
    deadline = setTimeout(() => ws.close(4000, "heartbeat timeout"), timeoutMs);
  };
  ws.addEventListener("message", (ev) => {
    clearTimeout(deadline);                 // any inbound message proves liveness
    const msg = JSON.parse(ev.data);
    if (msg.type === "hb-ack") recordRtt(Date.now() - msg.t);
  });
  timer = setInterval(tick, intervalMs + Math.random() * 2000);   // jitter
  ws.addEventListener("close", () => { clearInterval(timer); clearTimeout(deadline); });
}

Treat any inbound frame as proof of life, not only the heartbeat reply, so busy connections never time out spuriously. Reconnection after a timeout needs backoff and jitter, as covered in the WebSocket guide.

HTTP/2 PING and gRPC keepalive

HTTP/2 has a PING frame with an 8-byte opaque payload; the receiver must reply with a PING carrying the same payload and the ACK flag. It measures round-trip time and liveness of the whole connection, across all its streams. gRPC builds its keepalive on it, and the grpc.io keepalive guide documents the C-core defaults: the client's keepalive time is disabled (INT_MAX), the keepalive timeout is 20 seconds, and pings without active calls are off. The server's own keepalive time is 2 hours, and by default it permits client pings no more often than every 5 minutes.

That last default is the classic trap. A client configured to ping every 30 seconds against a server left at defaults will be told off: the guide says the server eventually sends GOAWAY with the debug data too_many_pings, and the client's calls fail. Client and server settings must be changed together:

import grpc

client_opts = [
    ("grpc.keepalive_time_ms", 30_000),            # ping after 30 s of inactivity
    ("grpc.keepalive_timeout_ms", 10_000),         # wait 10 s for the ack
    ("grpc.keepalive_permit_without_calls", 1),    # ping even with no active calls
    ("grpc.http2.max_pings_without_data", 0),      # lift the default cap of 2 pings
]
server_opts = [
    ("grpc.http2.min_ping_interval_without_data_ms", 20_000),  # allow the client's rate
    ("grpc.keepalive_permit_without_calls", 1),
]
channel = grpc.insecure_channel("svc:50051", options=client_opts)
server = grpc.server(executor, options=server_opts)

The string keys above are the C-core channel argument names that Python passes through. Note the third client option: by default C-core lets a client send only two pings without any data or header frame in between, so an idle channel stops pinging unless that cap is lifted. On the server, grpc.http2.max_ping_strikes (default 2) is how many over-frequent pings it tolerates before the GOAWAY. Defaults and option names differ between language implementations, so check your runtime's documentation rather than copying values between them. Streaming RPC patterns are covered in bidirectional gRPC.

SSE comments and MQTT keep alive

Server-Sent Events has no ping frame, but any line beginning with a colon is a comment the client ignores. Sending a comment every 15 to 30 seconds keeps proxies from closing the response as idle. Many reverse proxies time out an upstream read after about a minute of silence by default, and nginx's proxy_read_timeout is 60 seconds. The browser's EventSource reconnects on its own when the stream drops, so detection on the client is built in; the server sees a dead client only when a write fails.

async def sse_stream(queue, send):
    while True:
        try:
            event = await asyncio.wait_for(queue.get(), timeout=20)
            await send(f"id: {event.id}\ndata: {event.json}\n\n")
        except asyncio.TimeoutError:
            await send(": keep-alive\n\n")        # comment line; resets proxy idle timers

MQTT builds keep alive into the session. The client declares a keep-alive period in CONNECT; if it has nothing else to send it sends PINGREQ and expects PINGRESP. The broker disconnects a client from which it receives no control packet within one and a half times the keep-alive period, and then publishes the client's will message if one was set, which makes keep alive the basis of presence in many IoT systems. Choose the period against the cellular NAT on the device side, not the broker; see MQTT for bidirectional messaging.

When you need an application heartbeat

Protocol-level pings stop at the first component that terminates the protocol. Add an application heartbeat when you need any of these: liveness of the application itself rather than its kernel or gateway; detection from a browser; presence semantics such as online, away and offline; or RTT and clock-offset measurements carried end to end. Give the message a sequence number and a send timestamp so the receiver can compute RTT and spot gaps, and let any application message count as a heartbeat so idle-only traffic stays small.

On mobile, heartbeats cost battery because each one can wake the cellular radio, which then stays in a high-power state for several seconds. Coalesce heartbeats with real traffic, lengthen the interval when the app is backgrounded, and for long background periods rely on the platform's push service rather than holding a socket open.

Worked example: a chat service

A chat service holds one million WebSocket connections behind an Application Load Balancer at its 60-second default, and the product wants presence accurate to about a minute. The plan above gives a 30-second interval and two tolerated misses: about 33,000 heartbeats per second across the fleet, roughly 1,700 per second on each of 20 gateway nodes, which is trivial for an event loop. The gateway sends protocol pings to clients; browsers send application heartbeats so they can show a reconnecting banner within about 40 seconds of the server disappearing.

Gateways keep TCP keepalive at 30, 10 and 3 seconds on their backend links. The first load test exposes the real constraint: when one gateway restarts, 50,000 clients detect it within seconds and reconnect at once. Jittered reconnect backoff, plus spreading heartbeat phases with the random offset in the code above, flattens the spike.

Failure modes

FailureSymptomPrevention
Interval above an idle timeoutConnections drop at a suspiciously regular ageMeasure the path; use half the smallest timeout
Relying on Linux keepalive defaultsDead peers held for over two hoursSet per-socket idle, interval, count and TCP_USER_TIMEOUT
gRPC client pings too oftenGOAWAY with too_many_pings, failed callsRaise the server's permitted ping rate together with the client's
Proxy answers the pingApp deadlocked but connection looks healthyApplication heartbeat end to end
Only heartbeats countBusy connections time out while streaming dataAny inbound message resets the deadline
Synchronised heartbeats and reconnectsLoad spikes and reconnect stormsJitter intervals and backoff
Aggressive mobile intervalBattery drain complaintsCoalesce, adapt when backgrounded, use push

Trade-offs

DecisionOption AOption B
IntervalShort: fast detection, more load and batteryLong: cheap, slow detection and risk of idle drops
LayerProtocol ping: free in most stacks, stops at the first proxyApplication heartbeat: end to end, costs bytes and code
Who drivesServer-driven: one policy for the fleet, no client changesClient-driven: client detects dead servers, needs client releases to tune

What to do next

  1. List every hop between client and application and record each idle timeout; measure the ones you cannot look up.
  2. Write down the detection time the product needs, and compute interval, misses and fleet-wide ping rate.
  3. Set TCP keepalive and TCP_USER_TIMEOUT per socket on server-side links instead of relying on system defaults.
  4. Configure the protocol's own mechanism: WebSocket pings, gRPC keepalive on both client and server, SSE comments, MQTT keep alive.
  5. Add an application heartbeat with sequence and timestamp wherever you need end-to-end or browser-side detection.
  6. Jitter intervals and reconnect backoff, then load-test a gateway restart.
  7. Alert on connection age distributions: a spike at one age means a timeout you missed.
Key takeaway: Heartbeats do two jobs: keeping middlebox state alive and detecting dead peers. Size the interval from the shortest idle timeout on the path and the detection time the product needs, take the tighter, and account for the linear cost. Use TCP keepalive as a tuned one-hop backstop, the protocol's own ping where it exists, matched settings on both ends for gRPC, and an application heartbeat whenever liveness must be proven end to end or detected from a browser.