HTTP/2 is usually introduced as a faster way to load web pages. For real-time systems it matters for a different reason: it turns one TCP connection into many independent streams, each of which can carry data in both directions at once. gRPC bidirectional streaming, WebSockets over HTTP/2, and long-lived server-sent event feeds all depend on that property, and they all inherit HTTP/2's limits along with it.

This article explains HTTP/2 as the machinery underneath those systems. It covers how a connection starts, what a frame looks like, how streams move through their states, how header compression and flow control are shared across a connection, how WebSockets run inside a stream, and what attacks and operational problems appear when many long-lived streams share one socket. The protocol is defined by RFC 9113, which replaced RFC 7540 in 2022, with header compression in RFC 7541.

Advertisement

Why HTTP/2 matters for bidirectional systems

HTTP/1.1 allows one request at a time per connection, and a response must finish before the next can start. A client that wants a live feed, a chat channel and ordinary requests needs several connections, each with its own TCP and TLS handshake and its own congestion state. HTTP/2 replaces this with streams. Each request-response exchange is a stream with its own id, frames from different streams interleave on one connection, and each stream is full duplex: the client can keep sending request data while the server is already sending response data.

That last property is what makes HTTP/2 a bidirectional transport. A gRPC bidirectional call is one stream in which both sides send messages until each closes its half. The general idea of multiplexing many logical channels over one connection is covered in stream multiplexing; this article is about how HTTP/2 specifically implements it.

One TCP connection, many independent bidirectional streamsApplicationgRPC call, WebSocket, fetchStream layerids, states, windowsFraming layer9-byte header + payloadTLS over TCPone ordered byte streamstream 1gRPC bidistream 3WebSocketstream 5GET /feedHPACK state + connection windowshared by every streamInterleaved frames on the wireH1 D1 D3 H5 D1 WU0 D3 PING ...A lost TCP segment stalls every stream behind it:multiplexing removes HTTP-level blocking, not transport-level blocking.
Applications open streams; every stream's frames share one HPACK state, one connection-level flow-control window and one TCP byte stream.

Connection start and the frame layer

Over TLS, the client and server agree on HTTP/2 during the handshake using ALPN with the identifier h2. The client then sends a fixed 24-byte connection preface, PRI * HTTP/2.0\r\n\r\nSM\r\n\r\n, followed by a SETTINGS frame. The server sends its own SETTINGS frame, and each side acknowledges the other's. From then on everything is frames.

Every frame begins with a 9-byte header: a 24-bit payload length, an 8-bit type, 8 bits of flags, one reserved bit and a 31-bit stream id. Stream 0 is the connection itself and carries SETTINGS, PING, GOAWAY and connection-level WINDOW_UPDATE. The maximum payload defaults to 16,384 bytes and can be raised by SETTINGS_MAX_FRAME_SIZE up to 16,777,215. A small parser shows how little there is to the format:

import struct

TYPES = {0x0: "DATA", 0x1: "HEADERS", 0x2: "PRIORITY", 0x3: "RST_STREAM",
         0x4: "SETTINGS", 0x5: "PUSH_PROMISE", 0x6: "PING", 0x7: "GOAWAY",
         0x8: "WINDOW_UPDATE", 0x9: "CONTINUATION"}
PREFACE = b"PRI * HTTP/2.0\r\n\r\nSM\r\n\r\n"

def parse_frames(buf):
    """Yield (type, flags, stream_id, payload) from a byte buffer after the preface."""
    i = 0
    while i + 9 <= len(buf):
        length = int.from_bytes(buf[i:i + 3], "big")          # 24-bit payload length
        ftype, flags = buf[i + 3], buf[i + 4]
        stream_id = struct.unpack(">I", buf[i + 5:i + 9])[0] & 0x7FFFFFFF   # drop reserved bit
        if i + 9 + length > len(buf):
            break                                              # incomplete frame
        yield TYPES.get(ftype, hex(ftype)), flags, stream_id, buf[i + 9:i + 9 + length]
        i += 9 + length

def settings(payload):
    """SETTINGS payload: repeated 16-bit id + 32-bit value."""
    return {struct.unpack(">H", payload[k:k + 2])[0]: struct.unpack(">I", payload[k + 2:k + 6])[0]
            for k in range(0, len(payload), 6)}
FrameCarriesNotes for streaming
HEADERS / CONTINUATIONCompressed header blockStarts a stream; trailers too. CONTINUATION frames must follow immediately
DATAMessage bytesOnly frame type subject to flow control
RST_STREAMError codeAborts one stream without closing the connection
SETTINGSParametersMust be acknowledged; values apply per direction
WINDOW_UPDATECredit incrementStream id 0 for the connection window
PING8 opaque bytesKeepalive and round-trip measurement
GOAWAYLast stream id, errorGraceful or fatal connection shutdown
PUSH_PROMISEPromised requestServer push; disabled by major browsers, not a bidi mechanism
Advertisement

Streams and their state machine

Clients open streams with odd ids and servers use even ids, so both sides can open streams without coordinating. Ids only increase and are never reused. A long-lived connection that opens streams quickly will eventually exhaust the 31-bit space and must be replaced, which clients handle by opening a new connection.

A stream starts idle. Sending or receiving HEADERS makes it open. The END_STREAM flag on a HEADERS or DATA frame closes that sender's direction, leaving the stream half-closed: half-closed (local) for the side that sent END_STREAM, half-closed (remote) for the other. When both directions have ended, or when either side sends RST_STREAM, the stream is closed. Streams reserved by server push have their own reserved states.

Half-closing is the protocol-level expression of bidirectional streaming. A gRPC client that finishes sending calls the equivalent of close-send, which sends END_STREAM, while it continues to read the server's messages. The server's final HEADERS frame, carrying gRPC status in trailers, closes the other half. The peer limits how many streams you can have open with SETTINGS_MAX_CONCURRENT_STREAMS. The specification recommends allowing at least 100. A client that holds many long-lived streams can reach this limit and see new calls queue or fail, even though the connection looks idle.

HPACK: header compression is connection state

Headers are compressed with HPACK. It combines a static table of 61 common header fields, a dynamic table of recently sent fields, and Huffman coding for literals. A repeated header such as a long authorization token costs one or two bytes after its first use, which matters when a streaming client sends many small messages with metadata.

The dynamic table defaults to 4,096 bytes, adjustable by SETTINGS_HEADER_TABLE_SIZE. It is shared by all streams in each direction, so header blocks must be decoded in exactly the order they were sent. That is why a HEADERS frame followed by CONTINUATION frames must be sent contiguously, with no frames from other streams in between. HPACK was designed after compression attacks on earlier schemes: it does not compress across fields in a way that lets an attacker guess secrets, and a sender can mark sensitive fields as never-indexed so intermediaries do not store them. Mark credentials that way.

Flow control: two windows, one of them shared

HTTP/2 flow control is credit-based and applies only to DATA frames. Each stream has a send window and so does the connection. A DATA frame can be sent only if both windows have enough credit, and it reduces both. The receiver restores credit with WINDOW_UPDATE frames: stream id 0 for the connection window, the stream's id for its own window.

Both windows start at 65,535 bytes. The asymmetry that implementations often get wrong is that SETTINGS_INITIAL_WINDOW_SIZE changes the initial size of stream windows only, including a retroactive adjustment for streams already open, which can drive them negative. The connection window can be raised only by WINDOW_UPDATE on stream 0. A server that raises the stream setting to several megabytes but never sends a connection-level update still lets its peer have only 64 KB of DATA in flight toward it, across all streams.

class SendWindows:
    """Sender-side flow control. Only DATA payloads consume window."""
    def __init__(self, initial_stream_window=65535):
        self.conn = 65535                      # never changed by SETTINGS
        self.initial = initial_stream_window
        self.streams = {}

    def open(self, sid):
        self.streams[sid] = self.initial

    def on_settings_initial_window(self, new):
        delta = new - self.initial             # applies to all open streams, may go negative
        self.initial = new
        for sid in self.streams:
            self.streams[sid] += delta

    def on_window_update(self, sid, inc):
        if sid == 0:
            self.conn += inc
        else:
            self.streams[sid] += inc

    def can_send(self, sid, n):
        return min(self.conn, self.streams[sid]) >= n

    def sent(self, sid, n):
        self.conn -= n
        self.streams[sid] -= n

For bidirectional streaming, flow control is how a slow reader pushes back on a fast writer. Its failure modes, including deadlocks when both sides stop reading and libraries that hide back-pressure by buffering without limit, are covered in gRPC bidirectional streaming and backpressure.

WebSockets over HTTP/2

HTTP/1.1 WebSockets use the Upgrade mechanism to take over a whole TCP connection. HTTP/2 has no Upgrade. RFC 8441 instead defines an extended CONNECT method that turns a single stream into a WebSocket tunnel. The server advertises support with SETTINGS_ENABLE_CONNECT_PROTOCOL. The client then opens a stream with a CONNECT request that carries a :protocol pseudo-header, and a 200 response establishes the tunnel:

# Server -> client, in its SETTINGS frame (RFC 8441)
SETTINGS_ENABLE_CONNECT_PROTOCOL (0x8) = 1

# Client -> server: HEADERS on a new stream, no END_STREAM
:method    = CONNECT
:protocol  = websocket
:scheme    = https
:path      = /chat
:authority = chat.example.com
sec-websocket-version = 13

# Server -> client: HEADERS
:status = 200          # not 101: there is no Upgrade in HTTP/2

# From here, DATA frames on this stream carry WebSocket frames in both directions.
# Closing the WebSocket ends the stream; other streams on the connection are unaffected.

Each WebSocket now costs one stream instead of one connection, so a page with several sockets shares a single TLS session, and WebSocket traffic follows HTTP/2's flow control. Support varies across browsers, servers and proxies. A proxy that does not implement extended CONNECT forces clients back to HTTP/1.1 WebSockets, so test the whole path. The WebSocket protocol itself is covered in WebSockets in depth.

Priorities and connection lifecycle

RFC 7540 defined a dependency tree for stream priorities. It was complex, inconsistently implemented, and deprecated in RFC 9113. Its replacement, RFC 9218, uses a priority header with an urgency from 0 to 7, defaulting to 3, and an incremental flag, plus a PRIORITY_UPDATE frame to change priority mid-stream. Priorities are hints; a server may ignore them. For a real-time app, give interactive streams a lower urgency number than bulk transfers.

Long-lived connections need explicit lifecycle handling. PING frames serve as keepalives and detect dead peers faster than TCP. GOAWAY shuts down gracefully: it carries the highest stream id the sender will process, and streams above it can be safely retried on a new connection. To drain a server for deployment, send a GOAWAY with the maximum stream id, wait about one round trip for requests already in flight, then send a second GOAWAY with the real last id, and close once open streams finish. Clients with long-lived streams must be written to reconnect and resume, because a drain eventually ends every stream.

Worked example: a chat client on one connection

A chat client opens one HTTP/2 connection. Stream 1 is a gRPC bidirectional call: the client sends typing and message events, and the server sends new messages and presence updates. Stream 3 fetches history with an ordinary request. Stream 5 uploads an image.

Suppose the image upload is 5 MB and the server's connection window is still at the 65,535-byte default. The upload consumes the whole connection window, and the chat stream's small DATA frames wait behind it until the server sends WINDOW_UPDATE on stream 0. The user sees chat lag during every upload. The fixes are for the server to raise the connection window with an early WINDOW_UPDATE on stream 0, and for the client's sender to schedule interactive streams ahead of the upload's DATA and limit how much any single stream takes from the shared window. If the lag persists on a lossy mobile network, the cause is TCP head-of-line blocking, which HTTP/2 cannot fix; that is the case for HTTP/3.

Failure modes and attacks

  • Rapid Reset (CVE-2023-44487). Disclosed in October 2023, the attack opens streams and immediately cancels them with RST_STREAM. Cancelled streams stop counting against the concurrency limit, but the server may already have started work for each one. Mitigate by limiting resets per connection over time and closing abusive connections with GOAWAY.
  • CONTINUATION flood. Disclosed in 2024 across many implementations: an endless series of CONTINUATION frames without END_HEADERS forces the server to buffer or decode headers indefinitely. Enforce SETTINGS_MAX_HEADER_LIST_SIZE and a time and size limit on header blocks.
  • Resource floods. Excessive SETTINGS, PING or empty frames can consume CPU. Count control frames and disconnect peers above a threshold.
  • Load imbalance. A layer-4 load balancer pins a long-lived connection, and all its streams, to one backend. Balance per stream at layer 7, or cap connection age so clients reconnect and redistribute.
  • Stream-id exhaustion and idle timeouts. Very long connections run out of ids, and middleboxes drop silent ones. Send PINGs and rotate connections deliberately.

Trade-offs

HTTP/2 is the right transport for bidirectional streaming when you want gRPC, many concurrent channels per client, and standard HTTP infrastructure. Its limit is TCP: one lost packet delays every stream. HTTP/3 moves streams onto QUIC so loss affects only one stream, at the cost of UDP deployment issues. Plain HTTP/1.1 WebSockets remain the most widely supported option for browser messaging, but each one needs its own connection.

What to do next

  1. Capture one session with a tool that decodes HTTP/2 frames and identify the preface, SETTINGS, HEADERS and WINDOW_UPDATE frames.
  2. Check your server's SETTINGS_MAX_CONCURRENT_STREAMS and initial window, and send a large connection-level WINDOW_UPDATE early.
  3. Confirm your server and proxies have Rapid Reset and CONTINUATION mitigations enabled and patched.
  4. Implement two-stage GOAWAY draining and test that clients reconnect and resume their streams.
  5. Load-test with a mix of bulk and interactive streams, and measure interactive latency during bulk transfers.
  6. If you use WebSockets, test extended CONNECT across the full proxy path before relying on it.
Key takeaway: HTTP/2 turns one TCP connection into many full-duplex streams, which is what gRPC bidirectional calls, WebSockets over HTTP/2 and long-lived feeds depend on. Frames carry a 9-byte header, streams move through open and half-closed states, and HPACK state and the connection flow-control window are shared by every stream. SETTINGS_INITIAL_WINDOW_SIZE changes only stream windows, so raise the connection window explicitly. Use extended CONNECT for WebSockets, RFC 9218 priorities for urgency, and two-stage GOAWAY for draining. Defend against Rapid Reset and CONTINUATION floods, and remember that TCP head-of-line blocking still applies to every stream.