HTTP/2 kept everything an application sees in HTTP, the methods, status codes, headers and caching rules, and replaced how those messages travel. Instead of one request at a time per TCP connection, it splits messages into binary frames and interleaves many requests, called streams, on a single connection. The semantics live in RFC 9110 and the HTTP/2 wire format in RFC 9113, which replaced the original RFC 7540 in 2022.

The frame types, stream state machine and HPACK internals are covered in HTTP/2 architecture for bidirectional systems. This article is for people who run HTTP/2: how a connection gets negotiated, what one request looks like on the wire, which defaults limit throughput, what changed with priorities and push, how to configure a proxy, why load balancing gets harder, which attacks forced every server to add limits, and how to see what is actually happening. Specifics were checked against RFC 9113 and current nginx documentation; throughput figures were computed, not estimated.

Advertisement

What HTTP/2 changes, and what it does not

HTTP/1.1 sends a request, waits for the whole response, then sends the next on that connection. Browsers worked around this by opening six or so connections per origin and sites worked around it with sprites, bundling and domain sharding. HTTP/2 makes the workarounds unnecessary: one connection carries many concurrent streams, each request gets an odd-numbered stream id chosen by the client, and frames from different streams interleave freely.

Headers are compressed with HPACK, which keeps a table of recently sent header fields on each side of the connection, so repeated cookies and user agents shrink to a few bytes. The initial table size is 4,096 bytes. What did not change: application semantics, the need for TLS in practice, and the fact that everything still rides on one TCP connection, which turns out to matter a great deal.

Three requests on one HTTP/2 connection over one TCP connectionStream 1: GET /HEADERS, then DATAStream 3: GET /app.jsHEADERS, then DATAStream 5: POST /apiHEADERS + DATAFrame interleavingH1 H3 H5 D5 D1 D3 D1 ...TCP byte streamin-order deliveryFlow controlper-stream and connection windowsOne lost segmentstalls every stream behind itMultiplexing removed HTTP/1.1's per-connection request queue; TCP's in-order delivery remains underneath.
Streams are independent at the HTTP layer and serialised again at the TCP layer.

Getting to HTTP/2: ALPN and prior knowledge

On the public web, HTTP/2 is negotiated inside the TLS handshake with ALPN: the client offers h2 and http/1.1, and the server picks one. No extra round trip is spent. If the server or a middlebox in front of it does not pick h2, the connection is plain HTTP/1.1, which is the most common reason a deployment that is supposed to be HTTP/2 is not.

Without TLS, RFC 7540 defined an upgrade path through an Upgrade: h2c request header and an HTTP2-Settings header. It was rarely deployed and RFC 9113 deprecates it. Cleartext HTTP/2 itself is still valid with prior knowledge, where the client simply starts speaking HTTP/2; that is how gRPC usually talks to services inside a trusted network.

Either way, the client then sends a fixed 24-byte preface and both sides exchange SETTINGS frames:

Client                                         Server
------                                         ------
TLS ClientHello  (ALPN: h2, http/1.1)  ------>
                                       <------ ServerHello (ALPN: h2) ... handshake done
"PRI * HTTP/2.0\r\n\r\nSM\r\n\r\n"  (24-byte preface)
SETTINGS (stream 0)                    ------>
                                       <------ SETTINGS (stream 0)
                                       <------ SETTINGS ACK
HEADERS stream 1 (:method GET :path / ...)  -->
SETTINGS ACK                           ------>
                                       <------ HEADERS stream 1 (:status 200)
                                       <------ DATA stream 1 ... DATA stream 1 (END_STREAM)

The preface is chosen so an HTTP/1.1 server reading it fails fast instead of misinterpreting it. SETTINGS carry each side's limits: maximum concurrent streams, initial stream window size, maximum frame size, which starts at 16,384 bytes and can be raised to 16,777,215, and header table size. Each side acknowledges the other's settings. Most surprising behaviour later is explained by what was in these first frames, so look at them first when debugging.

Advertisement

Flow-control windows and the bandwidth-delay product

HTTP/2 has its own flow control above TCP, on two levels: each stream has a send window, and the connection as a whole has another. Both start at 65,535 bytes. A sender may only send DATA frames while both windows have room; the receiver replenishes them with WINDOW_UPDATE frames as it consumes data. The SETTINGS_INITIAL_WINDOW_SIZE setting changes only the per-stream windows. The connection window can only grow through WINDOW_UPDATE frames on stream 0, a detail implementations still get wrong.

The default is small for modern networks. A window caps throughput at window divided by round-trip time, because the sender has to wait for updates once it has sent a window's worth. With 65,535 bytes: at 10 ms RTT that is about 52 Mbit/s per stream, at 50 ms about 10.5 Mbit/s, and at 100 ms about 5.2 Mbit/s. To fill a 100 Mbit/s path with 100 ms RTT, the bandwidth-delay product is 1.25 MB, so the window must be at least that.

Clients and servers that care, including browsers and most gRPC stacks, raise both windows right after the preface, and some grow them automatically by measuring the bandwidth-delay product. If you see a large download over HTTP/2 run at a fraction of the link speed while HTTP/1.1 on the same path is fast, check the windows before anything else. The opposite trade-off applies on servers handling many connections: generous windows mean each connection may buffer that much data in memory, so size them to the paths you serve rather than setting the maximum everywhere.

Head-of-line blocking moved, it did not disappear

HTTP/2 removed head-of-line blocking at the HTTP layer: a slow response no longer delays the requests queued behind it. But TCP delivers bytes in order. If one segment is lost, every byte after it waits in the kernel until the retransmission arrives, even bytes belonging to unrelated streams that arrived intact. On a clean network this is invisible. On a lossy mobile link, a single loss stalls every stream on the connection, and one HTTP/2 connection can perform worse than six HTTP/1.1 connections, where a loss stalls only one.

That is the main reason HTTP/3 exists: QUIC moves streams into the transport so a loss blocks only the stream it belongs to. The trade-off is discussed in HTTP/2 vs HTTP/3 streams and the QUIC side in HTTP/3 and QUIC. For HTTP/2 deployments, the practical mitigations are a modern congestion control and loss recovery stack on the server, as described in TCP Architecture, and offering HTTP/3 alongside HTTP/2 for clients on poor networks.

Coalescing and the 421 status code

Browsers will reuse one HTTP/2 connection for several hostnames if the certificate covers all of them and the names resolve to the same address. Coalescing saves handshakes, but it creates a failure that is hard to diagnose: a request for api.example.com arrives at a server that only answers for www.example.com, because both names were on a wildcard certificate and in the same DNS answer. The server is supposed to answer 421 Misdirected Request, defined in RFC 9110, after which the client retries on a fresh connection. Servers that answer 404 or serve the wrong virtual host instead produce intermittent, client-dependent errors.

If you terminate several hostnames with different backends behind one certificate, either make every edge node able to serve every name on it, or make sure misrouted requests get a 421.

Priorities and push: what is left

RFC 7540 let clients build a dependency tree of streams with weights. Implementations were inconsistent and many servers ignored the signals, so RFC 9113 deprecates that scheme. Its replacement, RFC 9218 Extensible Priorities, is a plain header: priority: u=0, i gives an urgency from 0, most urgent, to 7, with 3 as the default, and an incremental flag saying the response is useful in pieces, like a progressive image. It works the same way over HTTP/3. Servers and CDNs vary in how much they honour it, so treat it as a hint and measure.

Server push let a server send responses the client had not asked for. It was hard to use well and often sent resources already in the browser cache; Chrome disabled it by default in 2022 and other major clients have dropped it as well. Use the 103 Early Hints status with preload links instead, which lets the browser decide what to fetch.

Configuring a proxy

Most teams meet HTTP/2 through a reverse proxy or load balancer. An nginx example:

server {
    listen 443 ssl;
    http2 on;                              # nginx 1.25.1+; older: listen 443 ssl http2;
    ssl_certificate     /etc/ssl/site.pem;
    ssl_certificate_key /etc/ssl/site.key;

    http2_max_concurrent_streams 128;      # per connection; keep at or above 100
    client_header_timeout 10s;             # bound slow or never-ending header blocks

    location /api/ {
        proxy_pass http://app_backend;     # upstream leg is HTTP/1.1 unless you use grpc_pass
        proxy_http_version 1.1;
        proxy_set_header Connection "";    # allow upstream keepalive
    }
}

Note the version change: from nginx 1.25.1 HTTP/2 is enabled with the http2 directive, and the old http2 parameter on listen is deprecated but still accepted with a warning from nginx -t. Also note that HTTP/2 on the client side does not mean HTTP/2 to the backend: with proxy_pass nginx speaks HTTP/1.x upstream, and needs proxy_http_version 1.1 for upstream keepalive as above, which is usually fine and simplifies backends. RFC 9113 recommends that the concurrent-stream limit be no smaller than 100 so parallelism is not throttled.

From application code, check what you got rather than assuming it:

import httpx   # pip install "httpx[http2]"

with httpx.Client(http2=True, timeout=10) as client:
    for path in ("/", "/a", "/b"):
        r = client.get("https://example.com" + path)
        print(path, r.http_version, r.status_code)   # expect "HTTP/2" on every line

Load balancing long-lived connections

HTTP/1.1 clients open and close connections often, so a layer-4 load balancer that spreads connections spreads load. HTTP/2 clients open one connection and keep it for hours, sending everything over it. Behind a connection-level balancer, a new backend receives no traffic until clients reconnect, and one busy client pins all its load to one backend. This shows up most with gRPC, covered in gRPC Architecture.

Three fixes, usually combined: balance per request at layer 7 with a proxy that terminates HTTP/2 and spreads streams across backends; use client-side balancing with a connection per backend; and cap connection age on servers, which then send GOAWAY so clients reconnect gracefully and land somewhere new. GOAWAY carries the last stream id the server will process, so in-flight requests above it can be retried safely elsewhere.

Attacks that shaped server limits

Multiplexing means one connection can create a lot of work, and two disclosures forced every implementation to add limits. Rapid Reset (CVE-2023-44487), disclosed in October 2023 and used in record-sized DDoS attacks, opened streams and immediately cancelled them with RST_STREAM. Because cancelled streams do not count toward the concurrent-stream limit, a client could start far more requests than the limit implied, and servers did the work of each. Fixes count resets per connection and close connections that cancel too often.

The CONTINUATION flood, disclosed in April 2024 across many implementations, sent a header block as an endless series of CONTINUATION frames without ever ending it. Some servers buffered it until memory ran out; others burned CPU decoding it, and because no request completed, nothing appeared in access logs. Fixes cap the total header size and the time allowed to finish a header block. The operational lesson is the same for both: keep HTTP/2 libraries and proxies patched, keep the default limits on header size, stream count and reset rate rather than raising them blindly, and alert on connections closed for protocol errors.

Debugging toolkit

# What did we actually negotiate?
curl -s -o /dev/null -w '%{http_version}\n' https://example.com      # prints 2 if h2 was used
curl -sv --http2 https://example.com -o /dev/null 2>&1 | grep -i alpn

# Cleartext HTTP/2 with prior knowledge (internal services, gRPC without TLS)
curl -sv --http2-prior-knowledge http://localhost:8080/health

# Frame-level view: SETTINGS, WINDOW_UPDATE, HEADERS, DATA with stream ids
nghttp -nv https://example.com

# Load test with many streams per connection: 10 connections x up to 100 streams
h2load -n 20000 -c 10 -m 100 https://staging.example.com/api/ping

Use curl to confirm the negotiated version and ALPN, nghttp to see SETTINGS and WINDOW_UPDATE frames with stream ids, and h2load to load-test with realistic stream concurrency, since a load generator that opens a connection per request tests HTTP/1.1 behaviour even when it speaks HTTP/2. In browsers, enable the Protocol column in the developer tools network panel. When TLS is in the way, the details in TLS, in depth on ALPN and session keys apply.

Trade-offs

One connection per origin means fewer handshakes, better header compression and fair sharing of one congestion window, at the price of TCP head-of-line blocking on lossy paths and harder load balancing. Large flow-control windows give throughput on long paths and cost memory on busy servers. Terminating HTTP/2 at the edge and speaking HTTP/1.1 upstream keeps backends simple, while end-to-end HTTP/2 is needed for gRPC and long-lived streams. For browsers on mobile networks, add HTTP/3; for service-to-service traffic in a data centre, HTTP/2 remains an excellent default.

What to do next

  1. Check what your clients really negotiate with curl's http_version output, from outside your network and from behind your CDN.
  2. Dump the SETTINGS your servers send with nghttp and confirm stream limits of at least 100 and windows sized for your longest common round trip.
  3. On nginx 1.25.1 or later, move to the http2 directive and clear the deprecation warning.
  4. Confirm your HTTP/2 stack has the Rapid Reset and CONTINUATION fixes, and alert on connections closed for protocol errors.
  5. If you serve gRPC or other long-lived clients, set a maximum connection age or balance per request at layer 7.
  6. Replace any remaining server push with 103 Early Hints, and add the priority header where ordering matters.
  7. Test a lossy mobile profile and decide whether to offer HTTP/3 alongside HTTP/2.
Key takeaway: HTTP/2 keeps HTTP's meaning and changes its transport: many streams over one TLS connection, negotiated with ALPN, compressed with HPACK and governed by two levels of flow control. Most production problems trace back to a few things: default 65,535-byte windows on long round trips, TCP head-of-line blocking on lossy links, long-lived connections that defeat connection-level balancing, coalescing without 421, and unpatched stacks exposed to Rapid Reset and CONTINUATION floods. Look at the negotiated version and SETTINGS first, size windows to your paths, and keep the limits.