HTTP/2 and HTTP/3 expose the same thing to an application: many concurrent request and response exchanges over one connection, each carried on its own stream, with the same methods, status codes and header fields. The difference is underneath. HTTP/2 builds its streams on top of a single TCP byte stream, so every stream shares one ordering. HTTP/3 maps each request onto a QUIC stream, and QUIC orders bytes per stream rather than per connection. Almost every practical difference between the two, good and bad, follows from that one design choice.

This article compares the two stream models side by side, works through a lossy mobile link, and ends with failure modes and a rollout checklist. For the full mechanics of each protocol on its own, see HTTP/2 architecture in depth and QUIC bidi streams in depth.

Advertisement

Where the ordering lives

Same HTTP semantics, different place for stream orderingHTTP/2HTTP/3HTTP/2 framingstreams, HPACK, WINDOW_UPDATETLS 1.3 recordsencrypts one ordered byte streamTCPONE ordered byte stream, kernelIPHTTP/3 framingHEADERS/DATA per stream, QPACKQUICstreams, flow control, loss recoveryTLS 1.3 handshake inside, per-packet AEADuserspace, connection IDsUDPIPPacket carrying stream 5 is lost:HTTP/2 over TCPkernel holds bytes for streams 1, 3, 7 behind the gapevery stream waits for the retransmission(transport head-of-line blocking)HTTP/3 over QUICstreams 1, 3, 7 are delivered to the applicationonly stream 5 waits for its retransmitted frame(unless QPACK makes another stream wait for it)
HTTP/2 multiplexes streams inside one TCP byte stream, so one lost packet stalls every stream. HTTP/3 hands each request to a QUIC stream, and QUIC delivers each stream's bytes independently.

In HTTP/2 a stream is a logical sequence of frames that share a stream identifier. Frames from different streams are interleaved, encrypted into TLS records and delivered by TCP as one ordered byte stream. TCP knows nothing about HTTP/2 streams. If segment 40 is lost and segments 41 to 60 arrive, the kernel buffers 41 to 60 and the HTTP/2 library sees nothing until 40 is retransmitted, even if 41 to 60 belong entirely to other streams.

In HTTP/3 the stream is a transport object. QUIC packets carry STREAM frames, each tagged with a stream ID and a byte offset. When a packet is lost, QUIC retransmits the frames it carried, not the packet, and the receiver can deliver data for every other stream immediately because each stream has its own reassembly buffer. An HTTP/3 request stream carries HEADERS and DATA frames, and closing the QUIC stream ends the message.

Within a single stream both protocols still deliver bytes in order, so the benefit appears only when several streams are active and loss hits one of them.

Stream identifiers and lifecycles

HTTP/2 stream IDs are 31-bit integers. Clients open odd-numbered streams, servers use even numbers for server push, and stream 0 is reserved for connection-level frames such as SETTINGS, PING and GOAWAY. A stream moves through idle, open, half-closed and closed states, driven by the END_STREAM flag and by RST_STREAM. GOAWAY announces the last stream a server will process, so clients can retry later requests elsewhere.

QUIC stream IDs are 62-bit variable-length integers whose two low bits encode type: bit 0 says who opened the stream (client or server) and bit 1 says whether it is bidirectional or unidirectional. That gives four independent sequences, each counting up in steps of four. HTTP/3 uses client-initiated bidirectional streams for requests. It uses unidirectional streams, identified by a type byte at their start, for the control stream, the QPACK encoder and decoder streams, and server push. Server-initiated bidirectional streams are not used by HTTP/3 itself and are reserved for extensions such as WebTransport.

Cancellation differs in a way that matters for streaming. HTTP/2 has a single RST_STREAM that kills both directions. QUIC separates the two: RESET_STREAM abandons the sending direction and STOP_SENDING asks the peer to abandon its sending direction. A client can abandon a response while still finishing its upload. HTTP/3 carries application error codes in these frames, for example H3_REQUEST_CANCELLED.

Advertisement

Concurrency: a cap versus credit

HTTP/2 limits concurrency with SETTINGS_MAX_CONCURRENT_STREAMS, a cap on how many streams may be open at once. When a stream closes, a slot frees up automatically. QUIC uses cumulative credit instead. The peer advertises the highest stream number you may open, separately for bidirectional and unidirectional streams, through the initial transport parameters and later MAX_STREAMS frames. Closing a stream frees nothing until the peer sends more credit. If your peer is slow to raise MAX_STREAMS, you are stuck, and a well-behaved sender reports that with a STREAMS_BLOCKED frame.

This changes what you monitor. On HTTP/2 the question is how many streams are open. On HTTP/3 it is how quickly the server replenishes credit as requests finish, and whether clients spend time blocked waiting for it. A conservative setting shows up as queuing inside the client library, not as an error.

Flow control in two levels

Both protocols control flow per stream and per connection, and both make the receiver grant credit. In HTTP/2, every stream and the connection start with a 65,535-byte window. WINDOW_UPDATE frames add increments. SETTINGS_INITIAL_WINDOW_SIZE changes the starting window for streams but not for the connection, so a server that raises it and forgets to send a connection-level WINDOW_UPDATE still caps total throughput at 64 KiB per round trip, a classic cause of slow uploads on high-latency links.

QUIC sets initial limits through transport parameters such as initial_max_data and initial_max_stream_data_bidi_local, and raises them with MAX_DATA and MAX_STREAM_DATA frames. These carry absolute byte offsets rather than increments, which makes them safe to retransmit or reorder. Either way, the window must cover the bandwidth-delay product: a 100 Mbit/s path with 80 ms round trip needs about 1 MB in flight. A receiver that auto-tunes looks roughly like this:

# Receiver-side window auto-tuning, the same idea in HTTP/2 and QUIC.
def on_data_consumed(stream, n_bytes, now):
    stream.consumed += n_bytes
    remaining = stream.limit - stream.consumed
    if remaining < stream.window / 2:                 # half the window used
        elapsed = now - stream.last_update
        if elapsed < 2 * rtt_estimate():               # draining fast: grow
            stream.window = min(stream.window * 2, MAX_WINDOW)
        stream.limit = stream.consumed + stream.window
        stream.last_update = now
        send_credit(stream.id, stream.limit)          # QUIC: MAX_STREAM_DATA(limit)
                                                      # HTTP/2: WINDOW_UPDATE(limit - old)
    # repeat the same logic for the connection-level window

Flow control is also your backpressure. If the application stops reading a stream, credit stops flowing and the sender blocks on that stream only. That is the property bidirectional streaming relies on, and it is covered further in backpressure in streaming.

Header compression: why HPACK could not survive

HPACK, used by HTTP/2, keeps a dynamic table of recently sent header fields in both encoder and decoder. Every header block may insert entries and reference earlier ones, and that works only because TCP delivers header blocks in exactly the order they were encoded. Over QUIC there is no such cross-stream order. If stream 9's headers reference a table entry inserted by stream 5's headers, and stream 5's packet is lost, stream 9 cannot be decoded.

QPACK solves this by moving table updates onto a dedicated unidirectional encoder stream and acknowledgements onto a decoder stream. A header block may reference only entries the decoder already has, or the decoder must hold that stream until the insertion arrives. That second case is called a blocked stream, and the decoder caps how many it tolerates with SETTINGS_QPACK_BLOCKED_STREAMS. The dynamic table capacity, SETTINGS_QPACK_MAX_TABLE_CAPACITY, defaults to zero, so an implementation that never raises it gets only static-table and literal compression and no blocking at all.

This is the subtle point in any comparison: an aggressive QPACK encoder trades compression ratio for head-of-line blocking. If your encoder freely references unacknowledged entries, a loss on the encoder stream stalls every request that depends on it. If HTTP/3 helps less than expected under loss, check the blocked-streams metric first.

Priorities: the tree is gone

HTTP/2 originally defined a dependency tree with weights. Clients built it inconsistently, servers ignored it or implemented it badly, and RFC 9113 deprecated it. Its replacement, the extensible priority scheme in RFC 9218, works the same way on both protocols. A request carries a Priority header with an urgency from 0 to 7, where 3 is the default and lower is more urgent, and an incremental flag that says whether the response is useful in pieces. A client can change priority later with a PRIORITY_UPDATE frame.

GET /app.css   priority: u=0          # render-blocking, send first, whole
GET /hero.jpg  priority: u=2, i       # important, progressive is fine
GET /track.js  priority: u=5          # background

Long-lived bidirectional streams

gRPC runs bidirectional streaming over HTTP/2 streams, and so do WebSockets bootstrapped with the extended CONNECT method from RFC 8441. HTTP/3 has the matching extended CONNECT in RFC 9220, and WebTransport builds a richer session on HTTP/3 with additional streams and unreliable datagrams; see WebTransport. gRPC over HTTP/3 exists in some implementations, but support and maturity vary by language and library, so treat it as something to verify for your stack rather than assume. gRPC bidirectional streaming covers the HTTP/2 path in detail.

For a long-lived stream, the HTTP/3 advantage is narrower than for page loads. A single chat stream gains nothing from per-stream ordering. What it gains is independence from the other streams sharing its connection, and connection migration, discussed next. What it loses is that UDP NAT bindings often expire faster than TCP ones, so an idle stream needs QUIC keepalive pings, and the negotiated max_idle_timeout must be longer than your quietest period.

Connection setup, migration and discovery

A new HTTP/2 connection needs a TCP handshake and a TLS 1.3 handshake, two round trips before the first request. QUIC combines transport and TLS into one round trip, and resumed connections can send 0-RTT data. 0-RTT data can be replayed by an attacker, so servers should accept only idempotent requests in it, and clients and frameworks should not put state-changing requests there.

QUIC connections are identified by connection IDs rather than the address four-tuple, so a phone moving from Wi-Fi to cellular can keep its connection and every open stream. With TCP the connection dies and every stream must be re-established. This benefit only holds if your load balancer routes by connection ID; a layer 4 balancer that hashes the four-tuple sends the migrated packets to the wrong backend.

Clients do not try HTTP/3 blindly. They learn about it from an Alt-Svc response header or from an HTTPS DNS record (RFC 9460), and browsers typically race or fall back to TCP when UDP is blocked. A minimal nginx setup serving both looks like this:

server {
    listen 443 ssl;
    listen 443 quic reuseport;
    http2 on;
    ssl_certificate     /etc/ssl/site.pem;
    ssl_certificate_key /etc/ssl/site.key;
    add_header Alt-Svc 'h3=":443"; ma=86400' always;
}

Verify from a client with curl, which reports the negotiated version:

curl -sS -o /dev/null -w '%{http_version}\n' --http2 https://example.com/
curl -sS -o /dev/null -w '%{http_version}\n' --http3-only https://example.com/

Worked example: a mobile client on a lossy link

Take an app that opens one connection and keeps about 20 requests in flight: API calls, thumbnails and a streaming feed. The network has 80 ms round trip and 2 percent packet loss, typical of a congested cell. With HTTP/2, each loss stalls all 20 streams for at least one round trip, often more when the loss is detected by timeout rather than by duplicate acknowledgements. The user sees every image arrive late together.

With HTTP/3, the same loss stalls only the stream whose data was in the lost packet. Nineteen requests keep flowing. Small API responses that fit in one or two packets finish on time unless their own packet is lost. The improvement is largest for many small, independent responses and smallest for one large transfer. If the user leaves Wi-Fi, the QUIC connection migrates and the feed survives; HTTP/2 would reconnect from scratch.

Verify this for your traffic. Add delay and loss with Linux netem, replay a real request mix, and compare p95 completion times per request type.

Failure modes in production

  • UDP blocked or throttled. Some enterprise networks and middleboxes drop or rate-limit UDP 443. Clients fall back to TCP, so the site works, but your HTTP/3 share is lower than expected and latency for fallback users may include a failed attempt.
  • Path MTU problems. QUIC requires paths that carry 1,200-byte UDP payloads and probes upward. Tunnels and misconfigured MTUs cause silent drops of larger packets.
  • CPU cost. QUIC runs in userspace and encrypts every packet; without UDP GSO and GRO it costs noticeably more CPU than TCP.
  • Rapid reset. CVE-2023-44487 abused HTTP/2 by opening and immediately resetting streams so the concurrency cap never applied. Servers now count resets; make sure yours is patched and that HTTP/3 stream credit is replenished at a bounded rate too.
  • Credit and window starvation. A small connection window in HTTP/2 or slow MAX_STREAMS updates in HTTP/3 show up as low throughput with no errors.
  • QPACK blocking. Aggressive dynamic-table use reintroduces cross-stream stalls under loss.

Choosing

SituationPreferWhy
Browsers and mobile apps on lossy or changing networksOffer HTTP/3 with HTTP/2 fallbackPer-stream loss recovery, faster setup, migration
Service-to-service inside a data centreHTTP/2Low loss, mature tooling, cheaper CPU, gRPC support everywhere
Single large transfersEitherOne stream gains nothing from per-stream ordering
Long-lived bidirectional streamsHTTP/2 today, HTTP/3 where the stack supports itCheck library support and NAT idle timeouts
Unreliable real-time dataWebTransport datagrams over HTTP/3Neither protocol's streams drop stale data

What to do next

  1. Measure your current HTTP/2 traffic: streams per connection, connection-level window size and the share of users on mobile networks.
  2. Enable HTTP/3 at the edge with Alt-Svc or an HTTPS DNS record, keep HTTP/2 as fallback, and confirm with curl that both versions negotiate.
  3. Check that your load balancer routes QUIC by connection ID, or accept that migration will not work yet.
  4. Turn on GSO and GRO for the UDP sockets and compare CPU per request before and after.
  5. Replay real traffic under emulated delay and loss and compare p95 completion time per request type.
  6. Watch HTTP/3 share, fallback rate, QPACK blocked streams and STREAMS_BLOCKED counts after rollout.
  7. Send Priority headers for render-critical resources and verify the server honours urgency.
  8. Patch HTTP/2 rapid-reset protections and set limits on stream creation and reset rates for both protocols.
Key takeaway: HTTP/2 and HTTP/3 share HTTP semantics but place stream ordering at different layers. HTTP/2 multiplexes streams over one TCP byte stream, so a single lost packet stalls them all; HTTP/3 uses QUIC streams that recover from loss independently, add connection migration and shorten setup. In exchange QUIC needs stream credit replenishment, QPACK instead of HPACK, UDP that the network does not block, and more CPU. Offer HTTP/3 to users on lossy and mobile networks with HTTP/2 fallback, keep HTTP/2 where links are clean, and confirm the benefit with your own traffic under emulated loss.