QUIC is a transport protocol: it does the job TCP does, delivering reliable, ordered, congestion-controlled bytes, and adds the encryption TLS provides, in one layer that runs over UDP. It was standardised by the IETF in 2021 as RFC 9000, with TLS integration in RFC 9001 and loss recovery and congestion control in RFC 9002. HTTP/3 (RFC 9114) is HTTP mapped onto it, and most large web properties now serve a large share of traffic over it.

The point of this article is not the history but the mechanism. We build QUIC from the bottom up: what goes in a packet, how the handshake reaches application data in one round trip, how streams avoid TCP's head-of-line blocking, how loss detection works when packet numbers never repeat, and how a connection survives a change of IP address. Then we cover what changes when you run it in production, where UDP brings a new set of operational problems. If you want a refresher on the TCP baseline first, read TCP/IP, the protocol stack.

Advertisement

Why a new transport, and why on UDP

TCP has three problems that are hard to fix in place. Head-of-line blocking: TCP delivers one ordered byte stream, so when HTTP/2 multiplexes twenty requests over it, one lost segment stalls all twenty. Handshake latency: TCP plus TLS 1.3 costs two round trips before the first request byte; TCP Fast Open tried to remove one and was defeated by middleboxes. Ossification: TCP lives in kernels and its headers are visible to every middlebox, which depend on its exact behaviour, so new features take a decade to deploy.

QUIC answers all three by moving the transport into user space on top of UDP and encrypting nearly all of it. UDP passes almost everywhere; user space lets a browser ship a new congestion controller in a normal release; encrypted headers cannot be read or rewritten by middleboxes, so they cannot ossify. The cost is that offloads the kernel and NICs did for TCP need new mechanisms for UDP.

QUIC layering and the 1-RTT handshakeHTTP/3 or your protocolrequests on streamsQUIC framesSTREAM, ACK, CRYPTO, MAX_DATAQUIC packets3 packet number spacesUDP datagramskernel, GSO/GROTLS 1.3 runs inside CRYPTO frames;the transport is encrypted too.ClientServerInitial: ClientHello, padded to 1200 BInitial: ServerHelloHandshake: cert, Finished (3x limit)1-RTT: server data may startHandshake: Finished + 1-RTT request1-RTT: response, HANDSHAKE_DONEThe request leaves after one round trip; with 0-RTT it rides in the first flight.
Left: QUIC's layers. Right: the 1-RTT handshake, with packet types labelled. The server's first flight is capped at three times the bytes the client sent until the client's address is validated.

Packets, frames and packet number spaces

A UDP datagram carries one or more QUIC packets. Long-header packets (Initial, 0-RTT, Handshake, Retry) carry the version and both connection IDs; after the handshake, short-header 1-RTT packets carry only the destination connection ID, a packet number and the encrypted payload. The payload holds frames: STREAM for application bytes, ACK, CRYPTO for TLS messages, MAX_DATA, MAX_STREAM_DATA and MAX_STREAMS for flow control, plus PATH_CHALLENGE, NEW_CONNECTION_ID and CONNECTION_CLOSE.

Packets are numbered in three packet number spaces, Initial, Handshake and Application Data, each with its own keys. Within a space, packet numbers only increase and are never reused, even for retransmissions.

Most integers on the wire use a variable-length encoding whose two top bits give the size, and a stream ID's two lowest bits encode who opened it and whether it is bidirectional:

def encode_varint(v: int) -> bytes:
    """RFC 9000 variable-length integer: the top 2 bits give the length (1, 2, 4 or 8 bytes)."""
    if v < 2**6:
        return v.to_bytes(1, "big")
    if v < 2**14:
        return (v | 0x4000).to_bytes(2, "big")
    if v < 2**30:
        return (v | 0x8000_0000).to_bytes(4, "big")
    if v < 2**62:
        return (v | 0xC000_0000_0000_0000).to_bytes(8, "big")
    raise ValueError("varint out of range")


def decode_varint(buf: bytes, pos: int = 0) -> tuple[int, int]:
    length = 1 << (buf[pos] >> 6)
    value = buf[pos] & 0x3F
    for b in buf[pos + 1:pos + length]:
        value = (value << 8) | b
    return value, pos + length


def stream_kind(stream_id: int) -> str:
    """The two low bits of a stream ID encode initiator and direction."""
    who = "server" if stream_id & 0x1 else "client"
    way = "uni" if stream_id & 0x2 else "bidi"
    return f"{who}-{way}"

assert encode_varint(15293) == bytes.fromhex("7bbd")      # example from RFC 9000
assert [stream_kind(i) for i in (0, 1, 2, 3, 4)] == [
    "client-bidi", "server-bidi", "client-uni", "server-uni", "client-bidi"]
Advertisement

The handshake: TLS 1.3 inside QUIC

QUIC does not run TLS over itself as a byte stream; it uses TLS 1.3's handshake messages and key schedule directly. The client's first packet is an Initial carrying a ClientHello in a CRYPTO frame, with QUIC transport parameters in a TLS extension. Initial packets are protected with keys derived from the connection ID, which any observer can compute, so they provide integrity against accidents but not secrecy. The client must pad datagrams containing Initial packets to at least 1,200 bytes.

The server replies with an Initial (ServerHello), then Handshake packets carrying the certificate and Finished under real handshake keys. Until it has validated the client's address, the server may send at most three times the bytes received. This anti-amplification limit stops reflection attacks, and has a practical consequence: a 1,200-byte client Initial allows a 3,600-byte reply, so a larger certificate chain costs an extra round trip. Keep chains short. Under attack, a server can demand proof of address first with a Retry packet, also at one extra round trip.

Once the client has the server's Finished, it sends its own Finished and can immediately send its request in a 1-RTT packet in the same datagram. Application data flows after one round trip instead of TCP plus TLS's two.

0-RTT and the replay problem

A returning client can store a session ticket and send application data in 0-RTT packets alongside the ClientHello, before any round trip. But an attacker who captures 0-RTT data can replay it, and the server cannot tell the copy from the original.

So 0-RTT is safe only for requests harmless to repeat: servers typically accept only safe methods such as GET, and never payments or writes. A server may refuse 0-RTT, and the client resends after the handshake. In a custom protocol, make 0-RTT an opt-in per message type.

Streams and flow control

A stream is an ordered byte sequence carried in STREAM frames with offsets. Opening one needs no handshake, just a new ID within the peer's MAX_STREAMS limit. Each stream is ordered independently: if a packet carrying stream 8's bytes is lost, stream 12's bytes are still delivered. Head-of-line blocking remains only within a stream.

Flow control has two levels: per stream (MAX_STREAM_DATA) and per connection (MAX_DATA), raised by the receiver as the application consumes data. On long fat paths, a window below the bandwidth-delay product caps throughput, as with TCP. WebTransport exposes these streams, plus RFC 9221 datagrams, to browser applications.

Loss detection and congestion control

Because packet numbers are never reused, an ACK always identifies exactly which transmission arrived. TCP retransmits a segment with the same sequence number and cannot tell whether an ACK refers to the original or the retransmission, which pollutes RTT samples. QUIC's lost data is sent again in a new packet with a new number, so every RTT sample is clean. ACK frames also carry many ranges and an explicit ACK delay, so the sender knows how long the receiver held the ACK and can subtract it.

RFC 9002 declares a packet lost when a packet sent at least three packets later has been acknowledged, or when it was sent more than nine eighths of an RTT before an acknowledged packet. When nothing is acknowledged at all, a probe timeout (PTO) sends probe packets to elicit an ACK rather than immediately assuming disaster:

# Simplified from RFC 9002, one packet number space.
kPacketThreshold = 3            # reordering tolerance, in packets
kTimeThreshold   = 9 / 8        # reordering tolerance, in RTTs
kGranularity     = 0.001        # 1 ms timer granularity

def on_ack_received(ack, space):
    newly_acked = space.remove_acked(ack.ranges)
    if space.largest_acked in newly_acked and newly_acked[-1].ack_eliciting:
        # Packet numbers are never reused, so this sample is unambiguous.
        rtt.update(now() - newly_acked[-1].time_sent, ack.ack_delay)
    lost = detect_lost_packets(space)
    cc.on_packets_acked(newly_acked)
    cc.on_packets_lost(lost)        # lost frames are re-queued in NEW packets
    set_loss_detection_timer()

def detect_lost_packets(space):
    loss_delay = max(kTimeThreshold * max(rtt.latest, rtt.smoothed), kGranularity)
    lost = []
    for pkt in space.unacked():
        if pkt.number > space.largest_acked:
            continue
        if (space.largest_acked - pkt.number >= kPacketThreshold
                or now() - pkt.time_sent >= loss_delay):
            lost.append(pkt)
    return lost

def pto_interval():
    # Probe timeout: fires when nothing is acknowledged for too long; it sends probes,
    # it does not by itself declare anything lost or collapse the window.
    return rtt.smoothed + max(4 * rtt.var, kGranularity) + max_ack_delay

The RFC specifies a NewReno-style congestion controller with an initial window of min(10 x max_datagram_size, max(14,720 bytes, 2 x max_datagram_size)); with 1,200-byte datagrams that is 12,000 bytes. Because congestion control lives in user space, implementations commonly ship CUBIC or BBR instead, and can change it without a kernel upgrade. The flip side is that two QUIC stacks on the same path may behave very differently, so benchmark the implementation you actually deploy.

Connection IDs, migration and stateless reset

TCP names a connection by its four-tuple, so moving from Wi-Fi to cellular kills every connection. QUIC names it by connection IDs chosen by each endpoint, so packets from a new address are still recognised. Before trusting the new path, the server sends a PATH_CHALLENGE and waits for a matching PATH_RESPONSE, then resets congestion state for that path. Endpoints issue spare IDs via NEW_CONNECTION_ID and switch to a fresh one after migrating, so observers cannot link the paths.

A load balancer that reads routing bits in the connection ID can keep a connection on one backend across address changes; the IETF QUIC-LB work, still a draft at the time of writing, standardises this. A server that lost a connection's state answers with a stateless reset, ending in a token the peer received earlier, so the client closes at once. Compare Multipath TCP, which handles mobility with several TCP subflows.

A worked example: a page load on a lossy mobile link

Take a client with an 80 ms round-trip time and 2 percent packet loss, fetching an HTML page and then 30 assets from one origin.

StepTCP + TLS 1.3 + HTTP/2QUIC + HTTP/3
First request byte leavesafter 2 RTT = 160 msafter 1 RTT = 80 ms (0 ms with 0-RTT on a resumed connection)
One packet lost mid-transferevery stream stalls until retransmission, roughly 1 RTT or moreonly streams with bytes in that packet stall
30 assets, 2% loss, about 300 packetsabout 6 losses, each able to block all streamsabout 6 losses, each blocking one or a few streams
Phone switches Wi-Fi to cellularconnection reset, new handshakepath validation, about 1 RTT, streams continue

The gains concentrate where latency and loss are high, which is why mobile users see the biggest improvement. On a clean data-centre link with sub-millisecond RTTs, the handshake saving is negligible and QUIC's higher CPU cost per byte may dominate.

Operating QUIC in production

Several problems appear the day you turn QUIC on at scale:

  • CPU cost. A user-space stack does per-packet encryption and system calls that TCP offloaded. Use UDP GSO and GRO on Linux to send and receive batches of datagrams in one call, and expect higher CPU per gigabit than tuned TCP.
  • Socket buffers. Default UDP receive buffers are small, and a burst overflows them silently. Raise them and watch drop counters.
  • Load balancing. Four-tuple hashing breaks on migration and NAT rebinding. Route on connection IDs, or accept that migrated connections reset.
  • Fallback. Some networks block or rate-limit UDP. Clients discover HTTP/3 through Alt-Svc headers or HTTPS DNS records and must fall back to TCP quickly; always keep TCP serving.
  • Visibility. Encrypted headers defeat packet-capture tools. Use endpoint logging, such as qlog traces, and export per-connection RTT, loss and congestion window metrics.
# Is HTTP/3 actually negotiated end to end? (curl built with HTTP/3 support)
curl --http3 -sI https://example.com/ | head -1        # expect: HTTP/3 200

# Larger UDP socket buffers; quic-go's documentation suggests values of this order
sysctl -w net.core.rmem_max=7500000
sysctl -w net.core.wmem_max=7500000

# Watch UDP drops at the socket: rising RcvbufErrors means the server cannot keep up
nstat -az | grep -E 'UdpRcvbufErrors|UdpInErrors'

A CDN usually terminates QUIC for you at the edge; see how CDNs work for where that fits in the request path.

Failure modes

  • Silent fallback. HTTP/3 is advertised, but UDP is blocked for many users, who fall back to TCP while dashboards say QUIC is enabled.
  • Amplification-limited handshakes. Large certificate chains add a round trip for every new connection.
  • 0-RTT replay. A non-idempotent endpoint accepts early data and a replayed request executes twice.
  • Receive-buffer drops. Throughput collapses under bursts with no errors in the application log; only UDP drop counters show it.
  • Load balancer misrouting. After NAT rebinding, packets hash to a different backend, which answers with stateless resets.
  • Idle timeouts. NAT bindings for UDP often expire faster than for TCP; long-lived idle connections need keep-alive PINGs.

Trade-offs

AspectQUICTCP + TLS
Handshake1 RTT, 0-RTT on resumption2 RTT (TLS 1.3), 1 with TLS 0-RTT
Head-of-line blockingper stream onlywhole connection
Mobilityconnection migrationnew connection per path
Evolutionuser-space releases, encrypted headerskernel upgrades, ossified middleboxes
CPU per bytehigherlower, heavily offloaded
Network reachUDP sometimes blockeduniversal

What to do next

  1. Enable HTTP/3 at your CDN or edge proxy and keep TCP serving as the fallback.
  2. Measure the share of requests actually served over HTTP/3, per network and client, not just whether it is enabled.
  3. Trim certificate chains so the server's first flight fits within the anti-amplification limit.
  4. Decide which endpoints may accept 0-RTT and restrict it to idempotent requests.
  5. Raise UDP socket buffers on QUIC servers, enable GSO and GRO, and alert on UDP receive-buffer errors.
  6. Configure load balancers to route on connection IDs or confirm the edge terminates QUIC.
  7. Collect qlog or equivalent traces for a sample of connections so you can debug without packet captures.
  8. Compare p50 and p95 page-load or RPC latency on high-RTT mobile users before and after rollout.
Key takeaway: QUIC moves the transport into user space over UDP and folds TLS 1.3 into it. The result is a one-round-trip handshake, streams that do not block each other, unambiguous loss detection because packet numbers never repeat, and connections that survive address changes because they are named by connection IDs. The price is CPU, UDP-specific operational work and the need for a TCP fallback. Deploy it where latency and loss are high, measure real HTTP/3 usage, keep 0-RTT to idempotent requests, and give the servers the buffers and load balancing it needs.