QUIC, standardised in RFC 9000 in 2021, carries application data in streams: ordered byte streams that share one encrypted connection over UDP. A bidirectional stream carries bytes in both directions, like a miniature TCP connection. The difference is that a lost packet on one stream does not hold up the others. That property is why HTTP/3 runs on QUIC.
This article explains how streams actually work: how they are numbered and opened, the frames that carry and cancel them, the three independent flow-control limits that decide whether a sender may transmit, and how HTTP/3 uses them. It ends with code, the failure modes seen in production and a tuning checklist. The general idea of multiplexing logical streams is covered in stream multiplexing, and keeping a connection alive across network changes is covered in QUIC connection migration.
Why streams, and what they do not fix
Over TCP, HTTP/2 multiplexes many requests onto one ordered byte stream. If one TCP segment is lost, the kernel withholds every byte after it until the retransmission arrives, even bytes that belong to unrelated requests. This is transport-level head-of-line blocking.
QUIC moves ordering from the connection to the stream. Each STREAM frame carries a stream ID and a byte offset, so the receiver keeps a separate reassembly buffer per stream. A lost packet only delays the streams whose bytes it carried. Everything else is delivered as soon as it arrives.
Two limits remain. Within a single stream, bytes are still delivered strictly in order, so an application that pushes all its messages down one long-lived stream has rebuilt TCP's problem. And streams share the connection's congestion window and connection-level flow control, so a stream can still starve its neighbours of bandwidth or credit. Designing with QUIC is largely about choosing what gets its own stream.
Stream IDs: four types in two bits
A stream ID is a variable-length integer up to 2^62 - 1. Its two lowest bits encode the stream type, so either endpoint can tell who opened a stream and whether it is bidirectional just by looking at the number. Bit 0x01 is the initiator (0 means client) and bit 0x02 is the direction (0 means bidirectional).
| Low bits | Type | IDs in order | Typical use |
|---|---|---|---|
| 0x0 | Client-initiated, bidirectional | 0, 4, 8, 12 | HTTP/3 requests, RPC calls |
| 0x1 | Server-initiated, bidirectional | 1, 5, 9, 13 | Server-opened exchanges (WebTransport) |
| 0x2 | Client-initiated, unidirectional | 2, 6, 10, 14 | HTTP/3 control and QPACK streams |
| 0x3 | Server-initiated, unidirectional | 3, 7, 11, 15 | HTTP/3 control, push and QPACK streams |
There is no handshake to open a stream. The sender just sends a frame on the next unused ID of the right type, and the receiver creates the stream on first sight. Streams of each type must be used in order, and opening one implicitly opens every lower-numbered stream of the same type. If a client's first frame arrives on stream 8 because the packets for 0 and 4 were reordered, the server treats 0, 4 and 8 as open.
Either side may write to a bidi stream: the opener sends first and the peer replies on the same ID. That is the request-response pattern in QUIC, one bidi stream per exchange, closed by a FIN in each direction.
The frames that carry and end streams
Stream data travels in STREAM frames, types 0x08 to 0x0f. The three low bits of the type are flags: 0x04 means an offset field is present, 0x02 means a length field is present, and 0x01 is FIN, meaning this frame contains the last byte of the stream. Several STREAM frames for different streams can share one packet.
A stream ends cleanly when the sender's FIN is received and all bytes up to that final size are in. Each direction of a bidi stream closes independently, so a client can send a request with FIN (half-closing its side) and keep reading the response. Three more frames end streams abruptly:
RESET_STREAM(0x04): the sender abandons its sending part. It carries an application error code and the final size, the number of bytes the sender considers sent. The receiver may discard buffered data for that stream.STOP_SENDING(0x05): the receiver asks the peer to stop sending on a stream, with an error code. The peer is expected to answer with RESET_STREAM.- Connection close ends every stream at once.
The final size matters even after a reset, because flow control is accounted in bytes. Both sides must agree on how much credit the dead stream consumed, so the final size in RESET_STREAM is charged against the connection-level limit even if those bytes never arrived. A reset stream does not give credit back.
One gap is well known. A reset lets the receiver discard everything, including a header the application needs to understand which session was reset. An IETF draft, RESET_STREAM_AT, proposes a reset that still guarantees delivery up to a given offset. At the time of writing it is a draft, so check your stack's support before relying on it.
State machines on each side
RFC 9000 describes a stream as a sending part and a receiving part, each with its own state machine. A bidi stream has both on each endpoint; a unidirectional stream has only one. The states map directly to what an implementation must keep in memory:
| Part | States | What it means for memory |
|---|---|---|
| Sending | Ready, Send, Data Sent, Data Recvd | Unacknowledged bytes are kept for retransmission until Data Recvd |
| Sending (reset) | Reset Sent, Reset Recvd | Buffer can be dropped at once; only the RESET_STREAM must be delivered |
| Receiving | Recv, Size Known, Data Recvd, Data Read | Out-of-order bytes are buffered until the gap fills |
| Receiving (reset) | Reset Recvd, Reset Read | Discard buffer and tell the application |
A stream is fully closed only when both parts on both endpoints reach a terminal state. The practical consequence is that a stream the application forgets about, one where it never reads to the end or never sends FIN, stays open forever. It holds buffers and counts against the stream limit.
Three layers of flow control
A QUIC sender may only transmit stream data when three independent limits all allow it. Each is advertised by the receiver as an absolute value, not an increment, so a delayed or duplicated update cannot over-grant credit. Initial values are exchanged in the handshake as transport parameters, and frames raise them later.
| Limit | Initial transport parameter | Update frame | Blocked signal |
|---|---|---|---|
| Bytes on one stream | initial_max_stream_data_bidi_local, _bidi_remote, _uni | MAX_STREAM_DATA | STREAM_DATA_BLOCKED |
| Bytes across the connection | initial_max_data | MAX_DATA | DATA_BLOCKED |
| Number of streams opened | initial_max_streams_bidi, initial_max_streams_uni | MAX_STREAMS | STREAMS_BLOCKED |
The per-stream limit comes in three flavours because the receiver may want different windows for bidi streams it opened (local), bidi streams the peer opened (remote) and unidirectional streams. The stream-count limit is cumulative, not concurrent: MAX_STREAMS = 100 means the peer may open streams up to the hundredth of that type in total. As streams close, the receiver must keep raising the number, or the peer stalls after its hundredth request even if only two are open. The maximum value is 2^60.
The receiver runs a credit policy. A common one is to grant a fixed window beyond what the application has consumed, and to send an update when half of it has been used, so the sender never waits for a round trip:
# Receiver-side credit policy, per stream (RFC 9000 leaves the policy to the implementation)
on_bytes_consumed_by_app(stream, n):
stream.consumed += n
remaining = stream.max_offset_advertised - stream.consumed
if remaining < stream.window / 2: # half the window used up
stream.max_offset_advertised = stream.consumed + stream.window
queue_frame(MAX_STREAM_DATA(stream.id, stream.max_offset_advertised))
# same rule for MAX_DATA over all streams; raise MAX_STREAMS as streams finishWorked example: a path of 100 Mbit/s with a 50 ms round trip has a bandwidth-delay product of 100,000,000 / 8 x 0.05 = 625,000 bytes. If one stream must fill that path, its window must be at least about 625 KB, and the connection window larger again if several streams run at once. With the aioquic defaults of 1 MiB per stream and 1 MiB per connection, two busy streams already share one connection window, so one of them will sit in DATA_BLOCKED. Windows are also a memory commitment: 8 MiB of connection credit across 10,000 connections is 80 GB of potential buffering. The same pressure seen from the application side is covered in backpressure architecture.
Code: a bidi echo server and client
The Python library aioquic exposes the stream model almost directly. The server below buffers each bidi stream until FIN, echoes it back with FIN, and cancels oversized requests in both directions. Note the sid & 0x2 test, which is the stream-type bit from the table above.
import asyncio
from aioquic.asyncio import QuicConnectionProtocol, serve
from aioquic.quic.configuration import QuicConfiguration
from aioquic.quic.events import StreamDataReceived, StreamReset, StopSendingReceived
APP_CANCELLED = 0x101 # application-defined error code space
class EchoProtocol(QuicConnectionProtocol):
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
self.buffers = {} # stream_id -> bytearray
def quic_event_received(self, event):
if isinstance(event, StreamDataReceived):
sid = event.stream_id
if sid & 0x2: # unidirectional: we cannot reply on it
return
buf = self.buffers.setdefault(sid, bytearray())
buf.extend(event.data)
if len(buf) > 1_000_000: # refuse oversized requests
self._quic.stop_stream(sid, APP_CANCELLED) # sends STOP_SENDING
self._quic.reset_stream(sid, APP_CANCELLED) # sends RESET_STREAM
self.buffers.pop(sid, None)
elif event.end_stream: # peer sent FIN: request complete
reply = bytes(self.buffers.pop(sid))
self._quic.send_stream_data(sid, reply, end_stream=True)
self.transmit()
elif isinstance(event, (StreamReset, StopSendingReceived)):
self.buffers.pop(event.stream_id, None) # free per-stream state promptly
async def main():
config = QuicConfiguration(is_client=False, alpn_protocols=["echo"])
config.load_cert_chain("cert.pem", "key.pem")
config.max_data = 8 * 1024 * 1024 # connection-wide credit
config.max_stream_data = 1024 * 1024 # per-stream credit
await serve("0.0.0.0", 4433, configuration=config, create_protocol=EchoProtocol)
await asyncio.Future()
asyncio.run(main())In aioquic, QuicConfiguration exposes one max_data and one max_stream_data value, both 1 MiB by default. The finer-grained transport parameters in the table are protocol concepts that other stacks, such as quiche, msquic or quic-go, expose under their own names. The client side opens a stream by asking for the next free ID:
# Inside a connected client protocol (aioquic.asyncio.connect gives you one)
def send_request(self, payload: bytes) -> int:
sid = self._quic.get_next_available_stream_id(is_unidirectional=False) # 0, 4, 8, ...
self._quic.send_stream_data(sid, payload, end_stream=True) # data + FIN: half-close
self.transmit()
return sid # the reply arrives as StreamDataReceived events on the same sidEach call uses a fresh stream, so a slow response never blocks another request, at the cost of one stream from the MAX_STREAMS budget.
How HTTP/3 uses streams
HTTP/3 (RFC 9114) is the best-known client of QUIC streams and a good template for your own protocols. Every request and response uses one client-initiated bidi stream. The client sends HEADERS and optional DATA frames and then FIN; the server replies on the same stream and sends FIN. RFC 9114 says at least 100 request streams SHOULD be permitted at a time, so servers advertise at least that in initial_max_streams_bidi.
Connection-wide state lives on unidirectional streams, identified by a type byte at the start: 0x00 for the control stream carrying SETTINGS, 0x01 for server push, 0x02 and 0x03 for the QPACK encoder and decoder streams. Each endpoint must allow the peer at least three unidirectional streams and SHOULD give each at least 1,024 bytes of credit. If a client cancels a request, it resets the stream and sends STOP_SENDING with H3_REQUEST_CANCELLED, and the server stops work on it.
QUIC itself defines no priorities; which stream gets the next packet is the sender's choice. HTTP/3 carries hints using RFC 9218 extensible priorities: an urgency from 0 to 7 (default 3, lower is more urgent) and an incremental flag that says whether a response is useful in pieces. A long-lived session protocol on top of HTTP/3 is described in WebTransport architecture, and the HTTP/2 equivalent in gRPC bidi streaming.
Failure modes seen in production
- Stream budget exhaustion. A receiver that never raises MAX_STREAMS, or an application that leaks streams without FIN, eventually leaves the peer in STREAMS_BLOCKED. Requests hang after a fixed count and work again after reconnecting.
- One slow reader starves the connection. Stream credit is per stream but connection credit is shared. If the application stops reading one large stream, its unconsumed bytes pin connection credit and other streams stall in DATA_BLOCKED.
- Resets without STOP_SENDING. Resetting only your sending part leaves the peer still sending into a stream you no longer read. Cancel both directions, as the server code does.
- 0-RTT replay. Stream data sent in 0-RTT can be replayed by an attacker. Only allow idempotent requests in early data.
- UDP blocked or throttled. Some networks drop or rate-limit UDP. Browsers fall back to HTTP/2 over TCP; your own clients need the same fallback.
For debugging, enable qlog and inspect traces in a visualiser such as qvis; BLOCKED frames point at the limit to raise.
Design trade-offs
| Choice | Gains | Costs |
|---|---|---|
| One bidi stream per request | No blocking between requests; clean cancellation | Stream budget and per-stream credit to manage |
| One long-lived stream per session | Simple ordering, one set of buffers | Head-of-line blocking returns inside the stream |
| Datagrams (RFC 9221) instead of streams | No retransmission delay | No reliability or ordering; the app handles loss |
A good default is one bidi stream per request or independent object, a few long-lived unidirectional streams for control, and datagrams only for data that is worthless once late.
What to do next
- List which messages in your protocol need independent delivery and give each its own stream.
- Compute the bandwidth-delay product of your real paths and set stream and connection windows from it.
- Check that your server raises MAX_STREAMS as streams close.
- Make every code path either read to FIN or send STOP_SENDING.
- Enable qlog in staging and look for any BLOCKED frames.
- Keep a TCP fallback for networks that block UDP.