Real-time applications care about a different property than file downloads do. A download wants every byte, in order, eventually. A multiplayer game, a collaborative cursor, a live sports scoreboard or a telemetry dashboard wants the newest state as soon as possible, and is often better off never receiving an old update than receiving it late. TCP cannot express that difference: one lost packet holds up everything behind it. QUIC can, but only if you design your application protocol to use its pieces deliberately.

This article is about that design. The mechanics of QUIC streams themselves, such as stream id types, state machines and flow control, are covered in QUIC bidi streams architecture. Here the question is: given a stream of application messages with different freshness and ordering needs, which go on long-lived streams, which get a stream each, which go in datagrams, and what do you do when an update becomes stale before it arrives? The answer is worked through for a 30 Hz multiplayer game, with Python code using the aioquic library, and the reasoning carries over to collaboration tools and live dashboards.

Advertisement

What streams fix, and what they do not

QUIC (RFC 9000) runs many independent byte streams inside one encrypted connection over UDP. Each stream is delivered reliably and in order, but streams are ordered only relative to themselves. If a packet carrying data for stream 8 is lost, stream 8 waits for the retransmission while streams 4 and 12 keep delivering. That removes the cross-message head-of-line blocking that hurts WebSocket and other single-pipe protocols over TCP.

It does not remove head-of-line blocking within a stream. If you put every game update onto one long-lived stream, a single loss stalls all later updates on it until the retransmission arrives, typically one round-trip plus loss-detection time. You have rebuilt TCP inside QUIC. The unit of independence is the stream, so the first design decision is how much data shares one.

Two other properties matter. All streams share one congestion controller, so a bulk transfer on one stream still slows real-time traffic on another. And QUIC's transport has no priority field: when several streams have data ready, the order in which a sender packs them into packets is a local decision of the implementation, which you influence through whatever API your library offers.

Four ways to map messages onto QUIC

PatternOrderingOn lossGood for
One long-lived streamTotal order across all messagesEverything behind the loss waitsControl messages, chat, commands that must apply in sequence
Stream per object or per updateWithin one object onlyOnly that object waits; it can be resetSnapshots, keyframes, files, independent RPCs
Stream per entity or channelPer entityOnly that entity waitsDocuments in a collaboration app, rooms in a chat
Datagrams (RFC 9221)NoneThe datagram is gone; no retransmissionInputs, positions, cursors: small and superseded quickly

Most real systems combine three of these, as in the diagram. The test for each message type is two questions: does the receiver need every instance, and does it need them in order relative to other messages? Chat answers yes and yes, so it goes on an ordered stream. A full world snapshot answers "only the newest", so each snapshot gets its own stream that can be abandoned. A player's position answers no and no, so it goes in a datagram.

One QUIC connection, three lanes: pick the lane by what a lost packet should costGame clientone connectionGame serverper-session stateControl stream (bidi, long-lived)join, chat, purchases: orderedUnidirectional stream per snapshotreset when supersededDatagrams (RFC 9221)inputs, positions: never retransmittedkeyframe NLoss on one lane never blocks another lane: a stalled snapshot stream does not delay chat or inputs.All lanes still share one congestion controller, so a flood on one lane slows the others.
A hybrid layout for a game session: a long-lived control stream for ordered events, one unidirectional stream per snapshot that is reset when a newer one is sent, and datagrams for high-rate positional state.
Advertisement

Datagrams: unreliable, but not unmanaged

The unreliable datagram extension (RFC 9221) adds a DATAGRAM frame to QUIC. Both endpoints must advertise support with the max_datagram_frame_size transport parameter. Datagrams are encrypted and congestion controlled like everything else, but they are never retransmitted and are not part of any stream, so they cannot block anything.

Three constraints shape how you use them. First, a datagram must fit in one QUIC packet; it cannot be fragmented, so keep payloads well under the path MTU minus QUIC overhead, and plan for paths around 1,200 bytes, the minimum QUIC requires. Second, they can arrive out of order or duplicated, so carry a sequence number or tick in every one and drop anything older than what you already have. Third, a sender may drop a datagram locally when the congestion window is full; reliable delivery is not even attempted, so anything the receiver must eventually see does not belong here.

A useful pattern is redundancy instead of retransmission: each input datagram carries the last three inputs, so a single loss costs nothing and the server still never waits.

Resetting streams that have gone stale

Stream-per-snapshot only pays off if you abandon old snapshots. A sender ends a stream abruptly with RESET_STREAM, which tells the peer to stop expecting data and stops the sender from retransmitting lost frames for it. A receiver that no longer wants a stream sends STOP_SENDING, asking the sender to reset it. Both carry an application error code, which your protocol defines.

The code below is the server half. Each new snapshot resets the previous snapshot stream, then opens a new unidirectional stream. Positions go out as datagrams. RFC 9000 only lets a sender reset a stream whose data is not yet all acknowledged, so a production version tracks acknowledgement and skips the reset for a stream that has already been delivered; libraries differ in how they handle a late reset.

import asyncio, time
from aioquic.asyncio import QuicConnectionProtocol
from aioquic.quic.events import StreamDataReceived, DatagramFrameReceived, StreamReset

SUPERSEDED = 0x10      # application error code, defined by our protocol

class GameServerProtocol(QuicConnectionProtocol):
    def __init__(self, *args, **kwargs):
        super().__init__(*args, **kwargs)
        self.snapshot_stream = None    # stream id of the snapshot still in flight

    def send_snapshot(self, payload: bytes):
        # A newer snapshot makes the old one worthless: stop retransmitting it.
        # (Production code skips this when the old stream is already fully acked.)
        if self.snapshot_stream is not None:
            self._quic.reset_stream(self.snapshot_stream, SUPERSEDED)
        sid = self._quic.get_next_available_stream_id(is_unidirectional=True)
        self._quic.send_stream_data(sid, payload, end_stream=True)
        self.snapshot_stream = sid
        self.transmit()

    def send_position(self, entity_id: int, tick: int, xyz: bytes):
        # Fits in one packet, superseded within ~33 ms: a datagram, not a stream.
        self._quic.send_datagram_frame(entity_id.to_bytes(4, "big") +
                                       tick.to_bytes(4, "big") + xyz)
        self.transmit()

    def quic_event_received(self, event):
        if isinstance(event, DatagramFrameReceived):
            handle_input(event.data)                 # drop if tick is older than last seen
        elif isinstance(event, StreamDataReceived) and event.stream_id == 0:
            handle_control(event.data, event.end_stream)   # client's first bidi stream
        elif isinstance(event, StreamReset):
            pass                                     # peer abandoned a stream; nothing to repair

The receiver keeps only the newest complete snapshot. Streams can finish in a different order from the one they were opened in, so it compares ticks rather than trusting arrival order:

# Receiver side of the snapshot lane: keep only the newest complete snapshot.
latest_tick = -1
buffers = {}                                  # stream id -> bytearray

def on_stream_data(stream_id, data, end_stream):
    buf = buffers.setdefault(stream_id, bytearray())
    buf += data
    if not end_stream:
        return
    del buffers[stream_id]
    tick = int.from_bytes(buf[:4], "big")
    if tick > latest_tick:                    # streams can complete out of order
        apply_snapshot(tick, bytes(buf[4:]))
        set_latest(tick)

def on_stream_reset(stream_id):
    buffers.pop(stream_id, None)              # a newer snapshot is already on its way

One caution: a reset stream's data is discarded on the receiving side, including bytes that already arrived. If a partial delivery would be useful, the IETF has been working on a reliable-reset extension that lets the sender keep a prefix; treat it as a draft and check your library's support before designing around it.

The stream budget

Streams are cheap but not free. Each side limits how many streams the peer may open with MAX_STREAMS credits, counted separately for bidirectional and unidirectional streams, and an endpoint that runs out blocks until the peer grants more. Limits are cumulative counts of stream ids, not concurrent streams, so a peer extends them as streams close. Most libraries do that automatically, but the initial value is a transport parameter you set.

Do the arithmetic for a variant of the worked example below. If the server sent a full snapshot every tick at 30 Hz, it would open 30 unidirectional streams a second per client, 108,000 an hour. If the client's initial unidirectional limit is 100 and its library only grants new credit after the application reads streams to completion, a slow render loop can exhaust the credit in about three seconds, and snapshots then stall entirely: exactly the failure stream-per-object was meant to avoid. Set the receiving limit with headroom for at least a few seconds of traffic, consume reset streams promptly, and graph "stream blocked" events.

Flow-control windows matter too. A 40 KB snapshot on a stream whose initial window is 16 KB needs an extra round-trip for credit. Size per-stream initial windows to your largest routine object.

Scheduling: who goes first when the pipe is full

When the congestion window is smaller than the data waiting, the sender's scheduler decides what gets sent. That is where real-time quality is won or lost. A sensible order for the game is: control stream first (small and important), then datagrams for inputs and positions, then the newest snapshot stream, then bulk such as asset downloads.

How you express that depends on the layer you use. Raw QUIC libraries offer different knobs; some let you set a per-stream priority value, others send in round-robin, and some have no control at all, in which case you enforce order by not writing low-priority data until high-priority queues are empty. Extensible Priorities (RFC 9218) is an HTTP-level signal, a Priority header and a PRIORITY_UPDATE frame, used by HTTP/3; it is advice to a server about responses, not a QUIC transport feature. In browsers, the WebTransport API lets you set a send order on outgoing streams. See WebTransport architecture for the browser side.

Worked example: a 30 Hz match for 16 players

Each client sends inputs at 60 Hz as 40-byte datagrams, each carrying the last three inputs: about 2.4 KB/s up. The server sends positions for 16 entities at 30 Hz as one 600-byte datagram per tick, 18 KB/s down. Every second it sends a full snapshot of about 8 KB on a fresh unidirectional stream, and the control stream carries chat and match events at a few hundred bytes per second.

Now drop 2 percent of packets. Over TCP or a single stream, each loss stalls everything for about one round-trip; at 60 ms RTT, a 30 Hz feed loses roughly two frames per loss event and players see visible hitches. With this layout, a lost position datagram is replaced 33 ms later by the next one, a lost input is covered by redundancy, a lost snapshot packet delays only that snapshot, and if the next one is ready first the old stream is reset. Chat may wait one round-trip, which nobody notices.

Reconnects, 0-RTT and migration

Mobile real-time clients reconnect often. QUIC's connection migration lets a connection survive a change of client address, such as Wi-Fi to cellular, without a new handshake; connection migration explains the path-validation details. When the connection is truly gone, session resumption with 0-RTT lets the client send application data in its first flight.

0-RTT data can be replayed by an attacker who captures it, so only idempotent messages belong there: a resubscribe request or a fetch of the current snapshot, never a purchase, a chat message or a move that changes game state. Mark which message types are 0-RTT-safe in the protocol definition, and have the server reject anything else that arrives as early data.

Failure modes

  • One stream for everything. The most common mistake. Latency graphs look fine at zero loss and collapse at 1 percent. Test under loss with a network emulator before shipping.
  • Stream credit exhaustion. Per-object streams stop without errors; only "streams blocked" counters show it.
  • Oversized datagrams. Sends fail or are dropped silently on paths with a small MTU. Cap payloads in code and log rejections.
  • Assuming ordering across streams. A snapshot applied after a newer one rolls the world back. Version everything.
  • Bulk starving real-time data. An asset download on the same connection fills the congestion window. Schedule it last, rate-limit it, or move it to a separate connection.
  • UDP blocked. Some networks drop UDP entirely. Keep a WebSocket fallback and detect it fast; message ordering semantics for the fallback path are in message ordering.

Trade-offs

Stream-per-object gives the best freshness but costs per-stream overhead and credit management. Datagrams give the lowest latency but push sequencing, redundancy and loss tolerance into your code. Long-lived streams are the simplest and are correct for anything that must be complete and ordered. Media over QUIC, an IETF working-group draft, packages similar ideas for live media with tracks, groups and objects; watch it if you build media delivery, but design today against the RFCs your library implements.

What to do next

  1. List every message type in your app and answer the two questions for each: every instance needed, and ordered relative to what?
  2. Assign each type to a lane: long-lived stream, stream per object, stream per entity, or datagram.
  3. Version every update with a tick or sequence number and make receivers discard older ones.
  4. Reset superseded streams and give your protocol named application error codes.
  5. Size MAX_STREAMS and per-stream windows from your send rate and largest object, and alert on blocked counters.
  6. Decide the send order across lanes and implement it with your library's priority API or your own queues.
  7. Mark which messages are safe as 0-RTT data, and test the whole design under 1 to 5 percent loss with jitter.
Key takeaway: QUIC gives a real-time application independent streams, unreliable datagrams and the ability to abandon data that has gone stale, but only if the application protocol uses them on purpose. Put ordered, must-arrive messages on a long-lived stream. Give each superseding object its own stream and reset it when a newer one exists. Send small, high-rate state as sequenced datagrams. Budget stream credits, schedule real-time data ahead of bulk, keep 0-RTT for idempotent requests, and prove the design under packet loss before users find the hitches.