The textbook answer, TCP is reliable and UDP is fast, is wrong enough to cause bad designs. UDP is not faster at moving bytes; on the same path, both are limited by the same bandwidth and round-trip time. What differs is what happens when a packet is lost and who decides what to do about it. TCP decides for you: everything after the gap waits until the gap is filled. UDP decides nothing, so your application can choose to skip, resend, or replace the lost data, and must also take on all the jobs TCP was doing silently.
This page is about making that choice deliberately. For the kernel-level detail of TCP itself, read TCP architecture in depth; for the protocol that most new UDP-based designs should consider first, read QUIC architecture in depth.
What TCP gives you, and what it costs
TCP turns an unreliable packet network into a reliable, ordered byte stream between two endpoints. It sets up state with a three-way handshake, numbers every byte, retransmits what the receiver does not acknowledge, delivers bytes to the application only in order, slows the sender down when the receiver's buffer fills (flow control), and slows it down when the network drops or marks packets (congestion control). Every one of those features is something you would otherwise have to write.
The costs are specific. Setup takes one round trip before data, plus one more for TLS 1.3, so a fresh HTTPS connection on a 100 ms path spends 200 ms before the first request byte. The stream has no message boundaries, so the application must frame its own messages. And in-order delivery means head-of-line blocking: if one segment is lost, everything behind it sits in the kernel buffer, already received, until the retransmission arrives. Two small-packet interactions also matter: Nagle's algorithm holds small writes while data is unacknowledged, and the peer's delayed acknowledgement can hold the acknowledgement, so request-response protocols that write in two pieces can stall for tens of milliseconds. Setting TCP_NODELAY removes Nagle's delay; writing each message in one call avoids the pattern altogether.
What UDP gives you
UDP adds two things to IP: port numbers, so many applications can share an address, and a checksum. A datagram is delivered whole or not at all, it may arrive out of order or twice, and there is no connection, no retransmission, no flow control and no congestion control. Message boundaries are preserved, which is genuinely useful: one sendto is one recvfrom on the other side.
The important consequence is control. A lost UDP datagram delays nothing else, so the application can decide that a stale position update is worthless and drop it, while still retransmitting a chat message. That per-message decision is the only real reason to choose UDP over TCP, apart from multicast and broadcast, which TCP cannot do at all.
Head-of-line blocking in numbers
Consider a game server sending 60 state updates per second to a player over a 50 ms round-trip path with 1% packet loss. On TCP, a lost segment is usually repaired by fast retransmit after three duplicate acknowledgements, which costs roughly one round trip: about 50 ms, or three updates that arrive late and then all at once. If the loss hits the last packet in a burst there are no duplicate acknowledgements to trigger fast retransmit, so recovery waits for the retransmission timeout, whose minimum on Linux is 200 ms; that is twelve frames of frozen game. At 60 packets per second and 1% loss, a loss happens about every 1.7 seconds, so these stalls are a constant feature, not an edge case.
On UDP with a latest-state design, a lost update is simply superseded by the next one 16.7 ms later. The player sees one frame of slightly older interpolation instead of a freeze. The same arithmetic explains why voice and video calls use UDP: a 20 ms audio frame that arrives 200 ms late is worse than no frame at all, because the jitter buffer has already played past it.
What you must rebuild on UDP
Choosing UDP means accepting a list of responsibilities. Skipping any of them produces a protocol that works in the lab and fails in the field. RFC 8085 is the IETF's guidance for exactly this list and is worth reading in full.
- Datagram size. Keep each datagram under the path MTU to avoid IP fragmentation, which is frequently dropped by firewalls. IPv6 guarantees 1280 bytes; QUIC builds on a 1200-byte minimum. Staying near 1200 bytes of payload is a safe default; path MTU discovery explains how to probe for more.
- Sequencing. Put a sequence number in every datagram so the receiver can detect loss, reordering and duplicates, and discard anything older than what it has already applied.
- Selective reliability. Retransmit only what must arrive (chat, purchases, match results), using acknowledgements and timeouts, and never retransmit data that has been superseded.
- Congestion control. A sender with no back-off can flood a shared link and hurt every other flow on it. At minimum, pace sends and reduce the rate when loss rises; for bulk data, use a real algorithm or use QUIC.
- NAT keepalive. NATs and firewalls drop idle UDP mappings. RFC 4787 asks for at least two minutes, but many devices use far less, so send a small keepalive every 15 to 25 seconds while a session is idle. NAT traversal covers hole punching for peer-to-peer.
- Anti-spoofing and amplification. UDP source addresses are trivially forged. Never answer an unvalidated client with more bytes than it sent; require a cookie or token round trip first, as DTLS and QUIC do.
- Encryption and authentication. Use DTLS or QUIC rather than designing your own cryptography.
- A fallback. Some corporate and hotel networks block UDP except DNS. Any consumer product needs a TCP or WebSocket fallback on port 443.
Code: framing on TCP, latest-state on UDP
On TCP the most common bug is assuming one recv returns one message. The stream can split or merge writes arbitrarily, so prefix each message with its length and read exactly that many bytes.
import socket, struct
def send_msg(sock: socket.socket, payload: bytes) -> None:
# One sendall per message: header and body together, so Nagle and delayed ACK cannot split them.
sock.sendall(struct.pack("!I", len(payload)) + payload)
def recv_exact(sock: socket.socket, n: int) -> bytes:
buf = bytearray()
while len(buf) < n:
chunk = sock.recv(n - len(buf))
if not chunk:
raise ConnectionError("peer closed mid-message")
buf += chunk
return bytes(buf)
def recv_msg(sock: socket.socket, max_len: int = 1 << 20) -> bytes:
(length,) = struct.unpack("!I", recv_exact(sock, 4))
if length > max_len:
raise ValueError("frame too large") # never trust a length field from the network
return recv_exact(sock, length)On UDP, a latest-state receiver keeps the highest sequence number it has applied per sender and ignores anything older. Sequence numbers wrap, so compare them with serial-number arithmetic rather than a plain greater-than.
import socket, struct
HEADER = struct.Struct("!HI") # 16-bit sequence number, 32-bit session id
def newer(a: int, b: int) -> bool:
# True if sequence a is after b, modulo 2**16 (RFC 1982 style).
return a != b and ((a - b) & 0xFFFF) < 0x8000
sock = socket.socket(socket.AF_INET, socket.SOCK_DGRAM)
sock.setsockopt(socket.SOL_SOCKET, socket.SO_RCVBUF, 4 << 20)
sock.bind(("0.0.0.0", 7777))
last_seq = {} # (addr, session) -> last applied sequence
while True:
data, addr = sock.recvfrom(1500)
if len(data) < HEADER.size:
continue
seq, session = HEADER.unpack_from(data)
key = (addr, session)
if key in last_seq and not newer(seq, last_seq[key]):
continue # stale or duplicate: drop, never wait for it
last_seq[key] = seq
apply_state(key, data[HEADER.size:])
Decision matrix
| Workload | Pick | Why |
|---|---|---|
| APIs, databases, file transfer | TCP (or HTTP/2 over TCP) | Every byte matters and in-order is what you want |
| Web traffic on lossy mobile links | QUIC (HTTP/3) | Independent streams avoid cross-request blocking; faster setup |
| Real-time game state | UDP with your own protocol, or a game networking library | Latest-state data should be dropped, not waited for |
| Voice and video calls | UDP (RTP/SRTP via WebRTC) | Late media is useless; jitter buffers want loss, not delay |
| Metrics and logs | TCP, unless loss is acceptable | UDP-based metrics drop silently under load, exactly when you need them |
| DNS | UDP first, TCP on truncation | Small queries fit in one datagram; large answers fall back |
| Service discovery on a LAN | UDP multicast | TCP cannot multicast |
Worked example: a game with voice chat
A team ships a multiplayer game with in-game voice. They first send everything over one TCP connection per player, which is quick to build. In testing on mobile networks with 1 to 2% loss, players report rubber-banding and voice that speeds up and slows down. The cause is the head-of-line arithmetic above: every loss freezes both state and voice together.
The redesign uses three channels. Game state goes over UDP at 30 updates per second, each datagram carrying a full snapshot of nearby entities (under 1,100 bytes) with a sequence number, and the client interpolates between the two most recent snapshots, so a lost packet costs nothing visible. Inputs and events that must arrive, such as purchases and kill confirmations, go over the same UDP socket with acknowledgements and retransmission, but they are kept in their own reliable queue so they never delay snapshots. Voice uses WebRTC, which already handles jitter buffering, loss concealment and congestion control. The server sends a 20-byte keepalive every 15 seconds on idle sessions, rejects datagrams without a valid session token before replying, and the client falls back to a WebSocket over TCP 443 when no UDP reply arrives within two seconds. After the change, measured stalls over 100 ms fell from several per minute to almost none at the same loss rate.
Failure modes and operations
- Silent UDP drops at the receiver. If the application reads slower than packets arrive, the socket buffer overflows and the kernel discards datagrams. On Linux, watch
netstat -suornstatfor receive buffer errors, and raisenet.core.rmem_maxandSO_RCVBUFtogether. - Fragmentation black holes. Datagrams over the path MTU are fragmented, and lost fragments lose the whole datagram. Symptoms are that small messages work and large ones vanish.
- TCP stalls from Nagle and delayed ACK. Look for latency clustered near 40 ms on Linux peers; fix with single writes or
TCP_NODELAY. - Half-open TCP connections. A peer that vanishes without a FIN leaves a connection that looks alive. Use application heartbeats or TCP keepalive with short intervals.
- Ephemeral port and TIME_WAIT exhaustion. Clients that open a new TCP connection per request run out of ports under load; pool connections instead.
- UDP blocked. Track the fallback rate per network; a sudden rise usually means a carrier or enterprise firewall change, not a bug.
For TCP connections, ss -ti shows round-trip time, congestion window and retransmissions per socket, which is the fastest way to tell a slow network from a slow application.
What to do next
- For each message type, write down whether a late copy is still useful. If every message must arrive in order, use TCP or QUIC and stop.
- If you need independent streams or faster setup but still want reliability, use a QUIC library rather than building on raw UDP.
- If you build on UDP, implement the full list: size limit, sequence numbers, selective reliability, congestion back-off, keepalive, anti-amplification, DTLS or QUIC crypto and a TCP fallback.
- On TCP, frame messages with a length prefix, cap frame sizes, and write each message in a single call.
- Test under 1 to 3% loss and 100 ms of added latency with Linux
tc netembefore you ship. - Add the metrics: UDP receive buffer errors, TCP retransmissions, fallback rate, and stall counts above 100 ms.