UDP, the User Datagram Protocol defined in RFC 768, adds almost nothing to IP: port numbers so the kernel can deliver a datagram to the right socket, a length, and a checksum. There is no connection, no acknowledgement, no retransmission, no ordering, and no congestion control. DNS, NTP, syslog, StatsD, DHCP, VoIP, games, WireGuard and QUIC all run on it, precisely because it gets out of the way.
Getting out of the way also means every failure is silent. This article explains what UDP actually guarantees, how a datagram travels through a Linux host, where it gets dropped and how to see it, how to push millions of datagrams per second, and what NATs and attackers do to UDP traffic. Whether to use UDP at all is covered in TCP vs UDP in depth; this page assumes you already are.
The header and the checksum
Each UDP header has four 16-bit fields. The length counts header plus payload, so the smallest datagram is 8 bytes and the largest is 65,535. Subtracting the 20-byte IPv4 header leaves a maximum payload of 65,507 bytes over IPv4; over IPv6, whose payload length field excludes its own 40-byte header, the limit is 65,527 bytes without jumbograms.
The checksum covers the UDP header, the payload and a pseudo-header made from the source and destination IP addresses, the protocol number and the UDP length. Including the addresses means a datagram delivered to the wrong host fails the check. Over IPv4 a checksum of zero means "not computed" and is allowed; RFC 8085 says applications SHOULD enable checksums anyway. Over IPv6 the checksum MUST be used, because IPv6 has no header checksum of its own; RFC 6935 relaxes this only for specific tunnel encapsulations. A 16-bit one's-complement sum is weak, so protocols that care about integrity, such as QUIC, add their own cryptographic authentication on top.
What the socket API gives you
UDP sockets are message-oriented. One sendto() produces one datagram and one recvfrom() returns at most one datagram, so boundaries are preserved, unlike a TCP byte stream. Several details catch people out:
- Truncation. If your receive buffer is smaller than the datagram, the excess is discarded. On Linux,
MSG_TRUNCin the returned flags tells you it happened; size buffers for the largest message your protocol allows. - Connected UDP. Calling
connect()on a UDP socket sends nothing. It fixes the default peer, filters out datagrams from other addresses, and lets ICMP errors surface, so a latersend()orrecv()can fail withECONNREFUSEDafter the peer replied with port unreachable. Unconnected sockets never see these errors. - Sends rarely fail. A successful
sendto()means the kernel queued the datagram, not that anyone received it. - Duplicates and reordering are legal. The network may deliver a datagram twice or out of order. Protocols number their messages if it matters.
Size: why big datagrams are a bad idea
A datagram larger than the path MTU is fragmented by IP, and losing any one fragment loses the whole datagram. Fragments also confuse firewalls and load balancers that look at ports, which only the first fragment carries, and IPv6 routers do not fragment in transit at all. RFC 8085 says applications SHOULD NOT send datagrams that exceed the path MTU, and without path MTU discovery they should fall back to the default effective send MTU: the smaller of 576 bytes and the first-hop MTU for IPv4, and 1,280 bytes for IPv6.
Real protocols follow this. QUIC requires every path to carry 1,200-byte UDP payloads and pads its first packets to that size. The DNS community's 2020 flag day settled on an EDNS buffer size of 1,232 bytes so responses fit in a 1,280-byte IPv6 packet, falling back to TCP for anything bigger. Pick a payload ceiling around 1,200 bytes unless you control the path. The mechanics of MTUs and fragmentation are in MTU and fragmentation in depth and path MTU discovery.
Worked example: a metrics collector that loses data
A StatsD-style collector receives metrics over UDP on port 8125. Dashboards show gaps during traffic peaks, but the network team reports no packet loss. That combination nearly always means the receiving host drops datagrams because the application drains its socket too slowly. Check the kernel's UDP counters first:
$ nstat -az | grep -E 'Udp(InDatagrams|NoPorts|InErrors|RcvbufErrors|InCsumErrors)'
UdpInDatagrams 918442031
UdpNoPorts 12
UdpInErrors 4410288
UdpRcvbufErrors 4410288
UdpInCsumErrors 0
$ ss -uamn 'sport = :8125'
UNCONN 0 0 0.0.0.0:8125 0.0.0.0:*
skmem:(r0,rb212992,t0,tb212992,f4096,w0,o0,bl0,d4410288)The counters in this example are illustrative, but the pattern is the one to look for. RcvbufErrors equals InErrors, so every error is a full socket buffer, and the d field in ss pins the drops on this socket. rb212992 shows a receive buffer of about 208 KiB, a common distribution default, which holds only a few milliseconds of a busy stream. NoPorts counts datagrams for ports with no listener and InCsumErrors counts corruption; neither is the problem here.
The fix has two parts. First, give the socket a bigger buffer. On Linux, SO_RCVBUF requests are capped by net.core.rmem_max, and the kernel doubles the requested value to account for bookkeeping, so getsockopt reports twice what you asked for:
# sysctl -w net.core.rmem_max=26214400 (persist it in /etc/sysctl.d/)
import socket
sock = socket.socket(socket.AF_INET, socket.SOCK_DGRAM)
sock.setsockopt(socket.SOL_SOCKET, socket.SO_RCVBUF, 8 * 1024 * 1024)
print("effective rcvbuf", sock.getsockopt(socket.SOL_SOCKET, socket.SO_RCVBUF))
sock.bind(("0.0.0.0", 8125))
while True:
data, addr = sock.recvfrom(65535)
handle(data) # keep this fast: parse and hand off, never blockA larger buffer absorbs bursts; it does not fix a reader that is slower than the average arrival rate. The second part is to make the read loop cheap and parallel: hand parsed data to a queue rather than doing work in the loop, and run several receiving sockets bound to the same port with SO_REUSEPORT, which makes the kernel spread datagrams across them by a hash of the source and destination. After the change, RcvbufErrors should stop increasing during peaks; watch the rate, not the absolute total, which only resets at reboot.
High throughput: fewer system calls, bigger batches
At high packet rates the cost is per datagram, not per byte: one system call, one route lookup and one trip through the stack each time. Linux offers three ways to amortise it:
sendmmsg()andrecvmmsg()move many datagrams per system call, each with its own address and length.- UDP generic segmentation offload, the
UDP_SEGMENTsocket option (Linux 4.18 and later), lets you hand the kernel one large buffer and a segment size; it is split into equal-size datagrams as late as possible, in hardware where the NIC supports it. - UDP generic receive offload,
UDP_GRO(Linux 5.0 and later), works in reverse: consecutive datagrams of one flow arrive as a single large buffer, and a control message reports the segment size so the application can split it.
import socket
SOL_UDP = 17 # IPPROTO_UDP
UDP_SEGMENT = 103 # from linux/udp.h
sock = socket.socket(socket.AF_INET, socket.SOCK_DGRAM)
sock.connect(("198.51.100.7", 4433))
sock.setsockopt(SOL_UDP, UDP_SEGMENT, 1200)
# 24 KiB in one call becomes 20 datagrams of 1,200 bytes each on the wire.
sock.send(b"x" * (1200 * 20))QUIC implementations rely on exactly these features, which is why their performance depends heavily on kernel version and NIC drivers; QUIC architecture in depth explains the protocol built on top. Measure before and after: batching helps most when the per-packet stack cost, not the application, is the bottleneck.
NATs, firewalls and keep-alives
Because UDP has no handshake or teardown, a NAT or stateful firewall can only guess when a flow has ended, and it does so with an idle timer. When the timer expires, the mapping disappears and replies from the far side are dropped or delivered to nobody. RFC 4787 requires NAT UDP mapping timers of at least two minutes and recommends five minutes or more, but many devices and default configurations use less; Linux connection tracking, for instance, defaults to 30 seconds for UDP flows that have not seen replies.
Long-lived UDP sessions therefore send keep-alives. RFC 8085 says applications SHOULD NOT send them more often than once every 15 seconds and should use longer intervals when possible, because every keep-alive costs battery on mobile devices and load on servers. Make the interval configurable, start around 25 seconds for consumer networks, and treat a silent peer as possibly behind a rebound mapping with a new address and port.
Security: spoofing and amplification
UDP has no handshake, so the source address of a datagram is just a claim. An attacker can send a small request with a forged source address to a server that replies with a large response, and the response floods the victim. DNS, NTP's monlist command and memcached have all been used for reflection attacks with amplification factors far above 10x, and in the memcached case tens of thousands.
If you design a UDP protocol, do not let it become an amplifier. Before an address is validated, responses should be no larger than requests, or bounded by a small multiple. QUIC limits a server to sending three times the bytes it has received until the client's address is validated, and requires clients to pad their first packets so the budget is meaningful. Rate-limit responses per source prefix, and never expose debugging or statistics commands on public UDP ports.
Testing under realistic loss
Because UDP never retries, an application tested only on a clean LAN has never seen the network it will run on. Linux's netem queueing discipline injects delay, jitter, loss, duplication and reordering on an interface, so you can watch how your protocol behaves before users do. Run it in a test namespace or VM, not on a shared host:
# 40 ms +/- 10 ms delay, 2% loss, 0.5% duplicates, some reordering
tc qdisc add dev eth0 root netem delay 40ms 10ms loss 2% duplicate 0.5% reorder 5%
# capture what actually crossed the wire, including ICMP errors
tcpdump -ni eth0 'udp port 8125 or icmp'
# remove the impairment
tc qdisc del dev eth0 rootCheck that message loss produces the behaviour you designed, such as a retry, a gap marker or a degraded stream, rather than a hang or a crash, and that duplicated or reordered datagrams are handled idempotently.
Failure modes at a glance
| Symptom | Cause | Where to look |
|---|---|---|
| Gaps under load, network clean | Socket buffer full | RcvbufErrors, ss -uam d field |
| Large messages vanish, small ones work | Fragmentation dropped on path | Payload size versus path MTU |
| Session dies after a quiet period | NAT or firewall timer expired | Keep-alive interval, conntrack timeouts |
| Replies never arrive, no error | Unconnected socket hides ICMP | connect() the socket, tcpdump for ICMP |
| CPU saturated at modest bandwidth | Per-datagram syscall cost | sendmmsg, UDP_SEGMENT, UDP_GRO |
| Traffic spikes to a third party | Your service used as a reflector | Response size, per-source rate limits |
What to do next
- Graph
UdpRcvbufErrorsandUdpInErrorsfrom/proc/net/snmpon every host that receives UDP, as rates. - Raise
net.core.rmem_maxand setSO_RCVBUFexplicitly; confirm the effective value withgetsockopt. - Keep payloads at or below about 1,200 bytes unless you control the path.
- Make receive loops do nothing but read and hand off; scale with
SO_REUSEPORT. - Batch with
sendmmsg/recvmmsgor segmentation offload when per-packet CPU is the limit. - Send keep-alives no more often than every 15 seconds for long-lived flows through NATs.
- Audit every public UDP endpoint for response-to-request size ratio and add per-source limits.
- Use connected sockets for client-side flows so ICMP errors reach the application.