The maximum transmission unit (MTU) is the largest IP packet a link can carry in one frame. On Ethernet it is 1,500 bytes by default, a number that dates from the 1980s and still shapes almost every packet on the internet. When a packet is larger than the MTU of a link it must cross, one of two things happens: it is split into fragments, or it is dropped and the sender is told to send smaller packets. Both paths have costs, and when the 'told to send smaller' signal is lost you get the classic MTU bug: small requests work, large responses hang, and nobody can explain why.
This article builds the subject from the header bytes up. It covers what the MTU does and does not count, how IPv4 fragmentation works field by field, why IPv6 moved fragmentation to the sender, how TCP avoids fragmentation with the maximum segment size, what tunnels and overlays do to the budget, and how to measure and fix a path. The discovery mechanisms themselves, classic path MTU discovery and its packetization-layer successor, are covered in the path MTU discovery article; here they appear only where they explain a failure.
What the MTU counts
The MTU is measured at the IP layer: it is the largest IP packet, headers included, that fits in one link-layer frame. It does not include the link-layer framing. An Ethernet frame carrying a full 1,500-byte packet is 1,518 bytes on the wire with its 14-byte header and 4-byte frame check sequence, or 1,522 with an 802.1Q VLAN tag, plus preamble and gap that never show up in a capture. When a vendor says a switch supports 'jumbo frames of 9,216 bytes', check whether that figure is a frame size or an MTU; the two differ by the header bytes, and mismatches between devices are a common source of trouble.
Inside the MTU, headers take their share. An IPv4 header without options is 20 bytes and a TCP header without options is 20 bytes, so a full-size TCP segment over IPv4 carries 1,460 bytes of data. IPv6's fixed header is 40 bytes, leaving 1,440. TCP options such as timestamps take another 12 bytes from each segment in practice, which is why you will often see 1,448 bytes of payload per packet in captures. UDP's header is 8 bytes, so the largest unfragmented UDP payload on a 1,500-byte IPv4 path is 1,472 bytes, the number that appears in every ping-based MTU test.
Common MTU values
| Link or path | Typical MTU | Why |
|---|---|---|
| Ethernet | 1500 | The default for almost every interface |
| PPPoE (many DSL and fibre connections) | 1492 | 8 bytes of PPPoE and PPP header |
| IPv6 minimum | 1280 | Every IPv6 link must carry at least this |
| IPv4 minimum | 68 | Every IPv4 host must accept 576-byte datagrams after reassembly |
| VXLAN over a 1500 underlay | 1450 | 50 bytes of outer IPv4, UDP, VXLAN and inner Ethernet |
| WireGuard default | 1420 | Sized for the worst case of 80 bytes over an IPv6 outer header |
| Jumbo frames in a data centre | 9000 or close to it | Fewer packets per byte, lower CPU per gigabit |
| Within some cloud VPCs | Often 8,500 to 9,001 | Provider-specific; usually smaller across gateways, VPNs and the internet |
The path MTU between two hosts is the smallest MTU of every link on the path, and it can change when routing changes. Treat any value you have not measured as a guess, especially across clouds and VPNs. Cloud defaults differ by provider and by network type, so read your provider's documentation rather than assuming 1,500 or 9,001.
How IPv4 fragmentation works
When an IPv4 router must forward a packet larger than the outgoing link's MTU, and the packet's Don't Fragment (DF) bit is clear, it splits the payload into pieces that fit and gives each its own IPv4 header. Three header fields make this work. The 16-bit Identification field is the same in every fragment of one original packet. The 13-bit Fragment Offset says where this fragment's data starts in the original payload, in units of 8 bytes, which is why every fragment except the last carries a multiple of 8 data bytes. The More Fragments (MF) flag is set on every fragment except the last.
Worked example: a 4,000-byte IPv4 packet (20-byte header, 3,980 bytes of payload) meets a 1,500-byte link. Each fragment can carry at most 1,480 bytes of payload, which is a multiple of 8. The router emits three fragments:
| Fragment | Total length | Payload bytes | Offset field | MF |
|---|---|---|---|---|
| 1 | 1500 | 1480 (bytes 0 to 1479) | 0 | 1 |
| 2 | 1500 | 1480 (bytes 1480 to 2959) | 185 | 1 |
| 3 | 1040 | 1020 (bytes 2960 to 3979) | 370 | 0 |
Only the destination reassembles. It buffers fragments keyed by source, destination, protocol and Identification until it has every byte from offset zero to the fragment with MF clear, then hands the whole packet up. If any fragment is lost, the others wait in memory until a timer expires (30 seconds by default on Linux, the ipfrag_time setting) and the entire packet is discarded. The transport above must retransmit the whole thing, not just the missing piece.
Why fragmentation is considered harmful
- Loss amplification. Losing one fragment loses the whole packet, so a 1% fragment loss rate on a three-fragment packet becomes roughly 3% packet loss.
- Middleboxes. Only the first fragment carries the TCP or UDP header. Firewalls, load balancers and NAT devices that need port numbers either reassemble (expensive) or drop non-first fragments. Many simply drop all fragments.
- Identification wraparound. At high packet rates the 16-bit Identification field can wrap within the reassembly timeout, so fragments of different packets get stitched together. Checksums catch most, but not all, of these mis-assemblies. RFC 4963 describes the problem.
- Resource exhaustion. Reassembly buffers are a classic denial-of-service target, and overlapping-fragment tricks have been used to evade inspection.
- CPU cost. Fragmenting and reassembling are slow paths in most routers and kernels.
The practical conclusion is that every modern protocol tries never to fragment: TCP sizes segments to fit, QUIC sets DF and requires paths to carry at least 1,200-byte UDP payloads, and DNS over UDP has moved toward smaller responses with fallback to TCP.
IPv6 moved fragmentation to the sender
IPv6 routers never fragment. If a packet is too big for the next link, the router drops it and sends an ICMPv6 Packet Too Big message (type 2) back to the source, carrying the MTU of the constraining link. Only the source host may fragment, using an 8-byte Fragment extension header with the same offset-and-more-fragments logic and a 32-bit identification. IPv6 also guarantees a minimum link MTU of 1,280 bytes, so a sender that never exceeds 1,280 never needs discovery at all. That makes ICMPv6 Packet Too Big far more important than its IPv4 equivalent: filter it and IPv6 large-packet traffic breaks outright. The IPv6 article covers the rest of the header changes.
In IPv4 the equivalent signal is ICMP type 3, code 4, 'fragmentation needed and DF set', sent when a router meets a too-large packet with DF set. Since most operating systems set DF on TCP traffic to enable path MTU discovery, IPv4 in practice behaves much like IPv6: packets are dropped and the sender relies on the ICMP message to learn the right size.
MSS: how TCP stays under the MTU
TCP avoids fragmentation by advertising a maximum segment size in its SYN: the largest payload it is willing to receive, normally the interface MTU minus 40 bytes for IPv4 or 60 for IPv6. Each side sends segments no larger than the smaller of the peer's MSS and its own path MTU estimate. That handles the endpoints' own links, but not a narrower link in the middle of the path, which is where path MTU discovery and ICMP come in.
Routers and VPN gateways that know the path is narrower can clamp the MSS: rewrite the MSS option in SYN packets passing through so both ends start with a segment size that fits. It is a hack, it only helps TCP, and it is extremely effective, which is why almost every PPPoE router and tunnel endpoint does it. On Linux:
# Clamp the MSS of forwarded TCP SYNs to the outgoing route's path MTU.
iptables -t mangle -A FORWARD -p tcp --tcp-flags SYN,RST SYN \
-j TCPMSS --clamp-mss-to-pmtu
# nftables equivalent
nft add rule inet filter forward tcp flags syn tcp option maxseg size set rt mtuMSS clamping is the reason many MTU problems only show up for UDP-based protocols: QUIC, DNS with large responses, and VPNs carried over UDP do not benefit from it. The TCP/IP article covers the rest of TCP's segmentation behaviour.
Tunnels and overlays spend the budget
Every encapsulation adds headers inside the same physical MTU. VXLAN over IPv4 adds 50 bytes: 20 for the outer IPv4 header, 8 for UDP, 8 for the VXLAN header and 14 for the inner Ethernet header that the overlay carries. With a 1,500-byte underlay, the inner interfaces must use 1,450, or the underlay must be raised to at least 1,550; the VXLAN article covers the rest of the design. GRE adds 24 bytes, Geneve at least 50 plus options, and IPsec a variable amount depending on cipher and mode.
WireGuard adds 60 bytes over an IPv4 outer header (20 IP, 8 UDP, 32 WireGuard) and 80 bytes over IPv6, which is why its default interface MTU is 1,420: it fits the worst case on a 1,500-byte path. Over PPPoE you need 1,412 or less. See the WireGuard article for configuration. Kubernetes network plugins face the same arithmetic: each pod interface must be set to the node MTU minus the overlay overhead, and nesting overlays (a VPN inside a cloud overlay inside VXLAN) subtracts each layer in turn.
Raising the underlay MTU with jumbo frames is the cleaner fix when you control every switch and NIC on the path. Inside a data centre, 9,000-byte frames also cut per-packet CPU and interrupt load for bulk transfers such as storage replication and distributed training traffic, where fewer, larger packets mean fewer headers to process per gigabyte. The risk is a single device left at 1,500: then large packets are silently dropped for exactly the bulk flows that motivated the change.
Measuring the path
Send packets with DF set and shrink them until they pass. On Linux, ping -M do -s 1472 host sends a 1,500-byte IPv4 packet (1,472 bytes of data plus 8 of ICMP and 20 of IP) with DF set. On Windows the equivalent is ping -f -l 1472 host, and on macOS ping -D -s 1472 host. tracepath host walks the path and reports where the MTU drops, provided the routers send ICMP.
# Binary-search the largest IPv4 packet that crosses the path with DF set.
host=$1; lo=1200; hi=9000
while [ $((hi - lo)) -gt 1 ]; do
mid=$(( (lo + hi) / 2 ))
if ping -M do -c 1 -W 1 -s $((mid - 28)) "$host" >/dev/null 2>&1; then
lo=$mid
else
hi=$mid
fi
done
echo "path MTU to $host is about $lo bytes"A ping that times out without an error message at large sizes, while small sizes succeed, means a hop is dropping large packets without sending ICMP: a black hole. Check interface MTUs with ip link show, compare both ends of every tunnel, and look at the TCP connection's view with ss -ti, which shows the current path MTU and MSS.
Failure modes and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| SSH connects, then hangs on large output | PMTUD black hole: ICMP filtered on the path | Allow ICMP type 3 code 4 and ICMPv6 type 2; clamp MSS |
| HTTPS handshakes stall with large certificates | Same black hole, hit by the server's certificate flight | Same; also enable tcp_mtu_probing |
| Pod-to-pod works, cross-node large transfers stall | Overlay MTU not reduced by its overhead | Set CNI MTU to underlay minus overhead |
| VPN works for web, fails for QUIC or large DNS | MSS clamping hides the problem for TCP only | Lower the tunnel MTU itself |
| Throughput collapses after enabling jumbo frames | One device on the path still at 1500 | Audit every hop; test with DF-set pings at 9000 |
| Random corruption at very high UDP rates | Fragment Identification wraparound | Stop fragmenting; size datagrams under the path MTU |
Linux's net.ipv4.tcp_mtu_probing sysctl is a useful safety net for black holes: 0 disables probing, 1 enables it only after a black hole is detected, and 2 always probes. It implements packetization-layer path MTU discovery (RFC 4821), which finds the size by observing which segments are acknowledged rather than trusting ICMP. Its datagram counterpart, RFC 8899, is what QUIC implementations use. Mode 1 is a sensible default for servers that face the internet.
What to do next
- Measure the path MTU between your main pairs of hosts with DF-set pings or the script above, and write the results down.
- Check that firewalls and security groups allow ICMP type 3 code 4 and ICMPv6 Packet Too Big, in both directions.
- For every tunnel and overlay, subtract its overhead from the underlay MTU and confirm the interface MTUs match on both ends.
- Enable MSS clamping on gateways and set net.ipv4.tcp_mtu_probing to 1 on internet-facing Linux servers.
- If you adopt jumbo frames, audit every switch, NIC and virtual interface on the path, then test with 9,000-byte DF-set pings.
- Keep UDP application payloads under 1,200 to 1,400 bytes unless you control the path, and never rely on fragmentation.