A transfer between two well-provisioned servers on different continents crawls at a few megabits per second while every link on the path is nearly idle. The CPU is idle, the disks are idle, and no packets are lost. One frequent culprit is a single option byte in the first packet of the connection: the TCP window scale option, either missing or negotiated too small.
This page explains the option from first principles: why the original TCP header cannot express a large window, how the scale factor is negotiated and applied, how much window a given path actually needs, how Linux chooses the factor and how applications accidentally freeze it, and how to prove from a packet capture whether scaling is your problem. It assumes you know what a receive window is; for the wider Linux stack and buffer autotuning see TCP architecture, in depth.
Why a 16-bit window caps throughput
TCP flow control lets a receiver say how many more bytes it is willing to accept beyond what it has acknowledged. The sender may have at most that many unacknowledged bytes in flight. The field carrying this number in the TCP header is 16 bits, so the largest window it can express is 65,535 bytes.
A sender limited by the window can send one window of data per round trip and must then wait for acknowledgements. Throughput is therefore bounded by window divided by round-trip time:
max_throughput = window / RTT
65,535 bytes / 0.001 s (1 ms, same data centre) = 65.5 MB/s ~ 524 Mbit/s
65,535 bytes / 0.080 s (80 ms, transatlantic) = 819 KB/s ~ 6.6 Mbit/s
65,535 bytes / 0.250 s (250 ms, satellite) = 262 KB/s ~ 2.1 Mbit/sThe amount of data that must be in flight to fill a path is the bandwidth-delay product (BDP): link rate times round-trip time. A 1 Gbit/s path with 80 ms RTT has a BDP of 125 MB/s x 0.08 s = 10 MB. A 64 KB window fills well under one percent of that pipe. No amount of bandwidth helps, because the sender is idle waiting for acknowledgements most of the time.
The option: a shift count agreed in the handshake
Window scaling, defined originally in RFC 1323 and now in RFC 7323, adds a TCP option of kind 3 carrying one byte: a shift count. A side that sends a window value W with a negotiated shift S means W x 2S bytes. The rules are strict, and each one has caused real outages:
- Only in the handshake. The option may appear only in SYN and SYN-ACK segments. A server may include it in its SYN-ACK only if the client's SYN carried it. Scaling is enabled only if both sides sent it; if either omits it, neither direction scales and both windows top out at 65,535 bytes.
- Per direction. Each side picks its own shift for the windows it advertises. A client announcing 7 and a server announcing 9 is normal; the server multiplies the client's window values by 128, the client multiplies the server's by 512.
- Not applied to the handshake itself. The window field in the SYN and SYN-ACK is never scaled. That is why captures show a SYN with
win 64240and the next ACK with a tiny-lookingwin 502. - Capped at 14. The largest legal shift is 14, giving a maximum window of 65,535 x 16,384, just under 1 GiB. A receiver that sees a larger value uses 14. The cap keeps the window below half the 32-bit sequence space so old and new segments cannot be confused.
- Fixed for the connection. The shift cannot be renegotiated. Whatever was decided when the SYN went out bounds the window for the connection's whole life.
Scaling costs precision. With a shift of 9, the window is expressed in units of 512 bytes, so the receiver rounds what it advertises. RFC 7323 pairs scaling with the timestamps option because at high rates sequence numbers wrap quickly, and timestamps let the receiver reject stale duplicates (PAWS). Most stacks negotiate both together.
How Linux chooses the shift
On Linux, scaling is controlled by net.ipv4.tcp_window_scaling, which defaults to 1 and should stay that way. The shift a socket advertises is computed when its SYN or SYN-ACK is built: the kernel takes the largest receive buffer the socket could grow to and picks the smallest shift that lets that buffer be expressed in 16 bits. For an ordinary socket that maximum comes from the third value of net.ipv4.tcp_rmem (and net.core.rmem_max). A maximum of 6 MiB, common on current distributions, divided by 65,535 is about 96, which needs a shift of 7, so wscale:7 is what you most often see.
Applications change this in a way that surprises people. Calling setsockopt(SO_RCVBUF) locks the buffer at that size and switches off receive autotuning for the socket, and if done before the handshake it also sizes the shift for that locked buffer. Done after connect(), the shift is already decided. On a server, accepted sockets inherit their settings from the listening socket, and the SYN-ACK is sent by the kernel before your code calls accept(), so the buffer must be set on the listener before listen().
import math, socket
def bdp_bytes(rate_bit_s: float, rtt_s: float) -> int:
return int(rate_bit_s / 8 * rtt_s)
def shift_needed(window_bytes: int) -> int:
s = 0
while (65535 << s) < window_bytes and s < 14:
s += 1
return s
need = bdp_bytes(1e9, 0.080) # 10,000,000 bytes
print(need, shift_needed(need)) # 10000000 8
# Usually best: do NOT set SO_RCVBUF at all and let autotuning grow the buffer
# up to tcp_rmem[2]. If you must pin it, do it BEFORE connect() or listen().
s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
s.setsockopt(socket.SOL_SOCKET, socket.SO_RCVBUF, 2 * need) # Linux caps at rmem_max, then doubles
s.connect(("replica.example.net", 5432))Note the trap in the last lines: an explicit SO_RCVBUF is capped by net.core.rmem_max, which on many systems is only a few hundred kilobytes, while autotuning is capped by the larger tcp_rmem maximum. Code that sets a "large" buffer explicitly can end up slower than code that sets nothing.
Worked example: a slow cross-region copy
A nightly backup copies 2 TB from Frankfurt to Virginia over a dedicated 1 Gbit/s path. Ping shows 85 ms. The job takes about 18 hours, which works out to roughly 250 Mbit/s. Here is the diagnosis in order.
- Compute the target. BDP = 125 MB/s x 0.085 s = 10.6 MB. The receiver must be able to advertise about 10.6 MB, and the sender must be able to hold that much unacknowledged data.
- Check negotiation. On the receiver,
ss -tin dst <sender-ip>showswscale:7,7. Scaling is on, so this is not the 64 KB problem. With shift 7 the field can express up to 8 MiB, which is already short of 10.6 MB. - Check the buffer ceiling.
sysctl net.ipv4.tcp_rmemshows a maximum of 6,291,456. The advertised window is only part of the buffer, because the kernel reserves some for bookkeeping, so the effective window peaks below 6 MiB. 6 MiB / 0.085 s is about 590 Mbit/s as a hard ceiling, and autotuning ramps up gradually, so a real average near 250 Mbit/s is plausible. - Fix both ends. Raise the receiver's
tcp_rmemmaximum to 32 MiB and the sender'stcp_wmemmaximum likewise. New connections now negotiate a larger shift because the possible buffer is larger. Old connections keep their old shift. - Verify. Re-run the copy and watch
ss -tin: the shift is now 10 on the receiver's side,rcv_spacegrows over the first seconds, and throughput approaches the line rate if there is no loss.
# /etc/sysctl.d/90-long-fat-network.conf (values to test, not universal defaults)
net.ipv4.tcp_window_scaling = 1
net.ipv4.tcp_rmem = 4096 131072 33554432
net.ipv4.tcp_wmem = 4096 16384 33554432
net.core.rmem_max = 33554432
net.core.wmem_max = 33554432If throughput still falls short once the window is large enough, the limit has moved to congestion control: loss on the path shrinks the congestion window no matter what the receiver advertises. That is where a model-based algorithm such as BBR can help, and where fixing loss beats tuning buffers.
Reading the evidence in captures
Capture the handshake; without it no tool can interpret the window field.
$ tcpdump -ni eth0 'tcp[tcpflags] & tcp-syn != 0 and host 203.0.113.7'
IP 10.0.0.5.41022 > 203.0.113.7.443: Flags [S], seq 1, win 64240,
options [mss 1460,sackOK,TS val 1 ecr 0,nop,wscale 7], length 0
IP 203.0.113.7.443 > 10.0.0.5.41022: Flags [S.], seq 9, ack 2, win 65160,
options [mss 1460,sackOK,TS val 7 ecr 1,nop,wscale 9], length 0Both sides announced a shift, so scaling is on. If wscale appears in the SYN but not in the SYN-ACK, the server or something in front of it disabled scaling and the connection is capped at 64 KB.
Wireshark's TCP details show a "Window size scaling factor" for each segment. A value of -1 means Wireshark did not see the handshake and is showing the raw field, which is the most common way people misread a perfectly healthy connection as having a 500-byte window. A value of -2 means the handshake was seen and scaling was not negotiated. The "Calculated window size" field is the one to trust. The TCP stream graphs (window scaling and throughput) show whether the sender's in-flight data hugs the receiver's window line, which is the signature of a window-limited transfer.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Throughput stuck near 64 KB per RTT | One side, or a middlebox, removed the option from the SYN or SYN-ACK | Capture both ends of the handshake to find where it disappears; fix or bypass the device |
| Fine in one region, slow across oceans | Buffer maximum below the long path's BDP | Raise tcp_rmem and tcp_wmem maximums on both ends |
| Explicit large buffer made things slower | SO_RCVBUF capped by rmem_max and autotuning disabled | Remove the setsockopt, or raise rmem_max and set before connect |
| Long-lived connections stay slow after a sysctl change | Shift was fixed at their handshake | Reconnect; pools must be recycled to pick up new limits |
| Connections stall after a firewall upgrade | Device rewrites or normalises windows without tracking the shift | Upgrade or reconfigure the device to track scaling; do not disable scaling |
| Large memory use on busy servers | Tens of thousands of sockets each allowed huge buffers | Size maximums for the paths you serve; autotuning only grows buffers that need it |
On Windows the equivalent is receive window autotuning, shown by netsh interface tcp show global. Setting the autotuning level to disabled, a common folk remedy, caps the receive window at 64 KB and reproduces the original problem.
Trade-offs
Large maximums cost nothing for connections that never need them, because Linux grows buffers on demand, but they raise the worst case: a server with many slow readers can pin a lot of kernel memory. Buffers far larger than the BDP add queueing on the sender side without adding throughput. And window scaling only removes the flow-control ceiling; it does nothing about loss, small congestion windows after idle periods, or an application that writes in tiny chunks. For short request-response traffic the window rarely matters at all, which is why choosing between TCP and UDP usually turns on latency and head-of-line blocking rather than throughput. For a refresher on where the window sits in the header, see TCP/IP fundamentals.
What to do next
- Compute the BDP for your longest important path from link rate and measured RTT, and write it next to your sysctl settings.
- Run
ss -tinon a live long-distance connection and confirm bothwscalevalues are present and large enough for that BDP. - Check
tcp_rmem,tcp_wmemandrmem_maxon both ends, and raise the maximums only where long paths need them. - Search your code for
SO_RCVBUFandSO_SNDBUF; remove them unless a measurement justifies them, and set them before connect or listen if kept. - Capture a handshake through every firewall and load balancer on the path to prove the option survives.
- After any buffer change, recycle connection pools and re-measure throughput rather than assuming the change applied.