A transfer between two well-provisioned servers on different continents crawls at a few megabits per second while every link on the path is nearly idle. The CPU is idle, the disks are idle, and no packets are lost. One frequent culprit is a single option byte in the first packet of the connection: the TCP window scale option, either missing or negotiated too small.

This page explains the option from first principles: why the original TCP header cannot express a large window, how the scale factor is negotiated and applied, how much window a given path actually needs, how Linux chooses the factor and how applications accidentally freeze it, and how to prove from a packet capture whether scaling is your problem. It assumes you know what a receive window is; for the wider Linux stack and buffer autotuning see TCP architecture, in depth.

Advertisement

Why a 16-bit window caps throughput

TCP flow control lets a receiver say how many more bytes it is willing to accept beyond what it has acknowledged. The sender may have at most that many unacknowledged bytes in flight. The field carrying this number in the TCP header is 16 bits, so the largest window it can express is 65,535 bytes.

A sender limited by the window can send one window of data per round trip and must then wait for acknowledgements. Throughput is therefore bounded by window divided by round-trip time:

max_throughput = window / RTT

65,535 bytes / 0.001 s  (1 ms, same data centre)  = 65.5 MB/s  ~ 524 Mbit/s
65,535 bytes / 0.080 s  (80 ms, transatlantic)    = 819 KB/s   ~ 6.6 Mbit/s
65,535 bytes / 0.250 s  (250 ms, satellite)       = 262 KB/s   ~ 2.1 Mbit/s

The amount of data that must be in flight to fill a path is the bandwidth-delay product (BDP): link rate times round-trip time. A 1 Gbit/s path with 80 ms RTT has a BDP of 125 MB/s x 0.08 s = 10 MB. A 64 KB window fills well under one percent of that pipe. No amount of bandwidth helps, because the sender is idle waiting for acknowledgements most of the time.

The option: a shift count agreed in the handshake

Clientrcv buffer up to 6 MiBServerrcv buffer up to 16 MiBSYN win=64240 options: mss, sackOK, TS, wscale 7SYN-ACK win=65160 options: mss, sackOK, TS, wscale 9ACK win=502 (means 502 x 2^7 = 64,256 bytes)data, ACK win=4096 (means 4096 x 2^9 = 2 MiB)Window values in the SYN and SYN-ACK themselves are never scaled.Each side's shift applies to the windows THAT side advertises; both must send the option or neither direction scales.The shift is fixed for the life of the connection: buffers can grow later, the multiplier cannot.
Window scale negotiation. Each side announces its own shift in its SYN or SYN-ACK. After the handshake, every advertised 16-bit window is multiplied by 2 to the power of the sender's shift.

Window scaling, defined originally in RFC 1323 and now in RFC 7323, adds a TCP option of kind 3 carrying one byte: a shift count. A side that sends a window value W with a negotiated shift S means W x 2S bytes. The rules are strict, and each one has caused real outages:

  • Only in the handshake. The option may appear only in SYN and SYN-ACK segments. A server may include it in its SYN-ACK only if the client's SYN carried it. Scaling is enabled only if both sides sent it; if either omits it, neither direction scales and both windows top out at 65,535 bytes.
  • Per direction. Each side picks its own shift for the windows it advertises. A client announcing 7 and a server announcing 9 is normal; the server multiplies the client's window values by 128, the client multiplies the server's by 512.
  • Not applied to the handshake itself. The window field in the SYN and SYN-ACK is never scaled. That is why captures show a SYN with win 64240 and the next ACK with a tiny-looking win 502.
  • Capped at 14. The largest legal shift is 14, giving a maximum window of 65,535 x 16,384, just under 1 GiB. A receiver that sees a larger value uses 14. The cap keeps the window below half the 32-bit sequence space so old and new segments cannot be confused.
  • Fixed for the connection. The shift cannot be renegotiated. Whatever was decided when the SYN went out bounds the window for the connection's whole life.

Scaling costs precision. With a shift of 9, the window is expressed in units of 512 bytes, so the receiver rounds what it advertises. RFC 7323 pairs scaling with the timestamps option because at high rates sequence numbers wrap quickly, and timestamps let the receiver reject stale duplicates (PAWS). Most stacks negotiate both together.

Advertisement

How Linux chooses the shift

On Linux, scaling is controlled by net.ipv4.tcp_window_scaling, which defaults to 1 and should stay that way. The shift a socket advertises is computed when its SYN or SYN-ACK is built: the kernel takes the largest receive buffer the socket could grow to and picks the smallest shift that lets that buffer be expressed in 16 bits. For an ordinary socket that maximum comes from the third value of net.ipv4.tcp_rmem (and net.core.rmem_max). A maximum of 6 MiB, common on current distributions, divided by 65,535 is about 96, which needs a shift of 7, so wscale:7 is what you most often see.

Applications change this in a way that surprises people. Calling setsockopt(SO_RCVBUF) locks the buffer at that size and switches off receive autotuning for the socket, and if done before the handshake it also sizes the shift for that locked buffer. Done after connect(), the shift is already decided. On a server, accepted sockets inherit their settings from the listening socket, and the SYN-ACK is sent by the kernel before your code calls accept(), so the buffer must be set on the listener before listen().

import math, socket

def bdp_bytes(rate_bit_s: float, rtt_s: float) -> int:
    return int(rate_bit_s / 8 * rtt_s)

def shift_needed(window_bytes: int) -> int:
    s = 0
    while (65535 << s) < window_bytes and s < 14:
        s += 1
    return s

need = bdp_bytes(1e9, 0.080)             # 10,000,000 bytes
print(need, shift_needed(need))           # 10000000 8

# Usually best: do NOT set SO_RCVBUF at all and let autotuning grow the buffer
# up to tcp_rmem[2]. If you must pin it, do it BEFORE connect() or listen().
s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
s.setsockopt(socket.SOL_SOCKET, socket.SO_RCVBUF, 2 * need)   # Linux caps at rmem_max, then doubles
s.connect(("replica.example.net", 5432))

Note the trap in the last lines: an explicit SO_RCVBUF is capped by net.core.rmem_max, which on many systems is only a few hundred kilobytes, while autotuning is capped by the larger tcp_rmem maximum. Code that sets a "large" buffer explicitly can end up slower than code that sets nothing.

Worked example: a slow cross-region copy

A nightly backup copies 2 TB from Frankfurt to Virginia over a dedicated 1 Gbit/s path. Ping shows 85 ms. The job takes about 18 hours, which works out to roughly 250 Mbit/s. Here is the diagnosis in order.

  1. Compute the target. BDP = 125 MB/s x 0.085 s = 10.6 MB. The receiver must be able to advertise about 10.6 MB, and the sender must be able to hold that much unacknowledged data.
  2. Check negotiation. On the receiver, ss -tin dst <sender-ip> shows wscale:7,7. Scaling is on, so this is not the 64 KB problem. With shift 7 the field can express up to 8 MiB, which is already short of 10.6 MB.
  3. Check the buffer ceiling. sysctl net.ipv4.tcp_rmem shows a maximum of 6,291,456. The advertised window is only part of the buffer, because the kernel reserves some for bookkeeping, so the effective window peaks below 6 MiB. 6 MiB / 0.085 s is about 590 Mbit/s as a hard ceiling, and autotuning ramps up gradually, so a real average near 250 Mbit/s is plausible.
  4. Fix both ends. Raise the receiver's tcp_rmem maximum to 32 MiB and the sender's tcp_wmem maximum likewise. New connections now negotiate a larger shift because the possible buffer is larger. Old connections keep their old shift.
  5. Verify. Re-run the copy and watch ss -tin: the shift is now 10 on the receiver's side, rcv_space grows over the first seconds, and throughput approaches the line rate if there is no loss.
# /etc/sysctl.d/90-long-fat-network.conf  (values to test, not universal defaults)
net.ipv4.tcp_window_scaling = 1
net.ipv4.tcp_rmem = 4096 131072 33554432
net.ipv4.tcp_wmem = 4096 16384 33554432
net.core.rmem_max = 33554432
net.core.wmem_max = 33554432

If throughput still falls short once the window is large enough, the limit has moved to congestion control: loss on the path shrinks the congestion window no matter what the receiver advertises. That is where a model-based algorithm such as BBR can help, and where fixing loss beats tuning buffers.

Reading the evidence in captures

Capture the handshake; without it no tool can interpret the window field.

$ tcpdump -ni eth0 'tcp[tcpflags] & tcp-syn != 0 and host 203.0.113.7'
IP 10.0.0.5.41022 > 203.0.113.7.443: Flags [S], seq 1, win 64240,
    options [mss 1460,sackOK,TS val 1 ecr 0,nop,wscale 7], length 0
IP 203.0.113.7.443 > 10.0.0.5.41022: Flags [S.], seq 9, ack 2, win 65160,
    options [mss 1460,sackOK,TS val 7 ecr 1,nop,wscale 9], length 0

Both sides announced a shift, so scaling is on. If wscale appears in the SYN but not in the SYN-ACK, the server or something in front of it disabled scaling and the connection is capped at 64 KB.

Wireshark's TCP details show a "Window size scaling factor" for each segment. A value of -1 means Wireshark did not see the handshake and is showing the raw field, which is the most common way people misread a perfectly healthy connection as having a 500-byte window. A value of -2 means the handshake was seen and scaling was not negotiated. The "Calculated window size" field is the one to trust. The TCP stream graphs (window scaling and throughput) show whether the sender's in-flight data hugs the receiver's window line, which is the signature of a window-limited transfer.

Failure modes

SymptomCauseFix
Throughput stuck near 64 KB per RTTOne side, or a middlebox, removed the option from the SYN or SYN-ACKCapture both ends of the handshake to find where it disappears; fix or bypass the device
Fine in one region, slow across oceansBuffer maximum below the long path's BDPRaise tcp_rmem and tcp_wmem maximums on both ends
Explicit large buffer made things slowerSO_RCVBUF capped by rmem_max and autotuning disabledRemove the setsockopt, or raise rmem_max and set before connect
Long-lived connections stay slow after a sysctl changeShift was fixed at their handshakeReconnect; pools must be recycled to pick up new limits
Connections stall after a firewall upgradeDevice rewrites or normalises windows without tracking the shiftUpgrade or reconfigure the device to track scaling; do not disable scaling
Large memory use on busy serversTens of thousands of sockets each allowed huge buffersSize maximums for the paths you serve; autotuning only grows buffers that need it

On Windows the equivalent is receive window autotuning, shown by netsh interface tcp show global. Setting the autotuning level to disabled, a common folk remedy, caps the receive window at 64 KB and reproduces the original problem.

Trade-offs

Large maximums cost nothing for connections that never need them, because Linux grows buffers on demand, but they raise the worst case: a server with many slow readers can pin a lot of kernel memory. Buffers far larger than the BDP add queueing on the sender side without adding throughput. And window scaling only removes the flow-control ceiling; it does nothing about loss, small congestion windows after idle periods, or an application that writes in tiny chunks. For short request-response traffic the window rarely matters at all, which is why choosing between TCP and UDP usually turns on latency and head-of-line blocking rather than throughput. For a refresher on where the window sits in the header, see TCP/IP fundamentals.

What to do next

  1. Compute the BDP for your longest important path from link rate and measured RTT, and write it next to your sysctl settings.
  2. Run ss -tin on a live long-distance connection and confirm both wscale values are present and large enough for that BDP.
  3. Check tcp_rmem, tcp_wmem and rmem_max on both ends, and raise the maximums only where long paths need them.
  4. Search your code for SO_RCVBUF and SO_SNDBUF; remove them unless a measurement justifies them, and set them before connect or listen if kept.
  5. Capture a handshake through every firewall and load balancer on the path to prove the option survives.
  6. After any buffer change, recycle connection pools and re-measure throughput rather than assuming the change applied.
Key takeaway: The TCP header can only express a 64 KB window, which caps throughput at window divided by RTT and cripples long, fast paths. Window scaling fixes this with a shift count of up to 14 that each side announces only in its SYN or SYN-ACK and that stays fixed for the connection. Size receive and send buffer maximums from the bandwidth-delay product, avoid setting SO_RCVBUF late or without need, capture the handshake to prove the option survives the path, and reconnect after changing limits.