Most explanations of TCP describe the protocol: the three-way handshake, sequence numbers, sliding windows and slow start. Those are covered in the TCP/IP article. This article is about the other half, the implementation you actually operate. When a service drops connections under load, a cross-region copy runs at a fifth of the link speed, or a client hangs for fifteen minutes after a peer disappears, the explanation is almost always a specific queue, buffer or timer inside the kernel's TCP stack.

We follow a connection through the Linux implementation: how a listening socket queues new connections, how data moves from write() to the wire and back, how buffers are sized, how the stack decides that a segment is lost, how congestion control plugs in, and why closed connections linger. Each part ends with the counters and settings that let you see it in production, and a worked example ties them together.

Advertisement

The shape of a TCP socket in the kernel

Every TCP connection is a kernel socket object with a state (LISTEN, SYN_RECV, ESTABLISHED, FIN_WAIT_1 and so on), two byte queues, a set of congestion variables and a set of timers. The application sees a file descriptor; the kernel sees a send queue holding bytes not yet acknowledged, a receive queue holding in-order bytes not yet read, an out-of-order queue holding segments that arrived after a gap, and per-connection estimates of round-trip time and congestion window.

A TCP socket in the Linux kernel: queues, limits and timers between the app and the NICApplicationwrite() / read()Send bufferunsent + unackedtcp outputmin(cwnd, rwnd)qdisc + NICpacing, TSOcopysegmentsNIC + GROcoalescetcp inputACKs, SACK, orderReceive queue+ out-of-orderApplicationread()in ordercopyACK frees send buffer, grows cwndTimersRTO, TLP, delayed ACK, keepaliveCongestion controlCUBIC, BBR via tcp_congestion_opsListenerSYN queue, accept queue (somaxconn)Closed socketsTIME_WAIT for 60 s, FIN_WAIT_2Most production TCP problems are one of these queues filling, or one of these timers firing.
Data written by the application waits in the send buffer until the congestion window and the peer's receive window allow it out. Incoming segments are coalesced by GRO, ordered by tcp input and queued for read(). ACKs free send-buffer space and drive congestion control; timers handle loss and idle connections.

Hold this picture in mind: almost every TCP performance problem is one of the queues being too small or too full, and almost every reliability problem is one of the timers firing or not firing.

Accepting connections: two queues and SYN cookies

A listening socket has two queues. When a SYN arrives, the kernel creates a lightweight request socket in SYN_RECV and replies with a SYN-ACK. When the final ACK arrives, the connection is complete and moves to the accept queue, where it waits until the application calls accept(). The accept queue's length is the backlog argument to listen(), silently capped at net.core.somaxconn. The number of pending half-open requests is bounded separately, related to net.ipv4.tcp_max_syn_backlog.

If the application accepts too slowly, the accept queue fills. The kernel then ignores further completing handshakes; the client believes it is connected, sends its request, and waits for retransmissions. The symptom is latency spikes of one second or more on new connections with no errors in the application. With net.ipv4.tcp_syncookies enabled, a server under a SYN flood stops keeping state for half-open connections and encodes it in the SYN-ACK sequence number instead, at the price of losing some TCP options for those connections.

# Accept-queue overflow and drops since boot
nstat -az TcpExtListenOverflows TcpExtListenDrops

# For listening sockets, Recv-Q is the current accept-queue length, Send-Q the limit
ss -ltn 'sport = :8443'
# State   Recv-Q  Send-Q  Local Address:Port
# LISTEN  0       4096    0.0.0.0:8443

If the overflow counters grow, raise the backlog in the application and somaxconn together, and fix the underlying cause: an accept loop that is blocked, single-threaded or starved of CPU.

Advertisement

The send path

write() copies bytes into the socket's send buffer and returns; it does not mean the peer has anything. The buffer holds both unsent data and sent-but-unacknowledged data, because TCP must be able to retransmit. A segment leaves only when three limits allow it: the congestion window (how much the network is believed to carry), the peer's advertised receive window (how much the peer can buffer), and pacing (how fast the stack spreads packets over time). The amount in flight is bounded by the smaller window.

Below TCP, the stack builds large segments and lets the NIC cut them into MSS-sized packets (TSO, or GSO in software), which saves CPU. TCP Small Queues limits how many bytes one socket may have sitting in the qdisc and driver queues, so a single bulk flow cannot fill device queues and add latency to everyone else. Pacing, implemented in TCP itself or in the fq qdisc, spaces packets at a rate derived from the congestion window and RTT instead of sending bursts.

Two socket options matter for latency-sensitive applications. TCP_NODELAY disables Nagle's algorithm, which otherwise holds small writes while earlier data is unacknowledged; request-response protocols that write a header and a body separately need it. TCP_NOTSENT_LOWAT limits how much unsent data can sit in the send buffer, so an application that multiplexes streams, such as an HTTP/2 server, keeps its priority decisions in user space rather than committing megabytes to the kernel too early.

The receive path and buffer autotuning

On receive, the NIC driver and GRO coalesce consecutive packets of a flow into larger units before TCP sees them. TCP processes ACK information, puts in-order data on the receive queue and gaps on the out-of-order queue, and sends ACKs, delayed slightly so one ACK can cover two segments or ride on a response. The receive window it advertises is the free space it is willing to commit, scaled by the window-scale option negotiated in the handshake.

Linux grows each connection's receive buffer automatically as it measures the flow's throughput and RTT, within the limits in net.ipv4.tcp_rmem (minimum, default, maximum); the send side uses net.ipv4.tcp_wmem. Setting SO_RCVBUF or SO_SNDBUF explicitly turns autotuning off for that socket, which is a common way to make a fast long-distance transfer slower. Leave them unset unless you have measured a reason.

The number that governs bulk throughput is the bandwidth-delay product (BDP): link rate times round-trip time. A connection can never have more than one window in flight per RTT, so its throughput is at most window divided by RTT.

Loss detection and retransmission

TCP has two ways to learn that a segment was lost. The fast way uses evidence from later segments. With SACK, the receiver reports exactly which ranges arrived, and modern Linux uses RACK-TLP (RFC 8985): RACK marks a segment lost when a segment sent sufficiently later has been delivered, using time rather than a count of duplicate ACKs, and Tail Loss Probe sends a probe when the last segments of a burst get no response, so the loss is discovered by SACK instead of a timeout.

The slow way is the retransmission timeout. The stack keeps a smoothed RTT and an RTT variance per connection, and per RFC 6298 sets RTO = SRTT + max(G, 4 × RTTVAR), where G is the clock granularity, with an initial RTO of one second. Linux enforces a minimum RTO of 200 ms, so even in a datacenter with 100-microsecond RTTs a timeout costs at least 200 ms. On each successive timeout for the same segment the RTO doubles. When an RTO fires, the congestion window collapses to one segment, so timeouts are very expensive for throughput as well as latency.

How long the stack keeps retrying before giving up on an established connection is set by net.ipv4.tcp_retries2; with the default, a connection to a vanished peer with data outstanding can take on the order of fifteen minutes to fail. Applications that need faster failure should set TCP_USER_TIMEOUT (RFC 5482) per socket, which bounds how long transmitted data may remain unacknowledged before the connection is aborted. Idle connections send nothing and so never detect a dead peer; keepalive probes, by default after 7,200 seconds of idleness, 75 seconds apart, 9 times, address that and are almost always tuned shorter per socket.

import socket

s = socket.create_connection(("db.internal", 5432))
s.setsockopt(socket.IPPROTO_TCP, socket.TCP_NODELAY, 1)
# Abort if sent data stays unacknowledged for 30 s (milliseconds).
s.setsockopt(socket.IPPROTO_TCP, socket.TCP_USER_TIMEOUT, 30_000)
# Detect dead idle peers in about 60 s instead of over two hours.
s.setsockopt(socket.SOL_SOCKET, socket.SO_KEEPALIVE, 1)
s.setsockopt(socket.IPPROTO_TCP, socket.TCP_KEEPIDLE, 30)
s.setsockopt(socket.IPPROTO_TCP, socket.TCP_KEEPINTVL, 10)
s.setsockopt(socket.IPPROTO_TCP, socket.TCP_KEEPCNT, 3)
# Per-socket congestion control, if the module is loaded and allowed.
s.setsockopt(socket.IPPROTO_TCP, socket.TCP_CONGESTION, b"bbr")

Congestion control is a plug-in

Linux separates loss recovery from the decision of how fast to send. Congestion control algorithms are modules that implement a table of callbacks, struct tcp_congestion_ops, which the stack calls on every ACK, on loss, on entering recovery and on RTT samples. The module adjusts the congestion window and, for rate-based algorithms, the pacing rate. CUBIC is the default on most distributions; it is loss-based and grows its window as a cubic function of time since the last loss. BBR is model-based: it estimates the bottleneck bandwidth and minimum RTT and paces at that rate, which helps on long paths with random loss but can be unfair to loss-based flows on shallow buffers.

The system default is net.ipv4.tcp_congestion_control; net.ipv4.tcp_available_congestion_control lists loaded modules, and unprivileged sockets may only choose those in tcp_allowed_congestion_control. Because the choice is per socket, you can run BBR on a bulk-transfer service and leave everything else on CUBIC. ECN, which lets routers signal congestion without dropping packets, is a separate negotiation; see ECN.

Closing: FIN, TIME_WAIT and port exhaustion

The side that closes first passes through FIN_WAIT states and ends in TIME_WAIT, which Linux holds for a fixed 60 seconds. TIME_WAIT exists so delayed segments from an old connection cannot be mistaken for a new one with the same four-tuple, and so the final ACK can be resent if it was lost. It costs little memory, but it does occupy the four-tuple. A client that opens and closes many short connections to the same server address and port can exhaust its ephemeral port range, at which point connect() fails with EADDRNOTAVAIL.

The right fix is to reuse connections with keep-alive and pooling, covered in connection pooling. net.ipv4.tcp_tw_reuse lets outgoing connections reuse TIME_WAIT four-tuples when TCP timestamps make it safe. The old tcp_tw_recycle setting broke clients behind NAT and was removed from the kernel in 4.12; ignore advice that mentions it. Closing with SO_LINGER set to zero sends a RST and skips TIME_WAIT, but it also discards unsent data, so use it only when you mean to abort.

Worked example: a slow cross-region copy

A backup job copies data from Frankfurt to Virginia over a 1 Gbit/s path with an 80 ms RTT and reaches only about 400 Mbit/s. The BDP is 1 Gbit/s × 0.08 s = 80 Mbit, which is 10 MB. To fill the path, the connection needs about 10 MB in flight. The Linux default maximum for tcp_wmem is 4 MB on many systems, and 4 MB per 80 ms is 50 MB/s, or 400 Mbit/s, which is exactly the observed ceiling.

# illustrative output
$ ss -tin dst 203.0.113.40
ESTAB 0 3981312 10.0.1.7:51522 203.0.113.40:443
     cubic wscale:9,9 rto:284 rtt:80.6/1.2 mss:1448 cwnd:6950 ssthresh:7020
     bytes_acked:1734093824 retrans:0/112 rcv_space:14480
     send 998.9Mbps pacing_rate 1198.7Mbps delivery_rate 401.2Mbps
     unacked:2749 notsent:0

Read it field by field. The send queue holds about 3.9 MB, which is near the 4 MB limit, while the congestion window (6,950 segments, about 10 MB) would allow more. Retransmissions are low, so the network is not the limit. The delivery rate is 401 Mbit/s. The connection is send-buffer limited. Raising the maximum of tcp_wmem (and tcp_rmem on the receiver) to 16 MB, with no explicit SO_SNDBUF in the application, lets autotuning grow past the BDP; the transfer should then approach line rate. If after that the congestion window, not the buffer, is the limit and retransmissions climb, the path is lossy, and trying BBR on this one socket is a reasonable next experiment.

Failure modes and where they show up

SymptomLikely causeWhere to look
1 s or 3 s stalls on new connectionsAccept queue overflow or SYN dropsTcpExtListenOverflows, ss -lnt
Bulk throughput far below link rateBuffer below BDP, or explicit SO_SNDBUFss -ti: notsent, cwnd, buffer limits
Tail latency spikes in a datacenterRTOs at the 200 ms minimumTcpExtTCPTimeouts, retrans in ss -ti
Client hangs for minutes after peer diesNo TCP_USER_TIMEOUT or keepaliveSocket options in the client library
connect() fails with EADDRNOTAVAILEphemeral ports held by TIME_WAITss -s, connection reuse in the client
Large transfers hang, small ones workPath MTU black holeSee path MTU discovery

The last row is a path problem, not a TCP-buffer one, and is covered in path MTU discovery.

Tuning trade-offs

Larger buffer maximums raise throughput on long paths but let a few connections hold more memory, and the global net.ipv4.tcp_mem pressure thresholds then start trimming everyone; size them from your real BDP, not from a blog's largest number. Shorter keepalive and user timeouts detect failures faster but abort connections across brief network blips; set them per service to match its retry logic. A larger accept backlog absorbs bursts but hides a slow accept loop. And changing the default congestion control changes every flow on the host, while setting it per socket confines the experiment. Change one thing at a time and keep the counters from before.

What to do next

  1. Record a baseline: nstat -az for ListenOverflows, ListenDrops, TCPTimeouts and RetransSegs, plus ss -s, on each tier.
  2. For your busiest listener, compare the accept queue in ss -lnt with its limit and raise backlog and somaxconn together only if overflows grow.
  3. Compute the BDP of your longest bulk-transfer path and check that the tcp_rmem and tcp_wmem maximums exceed it, and that the application does not set SO_SNDBUF or SO_RCVBUF.
  4. Audit client libraries for TCP_USER_TIMEOUT, keepalive and TCP_NODELAY settings, and set them to match each service's failure-detection budget.
  5. Check for TIME_WAIT pressure on clients and fix it with connection reuse before touching tcp_tw_reuse.
  6. Run one controlled experiment with per-socket BBR on a lossy long path, and compare delivery rate and retransmissions against CUBIC.
Key takeaway: TCP in production is a set of kernel queues and timers. Accept queues explain connection stalls, send and receive buffers against the bandwidth-delay product explain slow transfers, RTO and RACK-TLP explain tail latency, user timeouts and keepalives explain hangs, and TIME_WAIT explains port exhaustion. Learn to read ss -ti and nstat, compute the BDP, and change one tunable at a time against a recorded baseline.