For decades TCP congestion control used packet loss as its signal for a full network. Reno and then CUBIC grow the sending window until a packet drops, back off, and grow again. That worked when router buffers were small and loss meant congestion. Modern paths break both assumptions: buffers are often large enough to hold hundreds of milliseconds of data, so loss-based senders fill them and inflate latency for everyone, and wireless links lose packets for reasons that have nothing to do with congestion, so loss-based senders slow down for no reason.
BBR, Bottleneck Bandwidth and Round-trip propagation time, was published by Google in 2016 with a different premise: estimate the path's bottleneck bandwidth and its minimum round-trip time directly, and send at that rate with about one bandwidth-delay product in flight. This article builds BBR from that model, walks through its state machine, explains what went wrong with the first version, and covers the third version described in the IETF draft, before turning to running and measuring it on Linux. The surrounding TCP machinery is covered in the Linux TCP architecture article.
The operating point loss-based control misses
Every path has a bottleneck: the slowest link the flow crosses. Call its rate BtlBw and the round-trip time with empty queues RTprop. Their product, the bandwidth-delay product or BDP, is the amount of data that fills the pipe without queueing. Send with less than one BDP in flight and the bottleneck idles some of the time. Send with exactly one BDP and you get full throughput at minimum delay. Send with more and the excess sits in the bottleneck's queue: throughput stays the same, delay rises, and when the buffer overflows packets drop.
Loss-based algorithms operate at the far edge of that range, because loss only happens once the buffer is full. On a 100 Mbit/s link with a 40 ms base RTT and a 1 MB buffer, a CUBIC flow regularly pushes queueing delay towards 80 ms, since 1 MB drains at 12.5 MB per second in 80 ms, tripling the RTT that a video call or game sharing the link experiences. BBR aims at the near edge: full rate, near-empty queue.
The difficulty is that the two quantities cannot be measured at the same time. Measuring the bottleneck rate requires enough data in flight to build a small queue; measuring the minimum RTT requires the queue to be empty. BBR resolves this by alternating: it probes each quantity in turn and keeps a filtered estimate of each.
The model: estimating bandwidth and minimum RTT
Every ACK yields a rate sample. The sender records, for each packet, how much data had been delivered when it was sent and when; when the ACK for that packet arrives, the delivery rate over that interval is the change in delivered data divided by the elapsed time. Using the longer of the send interval and the ACK interval protects the estimate from ACK compression, where ACKs bunch up and would otherwise suggest a rate higher than the link supports.
on_send(pkt):
pkt.delivered = C.delivered # bytes delivered when pkt left
pkt.delivered_time = C.delivered_time
pkt.first_sent = C.first_sent_time
pkt.app_limited = C.app_limited
on_ack(pkt, now):
C.delivered += pkt.size
C.delivered_time = now
send_elapsed = pkt.sent_time - pkt.first_sent
ack_elapsed = now - pkt.delivered_time
interval = max(send_elapsed, ack_elapsed)
rate = (C.delivered - pkt.delivered) / interval
rtt = now - pkt.sent_time
if rate >= max_bw.current() or not pkt.app_limited:
max_bw.update(rate) # windowed max over recent probing cycles
min_rtt.update(rtt) # windowed min over the last 10 secondsBandwidth uses a windowed maximum, because queueing and cross traffic can only make a sample look slower than the bottleneck, never faster. RTT uses a windowed minimum, because queueing can only add delay. Samples taken while the application had nothing to send are discarded unless they raise the estimate, so an idle chat connection does not conclude the network got slower. From the two estimates the sender computes BDP = max_bw x min_rtt and sets two controls: the pacing rate, pacing_gain x max_bw, which spaces packets in time, and the congestion window, cwnd_gain x BDP, which caps the data in flight.
Pacing is not optional
Loss-based TCP is ACK-clocked: each ACK releases a burst bounded by the window. BBR's primary control is the pacing rate, not the window, and without pacing it degenerates into bursts that build exactly the queue it is trying to avoid. On Linux, pacing is provided either by the fq queueing discipline or, since kernel 4.13, by TCP's internal pacing when another qdisc is in use. The window still matters as a safety cap for when ACKs stop arriving, for example during a burst of loss.
The state machine
BBR runs four modes, shown below using the IETF draft's names for the probing sub-states.
- Startup grows the sending rate exponentially, like slow start, using a pacing gain of about 2.89 (2/ln 2) in the original version and 4 ln 2, about 2.77, in the current draft. It exits when the bandwidth estimate fails to grow by 25 percent for three rounds in a row, meaning the pipe is full, or, in the newer version, when loss is excessive.
- Drain paces below the estimated rate to empty the queue Startup built, until data in flight is back down to about one BDP. The draft uses a gain of 0.5.
- ProbeBW is the steady state. The original version cycled through eight phases with pacing gains of 1.25, 0.75 and then six phases of 1: one round probing above the estimate, one draining the resulting queue, six cruising. The newer version splits this into DOWN, CRUISE, REFILL and UP sub-states and spaces bandwidth probes out in time, so that BBR flows sharing a link with loss-based flows probe less aggressively.
- ProbeRTT runs when the minimum RTT has not been refreshed for a while: originally after 10 seconds, cutting the window to four packets; in the draft every 5 seconds, cutting it to half a BDP, for 200 ms. The briefly emptied queue lets every flow on the link see the true propagation delay.
What went wrong with the first version
BBR v1, still the tcp_bbr module in mainline Linux at the time of writing, ignored loss as a signal altogether. That was the point on lossy wireless paths, but on paths with shallow buffers it produced very high retransmission rates, because a flow that believes the bandwidth estimate keeps sending through drops. Its interaction with CUBIC depended heavily on buffer depth: in shallow buffers BBR could take most of the link, in deep buffers CUBIC could fill the queue and starve BBR. Flows with longer RTTs received larger shares, since their windows scale with their longer min_rtt. And ProbeRTT's four-packet window produced a visible throughput dip every ten seconds.
BBR v3: bounded by loss and ECN
Google's later versions keep the model but add guard rails, and the third version is specified in the IETF draft draft-ietf-ccwg-bbr, published as experimental; revision 06 is dated July 2026. The main changes:
- Loss and ECN as bounds, not triggers. BBR tracks two limits on data in flight.
inflight_hiis the long-term ceiling learned when probing pushed loss above a threshold, 2 percent per round trip by default, or caused ECN marks;inflight_lois a short-term limit that backs off on recent loss, using a multiplicative decrease factor of 0.7. - Headroom. When cruising, the flow keeps inflight below the ceiling, leaving 15 percent headroom so new flows can enter.
- Gentler probing. ProbeBW waits between probes and raises the rate gradually during UP, which makes it coexist better with CUBIC.
- Shorter, shallower ProbeRTT, as described above, to reduce the periodic throughput dip.
- Shorter bandwidth memory. The max-bandwidth filter covers about two probing cycles, so the estimate decays faster when bandwidth drops.
The Linux implementation of this design is published on the v3 branch of Google's google/bbr repository rather than in mainline, so running it means building a patched kernel or module. QUIC stacks implement their own BBR variants in user space; check which version your library ships before assuming its behaviour matches either.
A worked example
Consider a flow over a 100 Mbit/s bottleneck with a 40 ms RTprop. The BDP is 12.5 MB/s x 0.04 s = 500 KB, about 342 segments of 1,460 bytes. Startup roughly doubles the delivery rate each round from an initial window of 10 segments, so it reaches the bottleneck in about five or six rounds, then needs three more rounds of flat estimates to declare the pipe full: around 350 ms in all. Because Startup's cwnd gain is 2, up to about one extra BDP may be queued at exit, adding up to 40 ms of delay, which Drain removes within a round or two.
In steady state, a v1 flow probing at 1.25 for one round adds a quarter of a BDP, about 125 KB or 10 ms of queue, then drains it in the next. Compared with the CUBIC flow above, which held up to 80 ms of queue, the median RTT stays near 40 ms. Now put the same flow behind a 64 KB shallow-buffer switch: v1's probe overflows the buffer every cycle and, ignoring the loss, retransmits heavily; v3 sees loss above 2 percent during UP, lowers inflight_hi, and stops probing that far until conditions change.
Running BBR on Linux
# Is it available, and what is the default?
sysctl net.ipv4.tcp_available_congestion_control
sysctl net.ipv4.tcp_congestion_control
# Enable system-wide (persist in /etc/sysctl.d/)
sysctl -w net.core.default_qdisc=fq
sysctl -w net.ipv4.tcp_congestion_control=bbr
# Inspect live connections: look for bbr:(bw:...,mrtt:...,pacing_gain:...,cwnd_gain:...)
ss -tin dst 10.0.0.0/8You can also choose per socket, which is the safer way to start, because it limits BBR to the traffic you are testing:
import socket
s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
s.setsockopt(socket.IPPROTO_TCP, socket.TCP_CONGESTION, b"bbr") # Linux only
s.connect(("storage.internal", 443))
print(s.getsockopt(socket.IPPROTO_TCP, socket.TCP_CONGESTION, 16))Congestion control is a sender-side choice; the receiver needs no change. That makes BBR most useful on servers that send a lot: CDNs, object stores, video origins and replication links between regions. It also means a fleet change affects every peer of those servers, so roll it out with the same care as any other change to production traffic.
Measuring whether it helped
Run an A/B test across hosts or connections, not a before-and-after comparison, because traffic and paths change daily. Measure throughput percentiles for bulk transfers, RTT percentiles, retransmission rate from nstat or ss, and an application metric such as time to first byte or rebuffer rate. Split the results by path type, such as datacentre, broadband and cellular, since BBR's gains concentrate on long, lossy or deeply buffered paths and may be absent or negative on short, clean ones. Look specifically at the retransmission rate: a large rise is the signature of v1 on shallow buffers. If your network supports ECN, the interaction is covered in the ECN architecture article.
Failure modes and trade-offs
- Fairness surprises: mixing BBR and CUBIC on a shared bottleneck may starve one or the other depending on buffer depth. Test on your own topology.
- Retransmission storms on shallow buffers with v1. Monitor retransmits per flow after enabling it.
- Policers: token-bucket policers that drop above a rate can mislead the bandwidth estimate; v1 in particular may keep sending at the burst rate.
- Missing pacing: offload or virtualisation layers that bunch packets back into bursts undo the pacing that BBR depends on.
- Short flows: most web objects finish within Startup, so for them congestion control choice matters less than initial window, connection reuse and the transport, discussed in the QUIC architecture article.
- The trade-off: BBR trades the simplicity and well-studied fairness of loss-based control for lower latency and better throughput on difficult paths. The newer versions narrow the fairness gap by reintroducing loss as a bound.
What to do next
- Check which congestion control and qdisc your senders use today, and which BBR version your kernel or QUIC library ships.
- Pick one high-volume sender class, such as cross-region replication, and enable BBR per socket on a test subset.
- Ensure pacing is in effect:
fqas the qdisc, or a kernel new enough for internal pacing. - A/B test throughput, RTT, retransmits and one user-facing metric, split by path type.
- Inspect
ss -tinon live flows to confirm the bandwidth and min-RTT estimates look like the path. - Test coexistence with CUBIC on a shared bottleneck that resembles your network.
- Write down the rollback (one sysctl) and the metric that would trigger it before widening the rollout.