BBR is usually explained as a model: estimate the bottleneck bandwidth and the minimum round-trip time, send at that rate, and keep about one bandwidth-delay product in flight. This site's BBR architecture article covers that model, the state machine, the problems of the first version and the changes in BBRv3. This article takes the other route. It reads BBR as Linux ships it today, in net/ipv4/tcp_bbr.c and net/ipv4/tcp_rate.c, and then builds a small lab where you can watch each mechanism happen.
One fact frames everything. Mainline Linux still ships BBR version 1, with fixes added over the years. BBRv2 and BBRv3 live in Google's out-of-tree google/bbr repository on the v3 branch and had not been merged as of October 2026. When you run sysctl net.ipv4.tcp_congestion_control=bbr on a stock kernel, you get the behaviour described here, including its known weaknesses with loss and shallow buffers. Constants below are quoted from mainline tcp_bbr.c; check your kernel's source if you depend on an exact value.
Where BBR plugs into the TCP stack
Linux congestion control is pluggable through struct tcp_congestion_ops. Classic algorithms such as CUBIC implement .cong_avoid and let the core stack handle recovery. BBR instead implements .cong_control, a hook the stack calls once per ACK with a rate sample, and takes over both the congestion window and the pacing rate. It also implements .ssthresh and .undo_cwnd so the core's loss-recovery code does not fight it, .set_state to react to entering loss recovery, and .get_info so ss can show its internal state.
Rate samples: measuring delivery, not sending
BBR's bandwidth estimate comes from tcp_rate.c, which the kernel maintains for every TCP socket. When a packet is sent, the stack records on it the connection's delivered count and the time of the most recent delivery. When an ACK confirms that packet, tcp_rate_gen() computes how much data was delivered between those two points and over what interval. It takes the longer of the send interval and the ACK interval, which stops ACK compression from inflating the estimate. The result is a struct rate_sample with fields including delivered, interval_us, rtt_us, losses and is_app_limited.
The app-limited flag matters more than it looks. If the application had nothing to send, a low delivery rate says nothing about the network. The stack marks such periods, and BBR uses an app-limited sample only if it would raise the current maximum. Without this rule, every idle pause in a chatty RPC connection would drag the bandwidth estimate down.
bbr_main(): the per-ACK update
Stripped of detail, the per-ACK path looks like this:
static void bbr_main(struct sock *sk, const struct rate_sample *rs)
{
struct bbr *bbr = inet_csk_ca(sk);
u32 bw;
bbr_update_model(sk, rs); /* bw filter, ACK aggregation, cycle phase,
full-pipe check, drain, min_rtt / PROBE_RTT */
bw = bbr_bw(sk); /* windowed max delivery rate (or policer estimate) */
bbr_set_pacing_rate(sk, bw, bbr->pacing_gain);
bbr_set_cwnd(sk, rs, rs->acked_sacked, bw, bbr->cwnd_gain);
}
/* The bandwidth filter, simplified from bbr_update_bw() */
if (rs->prior_delivered >= bbr->next_rtt_delivered) { /* a new round trip began */
bbr->next_rtt_delivered = tp->delivered;
bbr->rtt_cnt++;
}
bw = rs->delivered * BW_UNIT / rs->interval_us; /* packets per usec, scaled */
if (!rs->is_app_limited || bw >= bbr_max_bw(sk))
minmax_running_max(&bbr->bw, bbr_bw_rtts, bbr->rtt_cnt, bw);Note that BBR counts time in round trips, not wall-clock time, for the bandwidth filter. A round ends when a packet sent after the previous round began is acknowledged. The max filter (lib/minmax.c) keeps the best sample over the last bbr_bw_rtts rounds, which is the probe cycle length plus two, so ten rounds. The min-RTT filter instead uses wall-clock time: bbr_min_rtt_win_sec is 10 seconds.
The constants that define behaviour
| Constant (tcp_bbr.c) | Value | What it does |
|---|---|---|
bbr_high_gain | about 2.885 (2/ln 2) | Pacing and cwnd gain in STARTUP; doubles the sending rate each round |
bbr_drain_gain | about 0.35 (1/2.885) | Pacing gain in DRAIN, to empty the queue STARTUP built |
bbr_cwnd_gain | 2 | cwnd is twice the estimated BDP in PROBE_BW, so delayed and aggregated ACKs do not starve the pipe |
bbr_pacing_gain[] | 1.25, 0.75, then six phases of 1.0 | The PROBE_BW cycle; each phase lasts about one min_rtt |
bbr_full_bw_thresh, bbr_full_bw_cnt | 1.25 and 3 | Leave STARTUP when bandwidth grows less than 25% for 3 rounds |
bbr_min_rtt_win_sec | 10 | If min_rtt has not been refreshed for 10 s, enter PROBE_RTT |
bbr_probe_rtt_mode_ms | 200 | Stay in PROBE_RTT at least 200 ms and one round |
bbr_cwnd_min_target | 4 | cwnd in PROBE_RTT, in packets |
bbr_pacing_margin_percent | 1 | Pace 1% below the estimate to avoid building a queue |
Two details are easy to miss. PROBE_BW starts at a random phase, never the 0.75 phase, so flows sharing a bottleneck do not probe in lockstep. And the phases are asymmetric: the 1.25 phase lasts at least one min_rtt and can run longer, until inflight reaches 1.25 times the BDP or a loss occurs, while the 0.75 phase can end early as soon as inflight falls to the estimated BDP. Mainline v1 also has a long-term bandwidth mode, the lt_bw logic, that detects token-bucket policers from sustained loss and caps its rate to the policed rate. It is the main place v1 responds to sustained loss; otherwise loss only ends a probe-up phase and triggers packet conservation during recovery.
Pacing: how a rate becomes timing
BBR without pacing is not BBR. Its cwnd is deliberately twice the BDP, so if packets were sent as fast as the window allowed, they would leave in bursts and build exactly the queue BBR tries to avoid. bbr_set_pacing_rate() writes sk->sk_pacing_rate as gain times bandwidth times 0.99. Before the first bandwidth sample, BBR seeds it from the initial cwnd divided by the smoothed RTT, times the high gain.
Enforcement happens below the congestion module. Since Linux 4.20, TCP uses an earliest-departure-time model: each packet carries a departure timestamp in skb->tstamp computed from the pacing rate, and the fq qdisc holds it until that time. If the interface does not use fq, TCP falls back to internal pacing with high-resolution timers, available since 4.13. Both work. fq is usually cheaper at high connection counts, and it is why many guides pair net.core.default_qdisc=fq with BBR. The pacing rate also sets the TSO burst size through bbr_tso_segs_goal(), so low-rate flows send small bursts and high-rate flows larger ones.
A lab you can run in ten minutes
To see this behaviour, you need a real bottleneck with a known rate, delay and buffer. Do not put netem on the sending host's own interface, because it interferes with TCP small queues and pacing on the sender and distorts the results. Use three network namespaces and shape traffic on the middle one:
#!/bin/sh -e
# client <-> router <-> server, bottleneck on the router's egress toward the server
for ns in cli rtr srv; do ip netns add $ns; done
ip link add c0 netns cli type veth peer name r0 netns rtr
ip link add r1 netns rtr type veth peer name s0 netns srv
ip -n cli addr add 10.0.1.2/24 dev c0; ip -n rtr addr add 10.0.1.1/24 dev r0
ip -n rtr addr add 10.0.2.1/24 dev r1; ip -n srv addr add 10.0.2.2/24 dev s0
for l in "cli c0" "rtr r0" "rtr r1" "srv s0"; do set -- $l; ip -n $1 link set $2 up; done
ip -n cli route add default via 10.0.1.1; ip -n srv route add default via 10.0.2.1
ip netns exec rtr sysctl -qw net.ipv4.ip_forward=1
# 100 Mbit/s, 40 ms RTT (20 ms each way), buffer LIMIT packets
LIMIT=${LIMIT:-1300}
ip netns exec rtr tc qdisc add dev r1 root netem delay 20ms rate 100mbit limit $LIMIT
ip netns exec rtr tc qdisc add dev r0 root netem delay 20ms
ip netns exec cli tc qdisc replace dev c0 root fq
ip netns exec srv iperf3 -s -D
ip netns exec cli iperf3 -c 10.0.2.2 -t 30 -C bbr &
sleep 5; ip netns exec cli ss -tin dst 10.0.2.2 # shows bbr:(bw:...,mrtt:...,pacing_gain:...,cwnd_gain:...)The ss -tin output shows BBR's own view: estimated bandwidth, minimum RTT and the current gains. Sample it every 100 ms in a loop and you can watch STARTUP's gain of about 2.89, the brief drain, and then the 1.25, 0.75, 1, 1 pattern of PROBE_BW. For cwnd and smoothed RTT on every ACK, enable the tcp:tcp_probe tracepoint.
Worked example: reading the numbers
With 100 Mbit/s and 40 ms, the BDP is 100,000,000 / 8 x 0.040 = 500,000 bytes, about 345 full-size segments of 1,448 payload bytes. In steady state, expect ss to report a bandwidth near 100 Mbit/s and an mrtt near 40 ms. BBR's cwnd target is about twice the BDP, roughly 690 segments, but pacing keeps actual inflight near one BDP most of the time. Every ten seconds, unless the queue drains on its own, the flow enters PROBE_RTT: cwnd drops to 4 packets for at least 200 ms, and throughput dips briefly. That dip costs about 2% of a ten-second period, which is the price of refreshing the RTT estimate.
Now run the coexistence test. Start one BBR flow and one CUBIC flow (-C cubic) together, and repeat with LIMIT=30, a buffer of about a tenth of a BDP, and LIMIT=1300, about four BDPs. Record each flow's throughput and the retransmissions iperf3 reports. The published research on BBRv1 predicts the direction of the results, and your lab will show the size. In the shallow buffer, BBR takes most of the bandwidth and CUBIC collapses, because CUBIC backs off on every loss while BBR v1 keeps pacing at its estimate and retransmits heavily. In the deep buffer, CUBIC fills the queue, and the outcome depends on how BBR's cwnd cap of twice the BDP interacts with the inflated RTT. Do not trust a fixed fairness percentage from any article, including this one. Measure the buffer depths you actually have.
Failure modes and operational guidance
- No pacing in effect. A qdisc other than
fqon a kernel without internal pacing, or offload settings that coalesce bursts, makes BBR bursty. Checktc qdisc showand the kernel version. - High retransmission rates in shallow buffers. v1 largely ignores loss as a congestion signal, apart from recovery, ending a probe-up phase and policer detection. On lossy shallow-buffer paths, expect retransmit rates well above CUBIC's. Watch
TcpRetransSegsinnstatbefore and after rollout. - Unfairness to loss-based flows. Shared bottlenecks you do not control, such as office uplinks, may see BBR flows crowd out CUBIC flows. Many teams enable BBR on CDN and server egress, where it helps most, and leave internal bulk links alone. See ECN for the signal BBRv3 uses to improve this.
- min_rtt stuck high. On paths with a persistent standing queue, BBR may never observe the true minimum, which makes its BDP estimate and its queue too large. PROBE_RTT is designed to fix this but needs every flow to drain together.
- Receiver-limited flows. A small receive window caps throughput whatever BBR estimates; check window scaling with this guide.
- Mixed fleets. Congestion control is per sender. Switching servers to BBR changes download behaviour only; uploads still use the clients' algorithms.
To roll out, set net.ipv4.tcp_congestion_control=bbr and net.core.default_qdisc=fq on a canary group, or choose it per socket with setsockopt(fd, IPPROTO_TCP, TCP_CONGESTION, "bbr", 3) where only some traffic should change. Compare throughput percentiles, retransmissions and tail latency per region against a control group for at least a week, because the benefit is uneven: large on long, lossy paths and small inside a data center. The TCP fundamentals article covers the recovery machinery BBR sits beside.
What to do next
- Check which algorithm and qdisc your senders use:
sysctl net.ipv4.tcp_congestion_control net.core.default_qdiscandtc qdisc show. - Build the three-namespace lab and reproduce STARTUP, DRAIN, PROBE_BW and PROBE_RTT in
ss -tinoutput. - Run the BBR versus CUBIC test at the buffer depths your real bottlenecks have, and record throughput and retransmissions for each.
- Read
bbr_main(),bbr_update_bw()andtcp_rate_gen()in your kernel's source, alongside the constants table above. - Canary BBR on egress-heavy servers with
fq, and compare against a control group on throughput percentiles, retransmissions and tail latency. - If loss or fairness becomes a problem, evaluate BBRv3 from the
google/bbrv3 branch in the lab first, knowing it means running an out-of-tree kernel patch.