Adaptive bitrate (ABR) logic is usually explained as a family tree of algorithms. That is useful for choosing one, but it hides what actually happens inside a player: every few seconds, with incomplete information, a small piece of code picks one rung of the encoding ladder, and a bad pick either wastes bandwidth or empties the buffer. The best way to understand that code is to watch it make twenty decisions in a row.
This page builds a compact hybrid controller, runs it against a network trace that climbs, collapses and recovers, and walks through the resulting table row by row. Every number in the trace was produced by the simulator printed below. For the comparison of algorithm families (throughput rules, BBA, BOLA, MPC, learned policies) see adaptive bitrate selection algorithms; here we stay inside one controller.
The decision, stated precisely
Before requesting segment i, the player knows three things: a history of how fast past segments downloaded, how many seconds of media are already buffered, and which rung it used last. It must output one rung. After the request starts it learns one more thing, the speed of the download in progress, and it may change its mind by abandoning the request.
Two failure costs pull in opposite directions. Pick too high and the download takes longer than the buffer can cover, so playback stalls; a rebuffer is the most damaging event for viewer retention. Pick too low and the viewer sees a softer picture than the network could deliver. Oscillation is a third cost: a picture that changes sharpness every eight seconds looks broken.
| Rung | Bitrate | Typical resolution | Bytes per 4 s segment |
|---|---|---|---|
| 0 | 300 kbps | 360p | 150 KB |
| 1 | 750 kbps | 480p | 375 KB |
| 2 | 1,500 kbps | 720p | 750 KB |
| 3 | 3,000 kbps | 1080p | 1.5 MB |
| 4 | 6,000 kbps | 1080p high | 3 MB |
That is the ladder used throughout. Segments are 4 seconds and treated as constant bitrate, which real encodes are not; we return to that in the failure modes. Ladder design itself is covered in per-title encoding.
Step 1: turning downloads into an estimate
The raw measurement is simple: bits received divided by seconds taken. Using it directly is a mistake, because single samples are noisy. So the controller smooths samples with an exponentially weighted moving average, and it weights by time rather than by sample: a sample that took 6 seconds should move the average more than one that took 0.3 seconds.
The weight of the old value after a sample of duration d is alpha = 0.5 ** (d / half_life). With a 3-second half-life, a 3-second download replaces half of the old estimate. Two averages run in parallel, a fast one (3 s) that reacts to drops and a slow one (9 s) that remembers the longer past, and the estimate is the lower of the two. Taking the minimum is a deliberate asymmetry: drops are believed quickly, rises slowly.
Both averages start at zero, so they are divided by the total weight seen so far (zero-bias correction), otherwise the first estimate would be a fraction of the true rate. And before any sample exists the controller needs a default. These constants are modelled on the defaults hls.js ships at the time of writing: abrEwmaFastVoD: 3.0, abrEwmaSlowVoD: 9.0, abrEwmaDefaultEstimate: 500000, with bandwidth factors abrBandWidthFactor: 0.95 and abrBandWidthUpFactor: 0.7. The decision rule below is our own and is not how hls.js composes them.
Step 2: buffer zones and the choice rule
The estimate says what the network can do; the buffer says how much risk the player can afford. The controller splits the buffer into three zones and uses a different safety factor in each.
- Danger, under 8 s: choose the highest rung at or below 0.85 times the estimate, and never above the current rung. Stepping up while the buffer is thin is how stalls start.
- Steady, 8 to 20 s: target the highest rung at or below 0.7 times the estimate. The large margin absorbs estimator error.
- Comfort, 20 s and above: relax to 0.95 times the estimate, since the buffer itself can absorb a slow segment.
- Climb one rung at a time: if the target is above the current rung, move up by exactly one. Drops may skip rungs; climbs may not.
Step 3: abandoning a download in flight
The estimate is always about the past. When the network collapses mid-download, the request already sent may now need 30 seconds to finish. Waiting for it means draining the buffer for a segment you could have fetched at a lower rung. So after one second the watchdog projects the finish time from the measured in-flight rate. If the download would end less than 2 seconds before the buffer runs dry, it is cancelled, the second already spent is counted against the buffer and fed to the estimator, and the segment is fetched again at the highest rung at or below 0.85 times the measured rate. Abandonment wastes the bytes already received, so the threshold is conservative.
The simulator
The whole controller fits in one file. The network is a list of per-segment throughputs in kbps, download time is segment size divided by throughput, and the buffer is capped at 30 seconds: when it holds more than 26 seconds the player idles until there is room for one more segment. Playback starts once 8 seconds are buffered.
import math
LADDER = [300, 750, 1500, 3000, 6000] # kbps
SEG = 4.0 # seconds per segment
LOW, HIGH, MAX_BUF = 8.0, 20.0, 30.0 # buffer thresholds, seconds
class DualEwma:
"""Two time-weighted EWMAs (half-lives 3 s and 9 s); the estimate is the lower one."""
def __init__(self, default_kbps=500):
self.default = default_kbps
self.fast = self.slow = 0.0
self.wf = self.ws = 0.0 # total weight seen, for zero-bias correction
def sample(self, seconds, kbps):
for hl, attr, wattr in ((3.0, "fast", "wf"), (9.0, "slow", "ws")):
alpha = 0.5 ** (seconds / hl)
setattr(self, attr, kbps * (1 - alpha) + alpha * getattr(self, attr))
setattr(self, wattr, (1 - alpha) + alpha * getattr(self, wattr))
def estimate(self):
if self.ws == 0:
return self.default
return min(self.fast / self.wf, self.slow / self.ws)
def highest_fitting(kbps):
fits = [i for i, b in enumerate(LADDER) if b <= kbps]
return fits[-1] if fits else 0
def choose(est, buf, last):
if buf < LOW: # danger zone: rate rule, never step up
return min(highest_fitting(0.85 * est), last)
target = highest_fitting(0.95 * est if buf >= HIGH else 0.7 * est)
if target > last:
return last + 1 # climb one rung at a time
return target
def run(trace_kbps):
ewma, buf, last, playing, rows = DualEwma(), 0.0, 0, False, []
for i, net in enumerate(trace_kbps):
est = ewma.estimate()
rung = choose(est, buf, last) if i else highest_fitting(0.95 * est)
dl = LADDER[rung] * SEG / net
note = ""
if playing and rung > 0 and dl > buf - 2.0: # abandonment check at t = 1 s
spent = min(1.0, dl)
ewma.sample(spent, net)
buf = max(0.0, buf - spent)
old = rung
rung = min(highest_fitting(0.85 * net), rung - 1)
dl = LADDER[rung] * SEG / net
note = f"abandon {LADDER[old]}"
stall = max(0.0, dl - buf) if playing else 0.0
buf = (max(0.0, buf - dl) if playing else buf) + SEG
ewma.sample(dl, net)
if not playing and buf >= 8.0:
playing = True
idle = max(0.0, buf - MAX_BUF + SEG) if buf > MAX_BUF - SEG else 0.0
rows.append((i, net, round(est), LADDER[rung], round(dl, 2), round(stall, 2), round(buf, 1), note))
if idle:
buf -= idle # wait for room before the next request
last = rung
return rows
TRACE = [4000, 4500, 5000, 6000, 8000, 9000, 9000, 9000, 9000, 9000,
800, 1000, 1000, 1200, 2500, 4000, 5000, 5000, 5000, 5000]
if __name__ == "__main__":
for r in run(TRACE):
print(r)
The trace, segment by segment
| Seg | Link kbps | Estimate | Chosen | Download s | Buffer after s | What happened |
|---|---|---|---|---|---|---|
| 0 | 4,000 | 500 | 300 | 0.30 | 4.0 | No samples: default estimate |
| 1 | 4,500 | 4,000 | 300 | 0.27 | 8.0 | Danger zone, no climb; playback starts |
| 2 | 5,000 | 4,238 | 750 | 0.60 | 11.4 | Steady zone, target 1,500, climb one |
| 3 | 6,000 | 4,638 | 1,500 | 1.00 | 14.4 | Target 3,000, climb one |
| 4 | 8,000 | 5,295 | 3,000 | 1.50 | 16.9 | 0.7 x 5,295 fits 3,000 |
| 5 | 9,000 | 6,495 | 3,000 | 1.33 | 19.6 | 0.7 x 6,495 = 4,546: hold |
| 6 | 9,000 | 7,260 | 3,000 | 1.33 | 22.2 | Buffer 19.6, still steady: hold |
| 7 | 9,000 | 7,700 | 6,000 | 2.67 | 23.6 | Comfort zone, 0.95 x 7,700 fits 6,000 |
| 8 | 9,000 | 8,182 | 6,000 | 2.67 | 24.9 | Hold |
| 9 | 9,000 | 8,439 | 6,000 | 2.67 | 26.2 | Hold; buffer near the cap |
| 10 | 800 | 8,594 | 300 | 1.50 | 27.5 | 6,000 abandoned after 1 s |
| 11 | 1,000 | 5,263 | 750 | 3.00 | 27.0 | Estimate still high: climbs |
| 12 | 1,000 | 3,110 | 1,500 | 6.00 | 24.0 | Overshoot: 6 s for a 4 s segment |
| 13 | 1,200 | 1,523 | 750 | 2.50 | 25.5 | Estimate catches up, drops |
| 14 | 2,500 | 1,381 | 750 | 1.20 | 28.3 | Hold |
| 15 | 4,000 | 1,652 | 1,500 | 1.50 | 28.5 | Climb one |
| 16 | 5,000 | 2,341 | 1,500 | 1.20 | 28.8 | Hold |
| 17 | 5,000 | 2,985 | 1,500 | 1.20 | 28.8 | 0.95 x 2,985 = 2,836: hold |
| 18 | 5,000 | 3,409 | 3,000 | 2.40 | 27.6 | Climb one |
| 19 | 5,000 | 3,695 | 3,000 | 2.40 | 27.6 | Hold |
Two mechanics explain rows that otherwise do not add up. When the buffer exceeds 26 seconds after a segment, the player idles before the next request, so the buffer at the start of row 10 is 26.0, not 26.2. And in row 10 the abandoned request costs one second: 26.0 minus 1 is 25.0, minus the 1.5-second refetch is 23.5, plus 4 seconds of media is 27.5.
Read the first ten rows as startup and ramp. The default estimate puts segment 0 on the bottom rung, which is safe and fast (0.3 s). Segment 1 already has a 4,000 kbps estimate but sits in the danger zone, so it may not climb. From segment 2 the one-rung rule produces a staircase, and the high rung waits until the buffer passes 20 seconds at segment 7. That is a 28-second ramp to the top rung on a link that could carry it from segment 4, the price of never stalling.
Rows 10 to 13 are the collapse. The watchdog saves segment 10: the planned 6,000 kbps segment would have taken 30 seconds against a 26-second buffer. But the next two rows show the weakness. The abandoned second and the 1.5-second refetch are short samples, so the slow average still remembers 9 Mbps and the estimate reads 5,263 and then 3,110. The controller climbs to 1,500 kbps on a 1,000 kbps link and spends 6 seconds on a 4-second segment. With 26 seconds of buffer this is harmless; on a live stream holding 6 seconds it would have stalled.
Rows 14 to 19 are recovery, and the asymmetry shows again: the link reaches 5,000 kbps at segment 16 but the 3,000 rung waits until segment 18,.
What the trace teaches
- Abandonment is the real stall protection. No estimator reacts within one segment to a 90% drop; the watchdog did.
- Short samples barely move a time-weighted estimator. After a collapse, feed the in-flight rate into the decision directly, or cap the next choice at the abandonment rung for a few segments. Either fix would have prevented row 12.
- The comfort threshold sets the ramp time. Lowering it from 20 to 12 seconds reaches the top rung earlier on stable links and raises the stall risk on unstable ones.
- Live and VOD need different zones. A live player near the edge cannot hold 20 seconds, so its zones must scale with its target latency; see low-latency HLS for that budget.
Failure modes in real players
- Variable bitrate segments. A 3,000 kbps rung can contain a 6,000 kbps action scene. Use per-segment sizes from the manifest or byte-range index when available, not the nominal bitrate.
- Cache hits inflate estimates. A segment from the edge cache downloads far faster than a cache miss from origin; mixing them makes the estimate swing. Ignore samples that are too small to measure, and treat very short downloads with suspicion. The cache behaviour is covered in video CDN architecture.
- Chunked transfer in low-latency modes. When segments arrive as they are encoded, download time measures the encoder, not the network. Measure only the bursts, or the estimate collapses to the media bitrate and the player can never climb.
Operating an ABR controller
Never tune ABR constants on intuition. Record throughput traces from real sessions, replay them through the simulator for the old and new constants, and compare the distributions of rebuffer ratio, average delivered bitrate, startup time and switches per minute. A change that adds 4% bitrate and 0.1 points of rebuffering is usually a loss. Then ship behind an experiment, because the traces you recorded were shaped by the old controller. The player components around this logic are covered in ABR architecture.
What to do next
- Copy the simulator, run it and confirm you get the same twenty rows.
- Change one constant at a time (comfort threshold, 0.7 factor, half-lives) and note which rows move.
- Fix the row 12 overshoot by capping choices at the abandonment rung for three segments, and rerun.
- Export 50 real throughput traces from your player telemetry and replay them through both versions.
- Add per-segment sizes from your manifests instead of nominal bitrates.
- Track rebuffer ratio, startup time, average bitrate and switches per minute on one dashboard before any change.