Brotli is the compression format behind the br content encoding that almost every web response can now use. It was designed at Google by Jyrki Alakuijala and Zoltan Szabadka and specified in RFC 7932 in 2016. Like gzip it combines LZ77 string matching with Huffman-style prefix codes, but it adds three ideas that matter on the web: second-order context modelling for literals, a built-in static dictionary of common words and markup, and much larger windows. On this site's own HTML, measured below, the highest Brotli setting produces files 54 percent smaller than gzip -9.
This article explains the format from the bits up, so the knobs make sense: what a meta-block and a command are, how contexts pick a prefix code, how the dictionary is addressed, and what the quality levels actually buy, measured. It then covers serving Brotli safely: precompression, negotiation, caching headers, and the failure modes that bite in production. If prefix codes are new to you, read Huffman coding first.
The stream: windows, meta-blocks and commands
A Brotli stream starts with a header holding WBITS, a value from 10 to 24. The sliding window is (1 << WBITS) - 16 bytes, so from 1 KiB minus 16 bytes to 16 MiB minus 16 bytes. The decoder must keep that much recent output, which is the memory cost the encoder imposes on every client. The command-line tool and libraries default to lgwin 22, a 4 MiB window. A separate "large window" extension goes beyond 24 bits, but it is not part of RFC 7932 and browsers do not accept it.
After the header come meta-blocks, each decompressing to between 0 and 16,777,216 bytes. A meta-block can be stored uncompressed, which bounds the worst-case expansion on incompressible data. A compressed meta-block holds its own prefix codes and a sequence of commands. Each command is:
insert-and-copy length code, literal, literal, ..., literal, distance codeThe insert length says how many literal bytes follow, and the copy length says how many bytes to copy from an earlier position. gzip's DEFLATE has no insert length: each literal is its own symbol in an alphabet shared with match lengths. Brotli instead codes the insert and copy lengths jointly, in one symbol from a 704-symbol alphabet refined by extra bits, and codes the literals with separate codes. A combined symbol is cheaper because insert and copy lengths are correlated in real data.
Distances have a twist. Codes 0 to 15 do not carry a number at all: they refer to a ring buffer of the four most recent distances, plus small offsets such as last distance minus 1 or second-to-last plus 3. The ring starts at 16, 15, 11 and 4 at the beginning of the stream and is not reset at meta-block boundaries. Structured data such as tables, JSON arrays and repeated markup reuses the same distances constantly, so these codes are very cheap.
Block types and context modelling
gzip uses one Huffman code for all literals in a block. Brotli picks the code for each literal using two pieces of state. The first is the current block type. Each of the three symbol categories, literals, insert-and-copy codes and distances, is cut into runs, each tagged with one of up to 256 block types, and the stream contains explicit switch commands. An HTML file with inline JavaScript can therefore use one literal type for markup and another for script.
The second is the context ID, computed from the previous two output bytes p1 and p2. Each literal block type declares one of four context modes, all specified in RFC 7932:
def literal_context(mode, p1, p2, LUT0, LUT1, LUT2):
"""Context ID 0..63 from the two previous bytes (RFC 7932, section 7.1)."""
if mode == "LSB6": # binary data keyed on low bits
return p1 & 0x3F
if mode == "MSB6": # binary data keyed on high bits
return p1 >> 2
if mode == "UTF8": # text: classes such as letter, digit, space
return LUT0[p1] | LUT1[p2]
if mode == "Signed": # sequences of signed integers
return (LUT2[p1] << 3) | LUT2[p2]
def distance_context(copy_length):
return min(copy_length, 5) - 2 # 0, 1, 2, 3 for lengths 2, 3, 4, 5+A context map then sends each (block type, context ID) pair to one of the meta-block's prefix codes: 64 entries per literal block type and 4 per distance block type. Many contexts usually share a code, and the map itself is compressed with run-length coding and an optional move-to-front transform. The effect is second-order modelling at prefix-code speed: after a space, letters are likely; after a digit, digits are likely; after an opening angle bracket, a tag name is likely. Each situation gets its own code lengths without paying for arithmetic coding. For the theory of why this helps, see entropy coding, and for the coder Brotli chose not to use, arithmetic coding.
The static dictionary
Brotli ships a fixed dictionary of 122,784 bytes inside every encoder and decoder. It holds words of 4 to 24 bytes: common words in several languages and fragments of web markup. Each word can be used in 121 transformed forms, for example with a capitalised first letter, with a trailing space, or followed by punctuation or markup characters. A copy command points into the dictionary by using a distance larger than the maximum allowed distance at that position. The excess selects the word and the transform, and the copy length gives the word length.
The dictionary matters most where LZ77 has nothing to match against yet: the first few hundred bytes of every response, and small API payloads. On a 167-byte HTML snippet, gzip -9 produced 144 bytes and Brotli at quality 11 produced 84. On a 159-byte JSON object, gzip -9 gave 142 bytes and Brotli 104. For large files the window does most of the work and the dictionary's share shrinks.
Measured: quality levels against gzip
The benchmark used 200 of this site's HTML articles concatenated into one 6,088,481-byte file, compressed with the brotli 1.1.0 command-line tool and with Python's gzip module on a Windows laptop. Pages sharing one template are highly repetitive, so these ratios are better than you will see on unrelated files. Use the relative ordering. Speeds are rough: Brotli ran as a subprocess with its startup time subtracted.
| Setting | Output bytes | Ratio | Compress speed |
|---|---|---|---|
| gzip -6 | 1,356,526 | 22.3% | 21 MB/s |
| gzip -9 | 1,351,978 | 22.2% | 19 MB/s |
| br q1 | 1,013,427 | 16.6% | 217 MB/s |
| br q4 | 814,396 | 13.4% | 67 MB/s |
| br q5 | 758,830 | 12.5% | 42 MB/s |
| br q6 | 738,751 | 12.1% | 35 MB/s |
| br q9 | 697,191 | 11.5% | 13 MB/s |
| br q11 | 625,656 | 10.3% | 0.3 MB/s |
Three things stand out. Even quality 1 beats gzip -9 here, at ten times the speed. Quality 4 to 6 gives most of the gain at speeds fine for compressing responses on the fly. At quality 10 and 11 the reference encoder switches to a far more expensive optimal parse; quality 11 buys another 10 percent at less than a fortieth of the speed of quality 9, which is only acceptable when you compress once and serve many times. Window size also matters at quality 11: lgwin 16 gave 733,910 bytes, 20 gave 649,406 and 24 gave 624,076. Decompression through the same CLI pipe ran above 90 MB/s, a lower bound for the library; decoding cost is largely independent of the quality level used.
Serving Brotli
The client lists br in Accept-Encoding; the server picks an encoding and names it in Content-Encoding. Every response whose body depends on that negotiation must carry Vary: Accept-Encoding, or a cache may hand Brotli bytes to a client that cannot decode them. The standard deployment has two paths:
- Static assets: compress at build time with quality 11 and serve the stored file. CPU cost is paid once per deploy.
- Dynamic responses: compress on the fly at quality 4 to 6, or let the CDN do it at the edge.
A build step that precompresses and verifies looks like this. It uses the brotli command-line tool, decodes every output to prove it round-trips, and skips files where the saving is too small to justify a second variant:
import subprocess, sys
from pathlib import Path
TYPES = {".html", ".css", ".js", ".svg", ".json", ".xml", ".txt", ".wasm"}
MIN_BYTES = 1024 # below this, headers dominate
def brotli(data, *args):
return subprocess.run(["brotli", *args, "-c"], input=data,
capture_output=True, check=True).stdout
def precompress(root):
kept = skipped = 0
for f in sorted(Path(root).rglob("*")):
if f.suffix not in TYPES or not f.is_file():
continue
raw = f.read_bytes()
if len(raw) < MIN_BYTES:
skipped += 1
continue
packed = brotli(raw, "-q", "11", "-w", "22")
if brotli(packed, "-d") != raw:
raise RuntimeError(f"round trip failed for {f}")
if len(packed) >= len(raw) * 0.95:
skipped += 1
continue
f.with_name(f.name + ".br").write_bytes(packed)
kept += 1
print(f"wrote {kept} .br files, skipped {skipped}")
if __name__ == "__main__":
precompress(sys.argv[1])With nginx and the third-party ngx_brotli module, brotli_static on; serves those files, and brotli on; brotli_comp_level 5; compresses everything else on the fly. Restrict brotli_types to text formats.
One newer option is worth knowing. RFC 9842, Compression Dictionary Transport, published in September 2025, lets a previously downloaded response act as an external dictionary for a later one, using the dcb encoding for Brotli. The client announces a stored dictionary's SHA-256 hash in Available-Dictionary, and the server compresses against it. For a JavaScript bundle that changes a little each deploy, the new version becomes a small delta. Browser and CDN support is still uneven, so check before relying on it.
Failure modes
- Missing Vary. A shared cache stores the br variant and serves it to a client that never asked for it. The page renders as binary garbage.
- Double compression. An origin sends br and a proxy compresses again, or a precompressed file is served without
Content-Encoding. Check response headers in an end-to-end test. - Quality 11 on dynamic traffic. At 0.3 MB/s in the measurement above, one busy endpoint can saturate a CPU. Cap dynamic compression at 5 or 6.
- Compressing already compressed formats. JPEG, PNG, WOFF2, video and zip files gain nothing and cost CPU. Allow-list text types.
- Secrets next to attacker-controlled text. Compression over TLS leaks information through response length, the BREACH class of attacks. Do not compress responses that mix secrets such as CSRF tokens with reflected input, or mask the tokens per response.
- Decompression bombs. Services that accept br request bodies must cap decoded size, because a small input can expand enormously.
Trade-offs
| Concern | gzip | Brotli | Zstandard |
|---|---|---|---|
| Entropy coder | Huffman | Huffman with context maps | Huffman for literals, FSE (ANS) for sequences |
| Max window | 32 KiB | 16 MiB (RFC 7932) | configurable, large |
| Built-in dictionary | none | 122,784 bytes | none, trainable |
| Best fit | universal fallback | static web assets, small text | fast general purpose, logs, storage |
Brotli's strength is ratio at the top quality levels and on small text, with decoding cost that does not grow with quality. Its weakness is encode speed at those levels. Zstandard is generally faster to decompress and is preferred for storage and internal transport; check your own Accept-Encoding logs before betting on it for browser traffic. Keep gzip as the fallback for any client that advertises neither.
What to do next
- Measure your own assets: run brotli at quality 5 and 11 and gzip -9 on your ten largest text files and record the bytes.
- Add a precompression build step like the one above, with the round-trip check.
- Set dynamic compression to quality 4 to 6 and allow-list text content types only.
- Confirm
Vary: Accept-Encodingon every negotiated response, through the CDN as well as at the origin. - Audit endpoints that reflect user input next to secrets before turning compression on.
- Read RFC 7932 sections 2 and 7 with this article beside you; the format is small enough to follow in an afternoon.