GGUF · BYTES BEFORE QUALITY

GGUF Quantization Chooser

Size a pure GGUF quantization tensor by tensor from a published config, add the KV cache for your context, and find the largest type that fits your memory.

Direct tool · updates as you edit

Your experiment

Start with Llama 3.2 3B at Q4_K, 4 GiB of memory and 8,192 tokens. Compare the file with the KV cache, then raise the context, change the KV cache type, try Qwen2.5-0.5B and its fallbacks, and step through the types.

Every input recomputes the result immediately; there is no animation because nothing here unfolds over time. An input outside its allowed range is rejected with a message and the previous valid result stays on screen.

Computed data

Metrics

Block sizes are the static_asserts in ggml-common.h: Q4_K stores 256 weights in 144 bytes, Q8_0 32 weights in 34 bytes. Files are sized as llama-quantize --pure writes them: every two-dimensional weight at the chosen type, one-dimensional norms and biases in F32, and llama-quant.cpp's fallback (Q4_K to Q5_0, Q6_K to Q8_0 and so on) when a row length is not a multiple of 256. The default Q4_K_M and Q4_K_S mixtures raise some tensors to Q5_K or Q6_K and are a little larger. The KV cache follows llama.cpp: Gemma 2's alternate sliding-window layers keep 4,096 tokens (src/models/gemma2.cpp, swa_full off by default) and Phi-3's window is disabled there, so its full cache is counted. Metadata and the vocabulary strings add a few megabytes that are not counted.

Bytes per block

A GGUF tensor is stored in blocks. Q8_0 packs 32 weights into 34 bytes, 8.5 bits each; the K types pack 256 weights into one super-block: Q6_K 210 bytes, Q5_K 176, Q4_K 144, Q3_K 110 and Q2_K 84. Q4_K therefore costs 4.5 bits per weight, not 4, because each super-block also stores two fp16 scales and twelve bytes of 6-bit sub-block scales and minimums. The table sizes every weight tensor of the chosen model from these block sizes.

The opening model

With Llama 3.2 3B as the Published configuration and Q4_K as the Quantization type, the file is 1,807,773,696 bytes, 1.68 GiB, for 3,212,749,824 parameters: 4.501 bits per parameter, because only the 1-D norm weights stay in F32. At 8,192 tokens its f16 KV cache adds 0.88 GiB and the reserve 1 GiB, 3.56 GiB in total, so it FITS in 4 GiB of Memory available, and the largest type that fits is Q5_K. The Reserve for compute buffers and OS slider adds the same amount to every type's bar, so raising it can lower the largest type that fits. F16 would need a 5.98 GiB file and DOES NOT FIT; Q2_K shrinks the file to 0.98 GiB.

When the context does not fit

Raise the Context length to 32,768 tokens and the KV cache grows to 3.5 GiB, larger than the Q4_K file: DOES NOT FIT, and no type fits at all. Switching the KV cache type to q8_0, 34 bytes per 32 values, cuts it to 1.86 GiB, which still does not fit with Q4_K, though Q2_K would. Phi-3-mini-4k at 8,192 tokens is BEYOND TRAINED CONTEXT; at 4,096 tokens its full-attention cache is still 1.5 GiB and only Q2_K fits.

Fallbacks change the arithmetic

K types need a row length that is a multiple of 256. Qwen2.5-0.5B has a hidden size of 896, so every matrix that reads the hidden state falls back to Q5_0, as llama-quant.cpp does; only ffn_down, with rows of 4,864, stays Q4_K. That is 145 fallback matrices, and the file is 326,810,112 bytes, 5.292 bits per parameter instead of 4.5. Gemma 2 2B, with 2,304 = 9 x 256, needs no fallback and its Q4_K file is 1.37 GiB.

Reading the tool and its limits

The bars show file plus cache plus reserve for every type against your memory line, so the largest fitting type can be read off directly. The lanes give the nominal and effective bits, the fallback count and the memory terms; the table lists each tensor. This is the pure layout; the usual Q4_K_M mixture is somewhat larger. Size says nothing about quality, which you should measure with perplexity or task evaluations on the files you will ship.