Quantizing a model is quick. Knowing whether the quantized model is good enough is the hard part. The usual report gives a perplexity number and a benchmark score that both look close to the original, and teams ship on that basis. Then users find that the model has become worse at long documents, or at JSON, or at one language. Average scores hide those changes, because a model can get some answers newly wrong and others newly right and keep the same total.
This article treats quantization evaluation as a measurement problem with a specific question: how far has this model moved from its own baseline, and does the movement matter for the workload? It covers how to pin the setup so the comparison is fair, which metrics detect change (perplexity, KL divergence, top-1 agreement and flips), how to tell a real difference from noise, and how to arrange the checks into tiers so the expensive ones only run on candidates that deserve them. For the quantization methods themselves, see GPTQ and AWQ.
The question is relative, not absolute
A general model evaluation asks how good a model is. A quantization evaluation asks something narrower and easier to answer: how different is the quantized model from the one it was made from? The baseline is the same weights at BF16 or FP16. Every metric should be a comparison with it, on identical inputs, so the large uncertainty about the model's absolute quality cancels out.
Framing it this way has a practical consequence. You can use metrics that need no labels at all, such as the distance between the two models' output distributions, and run them on text from your own traffic. Labelled benchmarks are still useful, but they become one tier among several. Broader harness design is covered in LLM evaluation and small model evaluation. This page is about the relative comparison.
The evaluation pipeline
Pin the setup before measuring anything
Small differences in setup produce differences as large as the quantization effect, so fix everything except the weights:
- Same tokenizer and chat template. A converted model that silently uses a different tokenizer file or template changes every metric. Compare token IDs for a few prompts before you start.
- Same inputs, same order, same truncation. Store the evaluation set as token IDs, not raw text, so preprocessing cannot drift.
- Deterministic decoding. Greedy decoding for generation tests, fixed seeds and a fixed batch size. Batch size can change numerics on GPU kernels, so keep it constant across both models.
- The deployed kernels. Simulated (fake) quantization, where weights are rounded but computed in floating point, is useful during research but does not reproduce the real kernel's accumulation order, activation handling or KV cache format. Evaluate the artifact and runtime you will actually serve.
- Record everything. Model hashes, quantization config, calibration set (see calibration), runtime version, hardware and the exact command line.
Perplexity, done so it is comparable
Perplexity is the exponential of the mean negative log-likelihood of each next token: PPL = exp(-mean(log p(x_t | x_<t))). It needs no labels and reacts to quantization damage, so it is the usual first check. But it is easy to compute in ways that are not comparable.
The value depends on the tokenizer (perplexities across different tokenizers are not comparable at all), on the text, on context length and on how long texts are split into windows. With non-overlapping windows, tokens early in each window have little context and inflate the number. A sliding window with a stride shorter than the context gives every scored token more context and lowers it. Use the same choices for both models and report them alongside the result.
Perplexity is also a mean over tokens, which makes it insensitive in the way that matters most. A quantized model can match the baseline closely on common tokens and be badly wrong on the rare ones, which carry the facts, numbers and names. The mean barely moves. Treat a small perplexity increase as necessary but not sufficient.
KL divergence and top-1 agreement
A stronger logit-level test compares the full next-token distributions. For each position, compute the KL divergence from the baseline's distribution to the quantized model's: KL(P_base || P_quant) = sum_v P_base(v) * (log P_base(v) - log P_quant(v)). It is zero only when the distributions match and grows as they diverge, whether or not the change happens to raise or lower the probability of the actual next token. Report the mean and also a high quantile such as the 99th percentile, because damage is concentrated in a minority of positions. Alongside it, report top-1 agreement: the fraction of positions where both models' most likely token is the same, which relates directly to greedy decoding.
llama.cpp's llama-perplexity tool supports this as a two-step workflow, documented in its README. It first runs the baseline with --kl-divergence-base pointing at a file where it saves the logits. It then runs the quantized model with the same file and --kl-divergence, and reports perplexity for both models, KLD statistics and token agreement. The logits file is large, many gigabytes for a full Wikitext-2 pass with a large vocabulary, so budget disk space or use a smaller text sample from your own domain.
Task accuracy and flips
Labelled tasks show whether the change affects answers. The trap is comparing only totals. Dutta and colleagues, in "Accuracy is Not All You Need" (arXiv 2407.09141, NeurIPS 2024), measured flips: questions whose answer changes from correct to incorrect or the reverse between baseline and compressed model. They found substantial flip rates for many quantization schemes even where accuracy was close to the baseline's, and found that flips correlate well with KL divergence. Their conclusion was that accuracy and perplexity are necessary but not sufficient, and that distance metrics should be reported alongside them.
Flips matter for products because users experience individual answers, not averages. A model with the same accuracy but 5% of answers changed will break prompts, tests and user expectations that depended on the old behaviour. To count flips you need per-example results for both models in the same order. The lm-evaluation-harness writes them when run with --log_samples and an --output_path.
lm_eval --model hf \
--model_args pretrained=/models/base,dtype=bfloat16 \
--tasks mmlu --batch_size 8 --log_samples --output_path out/base
lm_eval --model hf \
--model_args pretrained=/models/w4-candidate \
--tasks mmlu --batch_size 8 --log_samples --output_path out/w4
Is the difference real? Paired statistics
The two models answered the same questions, so the comparison is paired, and the paired tests are much more sensitive than comparing two independent accuracies. Only the discordant questions carry information: b, right for the baseline and wrong for the quantized model, and c, the reverse. McNemar's test asks whether b and c are consistent with an even split. A paired bootstrap, resampling questions with replacement and recomputing the accuracy difference each time, gives a confidence interval that is easier to explain to a reviewer.
import math
import numpy as np
def log_softmax(z):
z = z - z.max(axis=-1, keepdims=True)
return z - np.log(np.exp(z).sum(axis=-1, keepdims=True))
def token_metrics(base_logits, q_logits, targets):
"""Logits are [tokens, vocab] from both models on the same inputs."""
lp_b, lp_q = log_softmax(base_logits), log_softmax(q_logits)
kld = (np.exp(lp_b) * (lp_b - lp_q)).sum(axis=-1) # KL(base || quant)
idx = np.arange(len(targets))
return {
"ppl_base": float(np.exp(-lp_b[idx, targets].mean())),
"ppl_quant": float(np.exp(-lp_q[idx, targets].mean())),
"kld_mean": float(kld.mean()),
"kld_p99": float(np.quantile(kld, 0.99)),
"top1_agree": float((lp_b.argmax(-1) == lp_q.argmax(-1)).mean()),
}
def flips(base_ok, quant_ok):
"""Boolean arrays, one entry per question, in the same order."""
b = int((base_ok & ~quant_ok).sum()) # right -> wrong
c = int((~base_ok & quant_ok).sum()) # wrong -> right
return b, c, (b + c) / len(base_ok)
def mcnemar_exact(b, c):
"""Two-sided exact McNemar: under H0 discordant pairs split 50/50."""
n, k = b + c, min(b, c)
return min(1.0, 2 * sum(math.comb(n, i) for i in range(k + 1)) / 2 ** n)
def paired_bootstrap(base_ok, quant_ok, iters=10000, seed=0):
rng = np.random.default_rng(seed)
d = quant_ok.astype(float) - base_ok.astype(float)
means = [d[rng.integers(0, len(d), len(d))].mean() for _ in range(iters)]
return np.quantile(means, [0.025, 0.975])Worked example. On 1,000 multiple-choice questions the baseline gets 687 right and a 4-bit candidate gets 675, a drop of 1.2 points. Underneath, 31 questions went from right to wrong and 19 from wrong to right: 50 flips, 5% of the set. McNemar's exact test on 31 against 19 gives p of about 0.12, and the paired bootstrap interval for the difference runs from about -2.6 to +0.2 points. So the accuracy drop is not clearly distinguishable from noise at this sample size, but one answer in twenty changed. Both statements belong in the report. Whether 5% flips is acceptable depends on the product, and the threshold should be set before the run, not after seeing the number.
Generation, long context and the edges
Logit metrics and multiple-choice tasks score single steps. Generation compounds errors: one changed token early in a greedy decode sends the rest of the output down a different path. Add checks that exercise the behaviours quantization tends to damage:
- Long context. Errors can grow with sequence length, especially if the KV cache is also quantized (see KV cache quantization). Measure KLD and retrieval accuracy by position bucket, not as one number.
- Structured output. Parse rate for JSON or tool calls on your real schemas. Formatting failures are cheap to detect and expensive in production.
- Code and arithmetic. Exact-answer tasks such as unit-tested code are sensitive to small numeric errors.
- Languages and rare tokens. Low-resource languages and rare vocabulary often degrade first. Slice every metric by language.
- Safety behaviour. Re-run refusal and policy tests. Quantization can change behaviour that was tuned on the baseline.
- Divergence point. For greedy outputs on the same prompts, record the first token where the two models differ. A falling median divergence position is an early warning.
Measure the gain you quantized for
Quantization trades quality for memory, throughput or latency, so the evaluation must measure the gain as carefully as the loss. Measure on the target hardware and runtime: weight and KV cache memory at the target context length, time to first token, decode tokens per second at the batch sizes you serve, and throughput at your latency target. A 4-bit model whose kernels are slow at large batch sizes may give no throughput gain on a busy server even though it fits in less memory. Quality and performance results together make the decision. Report them in the same table so neither is read alone.
Failure modes
| Failure | Effect | Prevention |
|---|---|---|
| Perplexity computed with different context or stride | Spurious gain or loss | Fix and record window, stride and text |
| Fake-quant evaluated, real kernel shipped | Shipped model differs from the tested one | Evaluate the deployed artifact and runtime |
| Only total accuracy compared | Flips hidden behind equal totals | Per-example logs, flips, McNemar |
| Calibration text overlaps eval text | Optimistic results | Disjoint calibration and evaluation sets |
| Mean KLD only | Rare severe errors hidden | Report p99 and per-slice KLD |
| Thresholds chosen after seeing results | Every candidate passes | Write the gate before the run |
| Short prompts only | Long-context damage missed | Position-bucketed metrics |
Trade-offs in the evaluation itself
Logit metrics are cheap, label-free and sensitive, but storing baseline logits costs disk, and a KLD threshold has no direct product meaning until you have calibrated it against task results. Task suites are interpretable but noisy at small sizes and may not resemble your traffic. Generation tests are closest to reality and the slowest to run. The tiered pipeline spends the budget where it is informative. Screen every candidate with tier 1 on a few hundred thousand tokens of in-domain text, confirm survivors with tier 2, and run tier 3 only on the one or two candidates you intend to ship.