Quantizing a model is quick. Knowing whether the quantized model is good enough is the hard part. The usual report gives a perplexity number and a benchmark score that both look close to the original, and teams ship on that basis. Then users find that the model has become worse at long documents, or at JSON, or at one language. Average scores hide those changes, because a model can get some answers newly wrong and others newly right and keep the same total.

This article treats quantization evaluation as a measurement problem with a specific question: how far has this model moved from its own baseline, and does the movement matter for the workload? It covers how to pin the setup so the comparison is fair, which metrics detect change (perplexity, KL divergence, top-1 agreement and flips), how to tell a real difference from noise, and how to arrange the checks into tiers so the expensive ones only run on candidates that deserve them. For the quantization methods themselves, see GPTQ and AWQ.

Advertisement

The question is relative, not absolute

A general model evaluation asks how good a model is. A quantization evaluation asks something narrower and easier to answer: how different is the quantized model from the one it was made from? The baseline is the same weights at BF16 or FP16. Every metric should be a comparison with it, on identical inputs, so the large uncertainty about the model's absolute quality cancels out.

Framing it this way has a practical consequence. You can use metrics that need no labels at all, such as the distance between the two models' output distributions, and run them on text from your own traffic. Labelled benchmarks are still useful, but they become one tier among several. Broader harness design is covered in LLM evaluation and small model evaluation. This page is about the relative comparison.

The evaluation pipeline

Compare the quantized model with its own baseline, on identical inputs, in tiersBaseline modelBF16 or FP16, pinnedQuantized modelreal deployed kernelsFixed inputssame text, same tokenizerTier 1: logitsPPL, KLD mean and p99, top-1 agreeTier 2: tasksaccuracy, flips, McNemarTier 3: generationlong context, format, code, safetyDecision gatethresholds set before the runPerformancememory, tokens/s, latencyCheap tiers run on every candidate; expensive tiers only on candidates that pass the cheap ones
Both models see the same inputs. Cheap logit-level metrics screen candidates; task and generation tiers confirm; the performance gain is measured on the same hardware the model will run on.
Advertisement

Pin the setup before measuring anything

Small differences in setup produce differences as large as the quantization effect, so fix everything except the weights:

  • Same tokenizer and chat template. A converted model that silently uses a different tokenizer file or template changes every metric. Compare token IDs for a few prompts before you start.
  • Same inputs, same order, same truncation. Store the evaluation set as token IDs, not raw text, so preprocessing cannot drift.
  • Deterministic decoding. Greedy decoding for generation tests, fixed seeds and a fixed batch size. Batch size can change numerics on GPU kernels, so keep it constant across both models.
  • The deployed kernels. Simulated (fake) quantization, where weights are rounded but computed in floating point, is useful during research but does not reproduce the real kernel's accumulation order, activation handling or KV cache format. Evaluate the artifact and runtime you will actually serve.
  • Record everything. Model hashes, quantization config, calibration set (see calibration), runtime version, hardware and the exact command line.

Perplexity, done so it is comparable

Perplexity is the exponential of the mean negative log-likelihood of each next token: PPL = exp(-mean(log p(x_t | x_<t))). It needs no labels and reacts to quantization damage, so it is the usual first check. But it is easy to compute in ways that are not comparable.

The value depends on the tokenizer (perplexities across different tokenizers are not comparable at all), on the text, on context length and on how long texts are split into windows. With non-overlapping windows, tokens early in each window have little context and inflate the number. A sliding window with a stride shorter than the context gives every scored token more context and lowers it. Use the same choices for both models and report them alongside the result.

Perplexity is also a mean over tokens, which makes it insensitive in the way that matters most. A quantized model can match the baseline closely on common tokens and be badly wrong on the rare ones, which carry the facts, numbers and names. The mean barely moves. Treat a small perplexity increase as necessary but not sufficient.

KL divergence and top-1 agreement

A stronger logit-level test compares the full next-token distributions. For each position, compute the KL divergence from the baseline's distribution to the quantized model's: KL(P_base || P_quant) = sum_v P_base(v) * (log P_base(v) - log P_quant(v)). It is zero only when the distributions match and grows as they diverge, whether or not the change happens to raise or lower the probability of the actual next token. Report the mean and also a high quantile such as the 99th percentile, because damage is concentrated in a minority of positions. Alongside it, report top-1 agreement: the fraction of positions where both models' most likely token is the same, which relates directly to greedy decoding.

llama.cpp's llama-perplexity tool supports this as a two-step workflow, documented in its README. It first runs the baseline with --kl-divergence-base pointing at a file where it saves the logits. It then runs the quantized model with the same file and --kl-divergence, and reports perplexity for both models, KLD statistics and token agreement. The logits file is large, many gigabytes for a full Wikitext-2 pass with a large vocabulary, so budget disk space or use a smaller text sample from your own domain.

Task accuracy and flips

Labelled tasks show whether the change affects answers. The trap is comparing only totals. Dutta and colleagues, in "Accuracy is Not All You Need" (arXiv 2407.09141, NeurIPS 2024), measured flips: questions whose answer changes from correct to incorrect or the reverse between baseline and compressed model. They found substantial flip rates for many quantization schemes even where accuracy was close to the baseline's, and found that flips correlate well with KL divergence. Their conclusion was that accuracy and perplexity are necessary but not sufficient, and that distance metrics should be reported alongside them.

Flips matter for products because users experience individual answers, not averages. A model with the same accuracy but 5% of answers changed will break prompts, tests and user expectations that depended on the old behaviour. To count flips you need per-example results for both models in the same order. The lm-evaluation-harness writes them when run with --log_samples and an --output_path.

lm_eval --model hf \
  --model_args pretrained=/models/base,dtype=bfloat16 \
  --tasks mmlu --batch_size 8 --log_samples --output_path out/base
lm_eval --model hf \
  --model_args pretrained=/models/w4-candidate \
  --tasks mmlu --batch_size 8 --log_samples --output_path out/w4

Is the difference real? Paired statistics

The two models answered the same questions, so the comparison is paired, and the paired tests are much more sensitive than comparing two independent accuracies. Only the discordant questions carry information: b, right for the baseline and wrong for the quantized model, and c, the reverse. McNemar's test asks whether b and c are consistent with an even split. A paired bootstrap, resampling questions with replacement and recomputing the accuracy difference each time, gives a confidence interval that is easier to explain to a reviewer.

import math
import numpy as np

def log_softmax(z):
    z = z - z.max(axis=-1, keepdims=True)
    return z - np.log(np.exp(z).sum(axis=-1, keepdims=True))

def token_metrics(base_logits, q_logits, targets):
    """Logits are [tokens, vocab] from both models on the same inputs."""
    lp_b, lp_q = log_softmax(base_logits), log_softmax(q_logits)
    kld = (np.exp(lp_b) * (lp_b - lp_q)).sum(axis=-1)       # KL(base || quant)
    idx = np.arange(len(targets))
    return {
        "ppl_base": float(np.exp(-lp_b[idx, targets].mean())),
        "ppl_quant": float(np.exp(-lp_q[idx, targets].mean())),
        "kld_mean": float(kld.mean()),
        "kld_p99": float(np.quantile(kld, 0.99)),
        "top1_agree": float((lp_b.argmax(-1) == lp_q.argmax(-1)).mean()),
    }

def flips(base_ok, quant_ok):
    """Boolean arrays, one entry per question, in the same order."""
    b = int((base_ok & ~quant_ok).sum())    # right -> wrong
    c = int((~base_ok & quant_ok).sum())    # wrong -> right
    return b, c, (b + c) / len(base_ok)

def mcnemar_exact(b, c):
    """Two-sided exact McNemar: under H0 discordant pairs split 50/50."""
    n, k = b + c, min(b, c)
    return min(1.0, 2 * sum(math.comb(n, i) for i in range(k + 1)) / 2 ** n)

def paired_bootstrap(base_ok, quant_ok, iters=10000, seed=0):
    rng = np.random.default_rng(seed)
    d = quant_ok.astype(float) - base_ok.astype(float)
    means = [d[rng.integers(0, len(d), len(d))].mean() for _ in range(iters)]
    return np.quantile(means, [0.025, 0.975])

Worked example. On 1,000 multiple-choice questions the baseline gets 687 right and a 4-bit candidate gets 675, a drop of 1.2 points. Underneath, 31 questions went from right to wrong and 19 from wrong to right: 50 flips, 5% of the set. McNemar's exact test on 31 against 19 gives p of about 0.12, and the paired bootstrap interval for the difference runs from about -2.6 to +0.2 points. So the accuracy drop is not clearly distinguishable from noise at this sample size, but one answer in twenty changed. Both statements belong in the report. Whether 5% flips is acceptable depends on the product, and the threshold should be set before the run, not after seeing the number.

Generation, long context and the edges

Logit metrics and multiple-choice tasks score single steps. Generation compounds errors: one changed token early in a greedy decode sends the rest of the output down a different path. Add checks that exercise the behaviours quantization tends to damage:

  • Long context. Errors can grow with sequence length, especially if the KV cache is also quantized (see KV cache quantization). Measure KLD and retrieval accuracy by position bucket, not as one number.
  • Structured output. Parse rate for JSON or tool calls on your real schemas. Formatting failures are cheap to detect and expensive in production.
  • Code and arithmetic. Exact-answer tasks such as unit-tested code are sensitive to small numeric errors.
  • Languages and rare tokens. Low-resource languages and rare vocabulary often degrade first. Slice every metric by language.
  • Safety behaviour. Re-run refusal and policy tests. Quantization can change behaviour that was tuned on the baseline.
  • Divergence point. For greedy outputs on the same prompts, record the first token where the two models differ. A falling median divergence position is an early warning.

Measure the gain you quantized for

Quantization trades quality for memory, throughput or latency, so the evaluation must measure the gain as carefully as the loss. Measure on the target hardware and runtime: weight and KV cache memory at the target context length, time to first token, decode tokens per second at the batch sizes you serve, and throughput at your latency target. A 4-bit model whose kernels are slow at large batch sizes may give no throughput gain on a busy server even though it fits in less memory. Quality and performance results together make the decision. Report them in the same table so neither is read alone.

Failure modes

FailureEffectPrevention
Perplexity computed with different context or strideSpurious gain or lossFix and record window, stride and text
Fake-quant evaluated, real kernel shippedShipped model differs from the tested oneEvaluate the deployed artifact and runtime
Only total accuracy comparedFlips hidden behind equal totalsPer-example logs, flips, McNemar
Calibration text overlaps eval textOptimistic resultsDisjoint calibration and evaluation sets
Mean KLD onlyRare severe errors hiddenReport p99 and per-slice KLD
Thresholds chosen after seeing resultsEvery candidate passesWrite the gate before the run
Short prompts onlyLong-context damage missedPosition-bucketed metrics

Trade-offs in the evaluation itself

Logit metrics are cheap, label-free and sensitive, but storing baseline logits costs disk, and a KLD threshold has no direct product meaning until you have calibrated it against task results. Task suites are interpretable but noisy at small sizes and may not resemble your traffic. Generation tests are closest to reality and the slowest to run. The tiered pipeline spends the budget where it is informative. Screen every candidate with tier 1 on a few hundred thousand tokens of in-domain text, confirm survivors with tier 2, and run tier 3 only on the one or two candidates you intend to ship.

Key takeaway: <p>Evaluate a quantized model as a change to a known baseline. Pin the setup, compute perplexity comparably, add KL divergence with its tail and top-1 agreement, count flips rather than only accuracy, test differences with paired statistics, and check the behaviours quantization tends to break: long context, structure, code, languages and safety. Then weigh what you measured against the memory and speed gain on the hardware you will actually run.</p><p><strong>What to do next:</strong></p><ol><li>Write down your quality gate before testing: maximum KLD mean and p99, maximum flip rate, and must-pass slices.</li><li>Freeze an evaluation set as token IDs, including a few hundred thousand tokens of your own domain text, disjoint from calibration data.</li><li>Capture baseline logits once and compute KLD and top-1 agreement for every candidate, using the code above or llama.cpp's two-step workflow.</li><li>Run task suites with per-sample logging and report flips, McNemar p-values and bootstrap intervals alongside accuracy.</li><li>Add long-context, structured-output and per-language slices, plus a greedy divergence-point check.</li><li>Benchmark memory, time to first token and throughput on the serving hardware, and put quality and performance in one decision table.</li></ol>