GPTQ is the most widely used way to turn a 16-bit language model into a 4-bit one without retraining. Running it is easy; knowing whether the checkpoint is any good is harder, and that is where teams lose time. A GPTQ model can match the original's perplexity and still break on code or in one runtime.

This article treats GPTQ as something to audit rather than something to run. It covers what the algorithm promises, a one-layer laboratory, Hessian arithmetic for a 7B model, calibration data, per-layer diagnostics and release gates. For the full pipeline, memory planning and packing details, read the GPTQ architecture deep dive; for the derivation of the update rule, read the GPTQ math walkthrough. This page assumes the outline and focuses on verification.

Advertisement

What GPTQ promises, and what it does not

GPTQ quantizes one linear layer at a time. For a layer with weights W and the inputs X it sees on calibration text, it chooses grid-restricted weights that keep the layer's output, W times X, close in squared error. It does this column by column: round one input column, then nudge the not-yet-quantized columns of each row to absorb the rounding error, weighted by how the inputs co-vary. The co-variation lives in one matrix per layer, H = 2XXT (averaged over tokens), whose size is input features by input features.

Three consequences follow, and every later check tests one of them.

  • The guarantee is per layer and on calibration inputs. GPTQ minimises output error on the activations it saw. If production text produces different activations, the error-compensating nudges may help less, or hurt.
  • Errors compound across depth. Each layer is quantized against inputs that already passed through quantized earlier layers (tools run block by block for this reason), but nothing optimises the final logits directly. A small error in one sensitive layer can matter more than a large one elsewhere.
  • H must be informative. If a column of X is always zero, or there are fewer distinct tokens than input features, H is singular and the update leans entirely on damping. The algorithm still runs; it just stops doing anything clever.
FP16 modelreference weightsCalibration setdomain-matched tokensCapture layer inputshooks, block by blockHessian checksdead cols, rank, dampingGPTQ per layerbits, group, act-orderLayer auditcalib vs held-out errorPack + saveformat flags, g_idxpassEnd-to-end gatesPPL, KL, tasksReleasetarget runtimeFix list16-bit, 8-bit, groupsfailGPTQ guarantees a small output error per layer on calibration inputs only. Everything after the GPTQ box exists to check that this transfers.
An audit-first GPTQ workflow: capture inputs, check the Hessians, quantize, measure each layer on calibration and held-out inputs, then gate the packed model end to end in the runtime you will ship.

A one-layer laboratory

The fastest way to build intuition, and to debug a bad checkpoint later, is to run GPTQ on a single layer yourself and compare it with round-to-nearest (RTN) at the same bit width and group size. The reference implementation below is unblocked and slow, but every line maps to the algorithm. It follows the original code's structure: dead channels zeroed, optional act-order, a damped Cholesky inverse, and group scales computed from already-updated weights.

import torch

def asym_params(w, bits):
    """Per-row asymmetric scale and zero point for a slice w of shape [rows, cols]."""
    maxq = 2 ** bits - 1
    lo = torch.clamp(w.min(dim=1, keepdim=True).values, max=0.0)
    hi = torch.clamp(w.max(dim=1, keepdim=True).values, min=0.0)
    scale = (hi - lo).clamp(min=1e-8) / maxq
    zero = torch.round(-lo / scale)
    return scale, zero, maxq

def fake_quant(w, scale, zero, maxq):
    q = torch.clamp(torch.round(w / scale) + zero, 0, maxq)
    return scale * (q - zero)

def rtn_layer(W, bits=4, group=128):
    Q = torch.empty_like(W)
    for g in range(0, W.shape[1], group):
        s, z, m = asym_params(W[:, g:g + group], bits)
        Q[:, g:g + group] = fake_quant(W[:, g:g + group], s, z, m)
    return Q

def gptq_layer(W, X, bits=4, group=128, damp=0.01, act_order=True):
    """W: [out, in] float32 weights. X: [tokens, in] float32 inputs seen by this layer.
    Unblocked reference version: clear, slow, and fine for one layer."""
    W = W.clone()
    H = 2.0 * X.T @ X / X.shape[0]
    dead = torch.diag(H) == 0            # input channels that were never active
    H[dead, dead] = 1.0
    W[:, dead] = 0.0
    perm = torch.argsort(torch.diag(H), descending=True) if act_order else torch.arange(W.shape[1])
    W, H = W[:, perm], H[perm][:, perm]
    H += damp * torch.mean(torch.diag(H)) * torch.eye(H.shape[0])
    L = torch.linalg.cholesky(H)         # raises if H is not positive definite
    U = torch.linalg.cholesky(torch.cholesky_inverse(L), upper=True)
    Q = torch.zeros_like(W)
    for j in range(W.shape[1]):
        if j % group == 0:               # group params from the already-updated weights
            s, z, m = asym_params(W[:, j:j + group], bits)
        q = fake_quant(W[:, j:j + 1], s, z, m)
        Q[:, j:j + 1] = q
        err = (W[:, j] - q[:, 0]) / U[j, j]
        W[:, j + 1:] -= err[:, None] * U[j, j + 1:][None, :]
    return Q[:, torch.argsort(perm)]

def rel_out_err(W, Wq, X):
    ref = X @ W.T
    return ((X @ Wq.T - ref).norm() / ref.norm()).item()

# X_cal and X_held: inputs captured with a forward pre-hook on one nn.Linear,
# from calibration text and from text the quantizer never saw.
# W = layer.weight.float()
# for name, Wq in [("rtn", rtn_layer(W)), ("gptq", gptq_layer(W, X_cal))]:
#     print(name, rel_out_err(W, Wq, X_cal), rel_out_err(W, Wq, X_held))

Capture X with a forward pre-hook on one nn.Linear, from a few hundred calibration sequences and, separately, from text the quantizer never sees. Then read the four numbers it prints. GPTQ's error on calibration inputs should be clearly below RTN's; if it is not, the inputs are wrong (the hook captured the wrong tensor, or the activations are all padding). GPTQ's error on held-out inputs should also be below RTN's but somewhat above its own calibration error. A large gap between the two is overfitting to the calibration set, the most useful early warning the lab gives.

Two experiments are worth running once. Toggle act-order: quantizing high-activation columns first leaves more columns to absorb their error, which usually helps layers with outlier channels. And vary the damping fraction from 0.001 to 0.1: too little and the Cholesky step can fail or amplify noise, too much and GPTQ drifts back towards RTN because the off-diagonal structure of H is drowned out.

Advertisement

Worked example: Hessian size and rank for a 7B model

Take the Llama 2 7B shape: hidden size 4,096, MLP intermediate size 11,008, 32 transformer blocks, a 32,000-token vocabulary. Inside each block the linear layers share inputs in four groups, so there are four distinct Hessians per block, not seven.

Input shared byInput featuresH size in FP32
q_proj, k_proj, v_proj4,0964,0962 x 4 B = 64 MiB
o_proj4,09664 MiB
gate_proj, up_proj4,09664 MiB
down_proj11,00811,0082 x 4 B = about 462 MiB

That is roughly 654 MiB of Hessians per block, small next to the activations a tool holds, but the rank arithmetic matters more than the memory. The rank of H is at most the number of calibration tokens. For down_proj, H is 11,008 by 11,008, so it cannot be full rank with fewer than 11,008 tokens, and in practice you want many times that, because consecutive tokens in a sequence are correlated. A common default of 128 sequences of 2,048 tokens gives 262,144 tokens, which is comfortable for a dense model.

Now consider a mixture-of-experts model. Each expert's down projection sees only the tokens routed to it. With 64 experts and top-2 routing, an average expert sees about 262,144 x 2 / 64 = 8,192 tokens, and rare experts far fewer. Their Hessians are rank-deficient, damping dominates and GPTQ behaves much like RTN, so MoE models need far more calibration tokens or forced routing, and per-expert audits.

The output size is easy to predict. The linear layers hold 32 x (4 x 4,0962 + 3 x 4,096 x 11,008), about 6.48 billion weights. At 4 bits with group size 128 and a 16-bit scale plus 4-bit zero point per group, each weight costs 4 + 20/128, about 4.16 bits, so about 3.36 GB. The embedding and output head stay in 16-bit at about 0.52 GB, for roughly 3.9 GB in total against about 13.5 GB in FP16. A much larger checkpoint means a layer was skipped.

Calibration data is a design decision

Because the guarantee only holds on calibration inputs, the calibration set is part of the model. Four properties matter.

  • Domain. Use text that looks like production traffic. A model that will write code should be calibrated with code; a multilingual model with every language it must serve.
  • Format. Apply the same chat template, system prompt and special tokens the runtime will send. Templated and untemplated text produce different activations, especially in the first layers.
  • Length. Use sequences at least as long as typical requests. Positional effects and attention sinks mean short snippets under-sample the activations seen deep into a long context.
  • Count and diversity. Enough tokens to make every Hessian well conditioned (see the rank arithmetic above), drawn from many documents rather than a few long ones.

Keep a held-out set from the same distribution and never calibrate on it. It feeds both the per-layer held-out error in the lab and the end-to-end gates below. For a general treatment of calibration across methods, see calibration for quantization.

Per-layer diagnostics

Run the lab's relative output error over every quantized layer of the real model, on calibration and held-out inputs, and plot it by depth and layer type. Healthy checkpoints show a smooth profile per layer type, often higher in down_proj and the first and last blocks; problems show up as spikes.

SymptomLikely causeFix
Cholesky fails on one layerH not positive definite: dead or duplicated channels, too few tokensRaise damping for that layer, add calibration tokens, check that dead channels are handled
One layer's error is 10x its neighboursOutlier activation channels or heavy-tailed weightsKeep it in 16-bit or 8-bit, use a smaller group size, enable act-order
Held-out error far above calibration errorCalibration set too small or off-domainMore and more representative calibration data
Rare MoE experts near RTN errorToo few routed tokens; rank-deficient HMore tokens, forced routing during calibration, or 8-bit experts
Fine in Python, garbage in the serverFormat mismatch: symmetric versus asymmetric, act-order flag, group index layoutLoad with the runtime's own loader and compare logits on one prompt

The last row deserves emphasis. With act-order, the quantized column order no longer matches the storage order, so the checkpoint carries a group index per input column, and the inference kernel must honour it. Some fast kernels require static groups or reorder at load time. A checkpoint that is numerically excellent can still produce nonsense in a runtime that ignores that metadata. The Marlin kernel article covers what fast 4-bit kernels expect.

End-to-end release gates

Per-layer error says where to look, not whether to ship. Three end-to-end gates, run in the target runtime, decide that. First, perplexity on the held-out set, compared with the FP16 model on identical tokens. Second, the mean per-token KL divergence between the FP16 and quantized next-token distributions, which is more sensitive than perplexity because it sees changes in the whole distribution, not only in the probability of the observed token. Third, the task evaluations that represent your product: exact-match answers, code that must pass tests, tool calls that must parse.

import torch, torch.nn.functional as F

@torch.no_grad()
def token_kl(ref_model, q_model, batches):
    """Mean KL(ref || quant) per token over held-out batches of input_ids."""
    total, n = 0.0, 0
    for ids in batches:
        lp_ref = F.log_softmax(ref_model(ids).logits.float(), dim=-1)
        lp_q = F.log_softmax(q_model(ids).logits.float(), dim=-1)
        kl = (lp_ref.exp() * (lp_ref - lp_q)).sum(-1)      # [batch, seq]
        total += kl.sum().item()
        n += kl.numel()
    return total / n

Set thresholds before you look at results, and track them per release. For example: perplexity within a fixed ratio of FP16, KL below a level set on a checkpoint you trusted, and no task metric outside its noise band. Also diff greedy outputs on a few hundred real prompts. Quantization evaluation methodology covers sample sizes and significance in detail.

Tooling in 2026

AutoGPTQ, the library that popularised GPTQ checkpoints on the Hugging Face hub, was archived in 2025, and its README points users to GPTQModel for continued support. llm-compressor, from the vLLM project, implements GPTQ as a GPTQModifier with knobs for block size, damping fraction, activation ordering, Hessian offloading and an ignore list, and applies it through a one-shot call. A minimal recipe looks like this:

# llm-compressor: check the installed version's docs; argument names have moved between releases.
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import GPTQModifier

recipe = GPTQModifier(
    targets="Linear",
    scheme="W4A16",          # 4-bit weights, 16-bit activations
    ignore=["lm_head"],      # plus any layer your audit flagged
)
oneshot(
    model=model,             # a loaded Hugging Face causal LM
    dataset=calib_dataset,   # your domain-matched, already templated text
    recipe=recipe,
    max_seq_length=2048,
    num_calibration_samples=512,
)

Record the recipe, tool version, calibration hash and audit numbers with the checkpoint; later, that is how you tell a quantization regression from a kernel one.

Trade-offs

Group size. Smaller groups (64, 32) lower error at the cost of more metadata and slower kernels; 128 is the common default.

Act-order. Usually lowers error, adds group-index metadata and constrains which kernels you can use. Measure both on your audit before choosing.

Mixed precision. Keeping a handful of sensitive layers in 8 or 16 bits costs little memory and often recovers most of the gap. Let the per-layer audit choose them, not folklore.

GPTQ or AWQ. GPTQ compensates errors using second-order input statistics; AWQ rescales salient channels before rounding and tends to need less calibration data. If GPTQ overfits your calibration set, AWQ is the natural comparison.

What to do next

  1. Run the one-layer lab on a down_proj layer of your model and compare RTN and GPTQ on calibration and held-out inputs.
  2. Do the Hessian rank arithmetic for your model's widest layer and, for MoE models, for your least-used experts.
  3. Build a calibration set from templated, production-like text, and a held-out set from the same source.
  4. Quantize, then plot per-layer relative output error by depth and type; move outlier layers to higher precision.
  5. Gate the packed checkpoint in the target runtime on perplexity, token KL and product tasks with thresholds set in advance.
  6. Store recipe, tool version, calibration hash and audit numbers with every checkpoint you release.
Key takeaway: GPTQ guarantees a small output error per layer, on the inputs it was calibrated with, and nothing more. Treat the calibration set as part of the model, check that every Hessian is well conditioned, measure each layer on held-out inputs as well as calibration ones, and gate the packed checkpoint in the runtime you ship with perplexity, KL and real tasks. Most bad GPTQ models come from data or format problems, not from the algorithm.