GPTQ is a post-training quantisation method that compresses the weights of a large language model to 4 or even 3 bits in a few GPU hours, with no retraining and only a small calibration set. Published by Frantar and colleagues in 2022, it remains one of the most widely used formats for open-weight models, and its core idea, quantise one weight and then adjust the remaining weights to cancel the error, appears in many later methods.
This article treats GPTQ as a system you build and run: the per-layer update in brief, the block-by-block pipeline, a working implementation, the memory and time it needs at scale. It then covers the practical knobs (group size, act-order, dampening, calibration data), a worked example on a 7B model, how the packed weights are stored and served, and the failure modes to watch for. For how GPTQ compares with activation-aware methods, see AWQ vs GPTQ.
Why rounding to nearest breaks at 4 bits
The simplest way to quantise a weight matrix is round-to-nearest: choose a scale per row or group, divide, round to the nearest of 16 levels, and store the integers. At 8 bits this is nearly lossless. At 4 bits each weight carries an error of up to half a step, and those errors add up across thousands of inputs to each output. Some inputs matter far more than others, because some activation channels are consistently large, so errors on the weights that multiply them do disproportionate damage.
GPTQ's answer is to stop treating weights independently. Rounding one weight up can be partly cancelled by nudging other weights in the same row that see correlated inputs. Doing that well requires knowing how inputs correlate, which is what the calibration data provides.
The algorithm on one page
GPTQ quantises one linear layer at a time. For weights W (rows are output features, columns are input features) and calibration inputs X, it looks for grid-restricted weights Ŵ that keep W X close to Ŵ X. The problem splits into independent rows that all share one matrix, H = 2XXᵀ, sized input features by input features. H depends only on the inputs: its diagonal says how strongly each input channel is used, and its off-diagonal entries say which channels move together. The full derivation, with a two-weight worked example, is in GPTQ math; here is what an implementation has to do.
for each column q, in one fixed order shared by every row:
round column q of W to the grid
err = (w_q - quant(w_q)) / [H^-1]_qq # per row, scaled by the inverse Hessian
W[:, later columns] -= err * H^-1[q, later] # let the rest of the row absorb the errorThis is the optimal brain surgeon update. Its predecessor, OBQ, chose a greedy order per row, which costs on the order of rows × columns³ and takes hours per layer. GPTQ makes three engineering changes. A fixed column order lets every row share one inverse Hessian. Lazy batch updates apply updates immediately inside blocks of 128 columns and push the accumulated error to later columns in one large matrix multiply, turning a bandwidth-bound loop into a compute-bound one. A Cholesky factorisation of H⁻¹, computed once after adding damping (1% of the mean diagonal by default), supplies every row of the inverse the loop needs without the error build-up of repeated inverse updates. The original paper quantised 175-billion-parameter models in about four GPU hours this way.
Running it at scale: memory, time and offloading
The pipeline's memory footprint is set by three things: one transformer block's weights in 16-bit, the calibration activations entering that block, and one Hessian per distinct layer input. Activations are usually the largest. 512 sequences of 2,048 tokens at hidden size 8,192 in FP16 is 512 × 2,048 × 8,192 × 2 bytes ≈ 17 GB, and the tool keeps both the block's inputs and its outputs. Halving the calibration set or sequence length halves this; so does keeping activations in CPU memory and streaming them through in micro-batches, which most tools support at some cost in speed.
Weights not being quantised stay on CPU or disk and are loaded one block at a time, so the model never needs to fit on the GPU. Time is dominated by the forward passes that collect inputs and, for wide layers, by the column loop; MLP down projections and the many small expert layers of mixture-of-experts models are the slow parts. Expect the run to scale roughly linearly with the number of blocks and calibration tokens, and checkpoint between blocks for very large models so a crash does not restart the job.
The pipeline, block by block
A full model is quantised sequentially. The calibration batch is pushed through the embedding layer, then each transformer block is processed in turn: hooks on every linear layer record its inputs and accumulate H, each layer is quantised, and then the block is run again with quantised weights to produce the inputs for the next block. Feeding later blocks the outputs of quantised earlier blocks lets each layer partly compensate for errors made upstream. Only one block and its activations need to be on the GPU at a time.
The algorithm in code
The core fits in about forty lines of PyTorch. This version uses asymmetric min-max quantisation per group and follows the structure of the reference implementation.
import torch
class HessianHook:
# Accumulates H = 2/n * sum(x x^T) over all calibration tokens seen by one Linear.
def __init__(self, cols, device):
self.H, self.n = torch.zeros(cols, cols, device=device), 0
def __call__(self, module, inp, out):
x = inp[0].reshape(-1, inp[0].shape[-1]).float() # [tokens, cols]
t = x.shape[0]
self.H *= self.n / (self.n + t)
self.n += t
x = (2 / self.n) ** 0.5 * x
self.H += x.T @ x
def fit_group(Wg, maxq):
lo = Wg.min(dim=1).values.clamp(max=0)
hi = Wg.max(dim=1).values.clamp(min=0)
scale = ((hi - lo) / maxq).clamp(min=1e-8)
return scale, torch.round(-lo / scale)
def gptq(W, H, bits=4, group=128, block=128, damp=0.01, act_order=True):
W, H = W.clone().float(), H.clone()
rows, cols = W.shape
maxq = 2 ** bits - 1
dead = torch.diag(H) == 0 # inputs never active in calibration
H[dead, dead] = 1
W[:, dead] = 0
perm = torch.argsort(torch.diag(H), descending=True) if act_order else torch.arange(cols)
W, H = W[:, perm], H[perm][:, perm]
H += damp * torch.mean(torch.diag(H)) * torch.eye(cols, device=H.device)
Hinv = torch.linalg.cholesky(torch.cholesky_inverse(torch.linalg.cholesky(H)), upper=True)
Q = torch.zeros_like(W)
for i in range(0, cols, block):
j2 = min(i + block, cols)
Err = torch.zeros(rows, j2 - i, device=W.device)
for j in range(i, j2):
if j % group == 0:
scale, zero = fit_group(W[:, j:j + group], maxq)
q = torch.clamp(torch.round(W[:, j] / scale) + zero, 0, maxq)
Q[:, j] = scale * (q - zero)
err = (W[:, j] - Q[:, j]) / Hinv[j, j]
W[:, j + 1:j2] -= err[:, None] * Hinv[j, j + 1:j2][None, :] # inside the block
Err[:, j - i] = err
W[:, j2:] -= Err @ Hinv[i:j2, j2:] # lazy update
return Q[:, torch.argsort(perm)] # undo act-order permutationIn a real tool the integer codes, scales and zero points are kept separately for packing rather than returned as dequantised floats, and the group index of each original column is recorded when act-order is on.
Groups, act-order and other knobs
- Group size. One scale and zero point per row is too coarse at 4 bits. Groups of 128 consecutive input columns per row are the common default; 32 or 64 improve accuracy at the cost of more metadata and slightly slower kernels. The general trade-off is covered in quantisation granularity.
- Act-order. Quantising the columns with the largest Hessian diagonal first, while there are still many unquantised columns left to absorb their error, usually improves accuracy noticeably. Because groups are then formed over the permuted order, each original column needs a group index (
g_idx), and the kernel must gather scales non-contiguously. Some tools offer a static-groups option that fixes groups in the original order before permuting, which keeps the kernel fast. - Symmetric or asymmetric. Asymmetric quantisation stores a zero point and fits skewed groups better; symmetric is simpler and required by some kernels.
- Damping. Raise it from 0.01 towards 0.1 if the Cholesky factorisation fails.
- Calibration data. The original paper used 128 random 2,048-token segments of C4. Data resembling your real traffic, including chat templates and code if you serve them, is better than generic web text. More guidance is in calibration.
Worked example: a 7B model
Take a 7B Llama-style model with hidden size 4,096 and MLP size 11,008. The attention projections and the MLP up and gate projections take 4,096 inputs, so each Hessian is 4,096 × 4,096 in FP32, or 64 MiB. The MLP down projection takes 11,008 inputs, so its Hessian is about 485 MB. The q, k and v projections see the same inputs and can share one Hessian. All of this fits easily on one GPU next to a single block.
Storage at 4 bits with groups of 128, an FP16 scale and a 4-bit zero point per group costs 4 + (16 + 4)/128 ≈ 4.16 bits per weight. Seven billion weights therefore take about 7 × 10⁹ × 4.16 / 8 ≈ 3.6 GB, compared with 14 GB in FP16. Embeddings and the output head are usually left in 16-bit, so the real file is somewhat larger. With 128 calibration sequences the whole run typically takes minutes to well under an hour on one modern GPU.
Packing, tooling and serving
A GPTQ checkpoint replaces each linear weight with packed tensors: qweight holding eight 4-bit codes per 32-bit integer, qzeros packed the same way, scales in FP16, and g_idx mapping columns to groups. Inference kernels unpack and dequantise tiles in registers just before the matrix multiply. Kernels such as Marlin are built for this format; vLLM, SGLang and Hugging Face Transformers can load GPTQ checkpoints.
The speed-up comes from reading a quarter of the bytes. Decoding at small batch sizes is limited by memory bandwidth, so it gains most. Large-batch prefill is compute-bound and gains little, because the multiplications still run in 16-bit after dequantisation.
# Quantising with llm-compressor (check the API of your installed version).
from transformers import AutoModelForCausalLM, AutoTokenizer
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import GPTQModifier
model_id = "your-org/your-7b-model"
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto")
tokenizer = AutoTokenizer.from_pretrained(model_id)
recipe = GPTQModifier(targets="Linear", scheme="W4A16", ignore=["lm_head"])
oneshot(model=model, dataset=calibration_dataset, recipe=recipe,
max_seq_length=2048, num_calibration_samples=512)
model.save_pretrained("model-w4a16", save_compressed=True)AutoGPTQ, the original library, is no longer maintained; GPTQModel and llm-compressor are the maintained options. They write different checkpoint layouts, so confirm your serving engine supports the one you produce.
Failure modes
- Cholesky failure. Too few calibration tokens, or input channels that never activate, leave H singular. Increase damping and calibration size; the dead-column handling above covers channels that are always zero.
- Starved MoE experts. In mixture-of-experts models, rarely routed experts see only a few calibration tokens, so their Hessians are poor. Use much more calibration data, check per-expert token counts, or leave rare experts at higher precision.
- Domain mismatch. Calibrating on English web text and serving code or another language increases errors exactly where you care. Calibrate on representative data.
- Perplexity-only evaluation. Perplexity can look fine while reasoning, maths or long-context accuracy drops. Run task evaluations against the 16-bit baseline.
- Slow act-order kernels. Act-order with non-static groups forces scattered scale reads; some kernels fall back to slower paths. Benchmark, or use static groups.
- Quantising sensitive layers. The output head and sometimes the first and last blocks are more sensitive. Exclude them first, then try quantising them and measure.
Trade-offs against other methods
| Method | Calibration | Cost | Typical 4-bit quality |
|---|---|---|---|
| Round-to-nearest | None | Seconds | Good with small groups; weak at 3 bits |
| GPTQ | Yes, second-order | Minutes to hours | Strong; sensitive to calibration data |
| AWQ (details) | Yes, activation scales | Minutes | Comparable; often more robust to domain shift |
| Quantisation-aware training | Full training data | Training run | Best, at much higher cost |
What to do next
- Measure the 16-bit model on your task evaluations and record decode tokens per second at your serving batch size.
- Build a calibration set of 256 to 512 sequences from real or representative traffic, formatted with your chat template.
- Quantise to 4 bits with group size 128, act-order on, damping 0.01 and the output head excluded.
- Compare task accuracy and speed with the baseline; if accuracy drops, try group size 64, more calibration data or excluding sensitive layers.
- For MoE models, log calibration tokens per expert before trusting the result.
- Confirm your serving engine loads the checkpoint with a fast kernel, and pin the tool versions used.