AWQ and GPTQ are the two calibrated methods people reach for when they want a large language model in 4-bit weights with 16-bit activations. Both produce the same kind of checkpoint: integer weights in groups, typically of 128, each group with a scale and possibly a zero point, served by the same fast kernels. The difference is entirely in how the integers are chosen, and that difference shows up only in quality, and only on some traffic.

So the useful question is not which method is better in general, but which is better for your model, your traffic and your constraints, and how to find out cheaply. This article compares what each method optimises, explains what each is sensitive to, and gives a bake-off procedure and decision rule. The internals are covered separately in AWQ architecture, in depth and GPTQ Architecture, in depth; here they are summarised only as far as the comparison needs.

Advertisement

Two objectives for one format

GPTQ treats each linear layer as a least-squares problem. Given calibration inputs X to the layer, it wants quantised weights whose outputs match the original outputs, and it builds an approximate Hessian from X multiplied by its transpose. It then quantises one column of weights at a time and, after rounding each column, updates the remaining unquantised columns to compensate for the error just introduced. The quantised weights can differ from plain rounding by a lot: they are chosen to cancel each other's errors on the calibration activations.

AWQ does no error compensation. It observes that a small fraction of input channels carry large activations, and that weight errors on those channels hurt the output most. It scales those weight channels up before quantisation and scales the matching activations down by the same factor, folded into the preceding operation, so the function is unchanged but the important weights get relatively finer resolution. The per-channel scale is the average activation magnitude raised to a power, and a small grid search per block picks the power that minimises the block's output error on calibration data. After scaling, the weights are rounded to the nearest level.

The consequence is the key to the whole comparison. GPTQ uses calibration data to fit individual weights, which is a much more expressive use of the data. AWQ uses calibration data only for per-channel averages and one exponent per block, which is far less expressive but also has far less room to overfit.

AWQGPTQ
What calibration data setschannel scales, one exponent per blockevery quantised weight, via error feedback
Statistics usedmean activation magnitudesecond-order: X times X transpose
Weight valuesrounded to nearest after scalingrounded with compensating updates
Main knobsgroup size, symmetric or asymmetric, scale gridgroup size, damping, act-order, block size
Cost to produceforward passes plus a small searchforward passes plus per-layer solves
Typical riskoutlier channels the scale cannot fixfitting the calibration distribution

What each method is sensitive to

Calibration distribution shift. The AWQ authors argued, and showed on their benchmarks, that AWQ depends less on the calibration set matching the evaluation domain, which follows from its low-dimensional use of the data. GPTQ's reconstruction can tilt the weights toward whatever the calibration set contains. The practical lesson for both is the same, and stronger for GPTQ: calibrate on templated text that looks like your traffic, with the languages, code and tool-call formats you serve, as described in quantization calibration.

Act-order and groups. GPTQ's act-order option quantises columns in decreasing order of Hessian diagonal, so the most sensitive columns are rounded first while there is the most freedom left to compensate. It often helps quality. Combined with grouping, it means a column's group no longer follows from its position, so the checkpoint carries a group index per column, and some kernels pay for the extra indirection. Check what your serving kernel supports before turning it on.

Symmetric versus asymmetric. The current AWQ example in llm-compressor uses an asymmetric scheme with zero points, while the GPTQ W4A16 example uses a symmetric one. That difference alone can move quality, so a fair comparison either runs both methods with the same scheme or treats scheme as a second, separately measured variable.

Group size. Smaller groups, such as 64 or 32, give each scale fewer weights to cover and reduce error for both methods, at the cost of more scale storage and sometimes slower kernels. Larger groups, or per-channel scales, make the method matter more, because rounding error per group grows and GPTQ's compensation or AWQ's rebalancing has more to repair. If the two methods are close at group size 128, they will usually be closer still at 64; if they differ sharply, try a smaller group before concluding that one method is unusable.

Architecture. Mixture-of-experts models route few calibration tokens to rare experts; both methods then estimate statistics from little data, and GPTQ's Hessian can become ill-conditioned, which damping only partly fixes. Models with extreme activation outliers can exceed what a per-channel scale can rebalance, and AWQ alone may lose more there.

Advertisement

The bake-off procedure

A fair comparison changes exactly one thing. Start from the BF16 model you will actually serve, after any fine-tune and adapter merge. Use one calibration set for both. Use the same group size, the same output-head exclusion and, ideally, the same symmetric or asymmetric scheme. Serve both checkpoints through the same engine on the same GPU, so kernel differences do not masquerade as quality differences. Then evaluate both against the BF16 model on held-out prompts that were not used for calibration.

An AWQ versus GPTQ bake-off: change one thingBF16 modelmerged, the exact releaseCalibration settemplated production promptsAWQ recipescale search, then W4A16 g128GPTQ recipeHessian updates, W4A16 g128Same engine and kernelvLLM, compressed-tensorsSame engine and kernelvLLM, compressed-tensorsHeld-out evaluation against BF16KL divergence, task scores, answer flips, latencyEverything except the algorithm is fixed, so the difference you measure belongs to the algorithm.
Fix the model, calibration set, scheme and serving path; vary only the algorithm; evaluate both against BF16 on held-out data.

The recipes below follow the llm-compressor examples on its main branch as of October 2026. Module paths have moved between releases, so check the examples directory of the version you install.

from transformers import AutoModelForCausalLM, AutoTokenizer
from llmcompressor import oneshot
from llmcompressor.modifiers.gptq import GPTQModifier                 # main, Oct 2026
from llmcompressor.modifiers.quantization import QuantizationModifier
from llmcompressor.modifiers.transform.awq import AWQModifier         # main, Oct 2026

MODEL = "./my-model-bf16"          # the merged model you will serve
CALIB = load_calibration_set()     # templated production-like prompts, a Dataset
SCHEME = "W4A16_ASYM"              # use the same scheme for both arms

def run(recipe, out_dir):
    model = AutoModelForCausalLM.from_pretrained(MODEL, torch_dtype="auto")
    tok = AutoTokenizer.from_pretrained(MODEL)
    oneshot(model=model, dataset=CALIB, recipe=recipe,
            max_seq_length=2048, num_calibration_samples=512)
    model.save_pretrained(out_dir, save_compressed=True)
    tok.save_pretrained(out_dir)

run([AWQModifier(),
     QuantizationModifier(targets="Linear", scheme=SCHEME, ignore=["lm_head"])],
    "./out-awq")
run(GPTQModifier(targets="Linear", scheme=SCHEME, ignore=["lm_head"]),
    "./out-gptq")
# serve each with: vllm serve ./out-awq   (format read from the saved config)

Measuring the difference

Benchmark averages hide what users notice. The most sensitive cheap signal is the KL divergence between the BF16 model's next-token distribution and each quantised model's, measured token by token on held-out prompts, together with the fraction of positions where the top token changes. Look at the tail, not only the mean: a method that is slightly better on average but much worse on the worst one percent of tokens will produce visibly broken answers on some prompts. Quantization evaluation methodology covers the full suite, including task scores and answer flips.

import torch
import torch.nn.functional as F

@torch.no_grad()
def kl_vs_reference(ref, quant, batches):
    """Per-token KL(ref || quant) and top-1 agreement on held-out token batches."""
    kls, agree = [], []
    for input_ids in batches:
        lp_ref = F.log_softmax(ref(input_ids).logits.float(), dim=-1)
        lp_q = F.log_softmax(quant(input_ids).logits.float(), dim=-1)
        kl = (lp_ref.exp() * (lp_ref - lp_q)).sum(-1)          # [batch, seq]
        kls.append(kl.flatten().cpu())
        agree.append((lp_ref.argmax(-1) == lp_q.argmax(-1)).flatten().cpu())
    kl = torch.cat(kls)
    return {"kl_mean": kl.mean().item(),
            "kl_p99": kl.quantile(0.99).item(),
            "top1_agreement": torch.cat(agree).float().mean().item()}

Run it per traffic slice: chat, code, each language you serve and long contexts. A method can win overall and lose on one slice, and that slice decides the choice if it carries revenue. Running both quantised models in one process with the reference needs memory for three models; for large models, store reference log-probabilities to disk in one pass and compare later.

Serving speed is usually a tie

Because both methods produce the same format, decode speed with the same scheme and kernel is normally the same; the kernels that make W4A16 fast are described in Marlin INT4 kernel architecture. Differences come from options, not methods: act-order with a per-column group index, zero points in asymmetric schemes, or a group size the kernel handles less efficiently. Measure latency and throughput for both checkpoints at your real batch sizes, but expect the decision to rest on quality.

Production cost differs more. GPTQ solves a problem per layer, and on large models it takes longer and needs more memory during quantisation than AWQ's search, though both are one-off jobs measured in minutes to hours. If you re-quantise often, for example after every fine-tune, that time counts.

Worked example: choosing for a fine-tuned assistant

A team serves a fine-tuned 8B model to an internal assistant used for both chat and code. They merge the adapters into BF16, build 512 calibration samples from logged prompts with the chat template applied, half chat and half code, and hold out 2,000 different prompts for evaluation. They run both recipes with the same asymmetric W4A16 scheme at group size 128.

Their decision rule, written before looking at results, is: prefer the checkpoint with lower 99th-percentile KL on the code slice, because code errors break builds; if the two are within the run-to-run noise measured by quantising twice with different calibration seeds, prefer AWQ, because it is faster to reproduce after each fine-tune; reject both if top-1 agreement on any slice falls below the threshold the team set from a blind review of sampled outputs. Writing the rule first keeps the team from rationalising whichever number came out ahead. Whichever wins, the procedure is rerun after every fine-tune, because the answer can change with the model.

Combining them

In llm-compressor, AWQ is expressed as a transform that rescales weights and is followed by a separate quantisation step, so it is structurally possible to follow the AWQ transform with the GPTQ modifier instead of plain rounding. The idea is sound, since the scaling and the error compensation address different problems, but it is not one of the documented examples. Treat it as a third arm of the bake-off and validate it with the same evaluation, rather than assuming it beats both.

Failure modes

  • Unequal comparison. Different schemes, group sizes or serving engines for the two arms. The result measures the setup, not the algorithm.
  • Calibration leakage. Evaluation prompts overlap calibration prompts and flatter GPTQ in particular. Keep the sets disjoint.
  • Untemplated calibration. Raw web text instead of chat-formatted prompts; quality drops on real traffic for both.
  • Quantising a 4-bit base. Starting from an NF4 QLoRA base rather than a merged BF16 model compounds two quantisation errors. See bitsandbytes vs AutoAWQ vs AutoGPTQ.
  • Kernel mismatch. Act-order or zero points that the serving kernel handles on a slow path or not at all. Test loading and latency before the quality run.
  • Single-number decision. Choosing on average perplexity while one slice degrades badly. Report per-slice tails.

What to do next

  1. Freeze the exact BF16 model you will serve and record its baseline on your evaluation suite.
  2. Build a templated calibration set from real traffic and a disjoint held-out set, sliced by task type and language.
  3. Write your decision rule, including a noise margin, before running anything.
  4. Run the AWQ and GPTQ recipes with the same scheme and group size, and serve both through the same engine.
  5. Measure per-slice KL mean, 99th-percentile KL and top-1 agreement against BF16, then latency at real batch sizes.
  6. Ship the winner, and rerun the bake-off whenever the model or the traffic mix changes.
Key takeaway: AWQ and GPTQ write the same 4-bit format and usually run at the same speed; they differ in how they use calibration data. GPTQ fits individual weights with second-order error feedback, which is powerful and can tilt toward the calibration set. AWQ sets only channel scales from activation averages, which leaves less to overfit and less to fix extreme outliers. Do not pick from a leaderboard: fix everything except the algorithm, measure per-slice KL tails against BF16, and apply a decision rule written in advance.