AWQ and GPTQ are the two most widely used methods for compressing a language model's weights to 4 bits after training, without any fine-tuning. Both produce the same kind of checkpoint, 4-bit integer weights with a 16-bit scale for each small group of weights, and both are served by the same fast kernels. Most write-ups test them on 7B to 70B models, where 4-bit weights cost little quality. Small language models, from about 1 to 4 billion parameters, behave differently: less of the model is quantizable, the savings are smaller than the headline ratio, and quality falls further.

This article explains both algorithms from first principles, then concentrates on what changes at small scale, with a worked byte budget, current tooling and an evaluation gate you can run before shipping. For full algorithm walkthroughs see GPTQ architecture and AWQ architecture; for a head-to-head bake-off method see AWQ vs GPTQ.

Advertisement

What W4A16 buys and why decoding gets faster

W4A16 means 4-bit weights and 16-bit activations. At inference time the kernel loads packed 4-bit weights, expands them to 16 bits in registers using the group's scale (and zero point, for asymmetric schemes), and multiplies with 16-bit activations. The arithmetic is unchanged; what shrinks is memory traffic. Generating one token at batch size one reads every weight once and does about two floating point operations per weight, so decoding is limited by memory bandwidth, not compute. Reading a quarter of the bytes can make decode close to proportionally faster, while prefill on long prompts, which is compute-bound, gains much less.

Groups are the key detail. Each run of 128 consecutive input weights in a row typically shares one scale. A 16-bit scale per 128 weights adds 0.125 bits per weight, and a 4-bit zero point adds about 0.03 more, so W4A16 at group size 128 costs roughly 4.13 to 4.16 bits per weight. Smaller groups track the weights more closely and cost more bits.

The two algorithms, compactly

Rounding each weight to the nearest grid point treats all errors as equal, but they are not: an error on a weight that multiplies a large, frequent activation changes the output much more. Both methods use a small set of calibration prompts to measure activations and reduce the error in the layer's output, not in the weights.

GPTQ (Frantar and colleagues, 2022) quantizes one layer at a time, one column at a time. It builds the matrix H = 2XXT from the layer's calibration inputs, which measures how strongly each pair of input channels co-varies. After rounding a column, it distributes the rounding error over the columns not yet quantized, weighted by the inverse of H, so later weights compensate for earlier errors. A Cholesky factorisation and processing columns in blocks make this fast enough that the paper quantized 175B-parameter models in about four GPU hours. Optional act order quantizes the columns with the largest activations first, while there are still many columns left to absorb their error.

AWQ (Lin and colleagues, 2023) starts from an observation: a small fraction of input channels, roughly 0.1 to 1%, matter far more than the rest, and they are found by activation size, not weight size. Keeping them in 16 bits helps, but mixed precision is awkward for kernels. Instead AWQ multiplies those channels' weights by a scale greater than one before quantizing, which gives them more effective resolution, and divides the activations by the same scale by folding it into the preceding operation. The scale is the channel's activation size raised to a power found by a small grid search. AWQ does no error feedback and no backpropagation, so it depends less on the details of the calibration set.

# Pseudocode. GPTQ, one linear layer W (out x in), calibration inputs X (in x tokens)
H = 2 * X @ X.T
H += damp * mean(diag(H)) * I                          # damp is typically 0.01
Hinv = cholesky_upper(inverse(H))
for each column j in order (optionally by descending diag(H), "act order"):
    q = quantize(W[:, j], scale_of_group(j))           # round to the 4-bit grid
    err = (W[:, j] - q) / Hinv[j, j]
    W[:, j+1:] -= outer(err, Hinv[j, j+1:])            # push the error onto unquantized columns
    W[:, j] = q

# AWQ, one linear layer and the operation feeding it
s_x = mean(abs(X), axis=tokens)                        # per input channel activation size
best = None
for alpha in grid(0, 1):
    s = s_x ** alpha                                    # larger scale for salient channels
    Wq = quantize(W * s) / s                            # divide back, folded into previous op
    loss = norm(W @ X - Wq @ X)
    best = min(best, (loss, s))
fold 1/s into the preceding norm or linear layer; quantize W * s
Advertisement

Worked example: where the bytes go in a 1.5B model

Take an illustrative configuration close to several current 1 to 2B models: 28 layers, hidden size 1536, MLP width 8960, grouped-query attention with 12 query heads and 2 key-value heads of dimension 128, a vocabulary of 151,936 and the input embedding tied to the output head. Each layer holds 1536 x 3584 attention weights (5.5M) and 3 x 1536 x 8960 MLP weights (41.3M), so 28 layers hold 1.31B linear weights. The embedding holds 151,936 x 1536 = 233M. The total is about 1.54B, 3.09 GB in BF16.

Quantize the linear layers to W4A16 at group size 128 and keep the tied embedding and head in BF16, as both published recipes do with ignore=["lm_head"]. The linear layers drop to about 0.68 GB, but the embedding stays at 0.47 GB, so the model is about 1.14 GB: 2.7 times smaller, not 4. In a 70B model the embedding and head are a few percent of the parameters; here it is 15% of the parameters and 41% of the quantized file.

The key-value cache does not shrink at all. Per token it stores 2 x 28 layers x 2 heads x 128 dimensions x 2 bytes = 28,672 bytes, so one 32k-token sequence needs about 0.94 GB, close to the whole quantized model. For long-context small models on small GPUs, the cache, not the weights, often sets the memory limit, and the remedy is cache quantization or shorter contexts, not a more aggressive weight scheme.

Where the bytes go in an illustrative 1.5B model, before and after 4-bit weight quantizationBF16 checkpoint: about 3.09 GBLinear layers 1.31B params: 2.62 GBEmbedding 0.47 GBW4A16, group 128, embedding kept in BF16: about 1.14 GBLinear 0.68 GBEmb 0.4741% of the quantized file is the unquantized embedding: a 2.7x reduction, not 4xGPTQerror feedback per layeruses H = 2 X X^T from calibrationAWQscale salient input channelsscales chosen from activation sizeSame output formatint4 weights + group scalesserved by the same W4A16 kernelsKV cache for one 32k-token sequence in the same model: about 0.94 GB in BF16,close to the size of the whole quantized model. Weight quantization does not shrink it.Dimensions are illustrative: 28 layers, hidden 1536, MLP 8960, 2 KV heads of 128, vocabulary 151,936, tied embeddings.
The byte budget of the illustrative model. Quantizing the linear layers cuts them by about 3.9x, but the 16-bit embedding and the key-value cache are untouched, so the whole-model saving is much smaller than the headline.

Why small models lose more at 4 bits

Three effects compound. First, less redundancy: a narrow model packs more distinct features into each weight, so the same relative rounding error removes more information. Benchmarks consistently show larger relative quality drops for smaller models at the same bit width, and the losses concentrate on multi-step reasoning, arithmetic, code and instruction following, while simple next-token perplexity can look almost unchanged. Second, outlier channels: a few input channels with very large activations exist at every scale, and with a hidden size of 1536 each row of the attention and up-projection weights has only 12 groups of 128, so one outlier influences a larger share of a row's scales. Third, fine-tuned behaviour is fragile: refusal style, output format and tool-call syntax learned in a small final fine-tuning step can break at 4 bits while general knowledge survives.

Practical consequences follow. Prefer GPTQ with act order or AWQ over plain round-to-nearest, always. Consider group size 64 or 32 for small models when your kernel supports it; the extra bits are cheap in absolute bytes. Leave the first and last layers or any layer with obvious outliers at higher precision if your format permits mixed schemes. Calibrate on your own traffic, including the system prompt and the output format, because that is the behaviour most likely to break.

4-bit bigger or 8-bit smaller

With a fixed memory budget the real question is which model to quantize. A 3B model at W4A16 and a 1.5B model at 8 bits use similar memory for weights. The larger model at 4 bits often wins on knowledge-heavy tasks; the smaller model at 8 bits often keeps formatting and tool calls more reliably and decodes faster because it reads fewer bytes in total and has fewer layers. There is no general answer, which is why the decision belongs in an evaluation, not a rule of thumb. On CPUs and phones, the GGUF formats used by llama.cpp are usually the better route; see GGUF and llama.cpp.

Running it: tooling as of October 2026

The original libraries are no longer maintained. AutoAWQ's README declares it deprecated and says it has been adopted by the vLLM project's llm-compressor; AutoGPTQ's README says it is unmaintained and points to GPTQModel. The code below follows the two llm-compressor example scripts on the main branch as published on 2026-10-02. The AWQ recipe pairs an AWQModifier with a separate quantization step using an asymmetric scheme; the GPTQ recipe is a single modifier with group size 128. Import paths have moved between releases, so pin the version you test with.

# Adapted from the vllm-project/llm-compressor AWQ and W4A16 examples (main branch, checked 2026-10-02).
from transformers import AutoModelForCausalLM, AutoTokenizer
from llmcompressor import oneshot
from llmcompressor.modifiers.gptq import GPTQModifier
from llmcompressor.modifiers.quantization import QuantizationModifier
from llmcompressor.modifiers.transform.awq import AWQModifier   # path has moved between releases

MODEL_ID = "your-org/your-1.5b-model"
model = AutoModelForCausalLM.from_pretrained(MODEL_ID)
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)

USE_AWQ = True
if USE_AWQ:
    recipe = [
        AWQModifier(duo_scaling="both"),
        QuantizationModifier(ignore=["lm_head"], scheme="W4A16_ASYM", targets=["Linear"]),
    ]
else:
    recipe = GPTQModifier(targets="Linear", scheme="W4A16", ignore=["lm_head"])

oneshot(
    model=model,
    dataset="perfectblend",          # replace with prompts that look like your traffic
    splits="train[:512]",
    recipe=recipe,
    max_seq_length=512 if USE_AWQ else 2048,
    num_calibration_samples=256 if USE_AWQ else 512,
)
SAVE_DIR = MODEL_ID.split("/")[-1] + ("-awq-asym" if USE_AWQ else "-W4A16-G128")
model.save_pretrained(SAVE_DIR, save_compressed=True)
tokenizer.save_pretrained(SAVE_DIR)
# Serve: vllm serve ./<SAVE_DIR>   (the quantization config travels in config.json)

The examples use 256 to 512 calibration samples. For a small domain model, replace the generic dataset with a few hundred real prompts and responses in your chat template; the run takes minutes on one GPU at this size.

An evaluation gate before shipping

Compare the quantized model with the original on held-out prompts from your own traffic, measuring three things. Token-level KL divergence between the two models' next-token distributions, and the share of tokens whose top prediction changed, catch damage that average perplexity hides. Task metrics on your real evaluation set catch functional regressions. Format checks, such as JSON validity and tool-call parse rates, catch the fragile fine-tuned behaviour.

import torch, torch.nn.functional as F

@torch.no_grad()
def token_kl(base, quant, batches):
    """Mean KL(base || quant) per token over held-out prompts; also the share of tokens
    whose top-1 prediction changed. Both models must share the tokenizer."""
    kl_sum, flips, n = 0.0, 0, 0
    for input_ids in batches:                       # LongTensor [batch, seq] on the right device
        lp_b = F.log_softmax(base(input_ids).logits.float(), dim=-1)
        lp_q = F.log_softmax(quant(input_ids).logits.float(), dim=-1)
        kl = (lp_b.exp() * (lp_b - lp_q)).sum(-1)   # [batch, seq]
        kl_sum += kl.sum().item()
        flips += (lp_b.argmax(-1) != lp_q.argmax(-1)).sum().item()
        n += kl.numel()
    return {"mean_kl": kl_sum / n, "top1_flip_rate": flips / n}

Set thresholds from a baseline: quantize with both methods, compute the metrics, and accept a checkpoint only if task scores stay within a tolerance you choose in advance, for example one point on your main metric and no drop in format validity. A rising top-1 flip rate on one category of prompt is the most useful early signal of where the model broke.

Failure modes

  • Perplexity says fine, users say broken. Arithmetic, code and tool calls degrade while perplexity barely moves. Gate on task and format metrics.
  • Calibration mismatch. Generic web text misses your chat template and output format, which then degrade most. Calibrate on your own traffic.
  • Expecting 4x. Tied embeddings and the key-value cache keep a large share of memory at 16 bits in small models.
  • Kernel fallback. An unusual group size or layer shape that the fast kernel does not support may run on a slower path. Benchmark tokens per second after loading, not just memory.
  • Unpinned tooling. Moved import paths or changed defaults silently produce a different checkpoint. Pin versions and record the recipe with the model.

What to do next

  1. Compute your model's byte budget: linear weights, embeddings and head, and key-value cache at your real context length.
  2. Build a held-out set of a few hundred real prompts with expected outputs and format checks.
  3. Quantize with both published llm-compressor recipes, replacing the dataset with your own traffic.
  4. Run the KL, flip rate, task and format gate against the BF16 model, and try a smaller group size if the result is marginal.
  5. Compare the best 4-bit checkpoint with an 8-bit smaller model at equal memory.
  6. Serve with a pinned vLLM version, measure decode tokens per second at batch one and at your real concurrency, and record the recipe, versions and gate results with the checkpoint. For device targets read SLM edge quantization.
Key takeaway: AWQ and GPTQ both produce 4-bit weights with per-group scales and run on the same W4A16 kernels, which speed up memory-bound decoding. GPTQ feeds each rounding error forward using second-order statistics from calibration data; AWQ scales the input channels with the largest activations before rounding. In 1-4B models the 16-bit embedding and the key-value cache keep savings well below 4x, and quality drops more, mostly in reasoning, code and output format. Calibrate on your own traffic, compare against an 8-bit smaller model, and ship only through a KL, task and format gate.