Activation-aware Weight Quantization (AWQ), by Ji Lin, Song Han and colleagues at MIT and published at MLSys 2024, compresses the weights of a large language model to 4 bits while keeping activations in 16-bit, with no retraining and only a small calibration set. Its central observation is that a small fraction of weights, around 1 percent, matter far more than the rest, and that you find them by looking at activation magnitudes, not weight magnitudes. Rather than keeping those weights in higher precision, AWQ scales their input channels so quantization treats them more gently.

The mathematics of why that scaling works is derived step by step in AWQ math, and its comparison with GPTQ is in AWQ vs GPTQ. This article covers AWQ as a system you run: the pipeline that walks the model block by block, exactly where scales are applied and absorbed, what the search does in code, what ends up on disk, which kernels execute it, and what goes wrong.

Advertisement

The idea in one paragraph

For a linear layer y = W x, multiplying input channel j of W by a factor s_j and dividing the matching activation x_j by s_j leaves the output unchanged in full precision. Under quantization it does not: scaling a column up gives its weights a larger share of the quantization grid, reducing their relative rounding error, while the matching activation shrinks. When channel j carries large activations, its weights' errors are amplified in the output, so protecting it pays. AWQ chooses one scale per input channel from activation statistics, searches how strongly to apply it, and absorbs the division into the operation that produces x, so inference pays nothing extra.

The pipeline end to end

AWQ is a post-training pass that needs the full-precision model, a calibration set of a few hundred text samples, and enough memory to run one transformer block at a time. The implementations in the original llm-awq code, AutoAWQ and now llm-compressor share the same structure:

AWQ as a pipeline: one transformer block at a timeFP16 modelon CPU or offloadedCalibration seta few hundred samplesRun block icapture input activationsScale search per pairgrid of 20 ratios, min MSEFold scalesinto norm or prior linearClip search (optional)shrink max rangeQuantize + packINT4, group 128, zerosBlock i outputfeeds block i + 1Checkpointpacked weights + scalesServing kerneldequant inside the GEMMnext blockload
AWQ processes the model sequentially. Each block is calibrated on the outputs of the already-processed blocks before it, then quantized and packed.
  1. Run the calibration samples through the embedding layer and capture the inputs to block 0.
  2. For block i, run it once in full precision and record the inputs to each group of linear layers that share an input.
  3. For each such group, search a per-channel scale that minimizes the error of the group's output after quantization.
  4. Fold the chosen scales into the weights and the preceding operation so the block's full-precision function is unchanged.
  5. Optionally search a clipping range per output channel that shrinks the quantization range to cut rounding error on outliers.
  6. Quantize the block's linear weights to INT4 in groups, record the scales and zero points, and compute the block's outputs as the inputs to block i + 1.

Only one block's activations and weights need to be on the GPU at once, which is why a single GPU can quantize a model far larger than it could serve in 16-bit. Calibration cost is roughly a few forward passes over the calibration set per block, times the grid size, so an 8B model typically quantizes in well under an hour on one modern GPU.

Advertisement

Where the scales go: smoothing pairs

A scale on the input channels of a linear layer must be cancelled by dividing the activations, and that division must be folded into whatever produced the activation. That constraint determines which layers AWQ can scale. For a Llama-style block the standard mappings are:

Scaled linears (balance layers)Absorbed into (smooth layer)Notes
q_proj, k_proj, v_projinput_layernorm weightOne shared scale, because all three read the same normalized input
o_projv_proj output rowsOnly when shapes match; with grouped-query attention, AutoAWQ's Llama definition skips it
gate_proj, up_projpost_attention_layernorm weightShared input again
down_projup_proj output rowsValid because SiLU(gate) x up is linear in up

Absorbing into an RMSNorm weight is exact because the norm multiplies each channel by a learned vector; dividing that vector by s is the same as dividing the activation. Absorbing into a preceding linear layer divides its output rows by s. This is why a new architecture needs a mapping definition before AWQ can run on it, and why unusual blocks, such as fused QKV projections, parallel attention and MLP, or MoE experts with a router in between, need architecture-specific handling. llm-compressor infers mappings for known architectures and accepts explicit ones otherwise.

The scale search as code

The search is a one-dimensional grid over an exponent. For each ratio r between 0 and 1, the candidate scale is the mean activation magnitude per channel raised to r; with duo scaling it is also divided by the mean normalized weight magnitude raised to 1 - r. The scale is normalized, applied, the weights fake-quantized, and the group output compared with the full-precision output. This sketch follows that structure:

import torch

def search_scale(x, linears, block_fn, quantize, n_grid=20, duo_scaling=True):
    """x: captured inputs [tokens, in_features] to a group of linears sharing one input.
    block_fn(x) runs the sub-module; quantize(w) fake-quantizes to INT4 per group."""
    x_mean = x.abs().float().mean(dim=0)                        # per input channel
    w_cat = torch.cat([l.weight for l in linears], dim=0)
    w_mean = (w_cat.abs() / w_cat.abs().amax(dim=1, keepdim=True)).mean(dim=0)
    ref = block_fn(x)                                           # fp16 output to match
    originals = [l.weight.data.clone() for l in linears]
    best = (float("inf"), None)
    for i in range(n_grid):
        r = i / n_grid
        s = (x_mean.pow(r) / (w_mean.pow(1 - r) + 1e-4)) if duo_scaling else x_mean.pow(r)
        s = s.clamp(min=1e-4)
        s = s / (s.max() * s.min()).sqrt()                      # keep scales centred near 1
        for l, w0 in zip(linears, originals):
            l.weight.data = quantize(w0 * s) / s                # scale up, quantize, scale back
        err = (block_fn(x) - ref).float().pow(2).mean().item()
        if err < best[0]:
            best = (err, s)
    for l, w0 in zip(linears, originals):
        l.weight.data = w0                                      # restore; caller folds best scale
    return best[1]

Two design choices are worth noticing. The objective is the output error of the whole group on real activations, not the weight reconstruction error, which is what makes the method activation-aware. And the search is over a single scalar per group, which keeps it robust with a small calibration set: there is little to overfit. llm-compressor's default grid has 20 points and duo scaling on.

A worked example makes the effect concrete. Suppose a group of 128 weights has values up to 0.5, so with asymmetric INT4 the step size is about 1 divided by 15, or 0.067, and the worst rounding error per weight is about 0.033. If one input channel's activations average 20 while the others average 1, that channel's error contributes up to 20 x 0.033 = 0.67 to each output. Scaling that weight up by 2 before quantization, without changing the group's maximum, halves its relative error, and dividing the activation by 2 keeps the product exact, so its contribution drops to about 0.33. Scaling too far enlarges the group's maximum and coarsens every other weight's step, which is what the grid search balances.

Quantization, storage and packing

After scaling, weights are quantized per group of consecutive input channels, typically 128, with one 16-bit scale and one 4-bit zero point per group and output channel in the asymmetric scheme. The storage cost is 4 + (16 + 4) / 128, about 4.16 bits per weight. For Llama 3 8B, whose linear layers hold about 7.0 billion of its 8.0 billion parameters, that is about 3.6 GB for the linears plus about 2.1 GB for the embedding and output head, which are usually left in 16-bit: roughly 5.7 GB against 16 GB for the bf16 model. The trade-offs between group sizes are discussed in quantization granularity.

Legacy AutoAWQ checkpoints store tensors named qweight, qzeros and scales, with eight 4-bit values packed into each 32-bit integer in an interleaved order chosen to make dequantization in the kernel cheap. llm-compressor instead saves in its compressed-tensors format, which vLLM reads directly. Either way, the layout is tied to the kernels that consume it, so treat the packing as an implementation detail of your serving engine rather than something to manipulate yourself.

Kernels: why W4A16 is fast at decode

AWQ keeps activations in 16-bit, so the matrix multiply still happens in fp16 or bf16 on tensor cores. The kernel loads packed 4-bit weights from memory, dequantizes them in registers using the group's scale and zero point, and feeds them to the multiply. During decode, a batch of a few tokens multiplies against every weight once, so the step is bound by memory bandwidth, and reading about a quarter of the bytes gives a large speedup, approaching 3 to 4 times on the linear layers at batch 1.

During prefill and at large batch, the multiply becomes compute-bound and the dequantization is extra work, so speedups shrink and a naive kernel can even be slower than fp16. Kernels such as Marlin are designed to keep W4A16 fast across batch sizes by overlapping loads and dequantization with tensor-core work; see the Marlin kernel. vLLM selects a Marlin-based path for 4-bit weight-only checkpoints on GPUs that support it.

Running it: llm-compressor and vLLM

AutoAWQ is deprecated and its functionality was adopted into vLLM's llm-compressor, which is now the recommended tool. This follows the AWQ example in the llm-compressor repository on its main branch in September 2026; the module path of AWQModifier differs in older releases:

from transformers import AutoModelForCausalLM, AutoTokenizer
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
from llmcompressor.modifiers.transform.awq import AWQModifier   # main branch, Sept 2026

MODEL_ID = "meta-llama/Meta-Llama-3-8B-Instruct"
model = AutoModelForCausalLM.from_pretrained(MODEL_ID)
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)

recipe = [
    AWQModifier(),                                   # mappings inferred from the architecture
    QuantizationModifier(ignore=["lm_head"], scheme="W4A16_ASYM", targets=["Linear"]),
]
oneshot(model=model, dataset="perfectblend", splits="train[:512]", recipe=recipe,
        max_seq_length=512, num_calibration_samples=256)

SAVE_DIR = "Meta-Llama-3-8B-Instruct-awq-asym"
model.save_pretrained(SAVE_DIR, save_compressed=True)
tokenizer.save_pretrained(SAVE_DIR)
# serve: vllm serve ./Meta-Llama-3-8B-Instruct-awq-asym   (format detected from the config)

The recipe runs the AWQ scale search and then applies W4A16 asymmetric quantization to every linear layer except the output head. Calibration data matters more than its size: choose text resembling your production prompts, including chat templates and any languages or code you serve. The effect of calibration choices on scales is covered in calibration.

Failure modes

  • Quality loss on one task family. Calibration data did not cover it, for example code or a second language, so the salient channels for that traffic were not protected. Recalibrate with a mixed set.
  • MoE experts degrade. Rarely routed experts receive few calibration tokens and get noisy statistics. Use more samples, check per-expert token counts, and keep the router in 16-bit.
  • Unsupported architecture or fused modules. Without correct mappings, scales are skipped or folded into the wrong place. Verify the mapping list before trusting the output.
  • Output head quantized. The head is sensitive and its logits drive sampling; leave it in 16-bit unless you have measured otherwise.
  • Slower than expected at high throughput. Prefill-heavy or large-batch workloads are compute-bound, where W4A16 gains little; profile with your real traffic, and consider FP8 or W8A8 schemes for those.
  • Out-of-memory during quantization. Long calibration sequences or too many samples on one GPU; reduce sequence length or offload cached activations.

Operational guidance and trade-offs

Treat a quantized checkpoint as a new model: run the same evaluation suite you run on the fp16 model, compare perplexity and task scores side by side, and version the recipe and calibration data with the checkpoint. AWQ's strengths are speed of quantization, robustness with small calibration sets and good behaviour on instruction-tuned models; GPTQ-style error compensation can edge it out at 3 bits, and activation quantization schemes win when you are compute-bound rather than bandwidth-bound.

What to do next

  1. Pick a model and write down the fp16 baseline: evaluation scores, memory and decode latency on your target GPU.
  2. Build a calibration set of 256 to 512 samples drawn from real or representative prompts, formatted with your chat template.
  3. Run the llm-compressor AWQ recipe with W4A16 and group size 128, leaving the output head in 16-bit.
  4. Serve the checkpoint in vLLM and measure decode latency and throughput at your production concurrency.
  5. Compare task scores against the baseline, per task family, and recalibrate if any family drops more than your tolerance.
  6. Record the recipe, calibration set hash and tool versions alongside the checkpoint.
Key takeaway: AWQ protects the weights that matter by scaling their input channels in proportion to activation magnitude. It chooses how strongly to scale with a small grid search against real outputs, and folds the scales into the preceding norm or linear layer so inference pays nothing. The pipeline walks the model one block at a time, quantizes to 4 bits in groups of 128 and packs the result for kernels that dequantize inside the matrix multiply. Its gains are largest in bandwidth-bound decode, and its quality depends on calibration data that matches your traffic.