SmoothQuant, published by Xiao, Lin, Seznec, Wu, Demouth and Han in 2022, is the technique that made 8-bit weights and 8-bit activations (W8A8) practical for large language models. The paper reports up to 1.56x speedup and 2x memory reduction with negligible accuracy loss, across OPT, BLOOM, GLM, MT-NLG, Llama, Falcon, Mistral and Mixtral. It does this without retraining. It applies an offline, mathematically equivalent rescaling that moves quantization difficulty out of the activations and into the weights.

The algebra is short and is derived in SmoothQuant math. This article is about the architecture around it: where the scales physically live in a transformer block, how calibration data turns into those scales, which quantization granularity the kernels then use, and how to build, evaluate and serve a W8A8 model. It ends with the ways it fails and a checklist.

Advertisement

The problem SmoothQuant solves

INT8 matrix multiplication is fast because both operands are 8-bit integers, the tensor cores accumulate in INT32, and one rescale at the end restores real units. Weights are easy to quantize: their values are roughly symmetric and well-behaved per output channel. Activations are not. In LLMs beyond a few billion parameters, a small number of hidden channels carry values tens to hundreds of times larger than the rest, and they do so consistently across tokens.

An activation scale has to cover the largest value it will see. With one scale per tensor, or even per token, a channel whose values reach 60 forces a step size of about 0.47 on an INT8 grid of 127 levels. Every ordinary channel, with values around 1, then gets only two or three usable levels, and the rounding error swamps the signal. You cannot give each activation channel its own scale either. Channel j of X is multiplied by row j of W and summed into every output, so a per-input-channel activation scale cannot be pulled out of the dot product as a single factor. The integer GEMM would need a different rescale inside every accumulation, which the hardware does not do.

Earlier work handled this at runtime. LLM.int8() splits outlier channels into a separate FP16 matmul. That preserves accuracy but adds a gather, a second GEMM and a merge, so it often runs slower than plain FP16. SmoothQuant moves the work offline instead.

The transform in one paragraph

For a linear layer Y = X W, pick a positive vector s with one entry per input channel and write Y = (X diag(s)^-1)(diag(s) W). The product is unchanged. Dividing activation channel j by s_j shrinks the outliers; multiplying weight row j by s_j grows the weights by the same factor. The paper chooses

s_j = max|X_j| ** alpha / max|W_j| ** (1 - alpha)

where max|X_j| comes from calibration and max|W_j| is the largest weight magnitude in row j. Alpha is the migration strength. At 0.5 both sides end up with the same per-channel maximum, sqrt(max|X_j| max|W_j|). Higher alpha pushes more of the difficulty into the weights. The paper used 0.5 for OPT and BLOOM, 0.75 for GLM-130B, whose outliers are more severe, 0.8 for LLaMA, and values between 0.6 and 0.9 for Llama-2, Falcon, Mistral and Mixtral. Weights tolerate the extra range because they are static and can be quantized per output channel, or corrected further by GPTQ.

Advertisement

Where the scales go in a decoder block

An equivalent transform is only free if the 1/s on the activation side costs nothing at inference. SmoothQuant gets this by folding 1/s into whatever operation produces X. In a pre-norm transformer, the input to the Q, K and V projections is the output of a LayerNorm or RMSNorm. The norm ends with an element-wise multiply by gamma, so dividing gamma (and beta, for LayerNorm) by s produces X/s with no extra kernel. Q, K and V all read the same normalized X, so they share one s, computed from the elementwise maximum of their three weight matrices. The MLP input after the second norm works the same way for gate_proj and up_proj, or fc1 in OPT-style models.

Where SmoothQuant scales live in one decoder block (default Llama-style mappings)Residual streamFP16 or BF16input_layernormgamma divided by sq, k, v projW rows times s, INT8Attentionsoftmax in FP16o_projnot smoothedpost_attn_layernormgamma divided by sgate, up projW rows times s, INT8SiLU times upelement-wise, FP16down_projnot smoothedX / sresidual addX / sOffline: calibration hooks record max |X_j|per input channel, then s is folded and the model is savedYellow boxes absorb 1/s for free; green boxes absorb s into their weights. Red boxes have no norm in front of them.
SmoothQuant folds scales into the two norms of each block. The smoothed layers read X/s, and their weights absorb s offline.

The current llm-compressor defaults encode exactly these two fold points: q, k and v balanced against input_layernorm, and gate and up balanced against post_attention_layernorm. Models with fused projections use the fused names, such as query_key_value for BLOOM or qkv_proj and gate_up_proj for Phi-3. Mixtral's mapping smooths attention only, because one norm feeds many experts and a shared s would be a compromise across all of them.

The o_proj and down_proj inputs have no norm directly in front of them. The o_proj input is the attention output and the down_proj input is SiLU(gate) times up. The default mappings do not smooth these two layers. They are still quantized to INT8, relying on per-token activation scales. If evaluation shows a regression concentrated there, exclude those layers from quantization or use a method such as rotation.

The calibration data flow

The only data-dependent quantity is max|X_j|. Calibration runs a few hundred representative sequences through the FP16 model and records, for every smoothed layer, the running maximum of the absolute activation per input channel. The paper used 512 random sentences from the Pile. The vLLM W8A8 guide uses 512 chat-formatted samples of up to 2,048 tokens, which matters for instruction-tuned models: the system prompt, role markers and chat template tokens produce their own activation patterns.

The code below is a minimal, self-contained version of the two steps. It records per-channel maxima with forward hooks, then folds s into a norm and the Linears that read it. It is useful for understanding and for small experiments. For production, use a maintained tool that also handles fused layers, offloading and the quantization step itself.

import torch

@torch.no_grad()
def collect_act_max(model, layer_names, batches):
    """Record max |x_j| per input channel for each named Linear, over calibration data."""
    stats, hooks = {}, []
    for name, mod in model.named_modules():
        if name in layer_names:
            def hook(m, inp, out, name=name):
                x = inp[0].detach().abs().reshape(-1, inp[0].shape[-1]).amax(dim=0).float()
                stats[name] = torch.maximum(stats[name], x) if name in stats else x
            hooks.append(mod.register_forward_hook(hook))
    for batch in batches:
        model(**batch)
    for h in hooks:
        h.remove()
    return stats

@torch.no_grad()
def smooth(norm, linears, act_max, alpha=0.5, eps=1e-5):
    """Fold s into a norm and the Linears that read its output. Output is unchanged."""
    w_max = torch.stack([l.weight.abs().amax(dim=0).float() for l in linears]).amax(dim=0)
    s = (act_max.clamp(min=eps) ** alpha) / (w_max.clamp(min=eps) ** (1 - alpha))
    s = s.clamp(min=eps)
    norm.weight.div_(s.to(norm.weight.dtype))            # X' = X / s
    if getattr(norm, "bias", None) is not None:
        norm.bias.div_(s.to(norm.bias.dtype))
    for l in linears:                                     # W' = W * s (per input column)
        l.weight.mul_(s.to(l.weight.dtype).view(1, -1))

The maximum is taken over every token of every calibration sequence, so one pathological sample, such as a long run of padding or repeated characters, can inflate a channel's maximum and waste its range. Filter calibration data as carefully as training data. For more on range estimation, see quantization calibration.

A worked example with two channels

Take one layer with two input channels. Calibration finds max|X_1| = 60 and max|X_2| = 1.0. The corresponding weight rows have max|W_1| = 0.2 and max|W_2| = 0.5.

QuantityChannel 1Channel 2Ratio of ranges
Before: activation max601.060x
Before: weight max0.20.52.5x
alpha 0.5: s17.31.41
alpha 0.5: activation max3.460.714.9x
alpha 0.5: weight max3.460.714.9x
alpha 0.8: s36.51.15
alpha 0.8: activation max1.640.871.9x
alpha 0.8: weight max7.300.5712.7x

Before smoothing, a per-token INT8 scale is 60/127, about 0.47, so channel 2 has about two levels. At alpha 0.5 the step size becomes 3.46/127, about 0.027, and channel 2 gets about 26 levels. At alpha 0.8 channel 2 gets about 67 levels, but the weight side now spans a 12.7x range across its input rows. That is the trade alpha controls. Past some point, the weight error that you create costs more than the activation error that you remove. The best alpha is model-specific, which is why you sweep it against an evaluation rather than trusting a default.

Granularity: O1, O2, O3 and what kernels run today

Smoothing produces a model that is easy to quantize. You still choose how scales are shared. The paper defines three levels, all with per-tensor weights:

LevelWeightsActivationsCost at runtime
O1per-tensorper-token, dynamica max reduction per token before each GEMM
O2per-tensorper-tensor, dynamicone reduction per tensor
O3per-tensorper-tensor, staticnone; scales are constants

O3 is fastest and least accurate. O1 is most accurate. Current serving stacks usually go one step beyond O1: per-output-channel weight scales plus per-token dynamic activation scales. The llm-compressor W8A8 example saves a model named for exactly that scheme, dynamic per-token. The GEMM then computes an INT32 accumulator and applies both scales in its epilogue, out = acc x scale_token[i] x scale_channel[j], before converting to FP16 or BF16. Because the scales form an outer product, this costs a few multiplies per output element, which the epilogue absorbs. See quantization granularity for how these choices trade accuracy against kernel complexity.

vLLM documents INT8 W8A8 support on NVIDIA GPUs from Turing onward. The paper also runs attention BMMs in INT8, but serving engines usually keep attention in 16-bit and quantize the KV cache separately.

Building and evaluating a W8A8 model

A practical pipeline is: smooth, then quantize weights with GPTQ, then save in a format your server loads. The recipe below follows the vLLM INT8 guide, with the calibration set built from chat data. Recent llm-compressor releases have been moving modules; the README imports oneshot from the package root, so check the docs for your installed version if an import fails.

from datasets import load_dataset
from transformers import AutoModelForCausalLM, AutoTokenizer
from llmcompressor import oneshot                      # older releases: llmcompressor.transformers
from llmcompressor.modifiers.quantization import GPTQModifier
from llmcompressor.modifiers.transform.smoothquant import SmoothQuantModifier  # older: modifiers.smoothquant

MODEL_ID = "meta-llama/Meta-Llama-3-8B-Instruct"
model = AutoModelForCausalLM.from_pretrained(MODEL_ID, device_map="auto", torch_dtype="auto")
tok = AutoTokenizer.from_pretrained(MODEL_ID)

ds = load_dataset("HuggingFaceH4/ultrachat_200k", split="train_sft").shuffle(seed=42).select(range(512))
ds = ds.map(lambda x: {"text": tok.apply_chat_template(x["messages"], tokenize=False)})
ds = ds.map(lambda x: tok(x["text"], max_length=2048, truncation=True, add_special_tokens=False),
            remove_columns=ds.column_names)

recipe = [
    SmoothQuantModifier(smoothing_strength=0.8),          # alpha; sweep 0.5-0.9 per model
    GPTQModifier(targets="Linear", scheme="W8A8", ignore=["lm_head"]),
]
oneshot(model=model, dataset=ds, recipe=recipe, max_seq_length=2048, num_calibration_samples=512)
model.save_pretrained("llama3-8b-w8a8", save_compressed=True)
tok.save_pretrained("llama3-8b-w8a8")

Evaluate the saved model with the same harness and prompts as the FP16 baseline. The guide uses lm-evaluation-harness on GSM8K through vLLM:

lm_eval --model vllm \
  --model_args pretrained="./llama3-8b-w8a8",add_bos_token=true \
  --tasks gsm8k --num_fewshot 5 --batch_size auto

Run the full task rather than a subset, add a task representative of your product, and check long-context behaviour separately.

Serving and operations

W8A8 roughly halves weight memory against FP16, so an 8B model drops from about 16 GB to about 8 GB of weights. Where the speedup comes from depends on the phase. Prefill is compute-bound, so INT8 tensor cores help directly. Decode at small batch sizes is memory-bound, so the gain there comes mostly from reading half the weight bytes. At large batch sizes decode becomes compute-bound again and INT8 GEMM helps again. Measure time to first token and inter-token latency at your real concurrency.

Treat a quantized checkpoint as a new model. Version it with its recipe, alpha and calibration dataset hash, and keep the FP16 model deployable for rollback.

Failure modes

  • Unrepresentative calibration. Plain web text calibrating a chat or code model leaves template and code-token outliers unmeasured, so they clip at runtime.
  • Alpha copied from another model. Too low leaves activation outliers; too high moves them into weight rows and GPTQ cannot fully recover. Sweep it.
  • Wrong mapping for the architecture. A norm that feeds a layer the mapping does not list, or a fused projection under a different name, means s is folded on one side only and the model is no longer equivalent. Test logits against FP16 right after smoothing, before quantizing.
  • Quantizing lm_head or embeddings. The output head is sensitive and small relative to the rest. Leave it in higher precision.
  • Static scales in production. O3-style static activation scales break on inputs outside the calibration distribution. Prefer dynamic per-token scales unless you have measured the difference.

Trade-offs against the alternatives

ApproachWhat it quantizesChoose it when
SmoothQuant W8A8 INT8weights and activationsprefill-heavy or high-batch serving on INT8 tensor cores
Weight-only INT4 (GPTQ, AWQ)weights onlymemory-bound decode at low batch; biggest model on least memory
FP8 W8A8weights and activationshardware with FP8 tensor cores; wider dynamic range and simpler calibration
Rotation (QuaRot, SpinQuant)weights, activations, KVyou need 4-bit activations; outliers are spread by an orthogonal rotation

SmoothQuant is a pre-processing step, not a format, so it combines with GPTQ in INT8 recipes. Rotation-based quantization goes further by spreading outliers out, which costs extra kernels but makes 4-bit activations feasible.

What to do next

  1. Measure the FP16 baseline on your evaluation suite, and record per-task scores with the prompts pinned.
  2. Build a calibration set of 256 to 512 samples from real traffic in the production chat template, with padding and degenerate samples filtered out.
  3. Confirm your model's norm-to-projection mappings, including any fused projection names. After smoothing, check that logits still match FP16 before you quantize.
  4. Sweep alpha at 0.5, 0.65, 0.8 and 0.9 with per-channel weights and dynamic per-token activations, and keep lm_head in high precision.
  5. Compare each candidate on full tasks with paired statistics, then benchmark time to first token and throughput at your real concurrency.
  6. Ship the winner as a versioned artifact with its recipe, alpha and calibration hash, and keep the FP16 model ready for rollback.
Key takeaway: SmoothQuant makes W8A8 work by dividing each activation channel by s_j and multiplying the matching weight row by s_j, so the layer's output is unchanged. In a transformer the 1/s folds into the preceding norm, which makes it free at inference, and the QKV and gate/up projections absorb s. Calibration supplies the per-channel activation maxima, and alpha decides how much difficulty moves into the weights. The main quality levers are the calibration data, the alpha sweep and correct layer mappings. Pair it with per-channel weights, dynamic per-token activations and GPTQ, and evaluate it as a new model.