SmoothQuant is a small idea with a large effect: before quantizing a linear layer to INT8 weights and INT8 activations, divide each input channel of the activations by a factor and multiply the matching row of the weights by the same factor. The layer computes exactly the same thing, but the activations become far easier to quantize. The method comes from Xiao et al., SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.

This article is about building the intuition so you can predict when it will work on your model, choose its one hyperparameter sensibly and recognise when you need something else. The full production pipeline, folding locations and granularity levels are covered in SmoothQuant, in depth. Here we go from the quantization step size, through the reason the obvious fix is impossible, to a runnable experiment you can modify.

Advertisement

One outlier channel ruins the whole token

Symmetric INT8 quantization maps a range of real values onto the integers from -127 to 127 with a single scale: step = max|value| / 127. Every value is rounded to the nearest multiple of the step, so the absolute rounding error is up to half a step regardless of how small the value is.

Now consider one token's activation vector entering a linear layer of a large language model. Most channels hold values of magnitude around 0.5, but a handful of channels hold values around 60, and they are the same channels on almost every token. This pattern is well documented in models above a few billion parameters; the LLM.int8() work called them emergent outlier features. With per-token quantization, that token's step is 60 / 127, about 0.47. A typical value of 0.3 rounds to either 0 or 0.47. Most of the channels now carry one or two bits of information. The outliers are quantized well; everything else is destroyed.

Weights are the opposite. Their magnitudes are fairly uniform across input channels, so INT8 weights with a scale per output channel lose very little. The difficulty sits almost entirely on the activation side.

Why the obvious fix is impossible

If a few channels are large, why not give each input channel its own activation scale? Because of where the scale has to go in the matrix multiply. For token i and output channel k, the layer computes Y[i,k] = sum over j of X[i,j] W[j,k], where j runs over input channels. An integer GEMM wants to compute the sum entirely in integers and multiply by real scales once, at the end.

# scales that can be pulled out of the sum: one per token (i), one per output channel (k)
Y[i, k] = sx[i] * sw[k] * sum_j( Xq[i, j] * Wq[j, k] )        # fine: integer accumulate

# a scale per input channel (j) sits inside the sum
Y[i, k] = sum_j( sx[j] * Xq[i, j] * sw[k] * Wq[j, k] )        # every term has its own factor

A per-input-channel activation scale multiplies each term of the sum by a different number, so the hardware would have to dequantize before accumulating, which throws away the speed of INT8 tensor cores. The allowed granularities are per tensor, per token and per output channel; the granularity article walks through each. The outliers live on exactly the axis that cannot have its own scale.

Advertisement

The seesaw identity

SmoothQuant's move is to change the data instead of the kernel. For any positive vector s with one entry per input channel, X W = (X diag(s)^-1)(diag(s) W). Dividing column j of X by s[j] and multiplying row j of W by s[j] cancels exactly. Choose s large for the outlier channels and the activation's per-token range collapses, while the corresponding weight rows grow.

Per-input-channel magnitude before and after smoothing (alpha 0.5)BeforeAfter: X divided by s, W multiplied by s|X||W|two outlier channels set the INT8 stepactivation range now spans about 2.5xweights are easyweights take part of the rangedivide by smultiply by sThe product X times W is unchanged; only the split of the difficulty moves.
Smoothing moves magnitude from activation channels into the matching weight rows. Activations become flat enough for per-token INT8; weights absorb some unevenness, which per-channel INT8 weights tolerate well.

Two things make this free at inference time. First, s is computed once from calibration data and frozen. Second, the division of X never runs as a separate operation: when the layer's input comes from a LayerNorm or RMSNorm, the division folds into the norm's per-channel gain (and its bias, for LayerNorm), and the multiplication folds into the stored weights. The model after smoothing is just a model with different numbers in it.

Alpha splits a fixed product

The paper sets s[j] = max|X[:,j]|^alpha / max|W[j,:]|^(1-alpha). Write a = max|X[:,j]| and b = max|W[j,:]| for one channel. After smoothing, the activation maximum for that channel is a / s = a^(1-alpha) b^(1-alpha), and the weight maximum is b s = a^alpha b^alpha. Multiply them and you get a b, the same as before. So the product of the activation and weight ranges per channel is invariant, and alpha only decides how that product is split.

At alpha = 0.5 both sides become the square root of a b: the difficulty is shared equally. At alpha = 1 the activations become perfectly flat at the cost of putting the whole range into the weights; at alpha = 0 the opposite. Because weights start out flat and activations start out wildly uneven, the right split is usually near the middle. The paper used 0.5 for most models and a larger value for some whose activation outliers were more severe. Treat any default as a starting point for a sweep.

This view also tells you what smoothing cannot do. If a channel is large in both X and W, the product a b is large and no alpha makes both sides small. Smoothing helps when activation unevenness is large and weight unevenness is small, which is exactly the typical LLM situation.

A runnable experiment

The script below builds a synthetic layer with three channels that are 60 times larger on every token, quantizes activations per token and weights per output channel to INT8, and measures the relative error of the output for different alphas. It runs in under a second with numpy.

import numpy as np

rng = np.random.default_rng(0)
T, C, O = 256, 512, 512
X = rng.normal(0, 1, (T, C))
outliers = [7, 100, 333]
X[:, outliers] *= 60            # a few channels are ~60x larger, on every token
W = rng.normal(0, 0.02, (C, O))
Y = X @ W

def q_per_row(A, bits=8):
    qmax = 2 ** (bits - 1) - 1
    s = np.abs(A).max(axis=1, keepdims=True) / qmax
    return np.round(A / s).clip(-qmax, qmax) * s

def q_per_col(A, bits=8):
    return q_per_row(A.T, bits).T

def rel_err(alpha):
    if alpha is None:
        Xs, Ws = X, W
    else:
        s = np.abs(X).max(axis=0) ** alpha / np.abs(W).max(axis=1) ** (1 - alpha)
        Xs, Ws = X / s, W * s[:, None]
    Yq = q_per_row(Xs) @ q_per_col(Ws)
    return np.linalg.norm(Yq - Y) / np.linalg.norm(Y)

print(f"no smoothing      rel err {rel_err(None):.4f}")
for a in [0.0, 0.25, 0.5, 0.75, 0.9, 1.0]:
    print(f"alpha={a:<4}        rel err {rel_err(a):.4f}")
SettingRelative output error, channel outliersRelative output error, scattered outliers
No smoothing0.04270.0418
alpha 0.00.04240.0424
alpha 0.250.01540.0332
alpha 0.50.00860.0301
alpha 0.750.01710.0318
alpha 0.90.03170.0369
alpha 1.00.04810.0445

The middle column is the script's output exactly as shown. The error curve is U-shaped: alpha 0.5 cuts the error about fivefold, and alpha 1.0 is worse than doing nothing, because the weight rows for the outlier channels now carry the full range and their per-output-channel scales are dominated by them. Alpha 0 is not the same as no smoothing because s = 1 / max|W| still rescales channels slightly.

The right column is a control. Replace the outlier line with a mask that scales random individual entries by 60 at the same overall rate (mask = rng.random((T, C)) less than 3 / C, then X[mask] *= 60). The outliers no longer live in consistent channels, so dividing a channel by a fixed factor cannot target them; the best alpha now only reduces error by about 28 percent. That is the premise of the method made visible: SmoothQuant works because outliers are a property of channels, not of individual tokens.

Read your own model before you trust the method

Collect the absolute maximum per input channel for each linear layer's input over a few hundred calibration samples, and check how concentrated it is. A simple diagnostic is the ratio of the largest channel max to the median channel max, plus the fraction of tokens on which the top channels are also the token's largest values.

import torch

def channel_stats(model, batches, layer_types=(torch.nn.Linear,)):
    stats = {}
    def hook(name):
        def fn(mod, inp, out):
            x = inp[0].detach().float().reshape(-1, inp[0].shape[-1])
            m = x.abs().amax(dim=0)
            stats[name] = torch.maximum(stats[name], m) if name in stats else m
        return fn
    hs = [m.register_forward_hook(hook(n)) for n, m in model.named_modules()
          if isinstance(m, layer_types)]
    with torch.no_grad():
        for b in batches:
            model(**b)
    for h in hs:
        h.remove()
    return {n: (m.max() / m.median()).item() for n, m in stats.items()}

Layers with ratios in the tens or hundreds are where smoothing pays off; layers with ratios of two or three need nothing. Calibration data should resemble production traffic, as discussed in quantization calibration, because the per-channel maxima are what s is built from.

Where the intuition breaks

  • Not every input has a place to fold. Inputs to the query, key, value and first MLP projections follow a norm, so the division folds into its gain. Other inputs need a preceding per-channel linear operation. In a SwiGLU MLP, the down projection's input is the product of the gate activation and the up projection, so the factor can fold into the up projection's output rows; where no fold exists, you insert an explicit multiply or leave that layer unsmoothed.
  • Too high an alpha creates weight outliers. The experiment shows it: past the minimum, error rises because the weight rows inherit the range. Weight quantization error then dominates.
  • Per-tensor static activation scales still hurt. Smoothing flattens channels, not tokens. If token magnitudes vary a lot, a single static scale wastes range; per-token dynamic scales recover it at a small runtime cost.
  • Calibration mismatch. If production prompts activate different channels than the calibration set, s is wrong where it matters. Recalibrate when the traffic changes.
  • Lower bit widths. At 4-bit activations the residual unevenness after smoothing is often too much; rotation methods spread outliers across all channels instead.

How it relates to the other mental models

MethodWhat it does with outliersRuntime cost
LLM.int8()Detects outlier channels and runs them in higher precisionMixed-precision matmul
SmoothQuantMoves activation range into weights with a fixed per-channel scaleNone after folding
AWQScales salient weight channels to protect them, for weight-only quantizationNone after folding
Rotations (QuaRot, SpinQuant)Rotates the basis so no channel stands outExtra transforms unless fused

All four exploit the same observation that a few channels carry most of the magnitude. SmoothQuant is the simplest of them and is the natural first try for W8A8. For 4-bit activations, see rotation-based quantization.

What to do next

  1. Run the experiment above, then change the outlier factor from 60 to 10 and to 200 and note how the best alpha and the gain move.
  2. Run the channel statistics on your model with a few hundred representative prompts and list the layers with the largest max-to-median ratios.
  3. Smooth only those layers first, sweep alpha from 0.3 to 0.9 in steps of 0.1, and evaluate perplexity plus a task benchmark at each value.
  4. Check every smoothed layer has a valid fold location; for any that does not, decide between an explicit scale and leaving it unsmoothed.
  5. Compare per-token dynamic activation scales against static per-tensor scales at your chosen alpha, measuring accuracy and latency.
  6. Re-run the channel statistics on fresh production samples periodically and recalibrate if the outlier channels change.
Key takeaway: Activation outliers in large language models live in a few consistent channels, and because per-input-channel scales cannot leave an integer matrix multiply, those channels inflate the INT8 step for every token. SmoothQuant divides each activation channel by s and multiplies the matching weight row by s, an exact identity that folds into norms and weights. For each channel the product of activation and weight range is fixed, and alpha only chooses the split. The error is U-shaped in alpha and the gain disappears when outliers are not channel-consistent, so measure your model's channel statistics and sweep alpha against a real evaluation.