Quantization stores and computes a model's numbers with fewer bits than it was trained with. Done well, a model shrinks by half or more, decodes faster and fits on cheaper hardware with almost no measurable quality loss. Done carelessly, it looks fine on a perplexity number and then fails on the long prompts or structured outputs your users actually send.

The field now has dozens of named methods and formats, and most pages explain one of them. This one is the map. It builds the core arithmetic from scratch, then answers four questions in order: what is your bottleneck, which tensors should you quantize, in which format, and with which method. It ends with how to evaluate the result and a checklist. Each method mentioned has its own deeper page.

Advertisement

The core idea, with real numbers

Quantization maps a real value x to an integer q using a scale s and, optionally, a zero point z: q = clamp(round(x / s) + z, qmin, qmax). Dequantization reverses it: x' = (q - z) * s. The difference between x and x' is quantization error, at most half a step for values inside the range and unbounded for values clipped at its edges.

Take five weights: -0.62, -0.10, 0.03, 0.41 and 0.88. For asymmetric 8-bit, the range is 1.50, so s = 1.50 / 255 = 0.00588 and z = round(0.62 / 0.00588) = 105. The value 0.41 maps to round(69.7) + 105 = 175 and comes back as (175 - 105) x 0.00588 = 0.412, an error of 0.002. Now try symmetric 4-bit with levels -8 to 7: s = 0.88 / 7 = 0.126. The value 0.41 becomes 3 and returns as 0.377; 0.03 becomes 0 and returns as 0; -0.10 becomes -1 and returns as -0.126. The errors are now around 0.03, fifteen times larger, and one outlier, 0.88, set the step for everyone.

Those two observations drive the whole field. Fewer bits mean coarser steps, and large values in a group make the steps coarse for every small value around them. Almost every method exists to manage outliers, either by isolating them in smaller groups, scaling them away, rotating them into a spread-out form, or compensating for the error they cause. INT8 quantization follows the integer arithmetic all the way into the matmul.

Why fewer bits make inference faster

For a large language model generating one token at a time, every weight is read from memory once per step, and the arithmetic per byte read is tiny. The step is limited by memory bandwidth, not compute. A 70-billion-parameter model is about 140 GB in 16-bit, 70 GB in 8-bit and about 35 GB plus scale overhead in 4-bit. Halving the bytes read roughly halves the time per decode step at small batch, and it can change how many GPUs you need at all.

At large batch sizes, during prompt prefill, or in training, the same weights are reused across many tokens and the matmuls become compute-bound. Then weight-only quantization helps little, because the weights are dequantized back to 16-bit before multiplying. To speed up compute-bound work, the activations must be low precision too, so the tensor cores run INT8 or FP8 math. Knowing which regime you are in is the first decision.

Advertisement

The decision map

The quantization decision map: four questions, in order1. Bottleneck?memory-bound or compute-bound2. What tensors?weights, activations, KV3. Which format?INT8, INT4, FP8, NF4, MX4. Which method?RTN, GPTQ, AWQ, rotation, QATSmall batch decodebytes per token dominateWeight-only: W4A16, W8A16Large batch, prefill, trainingmatmul FLOPs dominateWeights + activations: W8A8, FP8Long context, many usersKV cache fills memoryKV cache quantization: KV8, KV4Then: kernel exists on your hardware? eval against the same baseline?no kernel = no speedup; no eval = no evidence
Decide the bottleneck first; it determines which tensors to quantize, which in turn limits the formats and methods worth trying.

Shorthand helps. W4A16 means 4-bit weights with 16-bit activations; W8A8 means both 8-bit; KV8 means an 8-bit key-value cache. The three tensor groups behave very differently. Weights are fixed, known ahead of time and can be quantized offline with care. Activations change with every input and contain a few channels with values far larger than the rest, so they are harder. The KV cache grows with context length and concurrent users; at long context it can exceed the weights in size, which makes KV cache quantization the lever for capacity rather than speed.

Training adds gradients and optimizer states, which mixed-precision training and low-bit optimizers address. That is a separate topic; this page is about inference.

Granularity and what it costs

The worked example showed one scale for five values. Real tensors have millions, and how many values share a scale is the granularity. Per-tensor uses one scale for the whole matrix: cheap, and fragile against outliers. Per-channel gives each output row its own scale, the standard for INT8 weights. Per-group splits each row into groups, commonly 128 or 64 values, which is the norm for 4-bit weights. Block formats such as the OCP microscaling (MX) formats fix a block of 32 elements that share one 8-bit power-of-two scale.

Smaller groups track local ranges better but store more scales. With a 16-bit scale per group of 128, the overhead is 16 / 128 = 0.125 bits per weight, so INT4 with group 128 is really about 4.125 bits; at group 32 it is 4.5 bits. Asymmetric schemes also store zero points. When you compare file sizes or quality across methods, compare effective bits per weight, not the nominal width.

import torch

def quantize(x, bits=4, group=128, symmetric=True):
    """Fake-quantize the last dimension of x in groups; returns dequantized x."""
    shape = x.shape
    g = x.reshape(-1, group).float()
    if symmetric:
        qmax = 2 ** (bits - 1) - 1                     # 7 for INT4, 127 for INT8
        scale = g.abs().amax(dim=1, keepdim=True).clamp(min=1e-8) / qmax
        q = torch.clamp(torch.round(g / scale), -qmax - 1, qmax)
        deq = q * scale
    else:
        qmax = 2 ** bits - 1                           # 15 for UINT4
        lo, hi = g.amin(dim=1, keepdim=True), g.amax(dim=1, keepdim=True)
        scale = (hi - lo).clamp(min=1e-8) / qmax
        zero = torch.round(-lo / scale)
        q = torch.clamp(torch.round(g / scale) + zero, 0, qmax)
        deq = (q - zero) * scale
    return deq.reshape(shape).to(x.dtype)

def report(W, X):
    """Compare weight error and, more importantly, layer-output error."""
    ref = X @ W.T
    for bits, group in [(8, W.shape[1]), (4, W.shape[1]), (4, 128), (4, 32)]:
        Wq = quantize(W, bits, group)
        out_err = (X @ Wq.T - ref).norm() / ref.norm()
        extra = 16 / group                             # one fp16 scale per group
        print(f"INT{bits} group={group:5d}  bits/weight={bits + extra:5.3f}  "
              f"relative output error={out_err:.4f}")

torch.manual_seed(0)
W = torch.randn(4096, 4096) * 0.02
W[:, :8] *= 20                                         # a few outlier input channels
X = torch.randn(64, 4096)
report(W, X)

The script fake-quantizes a weight matrix with a few outlier input channels and reports the error in the layer's output, which is what matters, rather than the error in the weights. Expect INT8 per-row to be nearly lossless, INT4 per-row to be clearly worse, and smaller groups to recover much of the gap for a fraction of a bit. Try moving the outliers to a few output rows instead and see how per-row scales absorb them.

Number formats

FormatLayoutTypical useNotes
INT88-bit integer + scaleW8A8 and weight-only inferenceMature kernels nearly everywhere
INT44-bit integer + per-group scaleW4A16 weight-only LLM servingNeeds a method beyond round-to-nearest at scale
FP8 E4M3 / E5M21 sign, 4 or 5 exponent, 3 or 2 mantissa bitsW8A8 inference and training on GPUs with FP8 tensor coresE4M3 for weights and activations; E5M2's range suits gradients
NF44-bit codes mapped to normal-distribution quantilesQLoRA fine-tuning of frozen weightsStorage format; computed in 16-bit
MX (MXFP8, MXFP6, MXFP4, MXINT8)32-element block, shared 8-bit power-of-two scaleEmerging hardware-native low-bit formatsDefined by the OCP MX specification; check your hardware's support

Integers spread levels evenly; floating-point formats spend levels near zero, where most weights and activations live, and keep a wide range for the rest. That is why FP8 often needs less careful calibration than INT8 for activations. FP8 quantization goes deeper on the two FP8 variants and their scaling.

Methods, from simplest to strongest

Round to nearest. Compute scales from each group's range and round. For INT8 weights it is usually enough. At 4 bits on large language models it loses noticeable accuracy, and at 3 bits it often breaks the model.

Static versus dynamic activations. Static quantization fixes activation scales from a calibration set; dynamic quantization computes them per token at runtime, which costs a reduction but tracks outliers better. Calibration data should resemble real traffic.

Error-compensating PTQ. GPTQ quantizes a layer's weights column by column and updates the not-yet-quantized columns to cancel the error, using second-order information from calibration activations. It is the workhorse for 4-bit weights; see GPTQ.

Activation-aware scaling. AWQ observes that a small fraction of weight channels matter most because their input activations are large, and scales those channels up before quantization so they lose less precision. SmoothQuant uses a related per-channel rescaling to move activation outliers into the weights so W8A8 becomes feasible.

Rotations. QuaRot and SpinQuant multiply weights and activations by orthogonal matrices that leave the network's function unchanged but spread outliers across all channels, making 4-bit activations and KV caches practical.

Quantization-aware training. Simulate quantization in the forward pass during fine-tuning so the model adapts to it, passing gradients through the rounding with a straight-through estimator. It gives the best accuracy at the lowest widths and costs a training run.

Kernels decide whether you get a speedup

A quantized checkpoint is only a file. Speed comes from a kernel that reads packed low-bit weights, dequantizes them in registers and feeds the tensor cores, or that runs INT8 or FP8 matmuls directly. If your serving stack has no kernel for your exact format, group size and GPU, it may silently dequantize the whole layer to 16-bit and run slower than the original. Before choosing a method, check which formats your runtime accelerates on your hardware, then pick among the methods that produce those formats. Weight-only 4-bit kernels shine at small batch and fade at large batch; FP8 and INT8 W8A8 hold up as batch grows.

Proving it worked

Perplexity on held-out text is a quick smoke test and a weak acceptance test: it averages over easy tokens and can hide failures on reasoning, code, long context and strict output formats. Evaluate the quantized model and the original with the same harness, prompts, decoding settings and seeds, on the tasks you actually ship, and set a tolerance in advance. The shape of such a harness is below; quantization evaluation methodology covers suite design and statistical noise.

# Evaluation harness shape: same prompts, same decoding, same scorer for every variant.
variants = {"bf16": load("model", dtype="bf16"),
            "w8a8": load("model-w8a8"),
            "w4a16_g128": load("model-w4a16-g128")}
suites = ["perplexity_heldout", "task_suite_you_ship", "long_context_32k", "format_following"]

results = {}
for name, model in variants.items():
    for suite in suites:
        results[name, suite] = run_suite(model, suite, temperature=0.0, seed=0)

for name in variants:
    deltas = {s: results[name, s] - results["bf16", s] for s in suites}
    worst = min(deltas.values())
    print(name, deltas, "PASS" if worst >= -TOLERANCE else "FAIL")

Measure performance on the same footing: time to first token, decode tokens per second and memory at the batch sizes and context lengths you serve, not a single best-case number.

Failure modes

  • Quality cliff on long inputs. Errors accumulate over long contexts, especially with a quantized KV cache. Include long-context cases in evaluation.
  • Calibration mismatch. Calibrating on generic web text and serving code or another language shifts activation ranges. Calibrate on representative traffic.
  • No speedup or a slowdown. Missing kernel, unsupported group size, or a compute-bound workload where weight-only quantization cannot help.
  • Broken structured output. JSON or tool-call formats degrade before free-text quality does. Test them explicitly.
  • Sensitive layers. Embeddings, the output head and some early or late layers suffer most; keeping them at higher precision is common and cheap.
  • Re-quantizing a quantized model. Errors compound. Always quantize from the original high-precision weights.

What to do next

  1. Measure your workload's bottleneck: batch size, context length and whether decode is bandwidth-bound on your hardware.
  2. List the formats your serving runtime accelerates on your GPUs or CPUs, with supported group sizes.
  3. Run the fake-quantization script on one real layer of your model and compare output error across widths and group sizes.
  4. Try INT8 or FP8 first; move to 4-bit weights with GPTQ or AWQ only if memory or speed still fall short.
  5. Build the evaluation harness with your own tasks, a long-context case and a structured-output case, and fix a tolerance before looking at results.
  6. Keep the embeddings and output head in higher precision if evaluation shows they are sensitive.
  7. Add KV cache quantization only when cache memory, not weights, limits concurrency.
Key takeaway: Quantization replaces real numbers with a scale and a small integer or low-bit float, and almost every technique in the field exists to stop outliers from making those steps too coarse. Start from your bottleneck: weight-only quantization speeds memory-bound decode, weights plus activations speed compute-bound work, and KV cache quantization buys context and concurrency. Choose a format your hardware has kernels for, the simplest method that meets your tolerance, and judge the result with the same evaluation you would use for any model change.