INT4 quantization stores each weight of a neural network in four bits instead of sixteen. It is the reason a 70-billion-parameter model fits on hardware that could otherwise hold a quarter of it, and the reason many local and server deployments decode faster than they would in FP16. Four bits give only sixteen distinct values, though, and the whole craft of INT4 is choosing which sixteen values each small block of weights is allowed to take.
This article is about the format itself: how the integer grid, scales, zero points and groups fit together, what each choice costs in bits and in error, and what that means for memory, speed and quality. It works a group of weights by hand, gives runnable code to measure the trade-offs on your own matrices, and ends with a decision path. The algorithms that choose better values than plain rounding (GPTQ, AWQ) and the GPU kernels that run the result (INT4 kernels) have their own pages; this one gives you the vocabulary to read them.
Sixteen levels: the integer grid
Four bits encode the integers 0 to 15, or -8 to 7 in two's complement. Quantization maps a real weight w onto that grid with a scale s, the step between neighbouring levels, and optionally a zero point z, the integer that represents real zero. Asymmetric (affine) quantization uses q = clamp(round(w / s) + z, 0, 15) and reconstructs w-hat = s (q - z). Symmetric quantization fixes z at zero, uses the signed range and reconstructs w-hat = s q.
Asymmetric quantization fits the grid to the actual minimum and maximum: s = (max - min) / 15. It wastes no levels when a group's weights are skewed, at the price of storing z. Symmetric quantization uses s = max|w| / 7 (some implementations divide by 8 and clip the top value), is simpler in the kernel and wastes range when the weights are lopsided. Note the signed grid is itself lopsided, -8 to 7, so with s = max|w| / 7 the -8 code is never used and the grid effectively has fifteen levels. The detailed trade-off is in quantization granularity, which also covers per-tensor and per-channel options.
Rounding to the nearest level makes an error of at most s / 2. If weights are spread smoothly across the step, the error is roughly uniform, with mean squared error s2 / 12. Everything that follows is about keeping s small, which means keeping the range each scale must cover small.
Groups: why one scale per row is not enough
With eight bits, a single scale per output channel usually works. With four, it does not: one large weight in a 4,096-wide row stretches the range for all 4,096 weights. The fix is group-wise quantization: split each row into groups of g consecutive input weights and give every group its own scale (and zero). A large weight then damages only its own group.
Groups cost metadata. With a 16-bit scale per group the true storage is 4 + 16 / g bits per weight; adding a 16-bit zero point makes it 4 + 32 / g, and packing the zero point into 4 bits, as several formats do, makes it 4 + 20 / g.
| Group size | FP16 scale only | FP16 scale + 4-bit zero | FP16 scale + FP16 zero |
|---|---|---|---|
| 32 | 4.50 | 4.63 | 5.00 |
| 64 | 4.25 | 4.31 | 4.50 |
| 128 | 4.125 | 4.156 | 4.25 |
| Per row (4,096) | 4.004 | 4.005 | 4.008 |
Group size 128 is the most common default in published INT4 checkpoints because it captures most of the quality of smaller groups for about 4 percent extra storage. Smaller groups help small models and outlier-heavy layers; larger groups save a little memory and make kernels simpler.
A worked group, by hand
Take a group of eight weights: 0.12, -0.05, 0.31, -0.22, 0.02, 0.09, -0.40, 0.18. Quantize symmetrically: max|w| is 0.40, so s = 0.40 / 7 = 0.0571. Dividing and rounding gives the codes 2, -1, 5, -4, 0, 2, -7, 3. Multiplying back gives 0.114, -0.057, 0.286, -0.229, 0.000, 0.114, -0.400, 0.171. The worst error is 0.024, within the bound s / 2 = 0.029, and the weights keep their order and rough proportions.
Now replace 0.02 with an outlier, 1.60. The scale becomes 1.60 / 7 = 0.229, and the other seven weights round to 1, 0, 1, -1, 0, -2, 1. Weights that were 0.12 and 0.31 are now identical; -0.05 and 0.09 both vanish to zero. One value has destroyed the resolution of its seven neighbours. This single example explains most of INT4 engineering: smaller groups limit the blast radius, AWQ rescales channels so the outliers that matter are protected, GPTQ adjusts the remaining weights to compensate for rounding errors, and rotation methods spread outliers out before quantizing.
The outliers that matter most in large language models sit in a small number of input channels, driven by large activations. That is why per-input-channel effects, not just weight magnitudes, decide INT4 quality, and why methods that look at activations during calibration beat plain rounding.
Measuring it yourself
The script below implements asymmetric group-wise round-to-nearest (RTN), nibble packing and an error measurement across group sizes, on a synthetic matrix with one outlier channel. Swap in a real weight matrix from your model to see how its layers behave.
import numpy as np
def quantize_int4(w, group=128):
"""Asymmetric, group-wise round-to-nearest. w: (out, in) float array, in % group == 0."""
out_f, in_f = w.shape
g = w.reshape(out_f, in_f // group, group)
wmin = np.minimum(g.min(axis=2, keepdims=True), 0.0) # keep 0 exactly representable
wmax = np.maximum(g.max(axis=2, keepdims=True), 0.0)
scale = np.maximum((wmax - wmin) / 15.0, 1e-8)
zero = np.clip(np.round(-wmin / scale), 0, 15)
q = np.clip(np.round(g / scale) + zero, 0, 15).astype(np.uint8)
return q.reshape(out_f, in_f), scale.astype(np.float16), zero.astype(np.uint8)
def dequantize_int4(q, scale, zero, group=128):
out_f, in_f = q.shape
g = q.reshape(out_f, in_f // group, group).astype(np.float32)
return ((g - zero) * scale.astype(np.float32)).reshape(out_f, in_f)
def pack_nibbles(q):
"""Two 4-bit codes per byte, low nibble first. Real formats differ: check yours."""
flat = q.reshape(-1)
return (flat[0::2] | (flat[1::2] << 4)).astype(np.uint8)
def unpack_nibbles(packed, shape):
lo, hi = packed & 0x0F, packed >> 4
return np.stack([lo, hi], axis=1).reshape(shape)
rng = np.random.default_rng(0)
W = rng.normal(0, 0.02, size=(4096, 4096)).astype(np.float32)
W[:, 7] *= 25 # one outlier input channel
for group in (32, 128, 1024, 4096):
q, s, z = quantize_int4(W, group)
assert np.array_equal(unpack_nibbles(pack_nibbles(q), q.shape), q)
err = W - dequantize_int4(q, s, z, group)
bits = 4 + (16 + 4) / group # FP16 scale + 4-bit zero per group
print(f"group {group:5d}: {bits:.3f} bits/weight, relative RMSE {np.sqrt((err**2).mean() / (W**2).mean()):.4f}")Two things are worth noticing when you run it. Relative error rises steeply as the group grows past a few hundred, because the outlier channel inflates every group that contains it. And the round trip through pack_nibbles is lossless, but only because the same code packs and unpacks: real formats differ in nibble order, interleaving and whether zero points are stored with an offset, which is why loading a checkpoint with the wrong kernel produces fluent-looking garbage rather than an error. Weight error is also only a proxy; what matters is the error in each layer's output on real activations, which is what GPTQ and AWQ minimise.
Memory and speed arithmetic
A 7-billion-parameter model in FP16 needs 7e9 x 2 bytes = 14 GB for weights. At 4.156 bits per weight (group 128, FP16 scale, 4-bit zero) it needs 7e9 x 4.156 / 8 = 3.6 GB. A 70B model goes from 140 GB to about 36 GB, the difference between multi-GPU serving and a single large accelerator. In practice the result is somewhat larger because embedding and output layers are often left in 16-bit.
INT4 weights do not shrink the KV cache or activations. For long contexts and large batches the KV cache can exceed the weights, so INT4 weights alone may not deliver the capacity you expect; KV-cache quantization is a separate decision.
Speed depends on which resource is the bottleneck. In single-stream decoding every generated token reads every weight once, so decoding is memory-bandwidth bound and four times fewer weight bytes can approach a proportional speed-up, minus the cost of dequantization. In prefill or large-batch serving, the matrix multiplies are compute bound; weight-only INT4 does not reduce the FLOPs, and the dequantization work can make it no faster, or slower, than FP16. Benchmark at your real batch sizes rather than trusting single-stream numbers.
W4A16 versus W4A4: what the hardware actually computes
Almost all deployed INT4 is weight-only, written W4A16: weights stored in four bits, activations in FP16 or BF16, and the kernel dequantizes weight tiles into registers just before a 16-bit matrix multiply. Such kernels need no INT4 arithmetic units at all, which is why they run on any modern GPU and on CPUs.
W4A4 quantizes activations to four bits as well, so the multiply itself can run in low precision. That is much harder: activations have large, input-dependent outliers and are quantized on the fly with no calibration-time optimisation. It needs rotation or smoothing tricks and hardware with a fast low-bit integer path. Support varies by GPU generation; for example, the A100 datasheet lists INT4 tensor-core throughput while the H100 datasheet does not, and newer parts emphasise 4-bit floating-point formats instead. Check the datasheet of the exact part you deploy on before planning around INT4 arithmetic. W4A8 sits in between and is used by some serving stacks.
Integer INT4 is also not the only 4-bit format. NF4 places its sixteen levels at quantiles of a normal distribution instead of evenly; see NF4. FP4 variants are floating point with a tiny exponent. Uniform INT4 remains the most widely supported by fast kernels.
Choosing a method
| Method | Calibration data | Typical use | Note |
|---|---|---|---|
| RTN, group 128 | None | Quick baseline, larger models | Degrades most on small models |
| GPTQ | A few hundred sequences | Offline checkpoints for GPU serving | Order of quantizing columns matters |
| AWQ | Small calibration set | Fast, robust GPU deployments | Protects activation-salient channels |
| NF4 (bitsandbytes) | None | QLoRA fine-tuning, quick loading | Slower kernels than tuned INT4 |
| GGUF k-quants | Optional importance matrix | CPU and local inference | Mixed bit widths per tensor |
A sensible order: start with an existing, well-reviewed INT4 checkpoint of your model if one exists; otherwise run AWQ or GPTQ with calibration text similar to your traffic; fall back to RTN only for a quick capacity estimate. Keep the embedding, output head and any layer that measurably hurts quality in higher precision; mixed precision costs little memory.
Failure modes
- Small models suffer more. A 70B model often loses little at INT4; a 1B to 3B model can lose noticeably. Measure on your size.
- Perplexity hides damage. Perplexity can move slightly while multi-step reasoning, code, long-context retrieval or non-English text degrade more. Evaluate on your tasks; quantization evaluation methodology describes how.
- Calibration mismatch. Calibrating on generic web text and serving code or another language can leave the wrong channels protected.
- Format and kernel mismatch. Packing order, zero-point offset and act-order permutations differ across formats; the wrong loader yields garbage without an error.
- Silent slow paths. If no fast kernel supports your group size, shape or GPU, frameworks may fall back to dequantizing whole matrices, losing the speed benefit. Check the kernel actually used.
- Fine-tuning confusion. Training on INT4 weights directly does not work by ordinary gradient steps; QLoRA trains 16-bit adapters on a frozen 4-bit base, and merging adapters back requires requantizing.
What to do next
- Compute your memory budget: weights at about 4.16 bits per parameter, plus 16-bit embeddings, plus KV cache at your real context and batch.
- Run the error script on a few real layers to see where outliers sit and how group size changes error.
- Quantize with AWQ or GPTQ at group 128 using calibration text drawn from your own traffic.
- Evaluate against the FP16 model on task benchmarks and a sample of real prompts, not perplexity alone.
- Benchmark latency and throughput at production batch sizes and confirm the fast kernel is in use.
- Keep sensitive layers in 16-bit if quality loss is concentrated there, and record the exact format and loader version with the checkpoint.