INT4 quantization stores each weight of a neural network in four bits instead of sixteen. It is the reason a 70-billion-parameter model fits on hardware that could otherwise hold a quarter of it, and the reason many local and server deployments decode faster than they would in FP16. Four bits give only sixteen distinct values, though, and the whole craft of INT4 is choosing which sixteen values each small block of weights is allowed to take.

This article is about the format itself: how the integer grid, scales, zero points and groups fit together, what each choice costs in bits and in error, and what that means for memory, speed and quality. It works a group of weights by hand, gives runnable code to measure the trade-offs on your own matrices, and ends with a decision path. The algorithms that choose better values than plain rounding (GPTQ, AWQ) and the GPU kernels that run the result (INT4 kernels) have their own pages; this one gives you the vocabulary to read them.

Advertisement

Sixteen levels: the integer grid

Four bits encode the integers 0 to 15, or -8 to 7 in two's complement. Quantization maps a real weight w onto that grid with a scale s, the step between neighbouring levels, and optionally a zero point z, the integer that represents real zero. Asymmetric (affine) quantization uses q = clamp(round(w / s) + z, 0, 15) and reconstructs w-hat = s (q - z). Symmetric quantization fixes z at zero, uses the signed range and reconstructs w-hat = s q.

Asymmetric quantization fits the grid to the actual minimum and maximum: s = (max - min) / 15. It wastes no levels when a group's weights are skewed, at the price of storing z. Symmetric quantization uses s = max|w| / 7 (some implementations divide by 8 and clip the top value), is simpler in the kernel and wastes range when the weights are lopsided. Note the signed grid is itself lopsided, -8 to 7, so with s = max|w| / 7 the -8 code is never used and the grid effectively has fifteen levels. The detailed trade-off is in quantization granularity, which also covers per-tensor and per-channel options.

Rounding to the nearest level makes an error of at most s / 2. If weights are spread smoothly across the step, the error is roughly uniform, with mean squared error s2 / 12. Everything that follows is about keeping s small, which means keeping the range each scale must cover small.

Groups: why one scale per row is not enough

FP16 weight row4,096 input featuresSplit into groupse.g. 32 groups of 128Per group: scale s, zero zfrom min/max or an optimiserq = clamp(round(w / s) + z)integers 0..15Pack 8 codes per uint32plus FP16 scales, zerosStored modelabout 4.1 to 4.3 bits/weightKernel: load packed4x fewer bytes than FP16Dequantize in registersw = s (q - z)FP16 / BF16 matmulwith 16-bit activations (W4A16)Only storage and memory traffic are 4-bit. The arithmetic in a weight-only INT4 kernel is still 16-bit floating point.Bits per weight = 4 + (bits of scale + bits of zero point) / group size.
Group-wise weight-only INT4: each group of consecutive input weights gets its own scale and zero point; the kernel loads packed 4-bit codes and dequantizes to 16-bit before multiplying.

With eight bits, a single scale per output channel usually works. With four, it does not: one large weight in a 4,096-wide row stretches the range for all 4,096 weights. The fix is group-wise quantization: split each row into groups of g consecutive input weights and give every group its own scale (and zero). A large weight then damages only its own group.

Groups cost metadata. With a 16-bit scale per group the true storage is 4 + 16 / g bits per weight; adding a 16-bit zero point makes it 4 + 32 / g, and packing the zero point into 4 bits, as several formats do, makes it 4 + 20 / g.

Group sizeFP16 scale onlyFP16 scale + 4-bit zeroFP16 scale + FP16 zero
324.504.635.00
644.254.314.50
1284.1254.1564.25
Per row (4,096)4.0044.0054.008

Group size 128 is the most common default in published INT4 checkpoints because it captures most of the quality of smaller groups for about 4 percent extra storage. Smaller groups help small models and outlier-heavy layers; larger groups save a little memory and make kernels simpler.

Advertisement

A worked group, by hand

Take a group of eight weights: 0.12, -0.05, 0.31, -0.22, 0.02, 0.09, -0.40, 0.18. Quantize symmetrically: max|w| is 0.40, so s = 0.40 / 7 = 0.0571. Dividing and rounding gives the codes 2, -1, 5, -4, 0, 2, -7, 3. Multiplying back gives 0.114, -0.057, 0.286, -0.229, 0.000, 0.114, -0.400, 0.171. The worst error is 0.024, within the bound s / 2 = 0.029, and the weights keep their order and rough proportions.

Now replace 0.02 with an outlier, 1.60. The scale becomes 1.60 / 7 = 0.229, and the other seven weights round to 1, 0, 1, -1, 0, -2, 1. Weights that were 0.12 and 0.31 are now identical; -0.05 and 0.09 both vanish to zero. One value has destroyed the resolution of its seven neighbours. This single example explains most of INT4 engineering: smaller groups limit the blast radius, AWQ rescales channels so the outliers that matter are protected, GPTQ adjusts the remaining weights to compensate for rounding errors, and rotation methods spread outliers out before quantizing.

The outliers that matter most in large language models sit in a small number of input channels, driven by large activations. That is why per-input-channel effects, not just weight magnitudes, decide INT4 quality, and why methods that look at activations during calibration beat plain rounding.

Measuring it yourself

The script below implements asymmetric group-wise round-to-nearest (RTN), nibble packing and an error measurement across group sizes, on a synthetic matrix with one outlier channel. Swap in a real weight matrix from your model to see how its layers behave.

import numpy as np

def quantize_int4(w, group=128):
    """Asymmetric, group-wise round-to-nearest. w: (out, in) float array, in % group == 0."""
    out_f, in_f = w.shape
    g = w.reshape(out_f, in_f // group, group)
    wmin = np.minimum(g.min(axis=2, keepdims=True), 0.0)    # keep 0 exactly representable
    wmax = np.maximum(g.max(axis=2, keepdims=True), 0.0)
    scale = np.maximum((wmax - wmin) / 15.0, 1e-8)
    zero = np.clip(np.round(-wmin / scale), 0, 15)
    q = np.clip(np.round(g / scale) + zero, 0, 15).astype(np.uint8)
    return q.reshape(out_f, in_f), scale.astype(np.float16), zero.astype(np.uint8)

def dequantize_int4(q, scale, zero, group=128):
    out_f, in_f = q.shape
    g = q.reshape(out_f, in_f // group, group).astype(np.float32)
    return ((g - zero) * scale.astype(np.float32)).reshape(out_f, in_f)

def pack_nibbles(q):
    """Two 4-bit codes per byte, low nibble first. Real formats differ: check yours."""
    flat = q.reshape(-1)
    return (flat[0::2] | (flat[1::2] << 4)).astype(np.uint8)

def unpack_nibbles(packed, shape):
    lo, hi = packed & 0x0F, packed >> 4
    return np.stack([lo, hi], axis=1).reshape(shape)

rng = np.random.default_rng(0)
W = rng.normal(0, 0.02, size=(4096, 4096)).astype(np.float32)
W[:, 7] *= 25                                   # one outlier input channel
for group in (32, 128, 1024, 4096):
    q, s, z = quantize_int4(W, group)
    assert np.array_equal(unpack_nibbles(pack_nibbles(q), q.shape), q)
    err = W - dequantize_int4(q, s, z, group)
    bits = 4 + (16 + 4) / group                  # FP16 scale + 4-bit zero per group
    print(f"group {group:5d}: {bits:.3f} bits/weight, relative RMSE {np.sqrt((err**2).mean() / (W**2).mean()):.4f}")

Two things are worth noticing when you run it. Relative error rises steeply as the group grows past a few hundred, because the outlier channel inflates every group that contains it. And the round trip through pack_nibbles is lossless, but only because the same code packs and unpacks: real formats differ in nibble order, interleaving and whether zero points are stored with an offset, which is why loading a checkpoint with the wrong kernel produces fluent-looking garbage rather than an error. Weight error is also only a proxy; what matters is the error in each layer's output on real activations, which is what GPTQ and AWQ minimise.

Memory and speed arithmetic

A 7-billion-parameter model in FP16 needs 7e9 x 2 bytes = 14 GB for weights. At 4.156 bits per weight (group 128, FP16 scale, 4-bit zero) it needs 7e9 x 4.156 / 8 = 3.6 GB. A 70B model goes from 140 GB to about 36 GB, the difference between multi-GPU serving and a single large accelerator. In practice the result is somewhat larger because embedding and output layers are often left in 16-bit.

INT4 weights do not shrink the KV cache or activations. For long contexts and large batches the KV cache can exceed the weights, so INT4 weights alone may not deliver the capacity you expect; KV-cache quantization is a separate decision.

Speed depends on which resource is the bottleneck. In single-stream decoding every generated token reads every weight once, so decoding is memory-bandwidth bound and four times fewer weight bytes can approach a proportional speed-up, minus the cost of dequantization. In prefill or large-batch serving, the matrix multiplies are compute bound; weight-only INT4 does not reduce the FLOPs, and the dequantization work can make it no faster, or slower, than FP16. Benchmark at your real batch sizes rather than trusting single-stream numbers.

W4A16 versus W4A4: what the hardware actually computes

Almost all deployed INT4 is weight-only, written W4A16: weights stored in four bits, activations in FP16 or BF16, and the kernel dequantizes weight tiles into registers just before a 16-bit matrix multiply. Such kernels need no INT4 arithmetic units at all, which is why they run on any modern GPU and on CPUs.

W4A4 quantizes activations to four bits as well, so the multiply itself can run in low precision. That is much harder: activations have large, input-dependent outliers and are quantized on the fly with no calibration-time optimisation. It needs rotation or smoothing tricks and hardware with a fast low-bit integer path. Support varies by GPU generation; for example, the A100 datasheet lists INT4 tensor-core throughput while the H100 datasheet does not, and newer parts emphasise 4-bit floating-point formats instead. Check the datasheet of the exact part you deploy on before planning around INT4 arithmetic. W4A8 sits in between and is used by some serving stacks.

Integer INT4 is also not the only 4-bit format. NF4 places its sixteen levels at quantiles of a normal distribution instead of evenly; see NF4. FP4 variants are floating point with a tiny exponent. Uniform INT4 remains the most widely supported by fast kernels.

Choosing a method

MethodCalibration dataTypical useNote
RTN, group 128NoneQuick baseline, larger modelsDegrades most on small models
GPTQA few hundred sequencesOffline checkpoints for GPU servingOrder of quantizing columns matters
AWQSmall calibration setFast, robust GPU deploymentsProtects activation-salient channels
NF4 (bitsandbytes)NoneQLoRA fine-tuning, quick loadingSlower kernels than tuned INT4
GGUF k-quantsOptional importance matrixCPU and local inferenceMixed bit widths per tensor

A sensible order: start with an existing, well-reviewed INT4 checkpoint of your model if one exists; otherwise run AWQ or GPTQ with calibration text similar to your traffic; fall back to RTN only for a quick capacity estimate. Keep the embedding, output head and any layer that measurably hurts quality in higher precision; mixed precision costs little memory.

Failure modes

  • Small models suffer more. A 70B model often loses little at INT4; a 1B to 3B model can lose noticeably. Measure on your size.
  • Perplexity hides damage. Perplexity can move slightly while multi-step reasoning, code, long-context retrieval or non-English text degrade more. Evaluate on your tasks; quantization evaluation methodology describes how.
  • Calibration mismatch. Calibrating on generic web text and serving code or another language can leave the wrong channels protected.
  • Format and kernel mismatch. Packing order, zero-point offset and act-order permutations differ across formats; the wrong loader yields garbage without an error.
  • Silent slow paths. If no fast kernel supports your group size, shape or GPU, frameworks may fall back to dequantizing whole matrices, losing the speed benefit. Check the kernel actually used.
  • Fine-tuning confusion. Training on INT4 weights directly does not work by ordinary gradient steps; QLoRA trains 16-bit adapters on a frozen 4-bit base, and merging adapters back requires requantizing.

What to do next

  1. Compute your memory budget: weights at about 4.16 bits per parameter, plus 16-bit embeddings, plus KV cache at your real context and batch.
  2. Run the error script on a few real layers to see where outliers sit and how group size changes error.
  3. Quantize with AWQ or GPTQ at group 128 using calibration text drawn from your own traffic.
  4. Evaluate against the FP16 model on task benchmarks and a sample of real prompts, not perplexity alone.
  5. Benchmark latency and throughput at production batch sizes and confirm the fast kernel is in use.
  6. Keep sensitive layers in 16-bit if quality loss is concentrated there, and record the exact format and loader version with the checkpoint.
Key takeaway: INT4 stores weights on a sixteen-level grid whose scale is set per small group, so true storage is a little over four bits per weight. Outliers stretch the grid and are the root of most quality loss, which is why group size, AWQ and GPTQ matter. Weight-only INT4 cuts memory roughly fourfold and speeds memory-bound decoding, but not compute-bound prefill; evaluate on your own tasks and confirm the fast kernel is actually running.