Most explanations of INT8 quantization stop at the formula that maps a float to an 8-bit integer. That is the easy part. The speed and most of the bugs live in what happens next: how two int8 tensors are multiplied without ever turning back into floats, why the sum must be held in 32 bits, how a single integer multiply and shift replaces three floating-point scales, and which instructions in the CPU or GPU actually do the work. Understand that pipeline and you can predict where a model loses accuracy and why a slow INT8 model is slow.

This article walks through one quantized linear layer from first principles, gives a NumPy reference you can run and compare against, works a real example with the numbers it prints, and then covers the hardware paths, a deployment flow with ONNX Runtime, debugging, failure modes and trade-offs. Choosing scales is covered elsewhere on this site and is only summarised here.

Advertisement

What INT8 changes and what it does not

INT8 stores each value in one byte instead of four (FP32) or two (FP16 or BF16). That alone cuts weight memory and memory traffic by four or two times. The second benefit is compute: CPUs and GPUs have instructions that multiply groups of int8 values and add them into a 32-bit sum in one step, so an int8 matrix multiply can do several times more operations per cycle than the float equivalent on the same hardware.

Two deployments are both called INT8 and they behave differently. In W8A8, both weights and activations are int8 and the matrix multiply runs on integer units; this is what gives the compute speed-up and what this article is mostly about. In weight-only INT8 (often written W8A16), weights are stored as int8 but converted to FP16 just before the multiply. That saves memory and bandwidth, which is what matters for small-batch LLM decoding, but it never touches integer arithmetic. Knowing which one you are running explains most surprises in benchmarks.

The mapping, briefly

A real value r is represented by an integer q through r = s(q - z), where s is a positive float scale and z is an integer zero point. Symmetric quantization fixes z = 0; asymmetric quantization picks z so that the observed minimum and maximum land on the ends of the integer range. Weights are almost always symmetric, often restricted to [-127, 127] so the range is balanced, and get one scale per output channel. Activations after a ReLU or GELU are lopsided, so they are often asymmetric with one scale per tensor. The detailed trade-offs are in symmetric versus asymmetric quantization and quantization granularity, and choosing the ranges themselves is calibration.

One constraint is easy to forget: real zero must map exactly to an integer. Padding in convolutions and masked positions insert zeros, and if zero is represented as 0.4 of a step, every padded value carries a bias into the sum. Computing z by rounding guarantees exactness.

Advertisement

The integer GEMM and the zero-point algebra

Take one output of a linear layer: y = sum over k of x_k times w_k. Substituting the mapping for both operands gives a sum of s_x s_w (q_x - z_x)(q_w - z_w). The scales are constants, so they factor out of the sum, and what remains expands into four terms: the sum of q_x times q_w, minus z_w times the sum of q_x, minus z_x times the sum of q_w, plus K times z_x times z_w, where K is the reduction length.

Only the first term is a real matrix multiply, and it is exactly what int8 hardware computes. The third and fourth terms depend only on the weights and the activation zero point, so they are computed once when the model is loaded. The second term needs a row sum of the activations at run time, which is cheap but not free. This is the main reason weights are kept symmetric: with z_w = 0, the second and fourth terms vanish and the correction is a single precomputed vector per layer.

Each int8 by int8 product is at most 128 times 128, which is 16,384, or 2 to the power 14. A signed 32-bit accumulator holds up to about 2 to the power 31, so in the worst case it can absorb about 2 to the power 17, or 131,072, products before overflowing. Real data is far from the worst case, but check headroom when K is very large.

One quantized linear layer: int8 in, int32 inside, int8 out, with one fixed-point multiply per output channelActivations q_xint8, scale s_x, zp z_xWeights q_wint8, per-channel s_wInteger GEMMint8 x int8, int32 sumPrecomputed termsz_x x colsum(W), biasaddRequantizex m0, shift, roundint32Add z_y, clampfused ReLU hereOutput q_yint8 to next layerFloat values never appear between layers: the scales live only in the multiplier m0 and shiftBias is stored as int32 at scale s_x times s_w, so it adds straight into the accumulator
A W8A8 linear layer. The integer GEMM produces an int32 sum, the precomputed zero-point term and int32 bias are added, and a single fixed-point multiply and shift per output channel brings the result back to int8.

Requantization: three scales become one integer multiply

After accumulation, the int32 value acc represents s_x s_w acc in real units. The next layer wants an int8 with its own scale s_y and zero point z_y, so the output is q_y = z_y + round(M times acc), with M = s_x s_w / s_y. M is a positive real number, usually much smaller than one. Integer-only kernels do not multiply by a float. They write M as m0 times 2 to the power of minus shift, with m0 normalised into [0.5, 1) and stored as a 32-bit fixed-point integer, and then compute a 64-bit product followed by a rounding right shift. This is the scheme from Jacob and colleagues' 2018 paper on integer-arithmetic-only inference.

Two more operations fold in for free. The bias is quantized to int32 at the accumulator scale s_x s_w, so it adds directly into acc before requantization. A ReLU becomes a clamp: since real zero is z_y, clamping the output to at least z_y is the activation. A fused layer therefore reads int8, writes int8 and never materialises a float tensor.

A reference implementation you can run

The code below quantizes one linear layer with asymmetric per-tensor activations and symmetric per-channel weights, runs the integer pipeline exactly as described, and compares it with the float result. It is slow and meant for checking kernels and understanding, not for production.

import numpy as np

def quant_params_asym(lo, hi, qmin=-128, qmax=127):
    lo, hi = min(lo, 0.0), max(hi, 0.0)          # range must contain real zero
    scale = (hi - lo) / (qmax - qmin)
    zp = int(np.clip(round(qmin - lo / scale), qmin, qmax))
    return scale, zp

def quantize(x, scale, zp):
    return np.clip(np.round(x / scale) + zp, -128, 127).astype(np.int8)

def quantize_multiplier(m):                       # m = m0 * 2**-shift, m0 in Q31
    shift = 0
    while m < 0.5:
        m *= 2.0
        shift += 1
    m0 = int(round(m * (1 << 31)))
    if m0 == 1 << 31:                             # m rounded up to 1.0
        m0, shift = m0 // 2, shift - 1
    return m0, shift

def requantize(acc, m0, shift, zp_out):
    prod = acc.astype(np.int64) * m0
    total = 31 + shift
    out = (prod + (1 << (total - 1))) >> total    # rounding right shift
    return np.clip(out + zp_out, -128, 127).astype(np.int8)

rng = np.random.default_rng(0)
M_, K, N = 32, 256, 64
x = rng.uniform(-1.0, 5.0, size=(M_, K)).astype(np.float32)
w = (rng.standard_normal((K, N)) * 0.05).astype(np.float32)
b = (rng.standard_normal(N) * 0.1).astype(np.float32)

s_x, z_x = quant_params_asym(x.min(), x.max())
s_w = np.abs(w).max(axis=0) / 127.0               # per output channel
q_x = quantize(x, s_x, z_x)
q_w = np.clip(np.round(w / s_w), -127, 127).astype(np.int8)
q_b = np.round(b / (s_x * s_w)).astype(np.int32)  # bias at accumulator scale

y_ref = x @ w + b
s_y, z_y = quant_params_asym(y_ref.min(), y_ref.max())

acc = q_x.astype(np.int32) @ q_w.astype(np.int32)
acc -= z_x * q_w.astype(np.int32).sum(axis=0)     # precomputable at load time
acc += q_b
q_y = np.empty((M_, N), dtype=np.int8)
for n in range(N):
    m0, sh = quantize_multiplier(s_x * s_w[n] / s_y)
    q_y[:, n] = requantize(acc[:, n], m0, sh, z_y)

y_int8 = (q_y.astype(np.float32) - z_y) * s_y
print("max error in output steps:", np.abs(y_int8 - y_ref).max() / s_y)

Worked example: reading the numbers

Running the reference with that seed prints the following. The activations span -1 to 5, so s_x is about 0.0235 and the zero point is -86: real zero sits 86 steps below the middle of the range, which is what an asymmetric range looks like. The output range gives s_y of about 0.0503 and z_y of -12.

The largest accumulator magnitude is 245,864, which needs 18 bits, leaving 13 bits of headroom in int32. For the first output channel, M is about 0.000504, stored as m0 = 1,108,445,824 with a shift of 10. The integer result differs from the float result by at most 1.53 output steps and by 0.35 steps on average, and the signal-to-quantization-noise ratio (SQNR) is about 39.5 dB.

How should you read this? Requantization alone can only cost half a step. The rest comes from rounding the inputs and weights, which is unavoidable and is why the error exceeds one step in the worst case. An SQNR near 40 dB for a single layer is healthy; when a layer in a real model falls into the low twenties, it is usually the one hurting accuracy. If your production kernel disagrees with a reference like this by more than one step on any element, suspect the rounding mode, the shift or the zero-point term before you suspect the model.

How the hardware executes it

On x86, the VNNI extensions add VPDPBUSD, which multiplies unsigned 8-bit values by signed 8-bit values in groups of four and adds them into 32-bit lanes. The unsigned-times-signed shape is why x86 backends often store activations as uint8 with a zero point near 128 and weights as int8. Older AVX2 code uses VPMADDUBSW, whose intermediate pair sums saturate at 16 bits; some x86 backends therefore offer a reduced-range option that quantizes activations to 7 bits to avoid saturation, at a small accuracy cost.

On Arm, the dot-product extension adds SDOT and UDOT (four int8 products into each 32-bit lane), and the I8MM extension adds matrix-multiply instructions such as SMMLA that work on small 2 by 8 tiles. On NVIDIA GPUs, DP4A performs a four-way int8 dot product in the regular cores, and INT8 tensor cores perform whole tile multiplies with int32 accumulation. Support and throughput vary by generation and change with each one, so check the vendor documentation for your target rather than relying on a general ratio.

The software lesson is the same everywhere: integer speed only appears if the runtime selects a fused int8 kernel. A graph full of quantize, dequantize and float operations can be slower than the float model, because it pays for conversions and gets no integer GEMM.

A deployment flow with ONNX Runtime

ONNX Runtime ships a quantization tool that is a good reference workflow. You export a float model, write a small calibration reader that yields a few hundred representative inputs, and run static quantization in QDQ format, which inserts QuantizeLinear and DequantizeLinear pairs around each tensor. At load time the runtime recognises a pattern such as DequantizeLinear, MatMul, QuantizeLinear and replaces it with one integer kernel, so the QDQ graph is a description of the arithmetic above rather than something executed literally. TensorRT also consumes explicit Q and DQ nodes.

from onnxruntime.quantization import (CalibrationDataReader, QuantFormat,
                                      QuantType, quantize_static)

class Reader(CalibrationDataReader):
    def __init__(self, batches, input_name):
        self.it = iter([{input_name: b} for b in batches])
    def get_next(self):
        return next(self.it, None)               # None ends calibration

quantize_static(
    "model_fp32.onnx", "model_int8.onnx",
    Reader(calibration_batches, "input"),
    quant_format=QuantFormat.QDQ,
    per_channel=True,                             # one scale per output channel
    activation_type=QuantType.QInt8,
    weight_type=QuantType.QInt8,
)

Dynamic quantization is the other option: weights are quantized ahead of time and activation scales are computed per batch at run time. It needs no calibration and suits LSTMs and transformers on CPU, at the cost of a reduction over each activation tensor per call. For large language models, W8A8 runs into activation outliers that a single scale cannot cover; SmoothQuant moves them into the weights and LLM.int8() routes them through a float path.

Debugging accuracy layer by layer

When the quantized model loses more accuracy than expected, do not tune globally. Run the float and quantized models on the same batch, capture every layer output, dequantize, and compute SQNR per layer as in the example. A layer whose SQNR drops sharply relative to its neighbours is the culprit. Common ones are the first and last layers, depthwise convolutions with per-tensor weights, attention projections with outlier channels, and anything feeding a softmax. Fixes, cheapest first: per-channel scales, percentile or entropy calibration instead of min-max, keeping that layer in float, or quantization-aware training.

Failure modes

  • Silent float fallback. An operator has no int8 kernel, so the runtime dequantizes around it. Accuracy is fine; latency is worse than float. Profile the kernel list.
  • Calibration mismatch. Scales were computed on clean or short inputs, and production sends noisy or long ones, so activations clamp. Calibrate on traffic samples.
  • Wrong rounding. Round-half-to-even in one place and half-up in another produces one-step disagreements that accumulate across layers. Match the reference exactly.
  • Inexact zero. A hand-written zero point that does not map real zero to an integer adds a bias at every padded position.
  • Saturation on older x86. Full-range uint8 activations on an AVX2 path overflow 16-bit pair sums. Use the reduced-range option on such targets.

Trade-offs

ChoiceGainCost
W8A8 staticInteger GEMM speed, 4x smaller than FP32Needs calibration; outliers hurt
W8A8 dynamicNo calibration, robust rangesPer-call reduction; fewer fusions
Weight-only INT8Memory and bandwidth savings, high accuracyNo integer compute speed-up
Per-channel weightsMuch lower error for little costOne multiplier per channel
QATBest accuracy at 8 bitsTraining pipeline and time

What to do next

  1. Decide whether you need compute (W8A8) or memory (weight-only), and measure on your target hardware, not a different one.
  2. Run the NumPy reference against your runtime on one layer and confirm agreement within one output step.
  3. Quantize with per-channel symmetric weights and calibrate activations on a few hundred real inputs.
  4. Check the executed kernel list to confirm fused int8 kernels ran and no float islands remain.
  5. Compute per-layer SQNR, fix the worst layers one at a time, and keep a float fallback list.
  6. Re-check accuracy on your task metric, not only on SQNR, before shipping.
Key takeaway: INT8 inference is an integer pipeline: int8 operands, an int32 accumulator with a precomputed zero-point correction, an int32 bias, and one fixed-point multiply and shift per channel that replaces three float scales. Keep weights symmetric and per-channel, make sure fused integer kernels actually run, and debug with a per-layer comparison against a float reference rather than by global tuning.