INT8 and INT4 are the two quantization targets most teams weigh for serving large language models, and the question is usually framed as 'how much accuracy will INT4 cost me?'. That is half the question. The other half is whether INT4 is even faster for your workload, and the honest answer is: it depends on batch size, sequence lengths and which of the two formats actually reduces the work your GPU is waiting on.

This article assumes you know what quantization does to a tensor; INT8 quantization explains the integer arithmetic and INT4 quantization the 4-bit grid and group scales. Here we compare them as deployment choices: what each saves, where each wins, a worked roofline example, how accuracy behaves, a harness to measure your own case, and a procedure for deciding.

Advertisement

Name the formats precisely

'INT8' and 'INT4' each hide several different things, and comparisons go wrong when the labels are loose. The notation WxAy says how many bits the weights and activations use in the matrix multiply.

FormatWeightsActivations and mathWhat it saves
W8A16 (INT8 weight-only)8-bit, per-channel scalesFP16/BF16 math after dequantising weightsHalf the weight memory and bandwidth
W8A8 (INT8 full)8-bit8-bit activations, integer tensor-core GEMM, 32-bit accumulationHalf the bandwidth and higher peak math rate
W4A16 (INT4 weight-only)4-bit, group scales (often 128 weights per group)FP16/BF16 math after dequantising in registersAbout 3.5 to 4 times less weight bandwidth than FP16
W4A84-bit8-bit activations and integer mathINT4 memory plus INT8 math; needs specialised kernels

The asymmetry matters: in the common INT4 deployment, W4A16, the arithmetic still happens in 16-bit floating point. INT4 makes weights smaller to move, not cheaper to multiply. W8A8 makes them smaller to move and cheaper to multiply, if the hardware has fast integer matrix units and the runtime uses a fused kernel. Support differs between hardware generations, so check your vendor's documentation before assuming any particular integer throughput.

Two regimes: waiting for memory, waiting for math

Generating one token for one sequence requires reading every weight once and doing about two floating point operations per weight. A GPU can do hundreds of operations in the time it reads one byte, so at small batch sizes the matrix units sit idle waiting for weights. That is the memory-bound regime, and the time per step is roughly weight bytes divided by memory bandwidth. Shrinking weights is the direct fix, and INT4 shrinks them most.

Batching changes the arithmetic. With B sequences decoding together, the weights are read once and used B times, so the work grows with B and the memory traffic does not. Past some batch size the step becomes compute-bound, and only faster math helps. Prefill, which processes all prompt tokens at once, is compute-bound almost immediately. In this regime W4A16 gives nothing back and can cost a little, because dequantising weights adds instructions, while W8A8 doubles the math rate on hardware with fast INT8 paths.

Advertisement

Worked example: where the lines cross

Time per decode step on a hypothetical GPU: 8B model, roofline onlybatch size (concurrent sequences)ms1641281922560510W4A16: 2.1 ms floor, FP16-rate math, compute-bound above batch ~51W8A8: 4 ms floor until batch ~200FP16: 8 ms floor until batch ~200Lines cross near batch 100: below it INT4 weights win, above it INT8 arithmetic wins.
Idealised time per decode step for an 8-billion-parameter model on a hypothetical GPU with 2 TB/s bandwidth, 400 TFLOPS FP16 and 800 TOPS INT8. Real kernels reach a fraction of peak, and attention and KV-cache traffic are ignored.

Take an 8-billion-parameter model and a hypothetical GPU with 2 TB/s of memory bandwidth, 400 dense FP16 TFLOPS and twice that for INT8. The numbers are illustrative; substitute your hardware's datasheet values.

  • Weight bytes. FP16: 16 GB. INT8: 8 GB plus small per-channel scales. INT4 with one 16-bit scale per 128 weights: about 4.125 bits per weight, so about 4.1 GB, a little more if embeddings and the output layer stay in 16-bit.
  • Memory floor per decode step (bytes / bandwidth): FP16 8 ms, W8A8 4 ms, W4A16 about 2.1 ms. At batch 1 that is an upper bound of roughly 125, 250 and 480 tokens per second.
  • Compute per step: 2 x 8 x 109 x B operations. W4A16 does this in FP16 at 400 TFLOPS, so it takes 0.04 x B ms; W8A8 does it in INT8 at 800 TOPS, 0.02 x B ms.
  • Where each turns compute-bound: W4A16 when 0.04 x B exceeds 2.1, near B = 51. W8A8 when 0.02 x B exceeds 4, near B = 200.

Read the chart left to right. At batch 1 W4A16 is about twice as fast as W8A8. As batch grows, W4A16 hits its compute ceiling first and its step time climbs, while W8A8 is still waiting on memory at a flat 4 ms. The lines cross near batch 100. At batch 256 W4A16 needs about 10.2 ms per step and W8A8 about 5.1 ms, so INT8 is now twice as fast. Prefill behaves like a very large batch, so W8A8 also shortens time to first token for long prompts.

Real systems blur the picture. Kernels do not reach peak, the dequantisation in W4A16 kernels costs extra at large batch (see INT4 kernels), and attention reads the KV cache, whose size grows with batch and context and is not reduced by weight quantization at all. But the shape is robust: weight-only INT4 is a latency and memory tool for small batches; W8A8 is a throughput tool for busy servers.

Memory capacity is the other axis

Speed is not the only reason to quantize. A model that does not fit forces you onto more GPUs, with tensor parallelism and its communication cost. INT4 can turn a two-GPU deployment into one, or leave enough free memory for a much larger KV cache, which raises the batch size the server can hold and therefore its throughput. That interaction can flip the conclusion above: on a memory-constrained GPU, W4A16 may win on throughput simply because it leaves room for more concurrent sequences.

This is also where INT8 KV-cache or FP8 KV-cache quantization enters. At long contexts the cache can exceed the weights, and quantizing it attacks the traffic that weight quantization leaves untouched; see KV cache quantization. Combinations such as W4A16 weights with an 8-bit KV cache, or W8A8 with an 8-bit cache, are common, and W4A8 kernels try to get INT4 capacity with INT8 math at the price of harder accuracy work and narrower kernel support.

How accuracy behaves

INT8 weight-only quantization with per-channel scales is close to lossless for most transformer models. W8A8 is harder because activations in large language models have outlier channels whose values are far larger than the rest; a single per-tensor scale wastes most of the 8-bit range on them. Techniques such as SmoothQuant move that difficulty from activations into weights, and per-token activation scales help further.

INT4 weight-only has only 16 levels, so it relies on group scales and on error-compensating methods such as GPTQ and AWQ. Its loss is usually small on general benchmarks but larger in three situations: smaller models, which have less redundancy; tasks that need precise reasoning, such as arithmetic, code and long-context retrieval; and domains far from the calibration data. Average perplexity can hide all three. The only trustworthy number is your own task suite run against the 16-bit baseline, ideally with paired per-example comparisons, as described in quantization evaluation methodology.

Measure it: a harness for your workload

Because the answer depends on batch size, the benchmark must sweep concurrency rather than report one number. Serve each candidate checkpoint behind the same server software on the same GPU type, replay prompts sampled from real traffic, and record output tokens per second at several concurrency levels. Most serving stacks expose an OpenAI-compatible completions endpoint whose response includes a usage.completion_tokens field, which is all this needs:

import asyncio, json, time, httpx

PROMPTS = [json.loads(l)["prompt"] for l in open("traffic_sample.jsonl")]

async def one(client, url, model, prompt, max_tokens):
    r = await client.post(f"{url}/v1/completions", json={
        "model": model, "prompt": prompt, "max_tokens": max_tokens, "temperature": 0})
    r.raise_for_status()
    return r.json()["usage"]["completion_tokens"]

async def throughput(url, model, concurrency, max_tokens=256, rounds=4):
    assert len(PROMPTS) >= rounds * concurrency, "too few prompts: batches would run short"
    async with httpx.AsyncClient(timeout=600) as client:
        t0, total = time.perf_counter(), 0
        for k in range(rounds):
            batch = PROMPTS[k * concurrency:(k + 1) * concurrency]
            done = await asyncio.gather(*(one(client, url, model, p, max_tokens) for p in batch))
            total += sum(done)
        return total / (time.perf_counter() - t0)

ENDPOINTS = {"fp16": "http://gpu-a:8000", "w8a8": "http://gpu-b:8000", "w4a16": "http://gpu-c:8000"}
for name, url in ENDPOINTS.items():
    for conc in (1, 8, 32, 128):
        tps = asyncio.run(throughput(url, "model", conc))
        print(f"{name:6s} concurrency={conc:4d} output tokens/s={tps:9.1f}")

Add time to first token and per-request latency percentiles if you serve interactive users, because a throughput win at concurrency 128 is worthless if p95 latency breaks your target. Run the quality suite on the same checkpoints in the same week; a format that is fast but fails the gate is not a candidate.

Read the sweep as a curve, not a winner. Plot tokens per second against concurrency for each format and find where the curves cross; that measured crossover replaces the idealised one from the roofline. Then look at the concurrency your production servers actually run at, from their metrics rather than from capacity plans. If the typical operating point sits well on one side of the crossover, the choice is easy. If it straddles it, prefer the format that wins at your peak, because that is when capacity is scarce.

A decision procedure

SituationUsually betterWhy
Single user, local or edge, latency-sensitiveW4A16Memory-bound at batch 1; smallest weights
Model does not fit in INT8 on the targetW4A16Capacity beats everything else
Busy server, high concurrency, short contextsW8A8Compute-bound; integer math rate
Long prompts, time to first token mattersW8A8Prefill is compute-bound
Strict quality bar, reasoning or codeW8A16 or W8A8Smaller and more predictable loss
Long contexts, many sequencesEither, plus KV-cache quantizationCache traffic dominates

Written as code, the procedure gates on quality first, on memory second and on the measured crossover third. CROSSOVER_MEASURED comes from your harness, not from the hypothetical chart:

def choose(weights_gb_fp16, gpu_mem_gb, typical_batch, quality_ok):
    """quality_ok: dict format -> bool from YOUR eval, gated against the FP16 baseline."""
    fits = {"fp16": weights_gb_fp16, "w8a8": weights_gb_fp16 / 2, "w4a16": weights_gb_fp16 / 3.8}
    candidates = [f for f in ("w8a8", "w4a16")
                  if quality_ok[f] and fits[f] < 0.6 * gpu_mem_gb]   # leave room for KV cache
    if not candidates:
        return "fp16 on more GPUs, or try a better INT4 method / larger groups of FP16 layers"
    if len(candidates) == 1:
        return candidates[0]
    return "w4a16" if typical_batch < CROSSOVER_MEASURED else "w8a8"

Failure modes

SymptomLikely causeFix
INT4 slower than FP16No fused INT4 kernel; weights dequantised to a full FP16 copyUse a runtime with fused W4A16 kernels for your GPU
W8A8 no faster than FP16Runtime fell back to dequantise and float GEMMCheck the kernel actually selected; profile
INT4 fine on benchmarks, wrong on production tasksCalibration data unlike traffic; reasoning-heavy tasksCalibrate on traffic samples; gate on your own suite
W8A8 accuracy collapseActivation outliers with per-tensor scalesSmoothQuant-style migration, per-token scales
Throughput win, latency SLO missedBenchmark only at high concurrencySweep concurrency; track p95 and time to first token

What to do next

  1. Write down your serving regime: typical and peak concurrency, prompt and output lengths, and latency targets.
  2. Do the roofline arithmetic for your model and GPU to estimate where W4A16 and W8A8 cross.
  3. Produce W8A8 and W4A16 checkpoints with established methods, calibrated on samples of real traffic.
  4. Run your task suite against the 16-bit baseline with paired comparisons, and drop any format that fails the gate.
  5. Sweep concurrency with the harness, recording tokens per second, p95 latency and time to first token.
  6. Pick the format for your typical batch, consider KV-cache quantization for long contexts, and re-run the comparison when traffic, model or hardware changes.
Key takeaway: INT4 and INT8 solve different bottlenecks. Weight-only INT4 shrinks the bytes that small-batch decoding waits on, so it wins for latency, edge deployment and models that would not otherwise fit. W8A8 also makes the arithmetic cheaper, so it wins once batching or long prefills make the GPU compute-bound. Quality favours INT8, and INT4's loss concentrates in small models and precise tasks. Decide with your own numbers: a quality gate against the 16-bit baseline, then a concurrency sweep on your real traffic.