INT8 and INT4 are the two quantization targets most teams weigh for serving large language models, and the question is usually framed as 'how much accuracy will INT4 cost me?'. That is half the question. The other half is whether INT4 is even faster for your workload, and the honest answer is: it depends on batch size, sequence lengths and which of the two formats actually reduces the work your GPU is waiting on.
This article assumes you know what quantization does to a tensor; INT8 quantization explains the integer arithmetic and INT4 quantization the 4-bit grid and group scales. Here we compare them as deployment choices: what each saves, where each wins, a worked roofline example, how accuracy behaves, a harness to measure your own case, and a procedure for deciding.
Name the formats precisely
'INT8' and 'INT4' each hide several different things, and comparisons go wrong when the labels are loose. The notation WxAy says how many bits the weights and activations use in the matrix multiply.
| Format | Weights | Activations and math | What it saves |
|---|---|---|---|
| W8A16 (INT8 weight-only) | 8-bit, per-channel scales | FP16/BF16 math after dequantising weights | Half the weight memory and bandwidth |
| W8A8 (INT8 full) | 8-bit | 8-bit activations, integer tensor-core GEMM, 32-bit accumulation | Half the bandwidth and higher peak math rate |
| W4A16 (INT4 weight-only) | 4-bit, group scales (often 128 weights per group) | FP16/BF16 math after dequantising in registers | About 3.5 to 4 times less weight bandwidth than FP16 |
| W4A8 | 4-bit | 8-bit activations and integer math | INT4 memory plus INT8 math; needs specialised kernels |
The asymmetry matters: in the common INT4 deployment, W4A16, the arithmetic still happens in 16-bit floating point. INT4 makes weights smaller to move, not cheaper to multiply. W8A8 makes them smaller to move and cheaper to multiply, if the hardware has fast integer matrix units and the runtime uses a fused kernel. Support differs between hardware generations, so check your vendor's documentation before assuming any particular integer throughput.
Two regimes: waiting for memory, waiting for math
Generating one token for one sequence requires reading every weight once and doing about two floating point operations per weight. A GPU can do hundreds of operations in the time it reads one byte, so at small batch sizes the matrix units sit idle waiting for weights. That is the memory-bound regime, and the time per step is roughly weight bytes divided by memory bandwidth. Shrinking weights is the direct fix, and INT4 shrinks them most.
Batching changes the arithmetic. With B sequences decoding together, the weights are read once and used B times, so the work grows with B and the memory traffic does not. Past some batch size the step becomes compute-bound, and only faster math helps. Prefill, which processes all prompt tokens at once, is compute-bound almost immediately. In this regime W4A16 gives nothing back and can cost a little, because dequantising weights adds instructions, while W8A8 doubles the math rate on hardware with fast INT8 paths.
Worked example: where the lines cross
Take an 8-billion-parameter model and a hypothetical GPU with 2 TB/s of memory bandwidth, 400 dense FP16 TFLOPS and twice that for INT8. The numbers are illustrative; substitute your hardware's datasheet values.
- Weight bytes. FP16: 16 GB. INT8: 8 GB plus small per-channel scales. INT4 with one 16-bit scale per 128 weights: about 4.125 bits per weight, so about 4.1 GB, a little more if embeddings and the output layer stay in 16-bit.
- Memory floor per decode step (bytes / bandwidth): FP16 8 ms, W8A8 4 ms, W4A16 about 2.1 ms. At batch 1 that is an upper bound of roughly 125, 250 and 480 tokens per second.
- Compute per step: 2 x 8 x 109 x B operations. W4A16 does this in FP16 at 400 TFLOPS, so it takes 0.04 x B ms; W8A8 does it in INT8 at 800 TOPS, 0.02 x B ms.
- Where each turns compute-bound: W4A16 when 0.04 x B exceeds 2.1, near B = 51. W8A8 when 0.02 x B exceeds 4, near B = 200.
Read the chart left to right. At batch 1 W4A16 is about twice as fast as W8A8. As batch grows, W4A16 hits its compute ceiling first and its step time climbs, while W8A8 is still waiting on memory at a flat 4 ms. The lines cross near batch 100. At batch 256 W4A16 needs about 10.2 ms per step and W8A8 about 5.1 ms, so INT8 is now twice as fast. Prefill behaves like a very large batch, so W8A8 also shortens time to first token for long prompts.
Real systems blur the picture. Kernels do not reach peak, the dequantisation in W4A16 kernels costs extra at large batch (see INT4 kernels), and attention reads the KV cache, whose size grows with batch and context and is not reduced by weight quantization at all. But the shape is robust: weight-only INT4 is a latency and memory tool for small batches; W8A8 is a throughput tool for busy servers.
Memory capacity is the other axis
Speed is not the only reason to quantize. A model that does not fit forces you onto more GPUs, with tensor parallelism and its communication cost. INT4 can turn a two-GPU deployment into one, or leave enough free memory for a much larger KV cache, which raises the batch size the server can hold and therefore its throughput. That interaction can flip the conclusion above: on a memory-constrained GPU, W4A16 may win on throughput simply because it leaves room for more concurrent sequences.
This is also where INT8 KV-cache or FP8 KV-cache quantization enters. At long contexts the cache can exceed the weights, and quantizing it attacks the traffic that weight quantization leaves untouched; see KV cache quantization. Combinations such as W4A16 weights with an 8-bit KV cache, or W8A8 with an 8-bit cache, are common, and W4A8 kernels try to get INT4 capacity with INT8 math at the price of harder accuracy work and narrower kernel support.
How accuracy behaves
INT8 weight-only quantization with per-channel scales is close to lossless for most transformer models. W8A8 is harder because activations in large language models have outlier channels whose values are far larger than the rest; a single per-tensor scale wastes most of the 8-bit range on them. Techniques such as SmoothQuant move that difficulty from activations into weights, and per-token activation scales help further.
INT4 weight-only has only 16 levels, so it relies on group scales and on error-compensating methods such as GPTQ and AWQ. Its loss is usually small on general benchmarks but larger in three situations: smaller models, which have less redundancy; tasks that need precise reasoning, such as arithmetic, code and long-context retrieval; and domains far from the calibration data. Average perplexity can hide all three. The only trustworthy number is your own task suite run against the 16-bit baseline, ideally with paired per-example comparisons, as described in quantization evaluation methodology.
Measure it: a harness for your workload
Because the answer depends on batch size, the benchmark must sweep concurrency rather than report one number. Serve each candidate checkpoint behind the same server software on the same GPU type, replay prompts sampled from real traffic, and record output tokens per second at several concurrency levels. Most serving stacks expose an OpenAI-compatible completions endpoint whose response includes a usage.completion_tokens field, which is all this needs:
import asyncio, json, time, httpx
PROMPTS = [json.loads(l)["prompt"] for l in open("traffic_sample.jsonl")]
async def one(client, url, model, prompt, max_tokens):
r = await client.post(f"{url}/v1/completions", json={
"model": model, "prompt": prompt, "max_tokens": max_tokens, "temperature": 0})
r.raise_for_status()
return r.json()["usage"]["completion_tokens"]
async def throughput(url, model, concurrency, max_tokens=256, rounds=4):
assert len(PROMPTS) >= rounds * concurrency, "too few prompts: batches would run short"
async with httpx.AsyncClient(timeout=600) as client:
t0, total = time.perf_counter(), 0
for k in range(rounds):
batch = PROMPTS[k * concurrency:(k + 1) * concurrency]
done = await asyncio.gather(*(one(client, url, model, p, max_tokens) for p in batch))
total += sum(done)
return total / (time.perf_counter() - t0)
ENDPOINTS = {"fp16": "http://gpu-a:8000", "w8a8": "http://gpu-b:8000", "w4a16": "http://gpu-c:8000"}
for name, url in ENDPOINTS.items():
for conc in (1, 8, 32, 128):
tps = asyncio.run(throughput(url, "model", conc))
print(f"{name:6s} concurrency={conc:4d} output tokens/s={tps:9.1f}")Add time to first token and per-request latency percentiles if you serve interactive users, because a throughput win at concurrency 128 is worthless if p95 latency breaks your target. Run the quality suite on the same checkpoints in the same week; a format that is fast but fails the gate is not a candidate.
Read the sweep as a curve, not a winner. Plot tokens per second against concurrency for each format and find where the curves cross; that measured crossover replaces the idealised one from the roofline. Then look at the concurrency your production servers actually run at, from their metrics rather than from capacity plans. If the typical operating point sits well on one side of the crossover, the choice is easy. If it straddles it, prefer the format that wins at your peak, because that is when capacity is scarce.
A decision procedure
| Situation | Usually better | Why |
|---|---|---|
| Single user, local or edge, latency-sensitive | W4A16 | Memory-bound at batch 1; smallest weights |
| Model does not fit in INT8 on the target | W4A16 | Capacity beats everything else |
| Busy server, high concurrency, short contexts | W8A8 | Compute-bound; integer math rate |
| Long prompts, time to first token matters | W8A8 | Prefill is compute-bound |
| Strict quality bar, reasoning or code | W8A16 or W8A8 | Smaller and more predictable loss |
| Long contexts, many sequences | Either, plus KV-cache quantization | Cache traffic dominates |
Written as code, the procedure gates on quality first, on memory second and on the measured crossover third. CROSSOVER_MEASURED comes from your harness, not from the hypothetical chart:
def choose(weights_gb_fp16, gpu_mem_gb, typical_batch, quality_ok):
"""quality_ok: dict format -> bool from YOUR eval, gated against the FP16 baseline."""
fits = {"fp16": weights_gb_fp16, "w8a8": weights_gb_fp16 / 2, "w4a16": weights_gb_fp16 / 3.8}
candidates = [f for f in ("w8a8", "w4a16")
if quality_ok[f] and fits[f] < 0.6 * gpu_mem_gb] # leave room for KV cache
if not candidates:
return "fp16 on more GPUs, or try a better INT4 method / larger groups of FP16 layers"
if len(candidates) == 1:
return candidates[0]
return "w4a16" if typical_batch < CROSSOVER_MEASURED else "w8a8"
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| INT4 slower than FP16 | No fused INT4 kernel; weights dequantised to a full FP16 copy | Use a runtime with fused W4A16 kernels for your GPU |
| W8A8 no faster than FP16 | Runtime fell back to dequantise and float GEMM | Check the kernel actually selected; profile |
| INT4 fine on benchmarks, wrong on production tasks | Calibration data unlike traffic; reasoning-heavy tasks | Calibrate on traffic samples; gate on your own suite |
| W8A8 accuracy collapse | Activation outliers with per-tensor scales | SmoothQuant-style migration, per-token scales |
| Throughput win, latency SLO missed | Benchmark only at high concurrency | Sweep concurrency; track p95 and time to first token |
What to do next
- Write down your serving regime: typical and peak concurrency, prompt and output lengths, and latency targets.
- Do the roofline arithmetic for your model and GPU to estimate where W4A16 and W8A8 cross.
- Produce W8A8 and W4A16 checkpoints with established methods, calibrated on samples of real traffic.
- Run your task suite against the 16-bit baseline with paired comparisons, and drop any format that fails the gate.
- Sweep concurrency with the harness, recording tokens per second, p95 latency and time to first token.
- Pick the format for your typical batch, consider KV-cache quantization for long contexts, and re-run the comparison when traffic, model or hardware changes.