FP8 inference stores weights, and usually activations, as 8-bit floating-point numbers and runs the large matrix multiplications on tensor cores that consume FP8 directly. Done well it halves weight memory against BF16, speeds up the compute-heavy prefill phase, frees memory for KV cache, and costs little accuracy on most large language models. Done carelessly, it silently clips outliers or wastes the speedup on layers that never needed it.
The broader picture of the two FP8 formats and FP8 training is in FP8 for training and inference. This page is about serving: the numerics you are actually relying on, the scaling schemes inference engines use, what happens inside a quantized linear layer, and the steps to produce, serve and validate an FP8 model.
What FP8 buys during serving, and what it does not
LLM inference has two phases with different bottlenecks. During decode, each step reads every weight to produce one token per sequence, so at small batch sizes speed is set by memory bandwidth: fewer bytes per weight means more steps per second. During prefill, a long prompt turns each layer into a large matrix multiply, so speed is set by arithmetic throughput, and FP8 tensor cores deliver more operations per second than BF16 ones. FP8 helps both phases, which is its advantage over weight-only 4-bit formats that save more memory but do their arithmetic in 16-bit. INT8 vs INT4 works through the memory-bound versus compute-bound crossover in detail.
What FP8 does not do: it does not shrink attention's quadratic cost, it does not speed up the many small element-wise operations, and on hardware without FP8 tensor cores it saves memory but not arithmetic.
The numbers in one byte
Inference uses E4M3: one sign bit, four exponent bits with bias 7, three mantissa bits. The variant used on NVIDIA hardware and in PyTorch as float8_e4m3fn gives up infinities to extend range: the largest finite value is 448, the smallest normal is 2-6, and subnormals reach down to 2-9. E5M2 has more range (maximum 57,344) but only two mantissa bits; it is mainly used for gradients in training and appears in inference only as an optional KV-cache format.
Three mantissa bits mean that within any power-of-two interval there are only eight representable values, spaced 12.5 percent of the interval start apart, so the worst-case rounding error is about 6 percent of the value and the typical error a few percent. Crucially, that relative error is roughly the same for small and large numbers. INT8 has the opposite shape: evenly spaced steps, so small values lose relative precision once a large outlier sets the scale. This simulation rounds to E4M3 exactly as the hardware does, including saturation:
import math
E4M3_MAX = 448.0 # float8_e4m3fn: no infinities, max finite 448
def round_e4m3(x: float) -> float:
"""Round a float to the nearest E4M3 (fn) value, saturating at +-448."""
if x == 0 or math.isnan(x):
return x
s, a = math.copysign(1.0, x), min(abs(x), E4M3_MAX)
e = max(math.floor(math.log2(a)), -6) # -6 is the smallest normal exponent
step = 2.0 ** (e - 3) # 3 mantissa bits
q = round(a / step) * step # Python rounds half to even
return s * min(q, E4M3_MAX)
def fake_quant(row, scale):
"""Quantize to E4M3 with one scale, then dequantize (what the GEMM 'sees')."""
return [round_e4m3(v / scale) * scale for v in row]
row = [0.02, -0.5, 1.3, 7.9, -0.004, 3.1, 400.0] # one outlier
scale = max(abs(v) for v in row) / E4M3_MAX
print([round(v, 4) for v in fake_quant(row, scale)])
# [0.0192, -0.5022, 1.3393, 8.0357, -0.0035, 3.125, 400.0]Worked example. With one scale for the whole row, the outlier 400 maps to 448 and every small value keeps two or three significant digits: 0.02 becomes 0.0192, 1.3 becomes 1.339. Run the same row through symmetric INT8 with scale 400/127 and 0.02, -0.5, 1.3 and -0.004 all become 0, and 7.9 becomes 9.45. This is the main reason FP8 tolerates the outlier channels of transformer activations better than INT8 without the smoothing tricks described in INT8 calibration.
Scales and their granularity
448 is a small maximum, so values are mapped into range with a scale: xfp8 = round(x / s) and x is approximately s times xfp8, with s usually set to amax / 448, where amax is the largest absolute value the scale covers. The design question is how many values share a scale, and whether the scale is fixed offline (static) or computed from the live tensor (dynamic).
| Scheme | Weights | Activations | Notes |
|---|---|---|---|
| Per-tensor, static | one scale per matrix | one scale per tensor, from calibration | cheapest kernels; one outlier coarsens everything; clipping if live data exceeds calibration |
| Per-tensor, dynamic activations | one scale per matrix | amax of the live tensor each step | no calibration; the vLLM docs describe this as the on-load fp8 mode |
| Per-channel weights, per-token activations | one scale per output channel | one dynamic scale per token row | llm-compressor's FP8_DYNAMIC; scales factor out of the GEMM |
| Block-wise | 128x128 weight blocks | 1x128 activation tiles | used for DeepSeek-V3's FP8 training; needs kernels that apply scales inside the K loop |
| MXFP8 (OCP microscaling) | 32-element blocks | 32-element blocks | shared power-of-two (E8M0) scale per block; native on newer tensor cores such as Blackwell |
Finer granularity contains outliers to fewer values, at the price of more scale storage and more complex kernels. Dynamic activation scales track each input but cost an amax reduction per token per layer, which engines typically fuse into the preceding kernel. Static activation scales cost nothing at run time but depend on calibration data that resembles production traffic.
Inside a quantized linear layer
The diagram shows the per-token, per-channel W8A8 path. Weights were quantized offline and stored as E4M3 with one FP32 scale per output channel. At run time a small kernel finds each token's amax, divides and casts the activations to E4M3. The tensor cores multiply E4M3 by E4M3 and accumulate in higher precision, FP32 in principle, because summing thousands of 8-bit products in 8 or 16 bits would lose the result.
Why the scales are cheap: output element yij is the sum over k of aikwkj. If aik is sa,i times a quantized value and wkj is sw,j times a quantized value, both scales are constant along k and move outside the sum, so yij = sa,i sw,j times the FP8 dot product. Block-wise scales vary along k, so the kernel must multiply partial sums by each block's scales as it goes, which is why block-wise FP8 needs purpose-built GEMMs.
Hardware sets the floor. FP8 tensor cores start at NVIDIA compute capability 8.9 (Ada Lovelace, for example the L4 and L40S) and include Hopper and Blackwell. On Ampere and Turing, vLLM can still load FP8 checkpoints as weight-only W8A16 using Marlin kernels: memory savings, 16-bit arithmetic.
What stays in higher precision
Quantize the big linear layers (attention projections and MLPs), which hold almost all parameters and FLOPs, and leave the rest alone. Embeddings and the output head (lm_head, excluded in the recipe below) are sensitive and cheap to keep in BF16. Normalisation, softmax, rotary embeddings and residual additions stay in 16- or 32-bit. In mixture-of-experts models the router is small and its decisions are discrete, so engines and published checkpoints commonly leave it unquantized; check what your engine does before assuming. Attention itself runs in FP8 only when both the KV cache is FP8 and the attention backend supports FP8 maths.
FP8 KV cache
The KV cache often outgrows the weights at long context or high concurrency, and storing it in FP8 halves it. vLLM exposes this as a separate switch, kv_cache_dtype, accepting fp8, fp8_e4m3 or fp8_e5m2 (or auto to match the model), and supports per-tensor or per-attention-head scales for K and V. Its documentation notes that with the FlashAttention 3 backend the attention operations themselves run in FP8; other backends store FP8 and compute in higher precision. Keys are often more sensitive than values, so prefer checkpoints with calibrated KV scales over unit scales, and evaluate long-context tasks specifically. KV cache quantization covers the cache-side design in more depth.
Producing and serving an FP8 checkpoint
The simplest production path is to quantize offline with llm-compressor using the FP8_DYNAMIC scheme, which needs no calibration data because only weights get static scales:
# Offline: produce an FP8 checkpoint (from the vLLM / llm-compressor docs)
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
from transformers import AutoTokenizer, AutoModelForCausalLM
MODEL_ID = "meta-llama/Meta-Llama-3-8B-Instruct"
model = AutoModelForCausalLM.from_pretrained(MODEL_ID, device_map="auto", dtype="auto")
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
# Static per-channel weight scales, dynamic per-token activation scales.
recipe = QuantizationModifier(targets="Linear", scheme="FP8_DYNAMIC", ignore=["lm_head"])
oneshot(model=model, recipe=recipe) # no calibration data needed for this scheme
SAVE_DIR = MODEL_ID.split("/")[1] + "-FP8-Dynamic"
model.save_pretrained(SAVE_DIR)
tokenizer.save_pretrained(SAVE_DIR)# Serve the pre-quantized checkpoint, with an FP8 KV cache as a separate opt-in
vllm serve ./Meta-Llama-3-8B-Instruct-FP8-Dynamic --kv-cache-dtype fp8
# Or quantize a BF16 checkpoint on load (quick trial; scales computed at load/run time)
vllm serve meta-llama/Meta-Llama-3-8B-Instruct --quantization fp8Prefer the offline checkpoint for production: it loads faster, records exactly which layers were quantized, and is the artefact you evaluate. The on-load fp8 mode is useful for a quick comparison but uses coarser per-tensor scales.
Worked example: an 8B model on a 24 GB Ada GPU
Take a Llama-3-8B-class model (about 8 billion parameters, 32 layers, 8 KV heads of dimension 128, a vocabulary of about 128K and untied input and output embeddings) on a 24 GB L4. In BF16 the weights are about 16 GB. KV cache per token is 32 layers x 8 heads x 128 x 2 (K and V) x 2 bytes = 131,072 bytes, 128 KiB. After weights, activations and runtime overhead, perhaps 5-6 GB remain for cache: roughly 40,000-48,000 tokens across all sequences.
With FP8 linear layers, about 7 billion parameters at one byte, plus the embeddings and lm_head kept in BF16 (about 1 billion parameters, 2 GB), weights take roughly 9 GB. With an FP8 cache (64 KiB per token), perhaps 12-13 GB remain, roughly 190,000-200,000 tokens: four to five times the concurrent context. These are planning estimates; confirm with the engine's own report of KV blocks at startup.
Evaluate before you ship
Compare the FP8 model against its BF16 parent on your tasks, not only on perplexity: task accuracy on a held-out set, exact-match rate of greedy outputs, long-context retrieval, structured-output validity and, for code models, pass rates. Run the evaluations with the same engine and sampling settings for both. Quantization evaluation methodology describes how to size these comparisons so a real regression is distinguishable from noise. Then canary on a slice of traffic and watch user-facing metrics, not just latency.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Quality drops only on some prompts | Static activation scales clipped out-of-distribution inputs | Use dynamic per-token scales or recalibrate on production-like data |
| No speedup | GPU below compute capability 8.9, or decode at tiny batch | Check hardware; measure prefill and batched decode separately |
| Garbage output after conversion | Scales stored inverted or ignored by the loader | Load with the engine the checkpoint format targets; compare one layer's output to BF16 |
| Long-context answers degrade | FP8 KV cache with unit or poor scales | Use calibrated KV scales, or keep the cache in BF16 |
| MoE model worse than dense peers | Router or rarely used experts quantized badly | Keep the router in 16-bit; evaluate expert-heavy prompts |
Trade-offs
Against INT8 W8A8, FP8 usually needs less outlier handling and is the simpler path on GPUs that support it, while INT8 runs on older hardware. Against 4-bit weight-only formats, FP8 saves less memory but accelerates prefill and usually loses less accuracy. Against BF16, FP8 trades a small, measurable quality risk for roughly half the weight memory and higher throughput.