FP8 inference stores weights, and usually activations, as 8-bit floating-point numbers and runs the large matrix multiplications on tensor cores that consume FP8 directly. Done well it halves weight memory against BF16, speeds up the compute-heavy prefill phase, frees memory for KV cache, and costs little accuracy on most large language models. Done carelessly, it silently clips outliers or wastes the speedup on layers that never needed it.

The broader picture of the two FP8 formats and FP8 training is in FP8 for training and inference. This page is about serving: the numerics you are actually relying on, the scaling schemes inference engines use, what happens inside a quantized linear layer, and the steps to produce, serve and validate an FP8 model.

Advertisement

What FP8 buys during serving, and what it does not

LLM inference has two phases with different bottlenecks. During decode, each step reads every weight to produce one token per sequence, so at small batch sizes speed is set by memory bandwidth: fewer bytes per weight means more steps per second. During prefill, a long prompt turns each layer into a large matrix multiply, so speed is set by arithmetic throughput, and FP8 tensor cores deliver more operations per second than BF16 ones. FP8 helps both phases, which is its advantage over weight-only 4-bit formats that save more memory but do their arithmetic in 16-bit. INT8 vs INT4 works through the memory-bound versus compute-bound crossover in detail.

What FP8 does not do: it does not shrink attention's quadratic cost, it does not speed up the many small element-wise operations, and on hardware without FP8 tensor cores it saves memory but not arithmetic.

The numbers in one byte

Inference uses E4M3: one sign bit, four exponent bits with bias 7, three mantissa bits. The variant used on NVIDIA hardware and in PyTorch as float8_e4m3fn gives up infinities to extend range: the largest finite value is 448, the smallest normal is 2-6, and subnormals reach down to 2-9. E5M2 has more range (maximum 57,344) but only two mantissa bits; it is mainly used for gradients in training and appears in inference only as an optional KV-cache format.

Three mantissa bits mean that within any power-of-two interval there are only eight representable values, spaced 12.5 percent of the interval start apart, so the worst-case rounding error is about 6 percent of the value and the typical error a few percent. Crucially, that relative error is roughly the same for small and large numbers. INT8 has the opposite shape: evenly spaced steps, so small values lose relative precision once a large outlier sets the scale. This simulation rounds to E4M3 exactly as the hardware does, including saturation:

import math

E4M3_MAX = 448.0          # float8_e4m3fn: no infinities, max finite 448

def round_e4m3(x: float) -> float:
    """Round a float to the nearest E4M3 (fn) value, saturating at +-448."""
    if x == 0 or math.isnan(x):
        return x
    s, a = math.copysign(1.0, x), min(abs(x), E4M3_MAX)
    e = max(math.floor(math.log2(a)), -6)    # -6 is the smallest normal exponent
    step = 2.0 ** (e - 3)                    # 3 mantissa bits
    q = round(a / step) * step               # Python rounds half to even
    return s * min(q, E4M3_MAX)

def fake_quant(row, scale):
    """Quantize to E4M3 with one scale, then dequantize (what the GEMM 'sees')."""
    return [round_e4m3(v / scale) * scale for v in row]

row = [0.02, -0.5, 1.3, 7.9, -0.004, 3.1, 400.0]   # one outlier
scale = max(abs(v) for v in row) / E4M3_MAX
print([round(v, 4) for v in fake_quant(row, scale)])
# [0.0192, -0.5022, 1.3393, 8.0357, -0.0035, 3.125, 400.0]

Worked example. With one scale for the whole row, the outlier 400 maps to 448 and every small value keeps two or three significant digits: 0.02 becomes 0.0192, 1.3 becomes 1.339. Run the same row through symmetric INT8 with scale 400/127 and 0.02, -0.5, 1.3 and -0.004 all become 0, and 7.9 becomes 9.45. This is the main reason FP8 tolerates the outlier channels of transformer activations better than INT8 without the smoothing tricks described in INT8 calibration.

Advertisement

Scales and their granularity

448 is a small maximum, so values are mapped into range with a scale: xfp8 = round(x / s) and x is approximately s times xfp8, with s usually set to amax / 448, where amax is the largest absolute value the scale covers. The design question is how many values share a scale, and whether the scale is fixed offline (static) or computed from the live tensor (dynamic).

SchemeWeightsActivationsNotes
Per-tensor, staticone scale per matrixone scale per tensor, from calibrationcheapest kernels; one outlier coarsens everything; clipping if live data exceeds calibration
Per-tensor, dynamic activationsone scale per matrixamax of the live tensor each stepno calibration; the vLLM docs describe this as the on-load fp8 mode
Per-channel weights, per-token activationsone scale per output channelone dynamic scale per token rowllm-compressor's FP8_DYNAMIC; scales factor out of the GEMM
Block-wise128x128 weight blocks1x128 activation tilesused for DeepSeek-V3's FP8 training; needs kernels that apply scales inside the K loop
MXFP8 (OCP microscaling)32-element blocks32-element blocksshared power-of-two (E8M0) scale per block; native on newer tensor cores such as Blackwell

Finer granularity contains outliers to fewer values, at the price of more scale storage and more complex kernels. Dynamic activation scales track each input but cost an amax reduction per token per layer, which engines typically fuse into the preceding kernel. Static activation scales cost nothing at run time but depend on calibration data that resembles production traffic.

Inside a quantized linear layer

The diagram shows the per-token, per-channel W8A8 path. Weights were quantized offline and stored as E4M3 with one FP32 scale per output channel. At run time a small kernel finds each token's amax, divides and casts the activations to E4M3. The tensor cores multiply E4M3 by E4M3 and accumulate in higher precision, FP32 in principle, because summing thousands of 8-bit products in 8 or 16 bits would lose the result.

W8A8 FP8 linear layer: where the scales are computed and where they are appliedBF16 activationstokens x hiddenQuantize kernelamax per token, casts_a[token]FP32, one per rowFP8 GEMMtensor cores, FP32 accFP8 weightsE4M3, stored offlines_w[channel]FP32, one per columnEpilogueacc x s_a x s_wBF16 outputto next opE4M3Kept in BF16/FP32: embeddings, lm_head, norms, softmax, residual addsand usually MoE routers; attention runs in FP8 only with a backend that supports itPer-token and per-channel scales factor out of the dot product, so they cost one multiply per output element.
Per-token activation scales and per-channel weight scales are applied once per output element in the GEMM epilogue.

Why the scales are cheap: output element yij is the sum over k of aikwkj. If aik is sa,i times a quantized value and wkj is sw,j times a quantized value, both scales are constant along k and move outside the sum, so yij = sa,i sw,j times the FP8 dot product. Block-wise scales vary along k, so the kernel must multiply partial sums by each block's scales as it goes, which is why block-wise FP8 needs purpose-built GEMMs.

Hardware sets the floor. FP8 tensor cores start at NVIDIA compute capability 8.9 (Ada Lovelace, for example the L4 and L40S) and include Hopper and Blackwell. On Ampere and Turing, vLLM can still load FP8 checkpoints as weight-only W8A16 using Marlin kernels: memory savings, 16-bit arithmetic.

What stays in higher precision

Quantize the big linear layers (attention projections and MLPs), which hold almost all parameters and FLOPs, and leave the rest alone. Embeddings and the output head (lm_head, excluded in the recipe below) are sensitive and cheap to keep in BF16. Normalisation, softmax, rotary embeddings and residual additions stay in 16- or 32-bit. In mixture-of-experts models the router is small and its decisions are discrete, so engines and published checkpoints commonly leave it unquantized; check what your engine does before assuming. Attention itself runs in FP8 only when both the KV cache is FP8 and the attention backend supports FP8 maths.

FP8 KV cache

The KV cache often outgrows the weights at long context or high concurrency, and storing it in FP8 halves it. vLLM exposes this as a separate switch, kv_cache_dtype, accepting fp8, fp8_e4m3 or fp8_e5m2 (or auto to match the model), and supports per-tensor or per-attention-head scales for K and V. Its documentation notes that with the FlashAttention 3 backend the attention operations themselves run in FP8; other backends store FP8 and compute in higher precision. Keys are often more sensitive than values, so prefer checkpoints with calibrated KV scales over unit scales, and evaluate long-context tasks specifically. KV cache quantization covers the cache-side design in more depth.

Producing and serving an FP8 checkpoint

The simplest production path is to quantize offline with llm-compressor using the FP8_DYNAMIC scheme, which needs no calibration data because only weights get static scales:

# Offline: produce an FP8 checkpoint (from the vLLM / llm-compressor docs)
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
from transformers import AutoTokenizer, AutoModelForCausalLM

MODEL_ID = "meta-llama/Meta-Llama-3-8B-Instruct"
model = AutoModelForCausalLM.from_pretrained(MODEL_ID, device_map="auto", dtype="auto")
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)

# Static per-channel weight scales, dynamic per-token activation scales.
recipe = QuantizationModifier(targets="Linear", scheme="FP8_DYNAMIC", ignore=["lm_head"])
oneshot(model=model, recipe=recipe)          # no calibration data needed for this scheme

SAVE_DIR = MODEL_ID.split("/")[1] + "-FP8-Dynamic"
model.save_pretrained(SAVE_DIR)
tokenizer.save_pretrained(SAVE_DIR)
# Serve the pre-quantized checkpoint, with an FP8 KV cache as a separate opt-in
vllm serve ./Meta-Llama-3-8B-Instruct-FP8-Dynamic --kv-cache-dtype fp8

# Or quantize a BF16 checkpoint on load (quick trial; scales computed at load/run time)
vllm serve meta-llama/Meta-Llama-3-8B-Instruct --quantization fp8

Prefer the offline checkpoint for production: it loads faster, records exactly which layers were quantized, and is the artefact you evaluate. The on-load fp8 mode is useful for a quick comparison but uses coarser per-tensor scales.

Worked example: an 8B model on a 24 GB Ada GPU

Take a Llama-3-8B-class model (about 8 billion parameters, 32 layers, 8 KV heads of dimension 128, a vocabulary of about 128K and untied input and output embeddings) on a 24 GB L4. In BF16 the weights are about 16 GB. KV cache per token is 32 layers x 8 heads x 128 x 2 (K and V) x 2 bytes = 131,072 bytes, 128 KiB. After weights, activations and runtime overhead, perhaps 5-6 GB remain for cache: roughly 40,000-48,000 tokens across all sequences.

With FP8 linear layers, about 7 billion parameters at one byte, plus the embeddings and lm_head kept in BF16 (about 1 billion parameters, 2 GB), weights take roughly 9 GB. With an FP8 cache (64 KiB per token), perhaps 12-13 GB remain, roughly 190,000-200,000 tokens: four to five times the concurrent context. These are planning estimates; confirm with the engine's own report of KV blocks at startup.

Evaluate before you ship

Compare the FP8 model against its BF16 parent on your tasks, not only on perplexity: task accuracy on a held-out set, exact-match rate of greedy outputs, long-context retrieval, structured-output validity and, for code models, pass rates. Run the evaluations with the same engine and sampling settings for both. Quantization evaluation methodology describes how to size these comparisons so a real regression is distinguishable from noise. Then canary on a slice of traffic and watch user-facing metrics, not just latency.

Failure modes

SymptomCauseFix
Quality drops only on some promptsStatic activation scales clipped out-of-distribution inputsUse dynamic per-token scales or recalibrate on production-like data
No speedupGPU below compute capability 8.9, or decode at tiny batchCheck hardware; measure prefill and batched decode separately
Garbage output after conversionScales stored inverted or ignored by the loaderLoad with the engine the checkpoint format targets; compare one layer's output to BF16
Long-context answers degradeFP8 KV cache with unit or poor scalesUse calibrated KV scales, or keep the cache in BF16
MoE model worse than dense peersRouter or rarely used experts quantized badlyKeep the router in 16-bit; evaluate expert-heavy prompts

Trade-offs

Against INT8 W8A8, FP8 usually needs less outlier handling and is the simpler path on GPUs that support it, while INT8 runs on older hardware. Against 4-bit weight-only formats, FP8 saves less memory but accelerates prefill and usually loses less accuracy. Against BF16, FP8 trades a small, measurable quality risk for roughly half the weight memory and higher throughput.

Key takeaway: <p><strong>What to do next.</strong> FP8 inference is E4M3 values plus scales: get the scale granularity right, keep the sensitive layers in 16-bit, and verify on your own tasks.</p><ol><li>Confirm your GPUs are compute capability 8.9 or newer; otherwise expect weight-only savings.</li><li>Quantize offline with FP8_DYNAMIC, ignoring lm_head, and keep the checkpoint as the evaluated artefact.</li><li>Benchmark prefill and decode throughput separately against BF16 at your real batch sizes.</li><li>Add the FP8 KV cache as a separate change and test long-context tasks before and after.</li><li>Run task evals and greedy-output diffs against the BF16 parent, then canary with rollback ready.</li><li>Record the scheme, ignored layers, engine version and kernel backend alongside the model version.</li></ol>