FP16, the IEEE 754 binary16 format, stores a real number in 16 bits. It halves memory and bandwidth relative to FP32, and on GPUs with tensor cores it multiplies matrices many times faster. It was the format that made mixed-precision training practical, it remains the default storage type for many inference runtimes, mobile NPUs and graphics shaders, and it is the baseline that integer quantization is measured against.
It also has a hard ceiling. The largest finite FP16 value is 65,504, and anything larger becomes infinity, which turns into NaN a few operations later. Most FP16 incidents come from forgetting that ceiling, or from the matching floor where small gradients round to zero. This article works through the format bit by bit, shows exactly where values are lost, explains loss scaling for training and overflow auditing for inference, and ends with when to choose FP16 over BF16.
The format from first principles
An FP16 value has one sign bit, five exponent bits and ten mantissa (fraction) bits. The exponent is stored with a bias of 15. For a normal number the value is (-1)^sign x 2^(exponent - 15) x 1.fraction, where the leading 1 is implicit, so the significand has 11 bits of precision. An exponent field of all zeros means a subnormal number, with no implicit 1 and a fixed exponent of -14, which lets values fade gradually towards zero. An exponent field of all ones means infinity when the fraction is zero and NaN otherwise.
| Quantity | FP16 value | Bit pattern |
|---|---|---|
| 1.0 | 1.0 | 0x3C00 |
| Largest finite | 65,504 | 0x7BFF |
| Smallest positive normal | 2^-14, about 6.10e-5 | 0x0400 |
| Smallest positive subnormal | 2^-24, about 5.96e-8 | 0x0001 |
| Machine epsilon (gap after 1.0) | 2^-10, about 9.77e-4 | 0x3C01 minus 1.0 |
| Positive infinity | inf | 0x7C00 |
Two consequences are worth memorising. FP16 carries about three significant decimal digits, so 1,000.3 and 1,000.5 cannot both be represented exactly; between 1,024 and 2,048 the gap between neighbours is exactly 1, and between 2,048 and 4,096 it is 2. And the usable dynamic range spans roughly 6e-8 to 6.5e4, about twelve orders of magnitude, against about seventy-six for FP32 and BF16.
import numpy as np
def show(x):
h = np.float16(x)
print(f"{x!r:>12} -> {float(h)!r:<24} bits=0x{h.view(np.uint16):04X}")
for x in [1.0, 0.1, 65504, 65519, 65520, 70000, 2049, 1e-5, 6e-8, 2e-8]:
show(x)
# 1.0 -> 1.0 bits=0x3C00
# 0.1 -> 0.0999755859375 bits=0x2E66
# 65504 -> 65504.0 bits=0x7BFF
# 65519 -> 65504.0 bits=0x7BFF (rounds down)
# 65520 -> inf bits=0x7C00 (rounds up to infinity)
# 70000 -> inf bits=0x7C00
# 2049 -> 2048.0 bits=0x6800 (odd integers above 2048 are gone)
# 1e-05 -> 1.0013580322265625e-05 bits=0x00A8 (subnormal: only 8 significant bits)
# 6e-08 -> 5.960464477539063e-08 bits=0x0001
# 2e-08 -> 0.0 bits=0x0000 (underflow)
Rounding, conversion and where values disappear
Converting FP32 or BF16 to FP16 uses round-to-nearest, ties-to-even, by default. Three things can happen to a value. If its magnitude is at least 65,520, the halfway point between 65,504 and the next step that does not exist, it becomes infinity. If it lies in the subnormal range below 6.1e-5, it keeps fewer significant bits the smaller it gets. Below about 3e-8 it becomes zero. Converting from FP16 back to FP32 or BF16 never overflows, although BF16 loses three mantissa bits on the way.
Hardware usually converts with these semantics, but not always. Some kernels and accelerators flush subnormals to zero for speed, which moves the effective floor up to 6.1e-5. Some conversion paths saturate to 65,504 instead of producing infinity. When an FP16 result differs between two backends, check the subnormal and saturation behaviour before suspecting the model.
Accumulate in FP32
FP16 is a fine format for inputs to a multiplication and a poor one for long sums. A dot product of length 4,096 adds 4,096 products; if the running total is held in FP16, every addition is rounded to the total's own step size, and once the total is large, small terms stop contributing at all. The example below shows the effect at its most extreme: adding 1.0 to 2,048 in FP16 does nothing.
import numpy as np
acc = np.float16(2048)
print(float(acc + np.float16(1))) # 2048.0: the step at 2048 is 2, and 2049 ties to even
x = np.random.default_rng(0).standard_normal(4096).astype(np.float16)
w = np.random.default_rng(1).standard_normal(4096).astype(np.float16)
exact = np.dot(x.astype(np.float64), w.astype(np.float64))
fp16_acc = np.float16(0)
for a, b in zip(x, w):
fp16_acc = np.float16(fp16_acc + a * b) # FP16 running sum
fp32_acc = np.dot(x.astype(np.float32), w.astype(np.float32))
print(exact, float(fp16_acc), float(fp32_acc))
# -25.4100..., -25.203125, -25.41009: the FP16 running sum is off by about 0.2;
# FP32 accumulation matches the exact result to six digitsThat is why tensor cores and well-written kernels multiply FP16 inputs and accumulate in FP32, and why frameworks keep reductions such as softmax denominators, layer-norm and RMSNorm statistics, loss computation and large summations in FP32 even inside an FP16 region. If you write a custom kernel, the accumulator type is the decision that matters most; the input type is the easy part.
Training in FP16: master weights and loss scaling
Mixed-precision training, introduced by Micikevicius and colleagues in 2018, made FP16 training work with three ideas. Keep an FP32 master copy of the weights and update that, because an update smaller than about one thousandth of a weight vanishes when added in FP16. Run the forward and backward matrix multiplications in FP16 with FP32 accumulation. And scale the loss, because many gradient values are smaller than FP16's smallest subnormal and would round to zero.
Loss scaling multiplies the loss by a factor S before the backward pass. By the chain rule every gradient is multiplied by S too, which lifts small gradients into FP16's representable range. Before the optimizer step the gradients are divided by S in FP32. If S is too large some gradients overflow to infinity; the step is then skipped and S is halved. After a run of clean steps S is doubled. This dynamic scheme finds the largest safe scale automatically.
import torch
model = build_model().cuda() # parameters stay in FP32 (the master copy)
opt = torch.optim.AdamW(model.parameters(), lr=3e-4)
scaler = torch.amp.GradScaler("cuda") # defaults: init_scale=2**16, growth_factor=2.0,
# backoff_factor=0.5, growth_interval=2000
for x, y in loader:
opt.zero_grad(set_to_none=True)
with torch.autocast(device_type="cuda", dtype=torch.float16):
loss = loss_fn(model(x.cuda()), y.cuda()) # matmuls in FP16; autocast keeps
# softmax, norms and losses in FP32
scaler.scale(loss).backward() # gradients carry the factor S
scaler.unscale_(opt) # divide by S before anything reads them
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
scaler.step(opt) # skipped if any gradient is inf or NaN
scaler.update() # halve S on overflow, grow it after clean stepsTwo mistakes are common. Clipping gradients before calling unscale_ clips scaled values against an unscaled threshold, so clipping is effectively off. And ignoring skipped steps hides problems: a few early skips while the scale settles are normal, but a scale that keeps falling means something produces infinities every step. Log the scale value; if it collapses towards 1, find the overflowing layer instead of lowering the learning rate. Models that need a steadily falling scale are usually the ones that should train in BF16 instead.
Inference in FP16: why BF16-trained models overflow
Weights are rarely the problem. Trained weights are mostly well below 1 in magnitude, far inside FP16's range, and because FP16 has more mantissa bits than BF16, every BF16 weight between 6.1e-5 and 65,504 converts to FP16 exactly. Activations are the problem. A model trained in BF16 never had to keep activations under 65,504, and some do not: large language models develop outlier features of very large magnitude in specific channels, residual streams can grow through depth, and attention logits can be large before softmax.
Normalisation layers are the classic trap. RMSNorm computes the mean of x squared; if any element exceeds about 256 in magnitude, its square exceeds 65,504, and an FP16 sum of squares becomes infinity even though the normalised output would have been ordinary. The T5 family is the widely reported example of models that produce infinities or NaNs in FP16 while working in BF16 and FP32. The fixes, in order of preference, are to compute norms, softmax and residual additions in FP32, to keep the few offending layers in BF16 or FP32, and only as a last resort to clamp activations, which changes the model's outputs.
Worked example: auditing a checkpoint before converting it
Suppose you need to serve an 8-billion-parameter model, trained in BF16, on hardware whose fast path is FP16. Memory first: 8e9 parameters at 2 bytes is 16 GB of weights, the same as BF16. A KV cache for a model with 32 layers, 8 KV heads and a head dimension of 128 costs 2 (K and V) x 32 x 8 x 128 x 2 bytes, which is 131,072 bytes or 128 KiB per token, so an 8,192-token sequence holds 1 GiB of cache. FP16 does not change those numbers; it changes whether the values fit.
Before converting, measure. Run the model in BF16 on a few hundred representative prompts, including long ones, and record the largest absolute value each module outputs. Then check the weights directly.
import torch
FP16_MAX, HEADROOM = 65504.0, 8.0 # flag anything within 8x of the ceiling
peaks = {}
def track(name):
def hook(module, inputs, output):
out = output[0] if isinstance(output, tuple) else output
if torch.is_tensor(out):
peaks[name] = max(peaks.get(name, 0.0), out.detach().abs().max().float().item())
return hook
handles = [m.register_forward_hook(track(n)) for n, m in model.named_modules()]
with torch.no_grad():
for batch in calibration_batches: # model is still in BF16 here
model(**batch)
for h in handles:
h.remove()
risky = {n: v for n, v in peaks.items() if v > FP16_MAX / HEADROOM}
big_w = {n: p.abs().max().item() for n, p in model.named_parameters() if p.abs().max() > FP16_MAX}
tiny = sum(((p != 0) & (p.abs() < 6.1e-5)).sum().item() for p in model.parameters())
print("activations near the ceiling:", sorted(risky.items(), key=lambda kv: -kv[1])[:20])
print("weights that would become inf:", big_w)
print("weights that become subnormal or zero:", tiny)Hooks see module outputs, not values inside a module, so a norm whose input peaks above 256 is a risk even if its output is small; compare each norm's input peak against that threshold too. If nothing is flagged, convert, then compare FP16 and BF16 outputs on held-out prompts: logits should agree closely and greedy generations should mostly match. If a handful of layers are flagged, keep them in higher precision, as described in mixed-precision inference.
Hardware and software support
NVIDIA GPUs have had FP16 tensor cores since Volta; BF16 tensor cores arrived with Ampere, so on V100 and T4-class hardware FP16 is the fast 16-bit path and BF16 may not be accelerated at all. Many mobile GPUs and NPUs run FP16 natively and treat it as their highest-throughput floating-point type. On CPUs, x86 has long had F16C conversion instructions, and AVX-512 FP16 adds native arithmetic on processors that support it; Arm added half-precision arithmetic as an architecture extension in Armv8.2-A. Graphics APIs expose half floats in shaders, and WebGPU offers them behind the optional shader-f16 feature. Always check that the specific operator you need has an FP16 kernel on your target; a missing kernel silently falls back to FP32 or to a slow path.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| NaN logits after converting a BF16 model | An activation or a squared norm exceeded 65,504 | Run the audit; keep norms and flagged layers in FP32 or BF16 |
| Training loss stalls; loss scale keeps falling | Persistent overflow in a layer | Find the layer with inf gradients; consider BF16 |
| Training silently learns slower than FP32 | No loss scaling, so small gradients underflow | Use GradScaler or an equivalent |
| Clipping has no effect | Clipped before unscaling | Call unscale_ first |
| Different results on two backends | Subnormal flushing or saturating conversion | Compare conversion settings before debugging the model |
| Long reductions inaccurate | FP16 accumulator | Accumulate in FP32 |
Trade-offs: FP16 against its neighbours
Against FP32, FP16 halves memory and is far faster on tensor hardware, at the cost of range and precision you must manage. Against BF16 it has eight times finer steps but a ceiling of 65,504 instead of about 3e38, so it suits models whose activations are known to fit and hardware without fast BF16, and it is the riskier choice for large models trained in BF16. Against FP8 and integer formats it is the safe high-quality baseline: it needs no calibration of scales per tensor, and it is the reference that quantized models are evaluated against, as the quantization deep dive explains.
What to do next
- Memorise the four numbers: 65,504 maximum, 6.1e-5 smallest normal, 6e-8 smallest subnormal and about 0.001 relative step.
- Before serving a BF16-trained model in FP16, run the activation and weight audit above on representative prompts.
- Keep softmax, normalisation statistics, losses and long reductions in FP32 in any custom code.
- For FP16 training, use autocast with a dynamic loss scaler, unscale before clipping and log the scale every step.
- Compare FP16 and BF16 outputs on a held-out set before switching production traffic.
- If your hardware has fast BF16 and the model was trained in BF16, prefer BF16 unless you have measured a reason not to.