The B200 is the first Blackwell data-centre GPU most teams will actually run. In practice the gains arrive only if the software meets the hardware halfway. The headline FP4 number needs block-scaled 4-bit formats that your framework has to produce. The new tensor cores need kernels compiled for a new architecture. The extra memory changes which parallelism layout is optimal.
This article explains the B200 from the software side: what is on the package, how the tensor cores and block-scaled formats work, how Transformer Engine exposes them, how to reason about capacity and speed, and what bites in the first weeks of a migration. Numbers are NVIDIA's published HGX B200 figures. NVIDIA quotes FP8 and FP4 peaks with 2:4 structured sparsity; this article uses the dense figures, half of those, because dense is what real workloads get.
What is on the package
A B200 package contains two large dies, 208 billion transistors in total by NVIDIA's count, joined by a die-to-die interface NVIDIA calls NV-HBI at 10 TB/s. The key design decision is that software sees one GPU. CUDA reports one device, memory is one address space, and a kernel launch spreads its thread blocks across the SMs of both dies. The cost of the split is a small NUMA-like penalty on cross-die memory reads, which matters mainly to kernel authors.
Memory is HBM3e. NVIDIA announced the B200 with 192 GB and 8 TB/s; the HGX B200 boards that most servers and clouds ship expose 180 GB per GPU at about 7.7 TB/s. Plan with the number your instance actually reports in nvidia-smi, not the launch slide. Each GPU has fifth-generation NVLink at 1.8 TB/s of total bandwidth, twice the H100's 900 GB/s, and an HGX B200 board connects eight GPUs through NVSwitch into one all-to-all domain. Power is up to 1,000 W per GPU on HGX, against 700 W for an H100 SXM.
| Per GPU | H100 SXM | B200 on HGX B200 | What it changes for software |
|---|---|---|---|
| Memory | 80 GB HBM3 | 180 GB HBM3e | Larger models per GPU, much more KV cache, less sharding |
| Memory bandwidth | 3.35 TB/s | about 7.7 TB/s | Decode and other memory-bound kernels run roughly 2x faster |
| FP8 tensor (dense) | about 1.98 PFLOPS | about 4.5 PFLOPS | Compute-bound GEMMs roughly 2.3x faster at equal precision |
| FP4 tensor (dense) | not supported | about 9 PFLOPS | New precision tier, needs block-scaled formats |
| NVLink | 900 GB/s | 1.8 TB/s | Tensor parallel and all-to-all traffic inside the node |
| Power | 700 W | up to 1,000 W | Rack density, cooling and power capping decisions |
Fifth-generation tensor cores and tensor memory
The architecture is compute capability 10.0, compiled as sm_100. Its tensor cores are driven by a new family of PTX instructions, tcgen05, and they change how a matrix multiply is organised inside an SM. On Hopper, a warpgroup issued asynchronous wgmma instructions and the accumulator lived in registers. Blackwell adds a dedicated on-chip tensor memory, TMEM, of 256 KB per SM. The MMA reads its operands from shared memory, which the Tensor Memory Accelerator fills asynchronously from HBM, and accumulates into TMEM. Registers are freed for the epilogue: scaling, bias, activation and conversion.
Two practical consequences follow. First, a Hopper-tuned kernel does not use any of this; Hopper's wgmma path is not how Blackwell reaches peak, so a binary built only for sm_90 either runs through JIT-compiled PTX at reduced speed or fails to find a kernel image. Second, you mostly consume the new tensor cores through libraries rather than writing them yourself: cuBLAS and cuBLASLt, cuDNN, CUTLASS 3.x templates, and the kernels inside FlashAttention, Transformer Engine and inference engines.
For readers who do write kernels, the relevant background is the Tensor Memory Accelerator and the general tensor core programming model; the B200 extends both rather than replacing them.
Block-scaled low precision: FP8, MXFP8 and NVFP4
The difficulty with low precision is range. An FP4 value in E2M1 format has one sign bit, two exponent bits and one mantissa bit, so it can represent only a handful of magnitudes, from 0.5 to 6. A tensor whose values span five orders of magnitude cannot be stored in it directly. The solution is scaling: store a scale factor alongside the low-precision values, so each value is interpreted as its code times the scale.
Hopper-era FP8 training used one scale per tensor, which works because E4M3 and E5M2 have a reasonable range but breaks when one outlier forces a large scale on everything else. Blackwell tensor cores apply scales per small block of elements in hardware, so each block gets a scale fitted to its own values and an outlier only damages its own block.
| Format | Element type | Block size | Block scale |
|---|---|---|---|
| FP8 (per tensor) | E4M3 or E5M2 | whole tensor | FP32 |
| MXFP8 | E4M3 | 32 elements | E8M0 (power of two) |
| MXFP4 | E2M1 | 32 elements | E8M0 (power of two) |
| NVFP4 | E2M1 | 16 elements | E4M3, plus a per-tensor FP32 scale |
NVFP4's two differences from MXFP4 both improve accuracy. Blocks of 16 instead of 32 mean fewer values share a scale. An E4M3 scale can take values between powers of two, so the block's largest value maps close to the top of the FP4 range rather than being rounded down by up to a factor of two. The cost is storage: 8 bits of scale per 16 values is 0.5 extra bits per element, against 0.25 for MXFP4. A minimal reference quantiser makes the mechanics concrete:
import numpy as np
FP4_GRID = np.array([0, 0.5, 1, 1.5, 2, 3, 4, 6], dtype=np.float32) # E2M1 magnitudes
def quantize_nvfp4(x, block=16):
"""Reference NVFP4-style quantiser: per-block scale, then round to the E2M1 grid."""
x = x.reshape(-1, block)
amax = np.abs(x).max(axis=1, keepdims=True)
tensor_scale = amax.max() / (448.0 * 6.0) # keeps block scales inside E4M3 range
block_scale = amax / 6.0 / tensor_scale # would be stored as E4M3
scaled = x / (block_scale * tensor_scale + 1e-12)
idx = np.abs(np.abs(scaled)[..., None] - FP4_GRID).argmin(-1)
q = np.sign(scaled) * FP4_GRID[idx] # would be stored as 4-bit codes
return q, block_scale, tensor_scale
def dequantize(q, block_scale, tensor_scale):
return (q * block_scale * tensor_scale).reshape(-1)
w = np.random.randn(4096).astype(np.float32)
q, s, t = quantize_nvfp4(w)
err = np.abs(dequantize(q, s, t) - w).mean() / np.abs(w).mean()
print(f"mean relative error: {err:.3f}")This is an illustration only: real kernels pack two codes per byte and apply the scales inside the tensor core. Running it on your own weights is still a quick signal of which layers will tolerate 4-bit.
Using the formats from a training loop
Transformer Engine is the usual way to get block-scaled precision into a PyTorch model without writing kernels. Its recipes live in transformer_engine.common.recipe: DelayedScaling and Float8CurrentScaling for per-tensor FP8, Float8BlockScaling for block-scaled FP8, and MXFP8BlockScaling and NVFP4BlockScaling for the formats that need Blackwell. You replace nn.Linear and related layers with TE modules and wrap the forward pass in an autocast context:
import torch
import transformer_engine.pytorch as te
from transformer_engine.common.recipe import MXFP8BlockScaling
major, minor = torch.cuda.get_device_capability()
assert (major, minor) >= (10, 0), "block-scaled recipes need a Blackwell GPU"
model = torch.nn.Sequential(
te.LayerNormLinear(4096, 16384),
te.Linear(16384, 4096),
).cuda()
opt = torch.optim.AdamW(model.parameters(), lr=1e-4)
recipe = MXFP8BlockScaling()
for x, y in loader:
with te.autocast(enabled=True, recipe=recipe): # older TE: te.fp8_autocast(enabled=True, fp8_recipe=recipe)
out = model(x.cuda())
loss = torch.nn.functional.mse_loss(out.float(), y.cuda())
loss.backward() # backward runs outside the context
opt.step(); opt.zero_grad(set_to_none=True)Master weights and optimizer state stay in higher precision; only the GEMM inputs are quantised. For NVFP4 training, Transformer Engine's documentation describes extra machinery: 2D 16x16 block scaling for weights, stochastic rounding for gradients and random Hadamard transforms that spread outliers before quantisation. A sensible adoption order is BF16 baseline, then MXFP8 with a loss-curve comparison, and only then NVFP4 on selected layers, keeping the first and last layers and anything numerically sensitive in higher precision.
Worked example: a 70B model on one B200
Consider serving a 70B-parameter dense decoder with grouped-query attention: 80 layers, 8 KV heads, head dimension 128. The questions are whether it fits, how much context it can hold and how fast it can decode.
Weights in FP8 take 70 GB. Reserve about 10 GB for activations, the CUDA context and allocator slack, which leaves about 100 GB of the 180 GB for KV cache. Each token stores a key and a value per layer per KV head: 2 x 80 x 8 x 128 = 163,840 elements. In FP8 that is 160 KiB per token, so 100 GB holds roughly 610,000 tokens: about 18 concurrent sequences at a 32k context. On an 80 GB H100 the same weights leave almost no room for cache, forcing two-GPU tensor parallelism.
Decode speed at small batch is bound by memory bandwidth, because every generated token must read all weights once. 70 GB at 7.7 TB/s is about 9 ms per step, a ceiling of roughly 110 tokens per second for a single sequence; the H100's 3.35 TB/s gives about 21 ms. Batching amortises the weight read across sequences, so throughput rises almost linearly with batch until KV-cache reads and compute take over. More cache means bigger batches, which is where serving economics live. Quantising weights to NVFP4 halves the weight read again, to about 35 GB, which roughly doubles the batch-one ceiling if the accuracy holds.
Roofline: when the B200 is compute-bound
The ridge point of a roofline is peak FLOPs divided by peak bandwidth: the arithmetic intensity, in FLOPs per byte, above which a kernel is compute-bound. For dense FP8 on a B200 that is about 4.5e15 / 7.7e12, or roughly 580 FLOPs per byte; for dense FP4 it is about 1,170. A GEMM of shape M x K times K x N performs 2MNK FLOPs and moves at least (MK + KN + MN) elements, so large training GEMMs clear the ridge easily while decode GEMMs with M equal to the batch size of a few dozen sit far below it.
The practical lesson is that lower precision helps two different regimes for different reasons. In training and prefill, FP4 and FP8 raise the compute ceiling. In decode, they shrink the bytes moved. For decode kernels, judge achieved bandwidth against the roughly 7.7 TB/s ceiling, not achieved FLOPs. The HBM article covers why achieved bandwidth usually tops out below the spec figure.
NVLink domains and parallelism layout
An HGX B200 board is an eight-GPU NVLink domain with 1.8 TB/s per GPU, 14.4 TB/s across the board. Tensor parallelism and expert-parallel all-to-all belong inside that domain; pipeline and data parallelism cross the slower scale-out network between servers. The 72-GPU NVLink racks are GB200 products, not HGX B200.
The extra memory usually argues for less tensor parallelism, not more. A model that needed TP=8 on H100 for memory reasons may fit at TP=4 or TP=2 on B200, which cuts per-layer communication and frees GPUs for data parallelism. Re-run the layout search rather than porting the configuration from the H100. The NVLink and NVSwitch article explains the topology and the collective costs in detail.
Software prerequisites and bring-up
The sm_100 target needs CUDA 12.8 or later, a matching driver, and framework builds compiled for it; PyTorch wheels built against CUDA 12.8 or newer include it. HGX systems with NVSwitch also need the fabric manager service running before multi-GPU jobs will initialise. A short bring-up script catches most problems before a long job does:
import torch, subprocess
print(torch.__version__, torch.version.cuda)
print(torch.cuda.get_device_name(0), torch.cuda.get_device_capability(0)) # expect (10, 0)
print("arch list:", torch.cuda.get_arch_list()) # must include sm_100
a = torch.randn(8192, 8192, device="cuda", dtype=torch.bfloat16)
torch.cuda.synchronize()
start, end = torch.cuda.Event(enable_timing=True), torch.cuda.Event(enable_timing=True)
start.record()
for _ in range(50):
a @ a
end.record(); torch.cuda.synchronize()
tflops = 50 * 2 * 8192**3 / (start.elapsed_time(end) / 1e3) / 1e12
print(f"bf16 GEMM: {tflops:.0f} TFLOPS")
print(subprocess.run(["nvidia-smi", "topo", "-m"], capture_output=True, text=True).stdout)Record the GEMM rate, topology and an NCCL all-reduce test per node and compare nodes: a node 15 percent slower than its neighbours is usually power-capped, throttled or missing a link, and it sets the pace of every synchronous training step.
Failure modes in the first month
- No kernel image for the device. A container built with an arch list that stops at sm_90 fails on launch or falls back to JIT-compiled PTX with poor performance. Rebuild every extension, including custom CUDA ops, with sm_100 in the arch list.
- Silent slow paths. A library version that predates Blackwell runs, but on generic kernels. Profile with Nsight Systems and check that hot GEMMs are Blackwell kernels, not fallbacks.
- Accuracy drift from 4-bit. NVFP4 inference that looks fine on perplexity can regress on long-form reasoning or on rare tokens. Evaluate on task metrics, not only perplexity, and keep a BF16 or FP8 reference.
- Power and thermal throttling. At 1,000 W per GPU, an under-provisioned rack or a liquid-cooling fault reduces clocks. Watch clock and throttle-reason metrics, not just utilisation.
What to do next
- Run the bring-up script on one B200 node and confirm compute capability (10, 0), sm_100 in the arch list, the GEMM rate and the NVLink topology.
- Rebuild every custom CUDA extension with sm_100 and profile one training step to confirm the hot GEMMs use Blackwell kernels.
- Redo the capacity arithmetic for your model: weights, KV cache per token and target batch, then pick the smallest tensor-parallel degree that fits.
- Establish a BF16 baseline, then trial MXFP8 with a loss-curve or task-metric comparison before trying NVFP4 on selected layers.
- Add clock, power and throttle-reason metrics to your dashboards and compare nodes against each other every week.
- Price the result per token or per training step, not per GPU hour, and compare against your current fleet.