The B200 is the first Blackwell data-centre GPU most teams will actually run. In practice the gains arrive only if the software meets the hardware halfway. The headline FP4 number needs block-scaled 4-bit formats that your framework has to produce. The new tensor cores need kernels compiled for a new architecture. The extra memory changes which parallelism layout is optimal.

This article explains the B200 from the software side: what is on the package, how the tensor cores and block-scaled formats work, how Transformer Engine exposes them, how to reason about capacity and speed, and what bites in the first weeks of a migration. Numbers are NVIDIA's published HGX B200 figures. NVIDIA quotes FP8 and FP4 peaks with 2:4 structured sparsity; this article uses the dense figures, half of those, because dense is what real workloads get.

Advertisement

What is on the package

One B200 package: two dies presented to software as one CUDA deviceHBM3estacksHBM3estacksHBM3estacksHBM3estacksDie 0SMs + L2 + tensor coresDie 1SMs + L2 + tensor coresNV-HBI die-to-die link (10 TB/s per NVIDIA)Per SM: registers, shared memory, 256 KB TMEMtcgen05 MMA reads operands from SMEM, accumulates in TMEMNVLink 5: 1.8 TB/s per GPUto NVSwitch, 8-GPU domain on HGX B200PCIe to host CPU and NICsscale-out over InfiniBand or EthernetHGX B200 board: 180 GB HBM3e and about 7.7 TB/s per GPU; up to 1,000 W per GPU.
The B200 joins two reticle-sized dies with a high-bandwidth die-to-die link and exposes them as a single GPU. HBM3e stacks sit beside each die; NVLink 5 and PCIe leave the package.

A B200 package contains two large dies, 208 billion transistors in total by NVIDIA's count, joined by a die-to-die interface NVIDIA calls NV-HBI at 10 TB/s. The key design decision is that software sees one GPU. CUDA reports one device, memory is one address space, and a kernel launch spreads its thread blocks across the SMs of both dies. The cost of the split is a small NUMA-like penalty on cross-die memory reads, which matters mainly to kernel authors.

Memory is HBM3e. NVIDIA announced the B200 with 192 GB and 8 TB/s; the HGX B200 boards that most servers and clouds ship expose 180 GB per GPU at about 7.7 TB/s. Plan with the number your instance actually reports in nvidia-smi, not the launch slide. Each GPU has fifth-generation NVLink at 1.8 TB/s of total bandwidth, twice the H100's 900 GB/s, and an HGX B200 board connects eight GPUs through NVSwitch into one all-to-all domain. Power is up to 1,000 W per GPU on HGX, against 700 W for an H100 SXM.

Per GPUH100 SXMB200 on HGX B200What it changes for software
Memory80 GB HBM3180 GB HBM3eLarger models per GPU, much more KV cache, less sharding
Memory bandwidth3.35 TB/sabout 7.7 TB/sDecode and other memory-bound kernels run roughly 2x faster
FP8 tensor (dense)about 1.98 PFLOPSabout 4.5 PFLOPSCompute-bound GEMMs roughly 2.3x faster at equal precision
FP4 tensor (dense)not supportedabout 9 PFLOPSNew precision tier, needs block-scaled formats
NVLink900 GB/s1.8 TB/sTensor parallel and all-to-all traffic inside the node
Power700 Wup to 1,000 WRack density, cooling and power capping decisions

Fifth-generation tensor cores and tensor memory

The architecture is compute capability 10.0, compiled as sm_100. Its tensor cores are driven by a new family of PTX instructions, tcgen05, and they change how a matrix multiply is organised inside an SM. On Hopper, a warpgroup issued asynchronous wgmma instructions and the accumulator lived in registers. Blackwell adds a dedicated on-chip tensor memory, TMEM, of 256 KB per SM. The MMA reads its operands from shared memory, which the Tensor Memory Accelerator fills asynchronously from HBM, and accumulates into TMEM. Registers are freed for the epilogue: scaling, bias, activation and conversion.

Two practical consequences follow. First, a Hopper-tuned kernel does not use any of this; Hopper's wgmma path is not how Blackwell reaches peak, so a binary built only for sm_90 either runs through JIT-compiled PTX at reduced speed or fails to find a kernel image. Second, you mostly consume the new tensor cores through libraries rather than writing them yourself: cuBLAS and cuBLASLt, cuDNN, CUTLASS 3.x templates, and the kernels inside FlashAttention, Transformer Engine and inference engines.

For readers who do write kernels, the relevant background is the Tensor Memory Accelerator and the general tensor core programming model; the B200 extends both rather than replacing them.

Advertisement

Block-scaled low precision: FP8, MXFP8 and NVFP4

The difficulty with low precision is range. An FP4 value in E2M1 format has one sign bit, two exponent bits and one mantissa bit, so it can represent only a handful of magnitudes, from 0.5 to 6. A tensor whose values span five orders of magnitude cannot be stored in it directly. The solution is scaling: store a scale factor alongside the low-precision values, so each value is interpreted as its code times the scale.

Hopper-era FP8 training used one scale per tensor, which works because E4M3 and E5M2 have a reasonable range but breaks when one outlier forces a large scale on everything else. Blackwell tensor cores apply scales per small block of elements in hardware, so each block gets a scale fitted to its own values and an outlier only damages its own block.

FormatElement typeBlock sizeBlock scale
FP8 (per tensor)E4M3 or E5M2whole tensorFP32
MXFP8E4M332 elementsE8M0 (power of two)
MXFP4E2M132 elementsE8M0 (power of two)
NVFP4E2M116 elementsE4M3, plus a per-tensor FP32 scale

NVFP4's two differences from MXFP4 both improve accuracy. Blocks of 16 instead of 32 mean fewer values share a scale. An E4M3 scale can take values between powers of two, so the block's largest value maps close to the top of the FP4 range rather than being rounded down by up to a factor of two. The cost is storage: 8 bits of scale per 16 values is 0.5 extra bits per element, against 0.25 for MXFP4. A minimal reference quantiser makes the mechanics concrete:

import numpy as np

FP4_GRID = np.array([0, 0.5, 1, 1.5, 2, 3, 4, 6], dtype=np.float32)   # E2M1 magnitudes

def quantize_nvfp4(x, block=16):
    """Reference NVFP4-style quantiser: per-block scale, then round to the E2M1 grid."""
    x = x.reshape(-1, block)
    amax = np.abs(x).max(axis=1, keepdims=True)
    tensor_scale = amax.max() / (448.0 * 6.0)          # keeps block scales inside E4M3 range
    block_scale = amax / 6.0 / tensor_scale             # would be stored as E4M3
    scaled = x / (block_scale * tensor_scale + 1e-12)
    idx = np.abs(np.abs(scaled)[..., None] - FP4_GRID).argmin(-1)
    q = np.sign(scaled) * FP4_GRID[idx]                 # would be stored as 4-bit codes
    return q, block_scale, tensor_scale

def dequantize(q, block_scale, tensor_scale):
    return (q * block_scale * tensor_scale).reshape(-1)

w = np.random.randn(4096).astype(np.float32)
q, s, t = quantize_nvfp4(w)
err = np.abs(dequantize(q, s, t) - w).mean() / np.abs(w).mean()
print(f"mean relative error: {err:.3f}")

This is an illustration only: real kernels pack two codes per byte and apply the scales inside the tensor core. Running it on your own weights is still a quick signal of which layers will tolerate 4-bit.

Using the formats from a training loop

Transformer Engine is the usual way to get block-scaled precision into a PyTorch model without writing kernels. Its recipes live in transformer_engine.common.recipe: DelayedScaling and Float8CurrentScaling for per-tensor FP8, Float8BlockScaling for block-scaled FP8, and MXFP8BlockScaling and NVFP4BlockScaling for the formats that need Blackwell. You replace nn.Linear and related layers with TE modules and wrap the forward pass in an autocast context:

import torch
import transformer_engine.pytorch as te
from transformer_engine.common.recipe import MXFP8BlockScaling

major, minor = torch.cuda.get_device_capability()
assert (major, minor) >= (10, 0), "block-scaled recipes need a Blackwell GPU"

model = torch.nn.Sequential(
    te.LayerNormLinear(4096, 16384),
    te.Linear(16384, 4096),
).cuda()
opt = torch.optim.AdamW(model.parameters(), lr=1e-4)
recipe = MXFP8BlockScaling()

for x, y in loader:
    with te.autocast(enabled=True, recipe=recipe):   # older TE: te.fp8_autocast(enabled=True, fp8_recipe=recipe)
        out = model(x.cuda())
    loss = torch.nn.functional.mse_loss(out.float(), y.cuda())
    loss.backward()                                   # backward runs outside the context
    opt.step(); opt.zero_grad(set_to_none=True)

Master weights and optimizer state stay in higher precision; only the GEMM inputs are quantised. For NVFP4 training, Transformer Engine's documentation describes extra machinery: 2D 16x16 block scaling for weights, stochastic rounding for gradients and random Hadamard transforms that spread outliers before quantisation. A sensible adoption order is BF16 baseline, then MXFP8 with a loss-curve comparison, and only then NVFP4 on selected layers, keeping the first and last layers and anything numerically sensitive in higher precision.

Worked example: a 70B model on one B200

Consider serving a 70B-parameter dense decoder with grouped-query attention: 80 layers, 8 KV heads, head dimension 128. The questions are whether it fits, how much context it can hold and how fast it can decode.

Weights in FP8 take 70 GB. Reserve about 10 GB for activations, the CUDA context and allocator slack, which leaves about 100 GB of the 180 GB for KV cache. Each token stores a key and a value per layer per KV head: 2 x 80 x 8 x 128 = 163,840 elements. In FP8 that is 160 KiB per token, so 100 GB holds roughly 610,000 tokens: about 18 concurrent sequences at a 32k context. On an 80 GB H100 the same weights leave almost no room for cache, forcing two-GPU tensor parallelism.

Decode speed at small batch is bound by memory bandwidth, because every generated token must read all weights once. 70 GB at 7.7 TB/s is about 9 ms per step, a ceiling of roughly 110 tokens per second for a single sequence; the H100's 3.35 TB/s gives about 21 ms. Batching amortises the weight read across sequences, so throughput rises almost linearly with batch until KV-cache reads and compute take over. More cache means bigger batches, which is where serving economics live. Quantising weights to NVFP4 halves the weight read again, to about 35 GB, which roughly doubles the batch-one ceiling if the accuracy holds.

Roofline: when the B200 is compute-bound

The ridge point of a roofline is peak FLOPs divided by peak bandwidth: the arithmetic intensity, in FLOPs per byte, above which a kernel is compute-bound. For dense FP8 on a B200 that is about 4.5e15 / 7.7e12, or roughly 580 FLOPs per byte; for dense FP4 it is about 1,170. A GEMM of shape M x K times K x N performs 2MNK FLOPs and moves at least (MK + KN + MN) elements, so large training GEMMs clear the ridge easily while decode GEMMs with M equal to the batch size of a few dozen sit far below it.

The practical lesson is that lower precision helps two different regimes for different reasons. In training and prefill, FP4 and FP8 raise the compute ceiling. In decode, they shrink the bytes moved. For decode kernels, judge achieved bandwidth against the roughly 7.7 TB/s ceiling, not achieved FLOPs. The HBM article covers why achieved bandwidth usually tops out below the spec figure.

NVLink domains and parallelism layout

An HGX B200 board is an eight-GPU NVLink domain with 1.8 TB/s per GPU, 14.4 TB/s across the board. Tensor parallelism and expert-parallel all-to-all belong inside that domain; pipeline and data parallelism cross the slower scale-out network between servers. The 72-GPU NVLink racks are GB200 products, not HGX B200.

The extra memory usually argues for less tensor parallelism, not more. A model that needed TP=8 on H100 for memory reasons may fit at TP=4 or TP=2 on B200, which cuts per-layer communication and frees GPUs for data parallelism. Re-run the layout search rather than porting the configuration from the H100. The NVLink and NVSwitch article explains the topology and the collective costs in detail.

Software prerequisites and bring-up

The sm_100 target needs CUDA 12.8 or later, a matching driver, and framework builds compiled for it; PyTorch wheels built against CUDA 12.8 or newer include it. HGX systems with NVSwitch also need the fabric manager service running before multi-GPU jobs will initialise. A short bring-up script catches most problems before a long job does:

import torch, subprocess

print(torch.__version__, torch.version.cuda)
print(torch.cuda.get_device_name(0), torch.cuda.get_device_capability(0))   # expect (10, 0)
print("arch list:", torch.cuda.get_arch_list())                              # must include sm_100

a = torch.randn(8192, 8192, device="cuda", dtype=torch.bfloat16)
torch.cuda.synchronize()
start, end = torch.cuda.Event(enable_timing=True), torch.cuda.Event(enable_timing=True)
start.record()
for _ in range(50):
    a @ a
end.record(); torch.cuda.synchronize()
tflops = 50 * 2 * 8192**3 / (start.elapsed_time(end) / 1e3) / 1e12
print(f"bf16 GEMM: {tflops:.0f} TFLOPS")

print(subprocess.run(["nvidia-smi", "topo", "-m"], capture_output=True, text=True).stdout)

Record the GEMM rate, topology and an NCCL all-reduce test per node and compare nodes: a node 15 percent slower than its neighbours is usually power-capped, throttled or missing a link, and it sets the pace of every synchronous training step.

Failure modes in the first month

  • No kernel image for the device. A container built with an arch list that stops at sm_90 fails on launch or falls back to JIT-compiled PTX with poor performance. Rebuild every extension, including custom CUDA ops, with sm_100 in the arch list.
  • Silent slow paths. A library version that predates Blackwell runs, but on generic kernels. Profile with Nsight Systems and check that hot GEMMs are Blackwell kernels, not fallbacks.
  • Accuracy drift from 4-bit. NVFP4 inference that looks fine on perplexity can regress on long-form reasoning or on rare tokens. Evaluate on task metrics, not only perplexity, and keep a BF16 or FP8 reference.
  • Power and thermal throttling. At 1,000 W per GPU, an under-provisioned rack or a liquid-cooling fault reduces clocks. Watch clock and throttle-reason metrics, not just utilisation.

What to do next

  1. Run the bring-up script on one B200 node and confirm compute capability (10, 0), sm_100 in the arch list, the GEMM rate and the NVLink topology.
  2. Rebuild every custom CUDA extension with sm_100 and profile one training step to confirm the hot GEMMs use Blackwell kernels.
  3. Redo the capacity arithmetic for your model: weights, KV cache per token and target batch, then pick the smallest tensor-parallel degree that fits.
  4. Establish a BF16 baseline, then trial MXFP8 with a loss-curve or task-metric comparison before trying NVFP4 on selected layers.
  5. Add clock, power and throttle-reason metrics to your dashboards and compare nodes against each other every week.
  6. Price the result per token or per training step, not per GPU hour, and compare against your current fleet.
Key takeaway: The B200 is two dies acting as one GPU, with 180 GB of HBM3e at about 7.7 TB/s, NVLink 5 and tensor cores that accumulate into on-chip tensor memory. Its biggest gains come from block-scaled formats (MXFP8 and NVFP4), which need Blackwell-aware libraries and careful accuracy checks. Rebuild for sm_100, re-plan parallelism around the extra memory, reason about decode as a bandwidth problem and training as a compute problem, and measure cost per token rather than per GPU hour.