The GB200 is not a GPU. It is a superchip: one NVIDIA Grace CPU and two Blackwell GPUs on one board, joined by a cache-coherent link. Thirty-six of them, wired through NVLink switches, make the GB200 NVL72 rack, a 72-GPU NVLink domain. The interesting engineering questions are not about tensor-core peak numbers (the B200 article covers the GPU itself) but about the two things GB200 adds: a large, fast CPU memory tier that the GPUs can reach directly, and a scale-up domain nine times larger than an eight-GPU server.
This article explains both from first principles, shows how to measure them, maps a large mixture-of-experts model onto a rack, and covers the operational details that bite in the first weeks: Arm containers, IMEX, topology-aware scheduling and partial-domain failure. Specifications were checked against NVIDIA's GB200 NVL72 product page and its multi-node NVLink documentation on 2026-10-02; anything we derived ourselves is labelled as arithmetic.
What is on the board
Each superchip carries one Grace CPU with 72 Arm Neoverse V2 cores and two Blackwell GPUs. Each Blackwell GPU is itself two reticle-sized dies joined by a 10 TB/s die-to-die interconnect, 208 billion transistors in total, presented to software as one CUDA device. The CPU connects to both GPUs over NVLink-C2C (chip-to-chip) at 900 GB/s. The GPUs carry HBM3E; the CPU carries LPDDR5X.
| Quantity | Per superchip | Per NVL72 rack |
|---|---|---|
| Grace CPUs / Blackwell GPUs | 1 / 2 | 36 / 72 |
| Arm Neoverse V2 cores | 72 | 2,592 |
| HBM3E capacity and bandwidth | 372 GB, 16 TB/s | 13.4 TB, 576 TB/s |
| LPDDR5X capacity and bandwidth | up to 480 GB, up to 512 GB/s | up to 17 TB, 14 TB/s |
| NVLink bandwidth | 3.6 TB/s (1.8 TB/s per GPU) | 130 TB/s |
| NVFP4 tensor, dense | 20 PFLOPS | 720 PFLOPS |
Two derived numbers are worth writing down. Dividing 372 GB by two gives roughly 186 GB of HBM per GPU, which is what your memory planner should assume, not the larger figure sometimes quoted for standalone Blackwell parts. And the bandwidth ratios matter more than the absolute values: HBM is about 16 TB/s per superchip, NVLink-C2C 0.9 TB/s, LPDDR5X at most 0.5 TB/s. CPU memory is a capacity tier roughly thirty times slower than HBM, but still several times faster than a PCIe-attached host.
The architecture at a glance
NVIDIA's own description of the DGX GB200 rack lists ten compute trays above and eight below the switches, eighteen in all, connected by nine NVLink switch trays. Seventy-two GPUs over eighteen trays is four GPUs, or two superchips, per tray. Each tray runs its own operating system image, which is the crucial software fact: GPUs in the same NVLink domain are owned by different kernels. That is why the rack needs a component an eight-GPU server never did, the Internode Memory Exchange service (IMEX), described below.
Beyond the rack, nothing changes from earlier generations. Data-parallel replicas in different racks communicate over the scale-out network, so the job designer's task is to keep the bandwidth-hungry traffic (tensor and expert parallelism) inside one NVLink domain and push only the gradient reductions across racks.
Two memory tiers, one address space
On an x86 server the GPU reaches host memory through PCIe, and in practice software treats host memory as a staging area: copy in, compute, copy out. NVLink-C2C changes two things. It is several times faster than a PCIe host link, and it is cache coherent, so the CPU and GPUs see a consistent view of shared data without explicit cache flushes. Pinned host allocations and CUDA managed memory both move over this link instead of PCIe; check NVIDIA's CUDA documentation for your driver version before relying on any other way of sharing memory.
Coherent does not mean free. A kernel that streams its working set from LPDDR5X runs at LPDDR5X speed. The right mental model is a two-tier memory: HBM for anything touched every step (weights in use, activations, the hot part of a KV cache), LPDDR5X for state that is large and touched rarely or sequentially (optimizer moments, cold KV blocks, staged data). Do not assume a particular NUMA layout from earlier Grace Hopper material; inspect it on your own system with numactl -H and nvidia-smi topo -m before pinning processes.
Measure the tiers before you design around them. This microbenchmark reports effective copy bandwidth for each path; run it once per node type and keep the numbers next to your capacity plan.
import time
import torch
def gbps(src, dst, iters=20):
dst.copy_(src, non_blocking=True) # warm up
torch.cuda.synchronize()
t0 = time.perf_counter()
for _ in range(iters):
dst.copy_(src, non_blocking=True)
torch.cuda.synchronize()
return iters * src.numel() * src.element_size() / (time.perf_counter() - t0) / 1e9
n = 4 * 2**30 # 4 GiB buffers
hbm_a = torch.empty(n, dtype=torch.uint8, device="cuda")
hbm_b = torch.empty_like(hbm_a)
pinned = torch.empty(n, dtype=torch.uint8, pin_memory=True)
pageable = torch.empty(n, dtype=torch.uint8)
print("HBM -> HBM ", gbps(hbm_a, hbm_b)) # each byte is read and written
print("pinned host -> GPU", gbps(pinned, hbm_a))
print("GPU -> pinned host", gbps(hbm_a, pinned))
print("pageable -> GPU ", gbps(pageable, hbm_a))Expect pinned host copies to be far faster than pageable ones, and expect all host paths to be bounded by the lower of NVLink-C2C and LPDDR5X bandwidth. If the pinned number looks like a PCIe number, the process is probably bound to the wrong CPU or memory node.
Using the CPU tier: optimizer offload
The cleanest use of LPDDR5X in training is to hold the state that the optimizer needs once per step. Mixed-precision Adam keeps a 4-byte master weight and two 4-byte moments per parameter, 12 bytes that are read and written exactly once per step. Moving them off the GPU frees most of the HBM a training job uses for state, and the Grace cores can apply the update while the GPUs wait. The sketch below shows the mechanics; production code (DeepSpeed ZeRO-Offload and similar) overlaps the transfer with the backward pass and shards state across ranks.
import torch
class GraceOffloadAdam:
"""bf16 weights and grads stay in HBM; fp32 master copy and Adam moments live in LPDDR5X."""
def __init__(self, params, lr=1e-4, betas=(0.9, 0.95), eps=1e-8):
self.params = [p for p in params if p.requires_grad]
self.master = [p.detach().float().cpu().pin_memory() for p in self.params]
self.m = [torch.zeros_like(w) for w in self.master]
self.v = [torch.zeros_like(w) for w in self.master]
self.grad_buf = [torch.empty_like(w).pin_memory() for w in self.master]
self.lr, self.b1, self.b2, self.eps, self.t = lr, *betas, eps, 0
@torch.no_grad()
def step(self):
self.t += 1
for p, g in zip(self.params, self.grad_buf):
g.copy_(p.grad, non_blocking=True) # GPU -> Grace over NVLink-C2C
torch.cuda.synchronize()
c1, c2 = 1 - self.b1 ** self.t, 1 - self.b2 ** self.t
for p, w, m, v, g in zip(self.params, self.master, self.m, self.v, self.grad_buf):
m.mul_(self.b1).add_(g, alpha=1 - self.b1) # runs on the Arm cores
v.mul_(self.b2).addcmul_(g, g, value=1 - self.b2)
w.addcdiv_(m / c1, (v / c2).sqrt_().add_(self.eps), value=-self.lr)
p.copy_(w, non_blocking=True) # fp32 -> bf16 on the way backThe trade-off is step time. The update now costs at least the time to read and write 12 bytes per parameter at LPDDR5X speed, plus the gradient and weight transfers over C2C. If the step is long (large batch, long sequences) that cost hides behind compute; if it is short, it dominates. Offload is a capacity lever, not a speed lever.
The NVL72 domain: why 72 matters
Within an NVLink domain every GPU can reach every other at 1.8 TB/s through the switches, and collectives such as all-reduce and all-to-all run without touching the network. On an eight-GPU server, tensor parallelism and expert parallelism are confined to eight GPUs; anything wider crosses the network at a small fraction of that bandwidth. On NVL72 the same communication pattern can span 72 GPUs, which changes model-parallel layouts more than any single-GPU improvement does. The NVLink article explains the switch fabric, and NCCL collectives explains which algorithms the library picks.
Mixture-of-experts models benefit most. Each token is routed to a few experts, so expert parallelism produces an all-to-all exchange every MoE layer. That exchange is latency- and bandwidth-sensitive and scales badly over a network. Keeping all experts of a layer inside one NVLink domain makes the all-to-all a switch-local operation. Dense models benefit too, from wider tensor parallelism that keeps per-GPU memory low without leaving NVLink.
Worked example: sizing a 1.2T-parameter MoE on one rack
Take a mixture-of-experts model with 1.2 trillion total parameters. For inference with FP8 weights that is about 1.2 TB, against 13.4 TB of HBM in the rack: sharded across 72 GPUs, roughly 17 GB of weights per GPU, leaving most of each GPU's 186 GB for KV cache and activations. One rack serves the model with experts spread across all 72 GPUs and the all-to-all on NVLink.
Training is harder. A common mixed-precision budget is 16 bytes per parameter: 2 for bf16 weights, 2 for gradients, 12 for the fp32 master and Adam moments. That is 19.2 TB, more than the rack's HBM before a single activation is stored. Two options follow. Shard the state across several racks with data parallelism, which adds cross-rack traffic, or keep the 4 bytes of weights and gradients in HBM (4.8 TB) and put the 12 bytes of optimizer state, 14.4 TB, in LPDDR5X, which fits within the rack's 17 TB.
The offload option has a visible price. Reading and writing 14.4 TB at the rack's 14 TB/s aggregate LPDDR5X bandwidth costs about two seconds per step for the update alone, before compute. If your step takes thirty seconds, that is acceptable; if it takes three, it is not, and multi-rack sharding wins. This is the kind of arithmetic to do before buying, because it decides whether the CPU tier is useful for your workload.
Scheduling: IMEX, ComputeDomains and cliques
Because each tray runs its own OS, sharing GPU memory across trays needs a broker. IMEX is driver-level software that exports and imports GPU memory between nodes in the same NVLink domain, with access control per IMEX domain and channel. Without it, peer access across trays fails and libraries fall back to the network, often silently.
On Kubernetes, NVIDIA's DRA driver for GPUs (NVIDIA's announcement used version 25.8.0, requiring Kubernetes 1.32 or later with dynamic resource allocation enabled) exposes this as a ComputeDomain custom resource. Pods that claim the domain's channel are placed in a shared IMEX domain, and the nvidia.com/gpu.clique node label identifies which nodes share an NVLink partition, so affinity on that label keeps a job inside one rack.
apiVersion: resource.nvidia.com/v1beta1
kind: ComputeDomain
metadata:
name: train-cd
spec:
numNodes: 0
channel:
resourceClaimTemplate:
name: train-cd-channel
---
apiVersion: v1
kind: Pod
metadata:
name: trainer-0
labels: {job: moe-train}
spec:
affinity:
podAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector: {matchLabels: {job: moe-train}}
topologyKey: nvidia.com/gpu.clique # stay inside one NVLink partition
resourceClaims:
- name: imex
resourceClaimTemplateName: train-cd-channel
containers:
- name: trainer
image: registry.example.com/trainer:arm64 # Grace is aarch64
resources:
claims: [{name: imex}]
limits: {nvidia.com/gpu: 4}The same idea applies to Slurm or any other scheduler: treat the NVLink domain as a placement unit, and never let a tensor- or expert-parallel group straddle two cliques.
Failure modes
- Wrong architecture images. Grace is aarch64. x86 container images, pip wheels without arm64 builds and custom CUDA extensions compiled for x86 fail at start-up. Build multi-architecture images and test them before the hardware arrives.
- Silent network fallback. If IMEX is not running or the pod lacks the channel claim, collectives still complete, over the scale-out network, at a fraction of the speed. Run with
NCCL_DEBUG=INFOat job start and assert in code that the reported transport is NVLink-based for intra-rack groups. - A job sized to exactly 72 GPUs. One failed GPU or tray shrinks the domain, and a job that needs all 72 cannot restart until repair. Design layouts that fit in fewer GPUs, for example 64, and keep the rest as hot spares.
- Assumed memory placement. Copying a Grace Hopper NUMA recipe can bind a process to the far memory node and reduce host bandwidth. Measure with the benchmark above.
- Thermal and power events. The rack is liquid cooled; coolant and power faults throttle or drop whole trays. Export clock, power and throttle-reason telemetry per GPU and alert on sustained clock drops, not only on errors.
GB200 NVL72 or eight-GPU servers?
| Workload | Better fit | Why |
|---|---|---|
| Large MoE training and serving | NVL72 | All-to-all stays inside NVLink |
| Dense models that fit in 8 GPUs with TP=8 | 8-GPU servers | No benefit from the wider domain, simpler operations |
| Long-context inference with huge KV caches | NVL72 | HBM plus a fast CPU tier for cold blocks |
| Many small independent jobs | 8-GPU servers | Rack-scale domains are hard to share finely |
| x86-only software stack | 8-GPU x86 servers | Grace requires arm64 builds |
What to do next
- Port your containers to arm64 and run your full test suite on Grace before capacity arrives.
- Run the bandwidth benchmark on one node and record HBM, pinned host and pageable numbers.
- Write the memory budget for your model: bytes per parameter for weights, gradients and optimizer state, KV cache per token, and which tier each lives in.
- Pick a parallel layout that keeps tensor and expert groups inside one clique, and leaves spare GPUs for failures.
- Deploy the DRA driver, define ComputeDomains, and add clique affinity to every multi-node job.
- Gate job start on an NCCL transport check, and alert on clock throttling per GPU.