The GB200 is not a GPU. It is a superchip: one NVIDIA Grace CPU and two Blackwell GPUs on one board, joined by a cache-coherent link. Thirty-six of them, wired through NVLink switches, make the GB200 NVL72 rack, a 72-GPU NVLink domain. The interesting engineering questions are not about tensor-core peak numbers (the B200 article covers the GPU itself) but about the two things GB200 adds: a large, fast CPU memory tier that the GPUs can reach directly, and a scale-up domain nine times larger than an eight-GPU server.

This article explains both from first principles, shows how to measure them, maps a large mixture-of-experts model onto a rack, and covers the operational details that bite in the first weeks: Arm containers, IMEX, topology-aware scheduling and partial-domain failure. Specifications were checked against NVIDIA's GB200 NVL72 product page and its multi-node NVLink documentation on 2026-10-02; anything we derived ourselves is labelled as arithmetic.

Advertisement

What is on the board

Each superchip carries one Grace CPU with 72 Arm Neoverse V2 cores and two Blackwell GPUs. Each Blackwell GPU is itself two reticle-sized dies joined by a 10 TB/s die-to-die interconnect, 208 billion transistors in total, presented to software as one CUDA device. The CPU connects to both GPUs over NVLink-C2C (chip-to-chip) at 900 GB/s. The GPUs carry HBM3E; the CPU carries LPDDR5X.

QuantityPer superchipPer NVL72 rack
Grace CPUs / Blackwell GPUs1 / 236 / 72
Arm Neoverse V2 cores722,592
HBM3E capacity and bandwidth372 GB, 16 TB/s13.4 TB, 576 TB/s
LPDDR5X capacity and bandwidthup to 480 GB, up to 512 GB/sup to 17 TB, 14 TB/s
NVLink bandwidth3.6 TB/s (1.8 TB/s per GPU)130 TB/s
NVFP4 tensor, dense20 PFLOPS720 PFLOPS

Two derived numbers are worth writing down. Dividing 372 GB by two gives roughly 186 GB of HBM per GPU, which is what your memory planner should assume, not the larger figure sometimes quoted for standalone Blackwell parts. And the bandwidth ratios matter more than the absolute values: HBM is about 16 TB/s per superchip, NVLink-C2C 0.9 TB/s, LPDDR5X at most 0.5 TB/s. CPU memory is a capacity tier roughly thirty times slower than HBM, but still several times faster than a PCIe-attached host.

The architecture at a glance

GB200: superchip, compute tray and NVL72 rackOne GB200 superchipGrace CPU72 Arm Neoverse V2 coresBlackwell GPU2 dies, HBM3EBlackwell GPU2 dies, HBM3ELPDDR5X up to 480 GBC2CC2CNVLink-C2C 900 GB/s, coherentCompute tray (by arithmetic)2 superchips = 2 Grace + 4 GPUs, one OS image9 NVLink switch trays1.8 TB/s per GPU, 130 TB/s across the rackNVLinkGB200 NVL72 rack: 18 compute trays, 36 Grace CPUs, 72 GPUs, one NVLink domain13.4 TB HBM3E at 576 TB/s, up to 17 TB LPDDR5X, liquid cooledInside the rack, GPUs in different trays (different OS instances) talk over NVLink; IMEX brokers memory sharing.Between racks, traffic leaves the NVLink domain and uses the scale-out network (InfiniBand or Ethernet).
A superchip couples one Grace CPU to two Blackwell GPUs over NVLink-C2C. By our arithmetic each compute tray holds two superchips; eighteen trays and nine NVLink switch trays form the NVL72 rack.

NVIDIA's own description of the DGX GB200 rack lists ten compute trays above and eight below the switches, eighteen in all, connected by nine NVLink switch trays. Seventy-two GPUs over eighteen trays is four GPUs, or two superchips, per tray. Each tray runs its own operating system image, which is the crucial software fact: GPUs in the same NVLink domain are owned by different kernels. That is why the rack needs a component an eight-GPU server never did, the Internode Memory Exchange service (IMEX), described below.

Beyond the rack, nothing changes from earlier generations. Data-parallel replicas in different racks communicate over the scale-out network, so the job designer's task is to keep the bandwidth-hungry traffic (tensor and expert parallelism) inside one NVLink domain and push only the gradient reductions across racks.

Advertisement

Two memory tiers, one address space

On an x86 server the GPU reaches host memory through PCIe, and in practice software treats host memory as a staging area: copy in, compute, copy out. NVLink-C2C changes two things. It is several times faster than a PCIe host link, and it is cache coherent, so the CPU and GPUs see a consistent view of shared data without explicit cache flushes. Pinned host allocations and CUDA managed memory both move over this link instead of PCIe; check NVIDIA's CUDA documentation for your driver version before relying on any other way of sharing memory.

Coherent does not mean free. A kernel that streams its working set from LPDDR5X runs at LPDDR5X speed. The right mental model is a two-tier memory: HBM for anything touched every step (weights in use, activations, the hot part of a KV cache), LPDDR5X for state that is large and touched rarely or sequentially (optimizer moments, cold KV blocks, staged data). Do not assume a particular NUMA layout from earlier Grace Hopper material; inspect it on your own system with numactl -H and nvidia-smi topo -m before pinning processes.

Measure the tiers before you design around them. This microbenchmark reports effective copy bandwidth for each path; run it once per node type and keep the numbers next to your capacity plan.

import time
import torch

def gbps(src, dst, iters=20):
    dst.copy_(src, non_blocking=True)          # warm up
    torch.cuda.synchronize()
    t0 = time.perf_counter()
    for _ in range(iters):
        dst.copy_(src, non_blocking=True)
    torch.cuda.synchronize()
    return iters * src.numel() * src.element_size() / (time.perf_counter() - t0) / 1e9

n = 4 * 2**30                                      # 4 GiB buffers
hbm_a = torch.empty(n, dtype=torch.uint8, device="cuda")
hbm_b = torch.empty_like(hbm_a)
pinned = torch.empty(n, dtype=torch.uint8, pin_memory=True)
pageable = torch.empty(n, dtype=torch.uint8)

print("HBM -> HBM        ", gbps(hbm_a, hbm_b))     # each byte is read and written
print("pinned host -> GPU", gbps(pinned, hbm_a))
print("GPU -> pinned host", gbps(hbm_a, pinned))
print("pageable -> GPU   ", gbps(pageable, hbm_a))

Expect pinned host copies to be far faster than pageable ones, and expect all host paths to be bounded by the lower of NVLink-C2C and LPDDR5X bandwidth. If the pinned number looks like a PCIe number, the process is probably bound to the wrong CPU or memory node.

Using the CPU tier: optimizer offload

The cleanest use of LPDDR5X in training is to hold the state that the optimizer needs once per step. Mixed-precision Adam keeps a 4-byte master weight and two 4-byte moments per parameter, 12 bytes that are read and written exactly once per step. Moving them off the GPU frees most of the HBM a training job uses for state, and the Grace cores can apply the update while the GPUs wait. The sketch below shows the mechanics; production code (DeepSpeed ZeRO-Offload and similar) overlaps the transfer with the backward pass and shards state across ranks.

import torch

class GraceOffloadAdam:
    """bf16 weights and grads stay in HBM; fp32 master copy and Adam moments live in LPDDR5X."""

    def __init__(self, params, lr=1e-4, betas=(0.9, 0.95), eps=1e-8):
        self.params = [p for p in params if p.requires_grad]
        self.master = [p.detach().float().cpu().pin_memory() for p in self.params]
        self.m = [torch.zeros_like(w) for w in self.master]
        self.v = [torch.zeros_like(w) for w in self.master]
        self.grad_buf = [torch.empty_like(w).pin_memory() for w in self.master]
        self.lr, self.b1, self.b2, self.eps, self.t = lr, *betas, eps, 0

    @torch.no_grad()
    def step(self):
        self.t += 1
        for p, g in zip(self.params, self.grad_buf):
            g.copy_(p.grad, non_blocking=True)          # GPU -> Grace over NVLink-C2C
        torch.cuda.synchronize()
        c1, c2 = 1 - self.b1 ** self.t, 1 - self.b2 ** self.t
        for p, w, m, v, g in zip(self.params, self.master, self.m, self.v, self.grad_buf):
            m.mul_(self.b1).add_(g, alpha=1 - self.b1)  # runs on the Arm cores
            v.mul_(self.b2).addcmul_(g, g, value=1 - self.b2)
            w.addcdiv_(m / c1, (v / c2).sqrt_().add_(self.eps), value=-self.lr)
            p.copy_(w, non_blocking=True)               # fp32 -> bf16 on the way back

The trade-off is step time. The update now costs at least the time to read and write 12 bytes per parameter at LPDDR5X speed, plus the gradient and weight transfers over C2C. If the step is long (large batch, long sequences) that cost hides behind compute; if it is short, it dominates. Offload is a capacity lever, not a speed lever.

The NVL72 domain: why 72 matters

Within an NVLink domain every GPU can reach every other at 1.8 TB/s through the switches, and collectives such as all-reduce and all-to-all run without touching the network. On an eight-GPU server, tensor parallelism and expert parallelism are confined to eight GPUs; anything wider crosses the network at a small fraction of that bandwidth. On NVL72 the same communication pattern can span 72 GPUs, which changes model-parallel layouts more than any single-GPU improvement does. The NVLink article explains the switch fabric, and NCCL collectives explains which algorithms the library picks.

Mixture-of-experts models benefit most. Each token is routed to a few experts, so expert parallelism produces an all-to-all exchange every MoE layer. That exchange is latency- and bandwidth-sensitive and scales badly over a network. Keeping all experts of a layer inside one NVLink domain makes the all-to-all a switch-local operation. Dense models benefit too, from wider tensor parallelism that keeps per-GPU memory low without leaving NVLink.

Worked example: sizing a 1.2T-parameter MoE on one rack

Take a mixture-of-experts model with 1.2 trillion total parameters. For inference with FP8 weights that is about 1.2 TB, against 13.4 TB of HBM in the rack: sharded across 72 GPUs, roughly 17 GB of weights per GPU, leaving most of each GPU's 186 GB for KV cache and activations. One rack serves the model with experts spread across all 72 GPUs and the all-to-all on NVLink.

Training is harder. A common mixed-precision budget is 16 bytes per parameter: 2 for bf16 weights, 2 for gradients, 12 for the fp32 master and Adam moments. That is 19.2 TB, more than the rack's HBM before a single activation is stored. Two options follow. Shard the state across several racks with data parallelism, which adds cross-rack traffic, or keep the 4 bytes of weights and gradients in HBM (4.8 TB) and put the 12 bytes of optimizer state, 14.4 TB, in LPDDR5X, which fits within the rack's 17 TB.

The offload option has a visible price. Reading and writing 14.4 TB at the rack's 14 TB/s aggregate LPDDR5X bandwidth costs about two seconds per step for the update alone, before compute. If your step takes thirty seconds, that is acceptable; if it takes three, it is not, and multi-rack sharding wins. This is the kind of arithmetic to do before buying, because it decides whether the CPU tier is useful for your workload.

Scheduling: IMEX, ComputeDomains and cliques

Because each tray runs its own OS, sharing GPU memory across trays needs a broker. IMEX is driver-level software that exports and imports GPU memory between nodes in the same NVLink domain, with access control per IMEX domain and channel. Without it, peer access across trays fails and libraries fall back to the network, often silently.

On Kubernetes, NVIDIA's DRA driver for GPUs (NVIDIA's announcement used version 25.8.0, requiring Kubernetes 1.32 or later with dynamic resource allocation enabled) exposes this as a ComputeDomain custom resource. Pods that claim the domain's channel are placed in a shared IMEX domain, and the nvidia.com/gpu.clique node label identifies which nodes share an NVLink partition, so affinity on that label keeps a job inside one rack.

apiVersion: resource.nvidia.com/v1beta1
kind: ComputeDomain
metadata:
  name: train-cd
spec:
  numNodes: 0
  channel:
    resourceClaimTemplate:
      name: train-cd-channel
---
apiVersion: v1
kind: Pod
metadata:
  name: trainer-0
  labels: {job: moe-train}
spec:
  affinity:
    podAffinity:
      requiredDuringSchedulingIgnoredDuringExecution:
      - labelSelector: {matchLabels: {job: moe-train}}
        topologyKey: nvidia.com/gpu.clique       # stay inside one NVLink partition
  resourceClaims:
  - name: imex
    resourceClaimTemplateName: train-cd-channel
  containers:
  - name: trainer
    image: registry.example.com/trainer:arm64   # Grace is aarch64
    resources:
      claims: [{name: imex}]
      limits: {nvidia.com/gpu: 4}

The same idea applies to Slurm or any other scheduler: treat the NVLink domain as a placement unit, and never let a tensor- or expert-parallel group straddle two cliques.

Failure modes

  • Wrong architecture images. Grace is aarch64. x86 container images, pip wheels without arm64 builds and custom CUDA extensions compiled for x86 fail at start-up. Build multi-architecture images and test them before the hardware arrives.
  • Silent network fallback. If IMEX is not running or the pod lacks the channel claim, collectives still complete, over the scale-out network, at a fraction of the speed. Run with NCCL_DEBUG=INFO at job start and assert in code that the reported transport is NVLink-based for intra-rack groups.
  • A job sized to exactly 72 GPUs. One failed GPU or tray shrinks the domain, and a job that needs all 72 cannot restart until repair. Design layouts that fit in fewer GPUs, for example 64, and keep the rest as hot spares.
  • Assumed memory placement. Copying a Grace Hopper NUMA recipe can bind a process to the far memory node and reduce host bandwidth. Measure with the benchmark above.
  • Thermal and power events. The rack is liquid cooled; coolant and power faults throttle or drop whole trays. Export clock, power and throttle-reason telemetry per GPU and alert on sustained clock drops, not only on errors.

GB200 NVL72 or eight-GPU servers?

WorkloadBetter fitWhy
Large MoE training and servingNVL72All-to-all stays inside NVLink
Dense models that fit in 8 GPUs with TP=88-GPU serversNo benefit from the wider domain, simpler operations
Long-context inference with huge KV cachesNVL72HBM plus a fast CPU tier for cold blocks
Many small independent jobs8-GPU serversRack-scale domains are hard to share finely
x86-only software stack8-GPU x86 serversGrace requires arm64 builds

What to do next

  1. Port your containers to arm64 and run your full test suite on Grace before capacity arrives.
  2. Run the bandwidth benchmark on one node and record HBM, pinned host and pageable numbers.
  3. Write the memory budget for your model: bytes per parameter for weights, gradients and optimizer state, KV cache per token, and which tier each lives in.
  4. Pick a parallel layout that keeps tensor and expert groups inside one clique, and leaves spare GPUs for failures.
  5. Deploy the DRA driver, define ComputeDomains, and add clique affinity to every multi-node job.
  6. Gate job start on an NCCL transport check, and alert on clock throttling per GPU.
Key takeaway: GB200 pairs a Grace CPU with two Blackwell GPUs over coherent 900 GB/s NVLink-C2C and builds racks with a 72-GPU NVLink domain. Treat it as a machine with two memory tiers (186 GB of HBM per GPU, plus LPDDR5X that is large but about thirty times slower) and one very wide scale-up domain. Keep MoE all-to-all and tensor parallelism inside a clique, use the CPU tier for state touched once per step, build arm64 images, and verify with NCCL logs that traffic really is on NVLink.