Most writing about GB200 starts at the rack: 72 GPUs in one NVLink domain, liquid cooling, rack-scale power. That view is covered in the GB200 overview and the NVLink Switch guide. This page looks one level down, at the unit you actually log into: the compute tray. Each tray is an ordinary Linux host built from two GB200 superchips, and most first-month problems on GB200 clusters are host problems: an x86-only dependency, a library that breaks on 64K pages, ranks running on the wrong CPU, a data loader allocating memory on the wrong NUMA node, or one weak tray that slows every job it joins.

The goal is a tray you can trust. The page covers what the OS sees, porting a training stack to Arm, page size, NUMA and CPU affinity, host-to-GPU transfer paths, and an acceptance test that finds weak trays before users do, with code for each step and a worked debugging example.

What the operating system sees

One GB200 compute tray as the operating system sees it: two superchips, four GPUsGB200 superchip 0Grace CPU72 Arm coresLPDDR5XGPU 0HBM3eGPU 1HBM3eC2CC2CGB200 superchip 1Grace CPU72 Arm coresLPDDR5XGPU 2HBM3eGPU 3HBM3eC2CC2Cscale-out NICsone per GPU is typical; check your OEMNVLink to the switch traysthe 72-GPU domainLinux sees one host: CPUs and memory per Grace (NUMA nodes) plus four CUDA devices.Each rank should run on the cores of the Grace wired to its GPU, using that Grace's memory.
Figure: a compute tray. Each Grace CPU is wired to two Blackwell GPUs by NVLink-C2C and owns its own LPDDR5X; GPUs reach the rest of the rack over NVLink.

A GB200 superchip pairs one Grace CPU with two Blackwell GPUs. Grace has 72 Arm Neoverse V2 cores and up to 480 GB of LPDDR5X. In the NVL72 configuration each GPU has 186 GB of usable HBM3e at about 8 TB/s. The CPU and GPUs are joined by NVLink-C2C, which NVIDIA quotes at 900 GB/s bidirectional and which is cache coherent: the GPU can read and write CPU memory without the explicit staging a PCIe system needs.

A compute tray holds two superchips, so the OS sees two Graces and four GPUs. An NVL72 rack has 18 such trays plus nine NVLink switch trays, giving each GPU 1.8 TB/s of NVLink bandwidth to the other 71. The rack is fast, but it is not one giant GPU: software sees 72 CUDA devices spread across 18 operating system instances, and the usual distributed training machinery (one process per GPU, NCCL collectives, a launcher per host) still applies.

The practical model for the host: two halves. Each half is a Grace, its memory, and two GPUs. A process driving GPU 0 should run on Grace 0's cores and allocate host memory there; a process driving GPU 3 belongs on Grace 1. Getting this wrong does not crash anything. It makes some ranks slower, and the whole job runs at the speed of its slowest rank.

Porting the stack to Arm

Grace is aarch64. Your CUDA kernels need no change for that, but everything around them does: Python wheels, containers, native extensions and the odd hand-written SIMD routine. Check the following before the hardware arrives, on any Arm machine or in an emulated build.

  • Containers. Build multi-architecture images, for example with docker buildx build --platform linux/amd64,linux/arm64. NVIDIA's NGC framework containers publish Arm builds; start from them rather than from an x86 base.
  • Wheels. Every dependency needs an aarch64 wheel (tags such as manylinux_2_28_aarch64) or must build from source. A package silently compiling from source during install is a sign of trouble; pin it and cache the built wheel.
  • Native code. x86 intrinsics from immintrin.h do not compile. Use portable code, a translation layer such as sse2neon, or Arm's own SIMD (Neon, and SVE2 on Neoverse V2). Remove -march=native from build scripts that run on a different machine.
  • Semantics. Plain char is unsigned on Arm Linux and signed on x86. Arm's memory model is weaker than x86's, so lock-free code that relied on x86 ordering without proper atomics can fail rarely and only under load.
# Smoke test to run inside the training image on the target host
uname -m                          # expect: aarch64
getconf PAGESIZE                  # 4096 or 65536, see the next section
python - <<'EOF'
import platform, torch
print(platform.machine(), torch.__version__, torch.version.cuda)
print(torch.cuda.device_count(), [torch.cuda.get_device_name(i) for i in range(torch.cuda.device_count())])
x = torch.randn(4096, 4096, device="cuda", dtype=torch.bfloat16)
print((x @ x).float().abs().mean().item())     # forces a real kernel launch
EOF

Page size: 4 KB or 64 KB

Arm Linux kernels can be built with 4 KB or 64 KB base pages. NVIDIA's Grace tuning guidance recommends 64 KB as the default: large allocations take fewer page faults and the TLB covers more memory, which helps the large host buffers training uses. Distributions differ in which kernel they ship, so check getconf PAGESIZE on your image rather than assuming.

The risk is software that assumes 4 KB. Typical symptoms on a 64 KB kernel:

  • A bundled jemalloc built for 4 KB pages aborts at start with a message about an unsupported page size. Rebuild it with a larger page size setting or use the system allocator.
  • Code that rounds mmap offsets or lengths to 4096 gets EINVAL or silently wastes memory.
  • Memory use per process rises for programs that allocate many small, separately mapped regions.

NUMA and CPU affinity

Start by asking the host what it is, rather than hard-coding node numbers. Two commands tell you almost everything:

numactl -H            # NUMA nodes, their CPUs and memory sizes
nvidia-smi topo -m    # GPU and NIC matrix, with CPU Affinity and NUMA Affinity per GPU

On a tray you should see two nodes with CPUs, one per Grace. On coherent Grace platforms the driver can also expose GPU memory as extra NUMA nodes that have memory but no CPUs; how many appear depends on the driver and its configuration, so read the output instead of assuming a layout. The CPU Affinity column of the topology matrix says which cores belong with each GPU.

Then bind each rank before it allocates anything. The script below maps the local GPU to its PCI device and reads that device's CPU list from sysfs. If the platform reports every CPU for every GPU, fall back to the topology matrix.

import os, subprocess

def gpu_cpus(local_rank):
    out = subprocess.check_output(
        ["nvidia-smi", "--query-gpu=index,pci.bus_id", "--format=csv,noheader"], text=True)
    for line in out.strip().splitlines():
        idx, bus = [f.strip() for f in line.split(",")]
        if int(idx) == local_rank:
            dom, rest = bus.lower().split(":", 1)            # 00000009:01:00.0
            path = f"/sys/bus/pci/devices/{int(dom, 16):04x}:{rest}/local_cpulist"
            return parse_cpulist(open(path).read())
    raise RuntimeError(f"GPU {local_rank} not found")

def parse_cpulist(s):                                      # "0-71" or "72-143"
    cpus = set()
    for part in s.strip().split(","):
        lo, _, hi = part.partition("-")
        cpus.update(range(int(lo), int(hi or lo) + 1))
    return cpus

local_rank = int(os.environ["LOCAL_RANK"])
cpus = gpu_cpus(local_rank)
if len(cpus) < os.cpu_count():
    os.sched_setaffinity(0, cpus)        # children (loader workers) inherit this
else:
    raise RuntimeError("sysfs reports every CPU; take affinity from nvidia-smi topo -m")
workers = max(2, len(os.sched_getaffinity(0)) // 2 - 4)   # two ranks per Grace; spare OS cores

Thread affinity also steers memory: Linux places a page on the node of the CPU that first touches it, so binding before allocation keeps pinned buffers local. For stricter control, launch each rank under numactl --cpunodebind=N --membind=N with N read from the topology output. Bind network traffic the same way: use the NIC that the topology matrix shows closest to each GPU.

Host-to-GPU transfer paths

On a PCIe server, host-to-GPU copies go through pinned staging buffers over a PCIe link. On GB200 they cross NVLink-C2C, which changes three things for software:

  • Bandwidth. Host transfers are far faster than over PCIe, but two GPUs share one Grace, and the Grace's own LPDDR5X bandwidth is a ceiling too. Measure what one GPU gets alone and what each gets when both copy at once.
  • Pageable memory. With coherence, the GPU can access ordinary pageable host memory. CUDA reports this through the device attribute cudaDevAttrPageableMemoryAccess. It is convenient, but pinned buffers and explicit async copies are still the predictable path for data loading.
  • Offload economics. Optimizer state or KV cache parked in LPDDR5X costs much less step time than over PCIe, which makes CPU offload of optimizer state more attractive on GB200 than on PCIe servers.

A quick probe from PyTorch, before reaching for a full tool such as NVIDIA's nvbandwidth:

import torch

def h2d_gbps(nbytes=1 << 30, pinned=True, iters=20):
    host = torch.empty(nbytes, dtype=torch.uint8, pin_memory=pinned)
    dev = torch.empty(nbytes, dtype=torch.uint8, device="cuda")
    dev.copy_(host, non_blocking=pinned)
    torch.cuda.synchronize()
    t0, t1 = torch.cuda.Event(enable_timing=True), torch.cuda.Event(enable_timing=True)
    t0.record()
    for _ in range(iters):
        dev.copy_(host, non_blocking=pinned)
    t1.record()
    torch.cuda.synchronize()
    return nbytes * iters / (t0.elapsed_time(t1) / 1e3) / 1e9

print(f"pinned   {h2d_gbps(pinned=True):7.1f} GB/s")
print(f"pageable {h2d_gbps(pinned=False):7.1f} GB/s")

Run it on every GPU of a tray, alone and in pairs, with and without affinity binding. The absolute numbers depend on driver and firmware; what you are looking for is the gap between ranks.

Tray acceptance testing

A rack is only as fast as its weakest tray, so test trays before they join the scheduler and again after any repair. A practical acceptance sequence:

  1. Inventory. nvidia-smi -q shows four GPUs, the expected driver, no retired-page or remapping alerts; nvidia-smi topo -m matches the reference tray.
  2. Diagnostics. dcgmi diag -r 3 runs NVIDIA DCGM's longer diagnostic set; any failure blocks the tray.
  3. Host paths. The transfer probe above, or nvbandwidth, per GPU.
  4. Fabric. nccl-tests all_reduce_perf -b 8 -e 8G -f 2 -g 1 with one process per GPU, first within the tray, then across trays in the NVLink domain.
  5. Burn-in. A real training job for several hours, watching for Xid errors in the kernel log and for thermal throttling.

Judge results against the fleet median, not the datasheet: datasheet figures are peaks, while a tray 10% below its identical neighbours is a real defect.

import json, statistics

def outliers(results, floor=0.90):
    """results: {tray: {metric: value}} where higher is better."""
    bad = {}
    metrics = {m for r in results.values() for m in r}
    for m in metrics:
        vals = [r[m] for r in results.values() if m in r]
        med = statistics.median(vals)
        for tray, r in results.items():
            if m in r and r[m] < floor * med:
                bad.setdefault(tray, []).append(f"{m}={r[m]:.1f} vs median {med:.1f}")
    return bad

print(json.dumps(outliers(json.load(open("acceptance.json"))), indent=2))

Worked example: two slow ranks per tray

An illustrative case. A team moves a 4-GPU-per-node training job to GB200 trays. Throughput is below plan, and per-rank timing shows local ranks 2 and 3 on every tray spending noticeably longer in the data loader than ranks 0 and 1.

  1. Topology. nvidia-smi topo -m shows GPUs 2 and 3 attached to the second Grace.
  2. Placement. Checking taskset -cp on the rank processes shows all four ranks and their loader workers allowed on every core, and the scheduler has packed most of them onto the first Grace.
  3. Memory. numastat -p for ranks 2 and 3 shows most of their pinned buffers on the first Grace's node, so every batch they copy comes from memory that is remote to their GPUs.
  4. Fix. The affinity wrapper above runs at the top of the training script, before the dataset and loader are created, and loader workers are sized per Grace.
  5. Verify. Per-rank loader times converge, and the step time now tracks the fastest rank.

Nothing in this failure produced an error. The only signals were a per-rank timing spread and a topology check, which is why both belong in the acceptance test and in job dashboards.

Failure modes

  • x86 leftovers. A dependency without an aarch64 wheel builds from source at install with different flags or fails only in production images.
  • Page-size assumptions. Allocators or I/O code written for 4 KB pages abort or waste memory.
  • Unbound ranks. Ranks and loader workers float across both Graces; remote memory slows half the GPUs.
  • Over-subscribed loaders. Two ranks per Grace each sized as if they owned all 72 cores.
  • Datasheet acceptance. Passing trays that meet a loose absolute bar but trail their peers.
  • Treating the rack as one device. Code that assumes a single process can drive all 72 GPUs without the multi-node setup the NVLink domain needs.

Trade-offs

ChoiceGainCost
64 KB pagesFewer faults and TLB misses on big buffersBreaks 4 KB-assuming software
Strict NUMA bindingPredictable, local memory trafficLess flexibility for helper processes
Pageable memory accessSimpler code, no stagingLess predictable bandwidth
CPU offload to LPDDR5XFrees HBM for activations or KV cacheShares C2C and Grace bandwidth
Fleet-median acceptanceCatches weak traysNeeds a baseline and repeat runs

What to do next

  1. Build your training image for linux/arm64 and run the smoke test on any Arm host now.
  2. Check getconf PAGESIZE on the target image and test your allocator and I/O code on it.
  3. Record numactl -H and nvidia-smi topo -m for a reference tray and diff every new tray against it.
  4. Add the affinity wrapper to your launcher and size loader workers per Grace.
  5. Run the transfer probe per GPU and alone versus paired, and keep the results as a baseline.
  6. Automate dcgmi diagnostics, nccl-tests and a burn-in job, and gate trays on the fleet median.
  7. Keep learning: the Blackwell GPU and its number formats and liquid cooling and its failure modes.
Key takeaway: A GB200 compute tray is a Linux host with two Grace CPUs, each wired to two GPUs by a coherent link. Most early problems are host problems: x86 dependencies, page-size assumptions, ranks and memory on the wrong Grace, and weak trays hiding in a big rack. Port and test on Arm early, read the topology instead of assuming it, bind every rank to its own Grace, and accept trays against the fleet median.