With eight-GPU servers, the unit you buy, install, monitor and repair is a server. With the GB200 NVL72, it is a rack: 72 Blackwell GPUs and 36 Grace CPUs in 18 compute trays, joined by 9 NVLink switch trays into one NVLink domain, cooled by liquid and fed on the order of a hundred kilowatts. Software sees one 72-GPU scale-up domain, which is the whole point. Operations sees one large, tightly coupled machine whose parts fail independently but whose jobs fail together.

The GB200 article covers the superchip, the memory tiers and how frameworks map models onto the domain, and the NVLink Switch article covers the fabric. This article is about the rack as a physical and failure unit: what is in it, what it asks of the facility, how to accept one, what breaks, and how to size jobs and checkpoints so that a broken tray costs minutes instead of days.

What is in the rack

GB200 NVL72 rack, front elevation (schematic, not to scale)power and management (placement varies by build)10 compute trays: each 2 Grace + 4 Blackwell, one OS9 NVLink switch trays: 2 switch chips, 144 ports each8 compute traysNVLink spinecopper cartridges, rearcoolant manifoldssupply and return to CDUEvery GPU has one NVLink to each of the 18 switch chips: 72 x 18 = 1,296 links = 9 trays x 144 ports
Figure 1. The NVL72 rack as a set of replaceable parts. Compute trays sit above and below the switch trays; the copper NVLink spine runs down the back; liquid cooling reaches every compute and switch tray.
PartCountWhat it means for software
Compute tray18Two GB200 superchips: 4 GPUs, 2 Grace CPUs, one OS image, its own scale-out NICs
NVLink switch tray9Two NVLink Switch chips each, 18 chips in all; no OS for jobs, managed by the fabric software
NVLink spine1Copper cable cartridges joining every GPU to every switch chip
Liquid coolingrack loopCPUs, GPUs and switches are liquid cooled through a CDU; some components remain air cooled
Rack total72 GPUsOne NVLink domain, 1.8 TB/s per GPU, about 130 TB/s aggregate, about 13.4 TB HBM3E

The tray is the field-replaceable unit that matters most. Each compute tray runs its own Linux kernel (see the compute tray as a Linux host), so a 72-GPU job is always a multi-node job from the operating system's point of view, even though every GPU can load and store into every other GPU's memory over NVLink.

Link arithmetic: 18 planes

The link arithmetic explains both the bandwidth and the failure behaviour. Each Blackwell GPU has 18 fifth-generation NVLink links at 100 GB/s, which is the 1.8 TB/s per GPU in the specifications. Each switch tray exposes 144 NVLink ports. The rack therefore has 72 x 18 = 1,296 GPU links and 9 x 144 = 1,296 switch ports: every link lands on a switch, with nothing spare and nothing oversubscribed. Multiply 72 GPUs by 1.8 TB/s and you get the quoted 130 TB/s figure.

The wiring pattern matters more than the totals. Each GPU sends one link to each of the 18 switch chips, so the rack is 18 parallel planes and any GPU reaches any other in one switch hop on every plane. Topologically, losing one switch chip removes one eighteenth of every GPU's bandwidth rather than cutting anyone off, and losing a switch tray removes two planes. Whether the fabric keeps running degraded or the domain must be drained depends on the fabric management software and its release; treat it as a rack-level event until your vendor's runbook says otherwise.

NVIDIA states that the spine uses copper rather than optics, reported as roughly two miles of cable, and credits that choice with saving on the order of 20 kW per rack. The trade is reach: copper at these speeds only spans a rack, which is why the NVLink domain stops at the rack and everything beyond it travels over InfiniBand or Ethernet at a fraction of the bandwidth.

Power, cooling and what software feels

Reported figures put a DGX GB200 NVL72 at about 120 kW and about 1.36 tonnes. Exact numbers vary by build and workload, so plan from your vendor's site requirements, but the order of magnitude is the point: roughly ten times an air-cooled rack of a few years ago, delivered through busbars and power shelves rather than PDUs, and removed by liquid. The liquid cooling and power density articles cover the facility side; here are the parts that reach software.

  • Thermals become clocks. A warm coolant supply or a partly blocked cold plate shows up as one tray clocking lower, which in synchronous training makes every rank wait for it. Export per-GPU clocks, power and throttle reasons, not only temperatures.
  • Power swings are synchronous. A training step alternates compute-heavy and communication-heavy phases, and every GPU in the job does it at the same moment, so a large job's draw swings in lockstep. Facilities and grid operators care about that; power caps and smoothing features exist for it, and you should know which are enabled.
  • Power caps are a tuning knob. A modest per-GPU cap often costs little throughput and buys headroom when a site is power-limited. Measure tokens per second per kilowatt, not just tokens per second.

Accepting a rack

Acceptance for a rack is a ladder: each rung must pass before the next one means anything. The facility rungs (pressure and leak tests on the cooling loop, power-on under load) belong to the integrator and site team. The software rungs belong to you:

# 1. uniform software: same driver, firmware and fabric software on every tray
pdsh -w tray[01-18] 'nvidia-smi --query-gpu=driver_version,vbios_version --format=csv,noheader' | sort | uniq -c

# 2. every GPU sees all 18 links up (exact output text varies by driver release)
pdsh -w tray[01-18] 'nvidia-smi nvlink --status' > nvlink.txt
grep -ci inactive nvlink.txt            # expect 0

# 3. per-tray health and stress
pdsh -w tray[01-18] 'dcgmi diag -r 3'

# 4. whole-domain collectives over NVLink, 72 ranks, 4 per tray
NCCL_DEBUG=INFO mpirun -np 72 -N 4 --hostfile trays.txt \
    ./build/all_reduce_perf -b 8 -e 16G -f 2 -g 1 | tee ar.txt

Step 1 catches the most common silent problem, a tray that came back from repair with a different firmware or driver. Step 2 confirms the fabric trained. Step 3 finds weak GPUs and memory. Step 4 is the real test: NCCL must use NVLink across trays, which needs the IMEX service and correct container configuration; recent NCCL releases report multi-node NVLink in their INFO output. If it falls back to the scale-out network, the run still completes, just several times slower. Record the bus bandwidth curve of your first known-good rack and compare every later rack and every repaired tray against it; a uniform shortfall across all ranks usually points at configuration, a single slow rank points at hardware. The NCCL collectives article explains how to read those curves.

Failure domains

Think in failure domains: what is lost, how a running job notices, and what you do.

FailureLostJob impactResponse
GPU (uncorrectable memory error, fell off the bus)1 GPU, often its trayJob crashes or hangs in a collectiveDrain tray, restart from checkpoint elsewhere
Compute tray (board, NIC, OS)4 GPUsJob crashesReplace tray; re-run acceptance rungs 1 to 4
Switch chip or tray1 or 2 of 18 planesSlower collectives or a domain drain, by policyVendor runbook; often a rack-level window
Spine cable or connectorSome links on some GPUsLink errors, retrains, slow ranksReseat or replace cartridge; recheck link status
Coolant or power faultOne or more traysThrottling, then trays droppingFacility response; drain affected trays

The asymmetry is what drives planning: parts fail one at a time, but a synchronous job that spans the domain stops whenever any of its parts does.

Worked example: spare trays and checkpoint cadence

The cheapest reliability tool for a rack is to not need all of it. Suppose, purely as an illustrative assumption, that any given compute tray is unavailable 1 percent of the time (failed, being repaired or being re-qualified). The binomial distribution tells you how often a rack has enough healthy trays for a job of a given size, and Young and Daly's formula (interval = sqrt(2 x checkpoint cost x MTBF)) tells you how often to checkpoint.

import math
from math import comb

def p_short(trays=18, needed=16, p_down=0.01):
    """Chance that fewer than `needed` trays are healthy at a random moment."""
    return sum(comb(trays, k) * p_down**k * (1 - p_down)**(trays - k)
               for k in range(trays - needed + 1, trays + 1))

def young_daly(ckpt_minutes, mtbf_hours):
    c = ckpt_minutes / 60
    tau = math.sqrt(2 * c * mtbf_hours)           # optimal checkpoint interval
    return tau, c / tau + tau / (2 * mtbf_hours)  # interval, fraction of time lost

for needed in (18, 17, 16):
    print(needed, f"{p_short(needed=needed):.4%}")
for racks in (1, 8):
    tau, waste = young_daly(3, 5 * 24 / racks)
    print(racks, round(tau * 60), "min", f"{waste:.1%}")
Job needsGPUsTime the rack is short (1% per tray, illustrative)
18 trays7216.5%
17 trays681.4%
16 trays640.07%

A job sized for all 72 GPUs waits for a repair about one moment in six under this assumption. Size it for 64 and keep two trays as warm spares, and the rack is almost always able to restart it immediately on the remaining healthy trays. Sixty-four also divides cleanly into the power-of-two expert, tensor and pipeline groups that frameworks prefer; 72 often does not. The spare trays need not idle: give them preemptible work such as evaluation or small experiments.

Checkpoint cadence follows the same logic. Assume, again as an illustration, one job-interrupting event per rack every five days and a three-minute checkpoint. A one-rack job then has a 120-hour MTBF, an optimal interval of about 208 minutes and about 2.9 percent of time lost. Spread the job over eight racks and the MTBF falls to 15 hours, the interval to about 73 minutes and the loss to about 8.2 percent, before counting restart time. Faster checkpoints (asynchronous writes to local or Grace memory, then upload) pay off directly; plug in your own measured failure rates.

Placing jobs on racks

Scheduling follows from the topology. Keep every tensor, expert or pipeline group that exchanges large activations inside one rack, because crossing racks drops from NVLink to the scale-out network. Put data parallelism, whose traffic is one gradient reduction per step and overlaps with compute, across racks.

def place(job, racks):
    """Model-parallel group inside one NVLink domain; data parallel across racks."""
    group = job.model_parallel_size                  # e.g. 64 GPUs = 16 trays
    usable = []
    for rack in racks:
        healthy = [t for t in rack.trays if t.healthy and t.accepted]
        n = (len(healthy) * 4) // group              # replicas this rack can host
        usable += [(rack, healthy[i*group//4:(i+1)*group//4]) for i in range(n)]
    if len(usable) < job.min_replicas:
        return None                                  # wait, or shrink data parallel
    return usable[: job.max_replicas]

Two rules make this robust. A tray returns to the pool only after it passes acceptance again, which is what the accepted flag is for. And jobs should restart elastically: if a rack drops below one replica's worth of healthy trays, the job resumes with one fewer data parallel replica (adjusting gradient accumulation to keep the global batch) rather than waiting. In Kubernetes the IMEX domain is expressed with ComputeDomains, covered in the GB200 article; in Slurm, the topology plugin can describe each rack as a block so allocations do not straddle racks.

Trade-offs

ChoiceGainCost
Use all 72 GPUs per jobMaximum per-replica memory and bandwidthWaits for every repair; awkward group sizes
Use 64 plus 2 spare traysNear-immediate restarts, power-of-two groupsAbout 11% of GPUs on preemptible work
Model-parallel groups larger than a rackFits bigger models per replicaActivations cross the slow network; every rack failure hits every replica
Frequent checkpointsLess lost workStorage bandwidth and step stalls unless writes are asynchronous

What to do next

  1. Draw your rack's failure-domain table with your vendor's actual field-replaceable units.
  2. Record a baseline NCCL all-reduce curve on your first accepted rack, and make re-running it a condition for returning any repaired tray to service.
  3. Alert on per-GPU clock drops and throttle reasons, not only on errors.
  4. Measure your own tray unavailability and job-interrupt rate for a month, then redo the spare-tray and checkpoint calculations with real numbers.
  5. Size model-parallel groups to fit 16 trays and test an elastic restart that drops one data parallel replica.
  6. Make checkpoint writes asynchronous and measure their real cost per step.
Key takeaway: An NVL72 rack is one 72-GPU NVLink domain built from 27 trays, a copper spine and a liquid loop, and it fails one tray at a time while synchronous jobs stop as a whole. Accept racks with a repeatable ladder ending in a 72-rank NCCL run, keep model-parallel groups inside a rack, size jobs to leave spare trays, and set checkpoint intervals from measured failure rates.