With eight-GPU servers, the unit you buy, install, monitor and repair is a server. With the GB200 NVL72, it is a rack: 72 Blackwell GPUs and 36 Grace CPUs in 18 compute trays, joined by 9 NVLink switch trays into one NVLink domain, cooled by liquid and fed on the order of a hundred kilowatts. Software sees one 72-GPU scale-up domain, which is the whole point. Operations sees one large, tightly coupled machine whose parts fail independently but whose jobs fail together.
The GB200 article covers the superchip, the memory tiers and how frameworks map models onto the domain, and the NVLink Switch article covers the fabric. This article is about the rack as a physical and failure unit: what is in it, what it asks of the facility, how to accept one, what breaks, and how to size jobs and checkpoints so that a broken tray costs minutes instead of days.
What is in the rack
| Part | Count | What it means for software |
|---|---|---|
| Compute tray | 18 | Two GB200 superchips: 4 GPUs, 2 Grace CPUs, one OS image, its own scale-out NICs |
| NVLink switch tray | 9 | Two NVLink Switch chips each, 18 chips in all; no OS for jobs, managed by the fabric software |
| NVLink spine | 1 | Copper cable cartridges joining every GPU to every switch chip |
| Liquid cooling | rack loop | CPUs, GPUs and switches are liquid cooled through a CDU; some components remain air cooled |
| Rack total | 72 GPUs | One NVLink domain, 1.8 TB/s per GPU, about 130 TB/s aggregate, about 13.4 TB HBM3E |
The tray is the field-replaceable unit that matters most. Each compute tray runs its own Linux kernel (see the compute tray as a Linux host), so a 72-GPU job is always a multi-node job from the operating system's point of view, even though every GPU can load and store into every other GPU's memory over NVLink.
Link arithmetic: 18 planes
The link arithmetic explains both the bandwidth and the failure behaviour. Each Blackwell GPU has 18 fifth-generation NVLink links at 100 GB/s, which is the 1.8 TB/s per GPU in the specifications. Each switch tray exposes 144 NVLink ports. The rack therefore has 72 x 18 = 1,296 GPU links and 9 x 144 = 1,296 switch ports: every link lands on a switch, with nothing spare and nothing oversubscribed. Multiply 72 GPUs by 1.8 TB/s and you get the quoted 130 TB/s figure.
The wiring pattern matters more than the totals. Each GPU sends one link to each of the 18 switch chips, so the rack is 18 parallel planes and any GPU reaches any other in one switch hop on every plane. Topologically, losing one switch chip removes one eighteenth of every GPU's bandwidth rather than cutting anyone off, and losing a switch tray removes two planes. Whether the fabric keeps running degraded or the domain must be drained depends on the fabric management software and its release; treat it as a rack-level event until your vendor's runbook says otherwise.
NVIDIA states that the spine uses copper rather than optics, reported as roughly two miles of cable, and credits that choice with saving on the order of 20 kW per rack. The trade is reach: copper at these speeds only spans a rack, which is why the NVLink domain stops at the rack and everything beyond it travels over InfiniBand or Ethernet at a fraction of the bandwidth.
Power, cooling and what software feels
Reported figures put a DGX GB200 NVL72 at about 120 kW and about 1.36 tonnes. Exact numbers vary by build and workload, so plan from your vendor's site requirements, but the order of magnitude is the point: roughly ten times an air-cooled rack of a few years ago, delivered through busbars and power shelves rather than PDUs, and removed by liquid. The liquid cooling and power density articles cover the facility side; here are the parts that reach software.
- Thermals become clocks. A warm coolant supply or a partly blocked cold plate shows up as one tray clocking lower, which in synchronous training makes every rank wait for it. Export per-GPU clocks, power and throttle reasons, not only temperatures.
- Power swings are synchronous. A training step alternates compute-heavy and communication-heavy phases, and every GPU in the job does it at the same moment, so a large job's draw swings in lockstep. Facilities and grid operators care about that; power caps and smoothing features exist for it, and you should know which are enabled.
- Power caps are a tuning knob. A modest per-GPU cap often costs little throughput and buys headroom when a site is power-limited. Measure tokens per second per kilowatt, not just tokens per second.
Accepting a rack
Acceptance for a rack is a ladder: each rung must pass before the next one means anything. The facility rungs (pressure and leak tests on the cooling loop, power-on under load) belong to the integrator and site team. The software rungs belong to you:
# 1. uniform software: same driver, firmware and fabric software on every tray
pdsh -w tray[01-18] 'nvidia-smi --query-gpu=driver_version,vbios_version --format=csv,noheader' | sort | uniq -c
# 2. every GPU sees all 18 links up (exact output text varies by driver release)
pdsh -w tray[01-18] 'nvidia-smi nvlink --status' > nvlink.txt
grep -ci inactive nvlink.txt # expect 0
# 3. per-tray health and stress
pdsh -w tray[01-18] 'dcgmi diag -r 3'
# 4. whole-domain collectives over NVLink, 72 ranks, 4 per tray
NCCL_DEBUG=INFO mpirun -np 72 -N 4 --hostfile trays.txt \
./build/all_reduce_perf -b 8 -e 16G -f 2 -g 1 | tee ar.txtStep 1 catches the most common silent problem, a tray that came back from repair with a different firmware or driver. Step 2 confirms the fabric trained. Step 3 finds weak GPUs and memory. Step 4 is the real test: NCCL must use NVLink across trays, which needs the IMEX service and correct container configuration; recent NCCL releases report multi-node NVLink in their INFO output. If it falls back to the scale-out network, the run still completes, just several times slower. Record the bus bandwidth curve of your first known-good rack and compare every later rack and every repaired tray against it; a uniform shortfall across all ranks usually points at configuration, a single slow rank points at hardware. The NCCL collectives article explains how to read those curves.
Failure domains
Think in failure domains: what is lost, how a running job notices, and what you do.
| Failure | Lost | Job impact | Response |
|---|---|---|---|
| GPU (uncorrectable memory error, fell off the bus) | 1 GPU, often its tray | Job crashes or hangs in a collective | Drain tray, restart from checkpoint elsewhere |
| Compute tray (board, NIC, OS) | 4 GPUs | Job crashes | Replace tray; re-run acceptance rungs 1 to 4 |
| Switch chip or tray | 1 or 2 of 18 planes | Slower collectives or a domain drain, by policy | Vendor runbook; often a rack-level window |
| Spine cable or connector | Some links on some GPUs | Link errors, retrains, slow ranks | Reseat or replace cartridge; recheck link status |
| Coolant or power fault | One or more trays | Throttling, then trays dropping | Facility response; drain affected trays |
The asymmetry is what drives planning: parts fail one at a time, but a synchronous job that spans the domain stops whenever any of its parts does.
Worked example: spare trays and checkpoint cadence
The cheapest reliability tool for a rack is to not need all of it. Suppose, purely as an illustrative assumption, that any given compute tray is unavailable 1 percent of the time (failed, being repaired or being re-qualified). The binomial distribution tells you how often a rack has enough healthy trays for a job of a given size, and Young and Daly's formula (interval = sqrt(2 x checkpoint cost x MTBF)) tells you how often to checkpoint.
import math
from math import comb
def p_short(trays=18, needed=16, p_down=0.01):
"""Chance that fewer than `needed` trays are healthy at a random moment."""
return sum(comb(trays, k) * p_down**k * (1 - p_down)**(trays - k)
for k in range(trays - needed + 1, trays + 1))
def young_daly(ckpt_minutes, mtbf_hours):
c = ckpt_minutes / 60
tau = math.sqrt(2 * c * mtbf_hours) # optimal checkpoint interval
return tau, c / tau + tau / (2 * mtbf_hours) # interval, fraction of time lost
for needed in (18, 17, 16):
print(needed, f"{p_short(needed=needed):.4%}")
for racks in (1, 8):
tau, waste = young_daly(3, 5 * 24 / racks)
print(racks, round(tau * 60), "min", f"{waste:.1%}")| Job needs | GPUs | Time the rack is short (1% per tray, illustrative) |
|---|---|---|
| 18 trays | 72 | 16.5% |
| 17 trays | 68 | 1.4% |
| 16 trays | 64 | 0.07% |
A job sized for all 72 GPUs waits for a repair about one moment in six under this assumption. Size it for 64 and keep two trays as warm spares, and the rack is almost always able to restart it immediately on the remaining healthy trays. Sixty-four also divides cleanly into the power-of-two expert, tensor and pipeline groups that frameworks prefer; 72 often does not. The spare trays need not idle: give them preemptible work such as evaluation or small experiments.
Checkpoint cadence follows the same logic. Assume, again as an illustration, one job-interrupting event per rack every five days and a three-minute checkpoint. A one-rack job then has a 120-hour MTBF, an optimal interval of about 208 minutes and about 2.9 percent of time lost. Spread the job over eight racks and the MTBF falls to 15 hours, the interval to about 73 minutes and the loss to about 8.2 percent, before counting restart time. Faster checkpoints (asynchronous writes to local or Grace memory, then upload) pay off directly; plug in your own measured failure rates.
Placing jobs on racks
Scheduling follows from the topology. Keep every tensor, expert or pipeline group that exchanges large activations inside one rack, because crossing racks drops from NVLink to the scale-out network. Put data parallelism, whose traffic is one gradient reduction per step and overlaps with compute, across racks.
def place(job, racks):
"""Model-parallel group inside one NVLink domain; data parallel across racks."""
group = job.model_parallel_size # e.g. 64 GPUs = 16 trays
usable = []
for rack in racks:
healthy = [t for t in rack.trays if t.healthy and t.accepted]
n = (len(healthy) * 4) // group # replicas this rack can host
usable += [(rack, healthy[i*group//4:(i+1)*group//4]) for i in range(n)]
if len(usable) < job.min_replicas:
return None # wait, or shrink data parallel
return usable[: job.max_replicas]Two rules make this robust. A tray returns to the pool only after it passes acceptance again, which is what the accepted flag is for. And jobs should restart elastically: if a rack drops below one replica's worth of healthy trays, the job resumes with one fewer data parallel replica (adjusting gradient accumulation to keep the global batch) rather than waiting. In Kubernetes the IMEX domain is expressed with ComputeDomains, covered in the GB200 article; in Slurm, the topology plugin can describe each rack as a block so allocations do not straddle racks.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Use all 72 GPUs per job | Maximum per-replica memory and bandwidth | Waits for every repair; awkward group sizes |
| Use 64 plus 2 spare trays | Near-immediate restarts, power-of-two groups | About 11% of GPUs on preemptible work |
| Model-parallel groups larger than a rack | Fits bigger models per replica | Activations cross the slow network; every rack failure hits every replica |
| Frequent checkpoints | Less lost work | Storage bandwidth and step stalls unless writes are asynchronous |
What to do next
- Draw your rack's failure-domain table with your vendor's actual field-replaceable units.
- Record a baseline NCCL all-reduce curve on your first accepted rack, and make re-running it a condition for returning any repaired tray to service.
- Alert on per-GPU clock drops and throttle reasons, not only on errors.
- Measure your own tray unavailability and job-interrupt rate for a month, then redo the spare-tray and checkpoint calculations with real numbers.
- Size model-parallel groups to fit 16 trays and test an elastic restart that drops one data parallel replica.
- Make checkpoint writes asynchronous and measure their real cost per step.