An NVLink Switch turns NVLink from a point-to-point cable between a few GPUs into a switched fabric. Inside one HGX baseboard, NVSwitch chips have long let eight GPUs talk to each other at full NVLink bandwidth. The NVLink Switch takes the same idea out of the server: switch chips live in dedicated trays, and GPUs in different compute trays reach each other through them with NVLink semantics, so a whole rack behaves as one scale-up domain. In a GB200 NVL72 that domain is 72 Blackwell GPUs.
This article explains the NVLink Switch from the software side: what the topology looks like, why one hop through a switch is different from a ring of point-to-point links, what the in-switch reduction engine (NVLink SHARP, which NCCL exposes as the NVLS algorithm) does to collective traffic, which daemons and libraries have to be running for any of it to work, and how a training job should be shaped to use a 72-GPU domain rather than treating it as nine eight-GPU servers. Figures are taken from NVIDIA's published GB200 NVL72 material; where a detail is not public it is left out rather than guessed. For the link layer itself, read NVLink as a memory fabric first.
Why a switched scale-up domain matters
Training parallelism has two kinds of traffic. Data-parallel gradient reduction moves a lot of bytes, but once per step and overlappable with the backward pass, so it tolerates a slower network. Tensor parallelism and expert parallelism move activations inside every layer, on the critical path, many times per step. Those collectives are why the scale-up domain matters: tensor parallelism is normally confined to GPUs joined by NVLink, because pushing per-layer all-reduces over InfiniBand or Ethernet at a fraction of the bandwidth stalls the matrix units.
With eight GPUs per NVLink domain, the tensor-parallel degree was effectively capped at eight, and mixture-of-experts all-to-all traffic beyond eight GPUs fell onto the scale-out network. The NVLink Switch raises the cap. That is the whole value proposition: not a faster link, but a bigger set of GPUs that all see each other at NVLink speed, so model-parallel groups can be larger before they touch the slower fabric. The tensor parallelism guide covers why those groups are so sensitive to bandwidth.
Topology: 72 GPUs, 18 switch chips, one hop
In GB200 NVL72, each Blackwell GPU has 18 fifth-generation NVLink links at 100 GB/s each (bidirectional), which is where the 1.8 TB/s per-GPU figure comes from: 900 GB/s in each direction. The rack holds nine NVLink Switch trays with two switch chips each, 18 chips in total. Every GPU sends exactly one of its 18 links to each of the 18 chips. The arithmetic closes: 72 GPUs times 18 links is 1,296 GPU-side ports, and 18 chips times 72 ports is also 1,296. NVIDIA quotes about 130 TB/s of aggregate NVLink bandwidth for the rack, which is 72 times 1.8 TB/s.
Three properties follow from this wiring, and they drive everything later in the article. First, any GPU reaches any other GPU in one switch hop, so there is no notion of near and far neighbours inside the domain. Second, a single transfer between two GPUs is striped across all 18 chips, so point-to-point bandwidth between any pair can approach the full 900 GB/s per direction when nothing else competes. Third, the switch layer is a set of 18 parallel planes: losing one chip removes one plane, and every GPU degrades by one eighteenth rather than one GPU being cut off.
The physical layer between trays is a copper cable cartridge in the back of the rack, not optics. That keeps power and cost down and is why the domain stops at a rack: copper reach is short. Scaling beyond the rack goes back to InfiniBand or Ethernet, covered in the InfiniBand article.
In-switch reduction: NVLink SHARP and NVLS
A plain switch only forwards. NVLink switches since the third NVSwitch generation also contain an arithmetic engine that can perform reductions in the fabric, branded NVLink SHARP. The idea mirrors SHARP in InfiniBand switches: instead of GPUs exchanging partial sums with each other, they send their contribution into the switch once, the switch adds the contributions, and the result is multicast back to every participant.
Count the bytes to see why this matters. A ring all-reduce of S bytes across n GPUs makes each GPU send and receive 2(n-1)/n times S. With in-switch reduction each GPU sends roughly S into the fabric and receives roughly S back. For large n the ring moves nearly 2S per GPU per direction, the switched reduction about S, so the bandwidth demand on every GPU's links halves. NVLink SHARP also takes the reduction arithmetic off the GPU's SMs, which matters when the collective overlaps compute: fewer SMs are borrowed by communication kernels.
NCCL exposes this as the NVLS algorithm. When the hardware and drivers support it, NCCL can pick NVLS for all-reduce on its own; NCCL_NVLS_ENABLE=0 turns it off, which is the quickest A/B test when you suspect it. NVLS needs extra GPU memory for multicast buffers, so a job that runs out of memory at communicator creation on a switched system is worth retrying with NVLS disabled to confirm the cause. The NCCL collectives article explains how NCCL chooses between ring, tree and NVLS.
Worked example: what the domain buys a training step
Take a concrete step. A dense model with 8 billion parameters keeps bf16 gradients, so a full gradient all-reduce moves S = 16 GB. Run it as pure data parallelism over all 72 GPUs of one rack and compare the idealised transfer time per GPU at 900 GB/s per direction. These are bandwidth floors, not measurements: real collectives reach a fraction of link rate, and the fraction is what you measure next.
| Algorithm | Bytes per GPU per direction | Ideal time at 900 GB/s |
|---|---|---|
| Ring all-reduce, n = 72 | 2 x 71/72 x 16 GB = 31.6 GB | about 35 ms |
| In-switch reduction (NVLS) | about 16 GB | about 18 ms |
| Same ring over 400 Gb/s scale-out NIC (50 GB/s) | 31.6 GB | about 630 ms |
The last row is why the domain size matters more than the headline per-link number: the gap between NVLink and a single 400 Gb/s NIC is roughly eighteen times. Now apply it to tensor parallelism. A tensor-parallel group of 16 performs two all-reduces of the activation tensor per transformer layer in the forward pass and two in the backward. With a micro-batch of 8 sequences of 4,096 tokens and hidden size 8,192 in bf16, each activation tensor is 8 x 4096 x 8192 x 2 bytes, about 537 MB. Ring over 16 GPUs moves 2 x 15/16 of that, about 1 GB per GPU, roughly 1.1 ms ideal on NVLink; on the NIC it would be about 20 ms per all-reduce, four times per layer. Without a switched domain, TP=16 is not a design choice you get to make.
To find your real fraction, run the standard nccl-tests binary across the domain and read the bus bandwidth, which normalises for the algorithm so ring and NVLS are comparable:
# all-reduce across 72 GPUs (one rank per GPU, launched by your MPI or Slurm setup)
mpirun -np 72 ./build/all_reduce_perf -b 64M -e 8G -f 2 -g 1
# A/B the in-switch reduction
NCCL_NVLS_ENABLE=0 mpirun -np 72 ./build/all_reduce_perf -b 64M -e 8G -f 2 -g 1
# see which algorithm NCCL chose
NCCL_DEBUG=INFO NCCL_DEBUG_SUBSYS=INIT,TUNING mpirun -np 72 ./build/all_reduce_perf -b 1G -e 1G -g 1For all-reduce, nccl-tests reports busbw = algbw x 2(n-1)/n. Compare busbw at large message sizes against the 900 GB/s per-direction link rate; that ratio, not the datasheet, is what your parallelism plan should assume.
The software stack that has to be running
The switch is useless without its control plane, and most NVLink Switch incidents are control-plane incidents. Bottom up, the pieces that have to agree are:
- Driver and fabric management. On HGX systems the
nvidia-fabricmanagerservice programs the NVSwitch routing tables; until it has done so, CUDA reports the GPUs but peer access over NVLink fails or applications hang at initialisation. On rack-scale NVL72 systems, fabric management runs at the rack level and also defines NVLink partitions, the sets of GPUs allowed to talk to each other. - IMEX. GPUs in different compute trays sit under different operating systems. Sharing memory between them needs the
nvidia-imexdaemon on each node, which exports and imports GPU memory across the domain. In CUDA the mechanism is a fabric memory handle: allocations made with the virtual memory APIs and theCU_MEM_HANDLE_TYPE_FABRIChandle type can be exported and mapped by a process on another node. - NCCL multi-node NVLink. NCCL detects that ranks on different hosts share an NVLink domain and routes traffic over NVLink instead of the NIC. If detection fails, the job still runs, just over the scale-out network, which is the worst kind of failure because nothing errors.
- Scheduler. Slurm or Kubernetes must place a job's ranks inside one NVLink partition. A scheduler that sees 18 four-GPU nodes and packs a 64-GPU job across two racks has silently split the tensor-parallel group across the slow fabric.
A quick health check on a node before blaming the model:
nvidia-smi topo -m # NV# entries between GPUs, not SYS/PHB
nvidia-smi nvlink -s # per-link state and speed; all 18 links active
systemctl status nvidia-imex # on multi-node NVLink systems
systemctl status nvidia-fabricmanager # on HGX baseboard systems
Failure modes
Failures in a switched domain have characteristic shapes. Recognising the shape saves hours.
| Symptom | Likely cause | What to check |
|---|---|---|
| Collective bandwidth near NIC speed, no errors | NCCL fell back to the network: ranks outside one partition, IMEX not running, or MNNVL not detected | NCCL_DEBUG=INFO transport lines; scheduler placement |
| Every job on the rack about 5 percent slower | One switch chip or cable plane down; every GPU lost 1/18 | per-link status on all GPUs; link error counters |
| One GPU slower in every collective | Links on that GPU degraded or down | nvidia-smi nvlink -s on that GPU; DCGM NVLink counters |
| Hang at communicator init | Fabric manager not ready, partition mismatch, IMEX channel missing | service logs; whether all ranks see the same domain |
| OOM only on switched systems | NVLS multicast buffers | rerun with NCCL_NVLS_ENABLE=0 |
| Job fine at 64 GPUs, worse at 72 | No room for a spare; one straggler gates the collective | plan TP/EP sizes that leave slack |
Two failure modes deserve emphasis. The silent fallback is common because NCCL prefers working to failing. Assert on it: parse the NCCL init log in your launcher and fail the job if ranks that should share NVLink report a network transport between them. And the uniform slowdown from a lost plane is easy to misread as a software regression, because nothing is down from the scheduler's point of view. Export NVLink link state and error counters through DCGM into your monitoring and alert on any link not at full speed.
Planning parallelism and trade-offs
Treat the rack as the unit of model parallelism. The usual layout is tensor and expert parallelism inside the NVLink domain, pipeline and data parallelism across domains. Concretely:
- Size tensor-parallel and expert-parallel groups to divide the GPUs you actually schedule per partition, and keep them inside it. If the operations team holds GPUs in reserve for failures, plan for 64, not 72.
- Move mixture-of-experts all-to-all onto the domain first. All-to-all is the collective that benefits most from uniform one-hop bandwidth, and it is the one that hurts most on a scale-out fabric.
- Keep data-parallel gradient reduction hierarchical: reduce inside the rack over NVLink, then across racks over the NIC, which NCCL does by default when topology detection works.
- Record the domain size in experiment metadata. A throughput number measured at TP=8 on one system and TP=16 on another is not comparable, and the switch is usually why.
The trade-offs are real. A switched domain costs switch trays, power and cabling that a point-to-point design does not. Its failure blast radius is larger: a fabric-management fault can take down NVLink for every job in the partition, where a broken HGX baseboard takes down one server. And it couples scheduling to physical topology, which generic cluster schedulers handle poorly without topology-aware plugins. If your models fit comfortably in TP=8 with data parallelism over InfiniBand, an eight-GPU domain may be the cheaper, simpler answer; the switch earns its keep for large mixture-of-experts models and long-context training where model-parallel groups larger than eight are unavoidable.
What to do next
- Run
nvidia-smi topo -mandnvidia-smi nvlink -son one node and confirm every GPU pair shows NVLink and every link is up. - Run all_reduce_perf across your largest NVLink partition with NVLS on and off; record busbw at 1 GB and 8 GB messages as your baseline.
- Add a launcher check that fails the job if NCCL reports a network transport between ranks in the same domain.
- Export NVLink link state and error counters via DCGM and alert on any link below full speed.
- Re-plan parallelism: tensor and expert groups inside the domain, sized to the GPUs you can schedule, with data parallelism hierarchical across racks.
- Make the scheduler topology-aware so a job's model-parallel ranks never straddle two partitions.