An NVLink Switch turns NVLink from a point-to-point cable between a few GPUs into a switched fabric. Inside one HGX baseboard, NVSwitch chips have long let eight GPUs talk to each other at full NVLink bandwidth. The NVLink Switch takes the same idea out of the server: switch chips live in dedicated trays, and GPUs in different compute trays reach each other through them with NVLink semantics, so a whole rack behaves as one scale-up domain. In a GB200 NVL72 that domain is 72 Blackwell GPUs.

This article explains the NVLink Switch from the software side: what the topology looks like, why one hop through a switch is different from a ring of point-to-point links, what the in-switch reduction engine (NVLink SHARP, which NCCL exposes as the NVLS algorithm) does to collective traffic, which daemons and libraries have to be running for any of it to work, and how a training job should be shaped to use a 72-GPU domain rather than treating it as nine eight-GPU servers. Figures are taken from NVIDIA's published GB200 NVL72 material; where a detail is not public it is left out rather than guessed. For the link layer itself, read NVLink as a memory fabric first.

Why a switched scale-up domain matters

Training parallelism has two kinds of traffic. Data-parallel gradient reduction moves a lot of bytes, but once per step and overlappable with the backward pass, so it tolerates a slower network. Tensor parallelism and expert parallelism move activations inside every layer, on the critical path, many times per step. Those collectives are why the scale-up domain matters: tensor parallelism is normally confined to GPUs joined by NVLink, because pushing per-layer all-reduces over InfiniBand or Ethernet at a fraction of the bandwidth stalls the matrix units.

With eight GPUs per NVLink domain, the tensor-parallel degree was effectively capped at eight, and mixture-of-experts all-to-all traffic beyond eight GPUs fell onto the scale-out network. The NVLink Switch raises the cap. That is the whole value proposition: not a faster link, but a bigger set of GPUs that all see each other at NVLink speed, so model-parallel groups can be larger before they touch the slower fabric. The tensor parallelism guide covers why those groups are so sensitive to bandwidth.

Topology: 72 GPUs, 18 switch chips, one hop

In GB200 NVL72, each Blackwell GPU has 18 fifth-generation NVLink links at 100 GB/s each (bidirectional), which is where the 1.8 TB/s per-GPU figure comes from: 900 GB/s in each direction. The rack holds nine NVLink Switch trays with two switch chips each, 18 chips in total. Every GPU sends exactly one of its 18 links to each of the 18 chips. The arithmetic closes: 72 GPUs times 18 links is 1,296 GPU-side ports, and 18 chips times 72 ports is also 1,296. NVIDIA quotes about 130 TB/s of aggregate NVLink bandwidth for the rack, which is 72 times 1.8 TB/s.

Three properties follow from this wiring, and they drive everything later in the article. First, any GPU reaches any other GPU in one switch hop, so there is no notion of near and far neighbours inside the domain. Second, a single transfer between two GPUs is striped across all 18 chips, so point-to-point bandwidth between any pair can approach the full 900 GB/s per direction when nothing else competes. Third, the switch layer is a set of 18 parallel planes: losing one chip removes one plane, and every GPU degrades by one eighteenth rather than one GPU being cut off.

Rack-scale NVLink domain: every GPU has one link to every switch chipGPU 018 linksGPU 118 linksGPU 218 linksGPU 318 links...GPU 7018 linksGPU 7118 linksswitch chip 072 portsswitch chip 172 portsswitch chip 272 ports...switch chip 1772 ports9 switch trays x 2 chips = 18 chips; 72 GPUs x 18 links = 1,296 GPU ports = 18 chips x 72 portsFabric / NVLink managementroutes, partitions, healthIMEX daemons (per node)export/import GPU memoryNCCLring, tree, NVLS algorithmsAny GPU-to-GPU transfer is one hop: GPU -> switch chip -> GPU, spread across all 18 chips.Lose a chip and every GPU loses 1/18 of its NVLink bandwidth, not one neighbour.
GB200 NVL72 NVLink Switch topology, simplified. Each GPU's 18 links fan out one per switch chip; the control-plane pieces underneath must be healthy before NCCL can use the fabric.

The physical layer between trays is a copper cable cartridge in the back of the rack, not optics. That keeps power and cost down and is why the domain stops at a rack: copper reach is short. Scaling beyond the rack goes back to InfiniBand or Ethernet, covered in the InfiniBand article.

In-switch reduction: NVLink SHARP and NVLS

A plain switch only forwards. NVLink switches since the third NVSwitch generation also contain an arithmetic engine that can perform reductions in the fabric, branded NVLink SHARP. The idea mirrors SHARP in InfiniBand switches: instead of GPUs exchanging partial sums with each other, they send their contribution into the switch once, the switch adds the contributions, and the result is multicast back to every participant.

Count the bytes to see why this matters. A ring all-reduce of S bytes across n GPUs makes each GPU send and receive 2(n-1)/n times S. With in-switch reduction each GPU sends roughly S into the fabric and receives roughly S back. For large n the ring moves nearly 2S per GPU per direction, the switched reduction about S, so the bandwidth demand on every GPU's links halves. NVLink SHARP also takes the reduction arithmetic off the GPU's SMs, which matters when the collective overlaps compute: fewer SMs are borrowed by communication kernels.

NCCL exposes this as the NVLS algorithm. When the hardware and drivers support it, NCCL can pick NVLS for all-reduce on its own; NCCL_NVLS_ENABLE=0 turns it off, which is the quickest A/B test when you suspect it. NVLS needs extra GPU memory for multicast buffers, so a job that runs out of memory at communicator creation on a switched system is worth retrying with NVLS disabled to confirm the cause. The NCCL collectives article explains how NCCL chooses between ring, tree and NVLS.

Worked example: what the domain buys a training step

Take a concrete step. A dense model with 8 billion parameters keeps bf16 gradients, so a full gradient all-reduce moves S = 16 GB. Run it as pure data parallelism over all 72 GPUs of one rack and compare the idealised transfer time per GPU at 900 GB/s per direction. These are bandwidth floors, not measurements: real collectives reach a fraction of link rate, and the fraction is what you measure next.

AlgorithmBytes per GPU per directionIdeal time at 900 GB/s
Ring all-reduce, n = 722 x 71/72 x 16 GB = 31.6 GBabout 35 ms
In-switch reduction (NVLS)about 16 GBabout 18 ms
Same ring over 400 Gb/s scale-out NIC (50 GB/s)31.6 GBabout 630 ms

The last row is why the domain size matters more than the headline per-link number: the gap between NVLink and a single 400 Gb/s NIC is roughly eighteen times. Now apply it to tensor parallelism. A tensor-parallel group of 16 performs two all-reduces of the activation tensor per transformer layer in the forward pass and two in the backward. With a micro-batch of 8 sequences of 4,096 tokens and hidden size 8,192 in bf16, each activation tensor is 8 x 4096 x 8192 x 2 bytes, about 537 MB. Ring over 16 GPUs moves 2 x 15/16 of that, about 1 GB per GPU, roughly 1.1 ms ideal on NVLink; on the NIC it would be about 20 ms per all-reduce, four times per layer. Without a switched domain, TP=16 is not a design choice you get to make.

To find your real fraction, run the standard nccl-tests binary across the domain and read the bus bandwidth, which normalises for the algorithm so ring and NVLS are comparable:

# all-reduce across 72 GPUs (one rank per GPU, launched by your MPI or Slurm setup)
mpirun -np 72 ./build/all_reduce_perf -b 64M -e 8G -f 2 -g 1

# A/B the in-switch reduction
NCCL_NVLS_ENABLE=0 mpirun -np 72 ./build/all_reduce_perf -b 64M -e 8G -f 2 -g 1

# see which algorithm NCCL chose
NCCL_DEBUG=INFO NCCL_DEBUG_SUBSYS=INIT,TUNING mpirun -np 72 ./build/all_reduce_perf -b 1G -e 1G -g 1

For all-reduce, nccl-tests reports busbw = algbw x 2(n-1)/n. Compare busbw at large message sizes against the 900 GB/s per-direction link rate; that ratio, not the datasheet, is what your parallelism plan should assume.

The software stack that has to be running

The switch is useless without its control plane, and most NVLink Switch incidents are control-plane incidents. Bottom up, the pieces that have to agree are:

  • Driver and fabric management. On HGX systems the nvidia-fabricmanager service programs the NVSwitch routing tables; until it has done so, CUDA reports the GPUs but peer access over NVLink fails or applications hang at initialisation. On rack-scale NVL72 systems, fabric management runs at the rack level and also defines NVLink partitions, the sets of GPUs allowed to talk to each other.
  • IMEX. GPUs in different compute trays sit under different operating systems. Sharing memory between them needs the nvidia-imex daemon on each node, which exports and imports GPU memory across the domain. In CUDA the mechanism is a fabric memory handle: allocations made with the virtual memory APIs and the CU_MEM_HANDLE_TYPE_FABRIC handle type can be exported and mapped by a process on another node.
  • NCCL multi-node NVLink. NCCL detects that ranks on different hosts share an NVLink domain and routes traffic over NVLink instead of the NIC. If detection fails, the job still runs, just over the scale-out network, which is the worst kind of failure because nothing errors.
  • Scheduler. Slurm or Kubernetes must place a job's ranks inside one NVLink partition. A scheduler that sees 18 four-GPU nodes and packs a 64-GPU job across two racks has silently split the tensor-parallel group across the slow fabric.

A quick health check on a node before blaming the model:

nvidia-smi topo -m          # NV# entries between GPUs, not SYS/PHB
nvidia-smi nvlink -s        # per-link state and speed; all 18 links active
systemctl status nvidia-imex          # on multi-node NVLink systems
systemctl status nvidia-fabricmanager # on HGX baseboard systems

Failure modes

Failures in a switched domain have characteristic shapes. Recognising the shape saves hours.

SymptomLikely causeWhat to check
Collective bandwidth near NIC speed, no errorsNCCL fell back to the network: ranks outside one partition, IMEX not running, or MNNVL not detectedNCCL_DEBUG=INFO transport lines; scheduler placement
Every job on the rack about 5 percent slowerOne switch chip or cable plane down; every GPU lost 1/18per-link status on all GPUs; link error counters
One GPU slower in every collectiveLinks on that GPU degraded or downnvidia-smi nvlink -s on that GPU; DCGM NVLink counters
Hang at communicator initFabric manager not ready, partition mismatch, IMEX channel missingservice logs; whether all ranks see the same domain
OOM only on switched systemsNVLS multicast buffersrerun with NCCL_NVLS_ENABLE=0
Job fine at 64 GPUs, worse at 72No room for a spare; one straggler gates the collectiveplan TP/EP sizes that leave slack

Two failure modes deserve emphasis. The silent fallback is common because NCCL prefers working to failing. Assert on it: parse the NCCL init log in your launcher and fail the job if ranks that should share NVLink report a network transport between them. And the uniform slowdown from a lost plane is easy to misread as a software regression, because nothing is down from the scheduler's point of view. Export NVLink link state and error counters through DCGM into your monitoring and alert on any link not at full speed.

Planning parallelism and trade-offs

Treat the rack as the unit of model parallelism. The usual layout is tensor and expert parallelism inside the NVLink domain, pipeline and data parallelism across domains. Concretely:

  • Size tensor-parallel and expert-parallel groups to divide the GPUs you actually schedule per partition, and keep them inside it. If the operations team holds GPUs in reserve for failures, plan for 64, not 72.
  • Move mixture-of-experts all-to-all onto the domain first. All-to-all is the collective that benefits most from uniform one-hop bandwidth, and it is the one that hurts most on a scale-out fabric.
  • Keep data-parallel gradient reduction hierarchical: reduce inside the rack over NVLink, then across racks over the NIC, which NCCL does by default when topology detection works.
  • Record the domain size in experiment metadata. A throughput number measured at TP=8 on one system and TP=16 on another is not comparable, and the switch is usually why.

The trade-offs are real. A switched domain costs switch trays, power and cabling that a point-to-point design does not. Its failure blast radius is larger: a fabric-management fault can take down NVLink for every job in the partition, where a broken HGX baseboard takes down one server. And it couples scheduling to physical topology, which generic cluster schedulers handle poorly without topology-aware plugins. If your models fit comfortably in TP=8 with data parallelism over InfiniBand, an eight-GPU domain may be the cheaper, simpler answer; the switch earns its keep for large mixture-of-experts models and long-context training where model-parallel groups larger than eight are unavoidable.

What to do next

  1. Run nvidia-smi topo -m and nvidia-smi nvlink -s on one node and confirm every GPU pair shows NVLink and every link is up.
  2. Run all_reduce_perf across your largest NVLink partition with NVLS on and off; record busbw at 1 GB and 8 GB messages as your baseline.
  3. Add a launcher check that fails the job if NCCL reports a network transport between ranks in the same domain.
  4. Export NVLink link state and error counters via DCGM and alert on any link below full speed.
  5. Re-plan parallelism: tensor and expert groups inside the domain, sized to the GPUs you can schedule, with data parallelism hierarchical across racks.
  6. Make the scheduler topology-aware so a job's model-parallel ranks never straddle two partitions.
Key takeaway: The NVLink Switch does not make a link faster; it makes the set of GPUs that share NVLink bandwidth bigger, with one hop between any pair and in-switch reduction that halves all-reduce traffic. Use it by keeping tensor and expert parallelism inside the domain, verifying that NCCL really uses NVLink, and monitoring the 18 planes so a lost link shows up as an alert rather than a mystery slowdown.