Large training jobs spend a measurable fraction of every step moving gradients, activations or expert tokens between nodes. Inside a node, NVLink carries that traffic; between nodes, on most frontier-scale clusters, InfiniBand does. Two generations matter in 2026: NDR, at 400 Gb/s per port, built around NVIDIA's Quantum-2 switches and ConnectX-7 adapters, and XDR, at 800 Gb/s per port, built around Quantum-X800 switches and ConnectX-8 SuperNICs.
This page is about what actually changes between them and why it matters to the people who write and run training code: where the doubling comes from, what switch radix does to the size of fabric you can build, why the host's PCIe link becomes the bottleneck, how to estimate collective time from link speed, and how to bring up and verify a fabric. The fundamentals (RDMA, subnet management, fat trees, adaptive routing, SHARP) are covered in InfiniBand for GPU clusters; read that first if those words are new.
Lanes, SerDes and where 400G and 800G come from
An InfiniBand port is normally four lanes wide, written 4x. Each lane is a serial differential link driven by a SerDes (serialiser/deserialiser). A port's speed is simply lanes times per-lane rate. NDR runs each lane at 100 Gb/s using PAM4 signalling (four voltage levels, two bits per symbol), so a 4x port carries 400 Gb/s. XDR moves to 200 Gb/s per lane, so the same four lanes carry 800 Gb/s. The doubling is entirely in the SerDes: Quantum-X800 is described by NVIDIA as the first switch built on 200 Gb/s-per-lane SerDes.
Because the lane is the atom, ports can be split. On Quantum-2, a 400G port can be split into two 2-lane NDR200 ports at 200 Gb/s each, doubling the switch's radix from 64 to 128 at half the speed. This is how one switch generation serves both 400G GPU links and cheaper 200G links for storage or management nodes.
Converting units is where people go wrong. 400 Gb/s is 50 GB/s per direction; 800 Gb/s is 100 GB/s per direction. Usable bandwidth after encoding and protocol overhead is somewhat less, so measure it rather than assuming it.
The hardware of each generation
The hardware for each generation, as published by NVIDIA and its partners:
| NDR generation | XDR generation | |
|---|---|---|
| Port speed | 400 Gb/s (NDR200 split: 200 Gb/s) | 800 Gb/s |
| Per-lane rate | 100 Gb/s | 200 Gb/s |
| Switch | Quantum-2 (e.g. QM9700), 1U | Quantum-X800 Q3400, 4U |
| Switch radix | 64 x 400G, or 128 x 200G | 144 x 800G |
| Cages | 32 OSFP, two ports per cage | 72 OSFP, two ports per cage |
| Switch throughput | 51.2 Tb/s bidirectional | about 230 Tb/s bidirectional (144 x 800G x 2) |
| In-network reduction | SHARPv3 | SHARP 4th generation |
| Host adapter | ConnectX-7, PCIe Gen5 | ConnectX-8 SuperNIC, PCIe Gen6 (up to 48 lanes) with an integrated PCIe switch |
The twin-port cage matters operationally. On the switch side one OSFP module carries two ports, so cabling plans count modules and ports separately, and a single failed module takes down two links. Cables are passive copper for very short runs inside a rack, active copper slightly longer, and optics beyond that; optics dominate the bill of materials and the failure statistics in large fabrics, so check the supported-cable list for your exact switch and adapter rather than reusing a reach figure from an older generation.
Radix and fabric size
Radix (ports per switch) is the most consequential change. In a non-blocking fat tree, each leaf switch uses half its ports down to endpoints and half up to spines. With radix k, a two-tier tree has at most k leaves (each spine has k ports, one per leaf) of k/2 endpoints each, so it connects k²/2 endpoints. A three-tier tree reaches k³/4.
| Radix | Two tiers (k²/2) | Three tiers (k³/4) |
|---|---|---|
| 64 (NDR) | 2,048 | 65,536 |
| 128 (NDR200 split) | 8,192 at 200G | 524,288 at 200G |
| 144 (XDR) | 10,368 | 746,496 |
The XDR two-tier figure matches NVIDIA's published 10,368 endpoints. Why it matters to software: every extra tier adds switch hops, more links that can fail, more places for congestion, and more optics. A 4,096-GPU cluster with one NIC per GPU needs three tiers at NDR but fits in two at XDR. Collectives that cross the top tier see higher latency and more contention, which shows up in NCCL as lower bus bandwidth at scale. Most GPU fabrics are also rail-optimised (GPU i of every node lands on leaf group i) so that the common case of same-rank traffic stays one hop away; see rail-aligned topology for why.
The host side: PCIe and GPU-NIC placement
A network port is only as fast as the path from GPU memory to the wire. With GPUDirect RDMA the adapter reads and writes GPU memory directly over PCIe, so the PCIe link between the GPU, the adapter and any PCIe switch between them is in the data path.
PCIe Gen5 x16 carries roughly 63 GB/s per direction. That comfortably covers NDR's 50 GB/s, which is why ConnectX-7 on Gen5 works. It does not cover XDR's 100 GB/s. Gen6 doubles the per-lane rate, giving roughly 126 GB/s per direction on x16, which is why ConnectX-8 is a Gen6 device. ConnectX-8 also carries its own PCIe switch, so a GPU can sit behind the same switch as its NIC and the GPU-to-NIC path avoids the CPU root complex. Put an 800G adapter in a platform whose path to the GPU is Gen5 and you have paid for bandwidth the host cannot feed.
Two checks belong in every node bring-up. lspci -vv on the adapter shows LnkSta: the negotiated speed and width must match LnkCap; a slot that trained down to half width silently halves throughput. nvidia-smi topo -m shows the relationship between each GPU and each NIC; you want each GPU paired with a NIC under the same PCIe switch, and a path through the CPU interconnect is a sign of a misconfigured slot or NUMA mapping.
Does the fabric limit your job? Collective-time arithmetic
To judge whether the network matters for a job, estimate collective time from link speed. For a ring all-reduce of S bytes across N participants, each participant sends and receives about 2(N−1)/N × S bytes, so time is roughly 2S/B for large N, where B is per-participant bandwidth. Real multi-node collectives are hierarchical: reduce inside the node over NVLink, all-reduce across nodes over the rails, then broadcast inside the node. With eight GPUs and eight NICs per node, each NIC carries about S/8 of the cross-node traffic.
Worked example. Pure data parallelism on a 7-billion-parameter model with bf16 gradients: S = 14 GB per step. Each of the eight rails carries S/8 = 1.75 GB. The cross-node phase costs about 2 × 1.75 GB ÷ B:
def allreduce_seconds(grad_bytes, rails=8, link_gbps=400, efficiency=0.8):
per_rail = grad_bytes / rails # hierarchical: each NIC carries a shard
bw = link_gbps / 8 * 1e9 * efficiency # bytes/s actually achieved per NIC
return 2 * per_rail / bw # ring all-reduce, large-N approximation
S = 7e9 * 2 # 7B params, bf16 gradients
for gbps in (200, 400, 800):
print(gbps, round(allreduce_seconds(S, link_gbps=gbps), 3))
# about: 200 -> 0.18 s 400 -> 0.09 s 800 -> 0.04 sIf the compute part of the step takes two seconds and the framework overlaps gradient communication with the backward pass, even NDR hides this entirely, and XDR buys nothing for this job. The picture changes with sharded data parallelism (which adds parameter all-gathers every step), with mixture-of-experts all-to-all traffic (which does not shrink with hierarchy and is latency-sensitive), with smaller per-GPU batches that shorten compute, or with pipeline parallelism across nodes. Compute your ratio of communication to compute per step before deciding a faster fabric is the fix; how NCCL all-reduce works explains the algorithms NCCL picks, and SHARP in-network reductions how the switch can do the summing and roughly halve the bytes each host sends.
Bring-up and verification
Software sees the fabric through the verbs stack. Bring-up proceeds from the link outward, and each step has a tool:
# 1. Link: state Active, rate as expected (400 for NDR, 800 for XDR), full width
ibstat mlx5_0 # look for "State: Active" and "Rate: 400"
ibv_devinfo -v | grep -E "active_width|active_speed"
# 2. Fabric: errors, degraded links, topology vs plan
ibdiagnet # run from a management node; review the reported warnings
# 3. Point to point, GPU memory to GPU memory over RDMA (perftest built with CUDA)
ib_write_bw -d mlx5_0 --use_cuda=0 --report_gbits -a # server
ib_write_bw -d mlx5_0 --use_cuda=0 --report_gbits -a <server> # client
# 4. Collectives at job scale (nccl-tests)
mpirun -np 64 -N 8 ./all_reduce_perf -b 8M -e 4G -f 2 -g 1Read all_reduce_perf's bus bandwidth column and compare it with what the link should deliver; a single slow rail shows up as a ceiling on the whole job, because a ring moves at the speed of its slowest link. For NCCL itself, the variables that most often need setting on InfiniBand are NCCL_IB_HCA (which adapters to use, so management ports are excluded) and NCCL_SOCKET_IFNAME (the interface for the bootstrap connection); start from the system vendor's recommended settings and change one variable at a time with NCCL_DEBUG=INFO to confirm what NCCL actually chose.
Mixing generations
Few sites replace a fabric in one go. When generations meet, a link runs at a speed both ends support, so an XDR adapter attached to an NDR leaf delivers NDR speed at best, and the compatibility of specific switch, adapter, firmware and cable combinations is defined by NVIDIA's support matrices rather than by the speed names. Check them before you order. The practical patterns are to build new XDR pods as separate fabric islands and schedule jobs within an island, or to use NDR200 split ports on older switches for storage and service nodes while GPUs move to the newer fabric. A job that spans islands runs at the speed of the gateway between them, which is usually far slower than either.
Failure modes
Failure modes that cost training time:
- Degraded links. A link that comes up at reduced width or rate still passes traffic, so nothing fails; the job is just slower. Alert on rate and width, not only on link state.
- Symbol and error counters climbing. Marginal optics or dirty connectors cause retransmissions that look like random step-time jitter. Track port error counters and replace modules proactively.
- Miscabled rails. A cable plugged into the wrong leaf breaks the rail layout and sends same-rank traffic through the spine. Validate the cabling against the plan with fabric discovery after every maintenance window.
- Host bottlenecks. PCIe trained down, or GPU and NIC on different sockets, caps a node below its link speed. These are invisible to switch-side monitoring.
- One bad node. In synchronous training the slowest participant sets the pace. Run a short NCCL test on every node before admitting it to a large job, and drain nodes that fail.
- Module failures doubled. With two ports per OSFP module, one failure removes two links; plan spares and adaptive routing headroom accordingly.
Trade-offs
XDR doubles per-port bandwidth and more than doubles radix, which flattens large fabrics and shortens communication-heavy steps. It costs more per port, needs PCIe Gen6 hosts to be useful, and brings 200G-per-lane signalling with tighter cable and optics constraints. NDR is mature, widely deployed and enough for many dense data-parallel jobs whose communication already overlaps with compute. Ethernet fabrics tuned for AI are the other option, with different trade-offs in congestion control and operations; training cluster design places the network decision alongside the others. Choose by measuring your workload's communication fraction, not by headline speed.
What to do next
- Compute your job's per-step communication bytes and the communication-to-compute ratio with the formula above before buying or requesting a faster fabric.
- On every node, check
LnkStaagainstLnkCapfor each adapter and confirm GPU-NIC pairing withnvidia-smi topo -m. - Confirm each port's rate and width with
ibstatand alert on any deviation, not just on link down. - Run
ib_write_bw --use_cudaper rail andall_reduce_perfat job scale; record bus bandwidth as a baseline. - Gate admission of nodes into large jobs on a short NCCL health test.
- If mixing generations, read the vendor compatibility matrix and plan islands rather than assuming links negotiate cleanly.