A GPU cluster's scale-out network is usually drawn as a fat tree, but the most important decision in it fits in one sentence: the NIC attached to GPU i in every server plugs into the switch for rail i. That rule, called rail alignment or rail optimization, means that collectives between GPUs with the same local index on different servers cross a single switch. It also lets the spine layer be built for the rest of the traffic rather than for all of it.
This page explains rail alignment from the software side: why training traffic has the shape that makes rails work, how NCCL picks NICs and when it crosses rails, what PXN does, and how rank placement and parallelism layout either use the rails or quietly bypass them. It also covers how to audit the cabling and test each rail, the failures that only rail-aligned fabrics have, and when designs drop the spine entirely. Figures for specific systems come from NVIDIA's DGX H100 SuperPOD reference architecture. Treat them as an example, and check your own vendor's design for newer generations.
The wiring rule
Start with an eight-GPU server such as a DGX H100. Inside it, NVLink through NVSwitch gives every GPU a path to every other GPU that is many times faster than any network link. For the scale-out fabric, each GPU has its own 400 Gb/s ConnectX-7 NIC on a nearby PCIe switch, so GPU 3 and NIC 3 can move data by GPUDirect RDMA without crossing the CPU.
Rail alignment assigns NIC i of every node to rail i. In the SuperPOD reference design, a scalable unit (SU) is 32 nodes, and each rail of an SU is one 64-port leaf switch: 32 ports down to the NIC i of each node and 32 ports up to the spines, which is non-blocking. Eight rails therefore mean eight leaves per SU, and the reference's switch table lists exactly that: eight leaves and four spines for one SU. In NVIDIA's words, traffic per rail is always one hop away from the other 31 nodes in an SU, while traffic between SUs or between rails crosses the spine layer.
The alternative is a top-of-rack style design in which all eight NICs of a node go to the same leaf. That looks tidier, but it puts every collective between nodes in different racks through the spines, and a single leaf failure takes whole nodes offline instead of one rail.
Why training traffic fits rails
Rails work because data-parallel training mostly exchanges data between GPUs with the same local index. Consider a hierarchical all-reduce of S bytes of gradients per GPU across N nodes. First, inside each node, a reduce-scatter over NVLink leaves GPU i holding the sum of shard i, S/8 bytes. Second, the eight GPU-i shards, one per node, are all-reduced among the N GPUs with local index i, which is exactly the set of NICs on rail i. Third, an all-gather over NVLink restores the full result on every GPU.
In the second step, the eight rails work in parallel and never exchange a byte. With a ring, each NIC sends and receives 2(N − 1)/N × S/8 bytes. NCCL's default ring and tree channels get the same effect by entering and leaving each node through a GPU-NIC pair and chaining node-to-node hops between NICs of the same index. On a rail-aligned fabric, each of those hops is NIC i, leaf i, NIC i.
NCCL documents this design goal directly. The NCCL_CROSS_NIC variable accepts 0 (always use the same NIC for the same ring or tree, to avoid crossing rails), 1 (allow different NICs) and 2, the default (try the same NIC, but allow different ones if that performs better). Recent releases also let NCCL_IB_HCA entries carry an explicit rail and plane, in the form <hca>[:<port>[:<rail>[:<plane>]]].
Cross-rail traffic, PXN and rail-only networks
Not all traffic follows rails. Expert-parallel all-to-all in mixture-of-experts models sends from every GPU to every GPU, and pipeline sends between different local ranks also cross rails. Without help, GPU 2 on node A sending to GPU 5 on node B would leave through NIC 2 on rail 2 and need the spine to reach rail 5.
NCCL 2.12 introduced PXN for this case, combining NVLink and PCIe so a GPU can use a NIC that is not its own. GPU 2 writes the data over NVLink to GPU 5 on its own node, which sends it through NIC 5 on rail 5 to GPU 5 on node B. The cross-rail hop becomes an NVLink hop plus a same-rail network hop, and the intermediate GPU can aggregate messages bound for the same destination NIC, which matters for the many small messages of all-to-all. NCCL_PXN_DISABLE=1 turns it off, which is useful for A/B tests while debugging, and NCCL_P2P_PXN_LEVEL controls when it is used for point-to-point. Check your NCCL version's documentation before tuning it.
Taken to its limit, this idea removes the spine. A 2023 paper by Wang and co-authors proposed rail-only networks for LLM training: rails with no spine at all, and every cross-rail transfer forwarded over NVLink first. It works when the parallelism layout's cross-node traffic is either same-rail or forwardable, and it gives up the general any-to-any connectivity that schedulers and storage traffic often assume.
Placement: the software half of rail alignment
Rails only help if the software puts traffic on them. Three rules carry most of the value:
- Keep high-volume groups inside the node first, then on one rail. Tensor and expert parallelism go inside the NVLink domain. Data-parallel and pipeline groups then pair GPUs with the same local index on different nodes, which is what you get when ranks are numbered node-major and local rank equals GPU index.
- Pack each data-parallel group into one scalable unit. Same-rail traffic inside an SU crosses one leaf. Same-rail traffic between SUs crosses leaf, spine, leaf. Pipeline sends carry far fewer bytes than gradient all-reduce, so they are the group to stretch across SUs.
- Allocate whole nodes. A job using four GPUs per node uses four NICs, half the rails. If two jobs share a node, they share the PCIe switches and NICs, and the GPU-to-NIC pairing that NCCL detects may no longer match the rail assumptions.
In virtual machines or containers the PCIe topology NCCL sees can be synthetic, and GPU-NIC affinity may be wrong or missing. NCCL_TOPO_DUMP_FILE writes the topology NCCL detected, and NCCL_TOPO_FILE loads a corrected one. Compare the dump with nvidia-smi topo -m, where each GPU should be PIX or PXB to its own NIC.
Worked example: 512 GPUs, two scalable units
Take two SUs, 64 nodes and 512 GPUs, training a 70B-parameter dense model with tensor parallelism 8, pipeline parallelism 4 and data parallelism 16. TP = 8 fills each node, so all tensor-parallel traffic stays on NVLink. Each GPU holds about 70B / 32 = 2.2B parameters, so its bf16 gradients are about 4.4 GB per step.
Place the 16 nodes of each pipeline stage's data-parallel group in the same SU: stages 0 and 1 in SU 0, stages 2 and 3 in SU 1. GPU i of those 16 nodes all sit on the rail-i leaf of one SU, so the gradient all-reduce is 16-way and entirely within one switch. A ring sends 2 × 15/16 × 4.4 GB, about 8.2 GB per NIC, or roughly 165 ms at an ideal 50 GB/s. Real NCCL bus bandwidth is lower, and overlap with the backward pass hides much of it. Pipeline activations between stage 1 and stage 2 are the only traffic crossing SUs, and they stay same-rail: GPU i to GPU i.
Now place the job naively, scattering nodes across both SUs in scheduler order. Every data-parallel ring now hops between SUs several times. Each hop crosses leaf, spine and leaf, and shares spine uplinks with other jobs. If the spine tier is oversubscribed to save cost, which rails make tempting, those hops run at a fraction of line rate, and the slowest ring sets the step time. Same hardware, same model: placement alone decides whether the rails are used.
Auditing and testing the rails
Rail alignment is a cabling invariant, and cabling invariants decay. A single swapped pair of cables passes every link-up check and still sends one node's rail traffic through the spine. Audit it from data. Export which switch port each NIC is connected to (from LLDP or the InfiniBand fabric discovery tools) together with the GPU each NIC is local to, then check every leaf against the majority rail:
import collections, csv
def audit(path, rails=8):
"""cabling.csv columns: node, hca, gpu_index, switch (GPU-local NIC -> leaf)."""
rows = list(csv.DictReader(open(path)))
votes = collections.defaultdict(collections.Counter)
for r in rows:
votes[r["switch"]][int(r["gpu_index"])] += 1
rail_of = {sw: c.most_common(1)[0][0] for sw, c in votes.items()}
errors = [f"{r['node']} {r['hca']}: GPU {r['gpu_index']} NIC on {r['switch']} (rail {rail_of[r['switch']]} leaf)"
for r in rows if int(r["gpu_index"]) != rail_of[r["switch"]]]
per_node = collections.Counter(r["node"] for r in rows)
errors += [f"{n}: {k} fabric NICs, expected {rails}" for n, k in per_node.items() if k != rails]
by_rail = collections.Counter(rail_of.values())
errors += [f"rail {i}: {by_rail[i]} leaves" for i in range(rails) if by_rail[i] != by_rail[0]]
return errorsThen test performance one rail at a time. Rail i is the path from GPU i through its local NIC, so pin both: CUDA_VISIBLE_DEVICES=i selects the GPU, and NCCL_IB_HCA==<hca> (a leading = means exact match) selects the NIC that nvidia-smi topo -m shows as local to GPU i. Do not assume mlx5_i, because device numbering often differs from GPU numbering. Run nccl-tests all_reduce_perf -b 8M -e 4G -f 2 -g 1 with one rank per node across an SU, and repeat for each rail. Healthy rails produce nearly identical bus bandwidth curves. One slow rail points at a leaf, a cable or a NIC, and because every multi-node collective uses all rails, that one rail sets the pace for the whole job. Finish with an all-rail run and keep the curves as the baseline for acceptance after maintenance.
Failure modes
- Swapped cables. Traffic quietly crosses the spine; bandwidth drops only under load. Run the audit after every hardware change.
- A rail leaf fails. GPU i on every node in the SU loses its NIC. NCCL generally does not move an in-flight communicator to another NIC, so jobs typically hang until the NCCL or framework timeout fires. Restart from checkpoint on capacity that excludes the SU, and alert on leaf health, not only node health.
- Cross-NIC rings on asymmetric nodes. With the default NCCL_CROSS_NIC=2, NCCL may pair different NIC indices across nodes when a node looks different, for example with a NIC down or a different PCIe layout. NCCL sees only the node's internal topology, not the cabling, so it cannot detect swapped cables. Setting 0 during acceptance tests keeps rings on matching NIC indices, so a bad rail shows up as a slow rail rather than being routed around.
- Wrong GPU-NIC affinity. A VM image or container that hides the PCIe topology makes NCCL pair GPUs with distant NICs. Compare NCCL_TOPO_DUMP_FILE output with nvidia-smi topo -m.
- Fragmented scheduling. Partial nodes and scattered allocations bypass rails without any error. Make the scheduler topology-aware, at least at whole-node and SU granularity.
Trade-offs
| Design | Strength | Weakness |
|---|---|---|
| Rail-optimized fat tree | same-index traffic within an SU crosses one leaf; a leaf loss degrades one rail, not whole nodes | eight leaf ports per node, placement discipline required |
| Per-node top-of-rack | simple cabling, any-to-any symmetric | every inter-rack collective uses spines; leaf failure removes nodes |
| Rail-only (no spine) | fewest switches and optics | cross-rail traffic must forward over NVLink; rigid layouts, harder for storage and multi-tenant use |
What to do next
- Export your cabling and GPU-NIC affinity, run the majority-vote audit, and fix every flagged port.
- Run nccl-tests per rail and all-rail on one SU; archive the curves as the acceptance baseline.
- Check that your launcher numbers ranks node-major with local rank equal to GPU index, and that TP sits inside nodes.
- Make the scheduler allocate whole nodes and pack each data-parallel group into one SU.
- A/B test an MoE job with and without NCCL_PXN_DISABLE=1 to see what PXN is buying you.
- Keep learning: the four networks of a GPU pod, NCCL collectives, the NVLink Switch, InfiniBand, RoCE v2 and SHARP in-network reductions.