A GPU cluster's scale-out network is usually drawn as a fat tree, but the most important decision in it fits in one sentence: the NIC attached to GPU i in every server plugs into the switch for rail i. That rule, called rail alignment or rail optimization, means that collectives between GPUs with the same local index on different servers cross a single switch. It also lets the spine layer be built for the rest of the traffic rather than for all of it.

This page explains rail alignment from the software side: why training traffic has the shape that makes rails work, how NCCL picks NICs and when it crosses rails, what PXN does, and how rank placement and parallelism layout either use the rails or quietly bypass them. It also covers how to audit the cabling and test each rail, the failures that only rail-aligned fabrics have, and when designs drop the spine entirely. Figures for specific systems come from NVIDIA's DGX H100 SuperPOD reference architecture. Treat them as an example, and check your own vendor's design for newer generations.

The wiring rule

Rail alignment: NIC i of every node lands on the rail-i leaf; NVLink joins the rails inside a nodeRail 0 leafone switch per railRail 1 leafone switch per railRail 2 leafone switch per railRail 3 leafone switch per railNode 0NVSwitch insideGPU 0 + NIC 0PCIe neighboursGPU 1 + NIC 1PCIe neighboursGPU 2 + NIC 2PCIe neighboursGPU 3 + NIC 3PCIe neighboursNode 1NVSwitch insideGPU 0 + NIC 0PCIe neighboursGPU 1 + NIC 1PCIe neighboursGPU 2 + NIC 2PCIe neighboursGPU 3 + NIC 3PCIe neighboursNode 2NVSwitch insideGPU 0 + NIC 0PCIe neighboursGPU 1 + NIC 1PCIe neighboursGPU 2 + NIC 2PCIe neighboursGPU 3 + NIC 3PCIe neighboursSpines (not drawn) join leaves for cross-SU and cross-rail traffic; PXN turns many cross-rail sends into NVLink hop + same-rail hop.Four rails shown for space; an eight-GPU node has eight rails and eight leaves per scalable unit.
Figure: rail alignment wires NIC i of every node to the rail-i leaf, so same-index GPUs on different nodes are one switch apart.

Start with an eight-GPU server such as a DGX H100. Inside it, NVLink through NVSwitch gives every GPU a path to every other GPU that is many times faster than any network link. For the scale-out fabric, each GPU has its own 400 Gb/s ConnectX-7 NIC on a nearby PCIe switch, so GPU 3 and NIC 3 can move data by GPUDirect RDMA without crossing the CPU.

Rail alignment assigns NIC i of every node to rail i. In the SuperPOD reference design, a scalable unit (SU) is 32 nodes, and each rail of an SU is one 64-port leaf switch: 32 ports down to the NIC i of each node and 32 ports up to the spines, which is non-blocking. Eight rails therefore mean eight leaves per SU, and the reference's switch table lists exactly that: eight leaves and four spines for one SU. In NVIDIA's words, traffic per rail is always one hop away from the other 31 nodes in an SU, while traffic between SUs or between rails crosses the spine layer.

The alternative is a top-of-rack style design in which all eight NICs of a node go to the same leaf. That looks tidier, but it puts every collective between nodes in different racks through the spines, and a single leaf failure takes whole nodes offline instead of one rail.

Why training traffic fits rails

Rails work because data-parallel training mostly exchanges data between GPUs with the same local index. Consider a hierarchical all-reduce of S bytes of gradients per GPU across N nodes. First, inside each node, a reduce-scatter over NVLink leaves GPU i holding the sum of shard i, S/8 bytes. Second, the eight GPU-i shards, one per node, are all-reduced among the N GPUs with local index i, which is exactly the set of NICs on rail i. Third, an all-gather over NVLink restores the full result on every GPU.

In the second step, the eight rails work in parallel and never exchange a byte. With a ring, each NIC sends and receives 2(N − 1)/N × S/8 bytes. NCCL's default ring and tree channels get the same effect by entering and leaving each node through a GPU-NIC pair and chaining node-to-node hops between NICs of the same index. On a rail-aligned fabric, each of those hops is NIC i, leaf i, NIC i.

NCCL documents this design goal directly. The NCCL_CROSS_NIC variable accepts 0 (always use the same NIC for the same ring or tree, to avoid crossing rails), 1 (allow different NICs) and 2, the default (try the same NIC, but allow different ones if that performs better). Recent releases also let NCCL_IB_HCA entries carry an explicit rail and plane, in the form <hca>[:<port>[:<rail>[:<plane>]]].

Cross-rail traffic, PXN and rail-only networks

Not all traffic follows rails. Expert-parallel all-to-all in mixture-of-experts models sends from every GPU to every GPU, and pipeline sends between different local ranks also cross rails. Without help, GPU 2 on node A sending to GPU 5 on node B would leave through NIC 2 on rail 2 and need the spine to reach rail 5.

NCCL 2.12 introduced PXN for this case, combining NVLink and PCIe so a GPU can use a NIC that is not its own. GPU 2 writes the data over NVLink to GPU 5 on its own node, which sends it through NIC 5 on rail 5 to GPU 5 on node B. The cross-rail hop becomes an NVLink hop plus a same-rail network hop, and the intermediate GPU can aggregate messages bound for the same destination NIC, which matters for the many small messages of all-to-all. NCCL_PXN_DISABLE=1 turns it off, which is useful for A/B tests while debugging, and NCCL_P2P_PXN_LEVEL controls when it is used for point-to-point. Check your NCCL version's documentation before tuning it.

Taken to its limit, this idea removes the spine. A 2023 paper by Wang and co-authors proposed rail-only networks for LLM training: rails with no spine at all, and every cross-rail transfer forwarded over NVLink first. It works when the parallelism layout's cross-node traffic is either same-rail or forwardable, and it gives up the general any-to-any connectivity that schedulers and storage traffic often assume.

Placement: the software half of rail alignment

Rails only help if the software puts traffic on them. Three rules carry most of the value:

  • Keep high-volume groups inside the node first, then on one rail. Tensor and expert parallelism go inside the NVLink domain. Data-parallel and pipeline groups then pair GPUs with the same local index on different nodes, which is what you get when ranks are numbered node-major and local rank equals GPU index.
  • Pack each data-parallel group into one scalable unit. Same-rail traffic inside an SU crosses one leaf. Same-rail traffic between SUs crosses leaf, spine, leaf. Pipeline sends carry far fewer bytes than gradient all-reduce, so they are the group to stretch across SUs.
  • Allocate whole nodes. A job using four GPUs per node uses four NICs, half the rails. If two jobs share a node, they share the PCIe switches and NICs, and the GPU-to-NIC pairing that NCCL detects may no longer match the rail assumptions.

In virtual machines or containers the PCIe topology NCCL sees can be synthetic, and GPU-NIC affinity may be wrong or missing. NCCL_TOPO_DUMP_FILE writes the topology NCCL detected, and NCCL_TOPO_FILE loads a corrected one. Compare the dump with nvidia-smi topo -m, where each GPU should be PIX or PXB to its own NIC.

Worked example: 512 GPUs, two scalable units

Take two SUs, 64 nodes and 512 GPUs, training a 70B-parameter dense model with tensor parallelism 8, pipeline parallelism 4 and data parallelism 16. TP = 8 fills each node, so all tensor-parallel traffic stays on NVLink. Each GPU holds about 70B / 32 = 2.2B parameters, so its bf16 gradients are about 4.4 GB per step.

Place the 16 nodes of each pipeline stage's data-parallel group in the same SU: stages 0 and 1 in SU 0, stages 2 and 3 in SU 1. GPU i of those 16 nodes all sit on the rail-i leaf of one SU, so the gradient all-reduce is 16-way and entirely within one switch. A ring sends 2 × 15/16 × 4.4 GB, about 8.2 GB per NIC, or roughly 165 ms at an ideal 50 GB/s. Real NCCL bus bandwidth is lower, and overlap with the backward pass hides much of it. Pipeline activations between stage 1 and stage 2 are the only traffic crossing SUs, and they stay same-rail: GPU i to GPU i.

Now place the job naively, scattering nodes across both SUs in scheduler order. Every data-parallel ring now hops between SUs several times. Each hop crosses leaf, spine and leaf, and shares spine uplinks with other jobs. If the spine tier is oversubscribed to save cost, which rails make tempting, those hops run at a fraction of line rate, and the slowest ring sets the step time. Same hardware, same model: placement alone decides whether the rails are used.

Auditing and testing the rails

Rail alignment is a cabling invariant, and cabling invariants decay. A single swapped pair of cables passes every link-up check and still sends one node's rail traffic through the spine. Audit it from data. Export which switch port each NIC is connected to (from LLDP or the InfiniBand fabric discovery tools) together with the GPU each NIC is local to, then check every leaf against the majority rail:

import collections, csv

def audit(path, rails=8):
    """cabling.csv columns: node, hca, gpu_index, switch (GPU-local NIC -> leaf)."""
    rows = list(csv.DictReader(open(path)))
    votes = collections.defaultdict(collections.Counter)
    for r in rows:
        votes[r["switch"]][int(r["gpu_index"])] += 1
    rail_of = {sw: c.most_common(1)[0][0] for sw, c in votes.items()}
    errors = [f"{r['node']} {r['hca']}: GPU {r['gpu_index']} NIC on {r['switch']} (rail {rail_of[r['switch']]} leaf)"
              for r in rows if int(r["gpu_index"]) != rail_of[r["switch"]]]
    per_node = collections.Counter(r["node"] for r in rows)
    errors += [f"{n}: {k} fabric NICs, expected {rails}" for n, k in per_node.items() if k != rails]
    by_rail = collections.Counter(rail_of.values())
    errors += [f"rail {i}: {by_rail[i]} leaves" for i in range(rails) if by_rail[i] != by_rail[0]]
    return errors

Then test performance one rail at a time. Rail i is the path from GPU i through its local NIC, so pin both: CUDA_VISIBLE_DEVICES=i selects the GPU, and NCCL_IB_HCA==<hca> (a leading = means exact match) selects the NIC that nvidia-smi topo -m shows as local to GPU i. Do not assume mlx5_i, because device numbering often differs from GPU numbering. Run nccl-tests all_reduce_perf -b 8M -e 4G -f 2 -g 1 with one rank per node across an SU, and repeat for each rail. Healthy rails produce nearly identical bus bandwidth curves. One slow rail points at a leaf, a cable or a NIC, and because every multi-node collective uses all rails, that one rail sets the pace for the whole job. Finish with an all-rail run and keep the curves as the baseline for acceptance after maintenance.

Failure modes

  • Swapped cables. Traffic quietly crosses the spine; bandwidth drops only under load. Run the audit after every hardware change.
  • A rail leaf fails. GPU i on every node in the SU loses its NIC. NCCL generally does not move an in-flight communicator to another NIC, so jobs typically hang until the NCCL or framework timeout fires. Restart from checkpoint on capacity that excludes the SU, and alert on leaf health, not only node health.
  • Cross-NIC rings on asymmetric nodes. With the default NCCL_CROSS_NIC=2, NCCL may pair different NIC indices across nodes when a node looks different, for example with a NIC down or a different PCIe layout. NCCL sees only the node's internal topology, not the cabling, so it cannot detect swapped cables. Setting 0 during acceptance tests keeps rings on matching NIC indices, so a bad rail shows up as a slow rail rather than being routed around.
  • Wrong GPU-NIC affinity. A VM image or container that hides the PCIe topology makes NCCL pair GPUs with distant NICs. Compare NCCL_TOPO_DUMP_FILE output with nvidia-smi topo -m.
  • Fragmented scheduling. Partial nodes and scattered allocations bypass rails without any error. Make the scheduler topology-aware, at least at whole-node and SU granularity.

Trade-offs

DesignStrengthWeakness
Rail-optimized fat treesame-index traffic within an SU crosses one leaf; a leaf loss degrades one rail, not whole nodeseight leaf ports per node, placement discipline required
Per-node top-of-racksimple cabling, any-to-any symmetricevery inter-rack collective uses spines; leaf failure removes nodes
Rail-only (no spine)fewest switches and opticscross-rail traffic must forward over NVLink; rigid layouts, harder for storage and multi-tenant use

What to do next

  1. Export your cabling and GPU-NIC affinity, run the majority-vote audit, and fix every flagged port.
  2. Run nccl-tests per rail and all-rail on one SU; archive the curves as the acceptance baseline.
  3. Check that your launcher numbers ranks node-major with local rank equal to GPU index, and that TP sits inside nodes.
  4. Make the scheduler allocate whole nodes and pack each data-parallel group into one SU.
  5. A/B test an MoE job with and without NCCL_PXN_DISABLE=1 to see what PXN is buying you.
  6. Keep learning: the four networks of a GPU pod, NCCL collectives, the NVLink Switch, InfiniBand, RoCE v2 and SHARP in-network reductions.
Key takeaway: Rail alignment wires NIC i of every node to rail i, so the same-index exchanges that dominate data-parallel training cross one switch, and the spine only carries what is left. NCCL keeps rings on rails by default, and PXN converts many cross-rail sends into an NVLink hop plus a same-rail hop. The benefit depends on placement and cabling: number ranks node-major, keep data-parallel groups inside one scalable unit, allocate whole nodes, and audit and test rail by rail.