NCCL is fast on a well-described machine and quietly slow on a misdescribed one. At communicator creation it builds a graph of each node: GPUs, NICs, PCIe switches, CPU sockets, NVLink. It then labels every GPU-to-GPU and GPU-to-NIC route with a path type, and those labels decide which transports are allowed, whether GPUDirect RDMA is used, and how channels map to NICs. If the graph is wrong, every later decision is wrong too, and nothing crashes.
Our NCCL deep dive covers the communicator lifecycle and the search in outline, and the GPU topology guide covers reading a node and binding processes. This article is the operator's view of the topology engine itself: the path-type ladder and the knobs that gate on it, topology files for virtual machines and containers, NIC selection, and two small tools, one to summarise transports from logs and one to diff the topology two hosts report. Flag names and defaults below are from NVIDIA's NCCL environment variable documentation; check them against your NCCL version, because they change between releases.
The path-type ladder
NCCL classifies a route by the slowest kind of hop on it. The ladder, from closest to farthest, uses the same names the configuration variables accept:
| Path type | Meaning | Typical consequence |
|---|---|---|
| NVL | Connected through NVLink (directly or via NVSwitch) | P2P over NVLink, highest bandwidth |
| PIX | Same PCIe switch, one bridge | P2P and GDR allowed; traffic stays off the CPU |
| PXB | Through several PCIe switches, not the CPU | Usually still allowed |
| PHB | Through the CPU root complex, same NUMA node | GDR often disabled; host bridge limits |
| SYS | Across the socket interconnect (UPI, xGMI) | Slowest host path; P2P usually off |
Two variables gate directly on the ladder. NCCL_P2P_LEVEL sets the farthest distance at which GPUs use the peer-to-peer transport; NCCL_NET_GDR_LEVEL sets the farthest GPU-to-NIC distance at which the NIC reads and writes GPU memory directly. Both accept LOC, which disables the feature, and the ladder names; P2P_LEVEL also accepts NVL. Both have legacy integer forms (LOC 0, PIX 1, PXB 2, PHB 3, SYS 4) that the documentation discourages because path type numbering has changed over time. Use the strings. Unset, NCCL picks a value per platform, which is almost always what you want.
Some routes are synthesised by NCCL rather than read from hardware. NVB sends between two GPUs through an intermediate GPU over NVLink, disabled with NCCL_NVB_DISABLE=1. PXN lets a GPU send to the network through a NIC that is not local to it, by moving data over NVLink to a GPU that is close to that NIC; NCCL_PXN_DISABLE=1 turns it off, and NCCL_P2P_PXN_LEVEL (default 2, always use it) controls it for send and receive. On Grace-based systems, NCCL_NET_GDR_C2C allows GDR through a CPU-attached NIC over the chip-to-chip link, enabled by default since NCCL 2.27.
From graph to channels
With paths computed, NCCL searches for channel layouts: rings and trees, plus NVLink SHARP (NVLS) and CollNet variants where the hardware supports them. Conceptually the search starts with an ambitious per-channel bandwidth target and the strictest path rules, tries to find enough channels that use the same NIC on every node, and relaxes step by step when it fails: lower bandwidth, farther path types, then different NICs. The resulting channels are duplicated to fill the link, and each channel later maps to a CUDA block on the GPU. Because the search runs once per communicator, a wrong graph fixes a slow plan for the whole job.
The NIC rule is where rails meet topology. NCCL_CROSS_NIC has three values: 0 always uses the same NIC for a ring or tree on every node, for rail-optimised fabrics with slow inter-rail links; 1 allows different NICs, for fabrics where all NICs reach the same switch; 2, the default, prefers the same NIC but allows different ones if that performs better. The rail-aligned topology guide explains why same-NIC channels avoid flow collisions on rail networks.
Algorithm and protocol choice come after the search, from a cost model. NCCL_ALGO and NCCL_PROTO can restrict them; since NCCL 2.24 they take per-collective lists such as allreduce:^tree, and a list without ring no longer falls back to ring silently. NVIDIA discourages forcing protocols except to rule out a suspected bug, and warns that enabling LL128 where it is unsupported can corrupt data.
Virtual machines and topology files
On bare metal NCCL reads the PCIe tree from sysfs. In a virtual machine the hypervisor often presents a flattened or invented PCIe hierarchy, so NCCL cannot tell which NIC sits beside which GPU. Everything lands at PHB or SYS, GDR switches off, and inter-node bandwidth falls well short of the hardware with no error printed.
The fix is a topology file. NCCL_TOPO_FILE names an XML file loaded before detection; by default NCCL loads /var/run/nvidia-topologyd/virtualTopology.xml if it exists. Azure's HPC images ship such files for its GPU VM types. Here is the first block of Microsoft's published file for the NDv4 (A100) series, from the azhpc-images repository:
<system version="1">
<cpu numaid="0" affinity="00000000,00000000,00ffffff" arch="x86_64" vendor="AuthenticAMD" familyid="23" modelid="49">
<pci busid="ffff:ff:01.0" class="0x060400" link_speed="16 GT/s" link_width="16">
<pci busid="0003:00:00.0" class="0x030200" link_speed="16 GT/s" link_width="16"/>
<pci busid="0103:00:00.0" class="0x020700" link_speed="16 GT/s" link_width="16"/>
<pci busid="0004:00:00.0" class="0x030200" link_speed="16 GT/s" link_width="16"/>
<pci busid="0104:00:00.0" class="0x020700" link_speed="16 GT/s" link_width="16"/>
</pci>
</cpu>
<!-- three more cpu blocks, numaid 1 to 3, with the same shape -->
</system>Read it as a statement of locality. The bridge ffff:ff:01.0 (class 0x0604, a PCI bridge) is a stand-in switch. Under it sit two GPUs (class 0x0302, 3D controller) and two InfiniBand adapters (class 0x0207), all on NUMA node 0 with the listed CPU affinity. The file tells NCCL that those GPUs and NICs share a switch, so their paths become PIX, GDR is allowed and each GPU gets a local NIC.
Two rules from the documentation matter. A topology file must describe a single host, even on multi-node NVLink systems. And NCCL_TOPO_DUMP_FILE on such systems dumps the whole NVLink domain, so trim a dump to one host before reusing it as input. Containers have a milder version of the VM problem: if /sys is masked or devices are hidden, detection sees less than the hardware has. Dump the topology from inside the container, not from the host.
Choosing NICs exactly
NCCL_IB_HCA selects RDMA adapters. Entries follow <hca>[:<port>[:<rail>[:<plane>]]], a leading ^ makes it an exclude list, and a leading = makes names exact. Without =, names are prefixes: mlx5_1 also selects mlx5_10 to mlx5_19 if they exist, a classic way to pull a storage or management adapter into training traffic. Always write the exact form:
# exact names, port 1 on each training adapter
export NCCL_IB_HCA==mlx5_0:1,mlx5_1:1,mlx5_2:1,mlx5_3:1
# or exclude the storage adapters and keep everything else
export NCCL_IB_HCA=^=mlx5_8,mlx5_9The optional rail and plane fields assign identities to devices on multi-plane fabrics. NCCL_IB_MERGE_NICS (default 1) combines the two ports of a dual-port adapter into one logical device, and recent releases add NCCL_NET_MERGE_POLICY (since 2.30.5) to stop distance-based merging from joining adapters on different rails. NCCL supports at most 32 HCAs.
Reading what NCCL decided
Ask NCCL to explain itself with NCCL_DEBUG=INFO and NCCL_DEBUG_SUBSYS=INIT,GRAPH,TUNING (the default subsystem list is INIT,BOOTSTRAP,ENV). GRAPH prints the topology and the channels found; INIT prints, for each channel and peer, the transport chosen, in lines such as Channel 00/0 : 0[0] -> 1[1] via P2P/IPC or ... via NET/IB/0/GDRDMA. The exact format varies by version, so the summariser below matches loosely. It counts transports per rank and flags the ones that usually mean a topology problem.
import re, sys
from collections import Counter
LINE = re.compile(r"Channel \d+/\S+ : (\d+)\[[^\]]*\] -> (\d+)\[[^\]]*\].*?via (\S+)")
def summarise(path):
per_rank = {}
for line in open(path, errors="replace"):
m = LINE.search(line)
if m:
src, _dst, via = m.groups()
per_rank.setdefault(int(src), Counter())[via] += 1
for rank, seen in sorted(per_rank.items()):
warn = [v for v in seen if v.startswith("SHM") or (v.startswith("NET") and "GDR" not in v)]
flag = " <-- check topology" if warn else ""
print(rank, dict(seen), flag)
summarise(sys.argv[1])SHM between GPUs in one NVLink node means P2P was refused; NET without GDRDMA means the GPU and NIC were judged too far apart. Both are worth chasing before tuning anything.
Diffing topology between hosts
The most useful habit is comparing what NCCL saw on a slow host with a good one. Set NCCL_TOPO_DUMP_FILE on both, then reduce each dump to its shape: which kinds of device share each bridge under each NUMA node. Bus IDs differ between hosts; the shape should not.
import sys, xml.etree.ElementTree as ET
KIND = {"0x0302": "GPU", "0x0300": "GPU", "0x0207": "IB-NIC", "0x0200": "ETH-NIC"}
def shape(path):
groups = {}
def walk(el, numa, bridge):
if el.tag == "cpu":
numa = el.get("numaid")
if el.tag == "pci":
kind = KIND.get(el.get("class", "")[:6])
if kind:
groups.setdefault((numa, bridge), []).append(kind)
bridge = el.get("busid")
for child in el:
walk(child, numa, bridge)
walk(ET.parse(path).getroot(), None, None)
return sorted((numa, tuple(sorted(k))) for (numa, _), k in groups.items())
good, bad = shape(sys.argv[1]), shape(sys.argv[2])
print("same shape" if good == bad else f"good: {good}\nbad: {bad}")Run against the full NDv4 file, this prints four groups, one per NUMA node, each with two GPUs and two InfiniBand NICs. A host whose dump shows all GPUs under one group and the NICs elsewhere is the VM-without-topology-file case.
Worked example: slow all-reduce on cloud VMs
A team moves an all-reduce job from bare metal to cloud VMs with the same GPUs and adapters. nccl-tests reports bus bandwidth far below the bare-metal baseline, with no errors. The log summariser shows every inter-node connection as NET/IB without GDRDMA. The topology diff shows the VM dump with GPUs and NICs hanging off unrelated bridges, so every GPU-to-NIC path is PHB or SYS and NCCL's automatically chosen GDR level refuses them. Data now bounces through host memory on every send.
The fix is to point NCCL_TOPO_FILE at the provider's topology file for that VM type, or to place it at the default nvidia-topologyd path in the image. After the change, the summariser shows GDRDMA on every network connection, and bandwidth returns to the baseline. Forcing NCCL_NET_GDR_LEVEL=SYS would also have enabled GDR, but blindly: NCCL would still pair GPUs with arbitrary NICs, some across sockets. Describe the machine correctly, then let NCCL choose.
Failure modes
- Stale topology file. A file written for one VM size is reused on another, so NCCL trusts locality that no longer exists. Version files with the image and check the shape at boot.
- Prefix match on adapters.
NCCL_IB_HCA=mlx5_1silently includes mlx5_10 and up. Use = for exact names. - Overrides that outlive their reason. A forced algorithm, protocol or P2P level set to dodge an old bug blocks better choices in newer releases. Review every NCCL variable at each upgrade.
- Containers hiding hardware. A masked /sys or missing devices make detection see less than exists. Dump the topology inside the container.
- Mixed hosts in one job. Nodes with different shapes force the search toward the weakest common layout. Diff shapes across the allocation before launch.
Trade-offs
| Approach | Gain | Risk |
|---|---|---|
| Let NCCL detect | Adapts to hardware and new releases | Wrong when sysfs is wrong |
| Topology file | Correct locality in VMs | Goes stale when hardware changes |
| Path-level overrides | Quick unblock for one platform | Hides the real cause, blocks improvements |
| Forced ALGO or PROTO | Rules out a suspected bug | Slower in other message ranges; LL128 misuse can corrupt data |
| CROSS_NIC 0 | Strict rail discipline | Fewer valid layouts on irregular allocations |
The rule of thumb: fix the description of the machine, not the decisions NCCL makes from it. For how the chosen channels turn into ring steps, see the ring all-reduce walkthrough.
What to do next
- Set NCCL_TOPO_DUMP_FILE on one known-good host per node type and archive the dumps.
- Diff every new host or image against its archived shape before it joins the pool.
- Run one job with NCCL_DEBUG=INFO and NCCL_DEBUG_SUBSYS=INIT,GRAPH,TUNING and summarise transports; SHM inside an NVLink node or NET without GDR needs an explanation.
- On VMs, confirm that a provider topology file is present at NCCL_TOPO_FILE or the default nvidia-topologyd path.
- Rewrite NCCL_IB_HCA with the = prefix and list only training adapters.
- Choose NCCL_CROSS_NIC deliberately to match whether your fabric is rail-optimised.
- List every NCCL override in your launch scripts, record why each exists, and retest without it at each NCCL upgrade.