Two jobs ask for two GPUs on the same eight-GPU server. One gets a pair under the same PCIe switch; the other gets a pair on opposite CPU sockets. Both see torch.cuda.device_count() == 2, both run the same code, and the second one's gradient all-reduce can take several times longer. Nothing in the framework warns you. That gap is what topology awareness is about.
A server is not a flat bag of GPUs. It is a graph of links with very different bandwidth and latency: NVLink between some GPUs, PCIe switches, CPU root complexes, the socket-to-socket link, and NICs hanging off particular switches. This article shows how to read that graph, how NCCL reads it, how to bind processes and memory to match it, and how schedulers can enforce it, with a worked example and the failure modes that hide it from you. For a specific system's wiring see the DGX H100 architecture; for cross-node placement see rail-aligned topology.
The node is a graph
NVIDIA tools describe the path between two devices with a short label. Learn these seven and most topology output becomes readable:
| Label | Path between the two devices | Typical consequence |
|---|---|---|
| NV# | Direct NVLink, # links bonded | Fastest path; peer-to-peer copies at NVLink speed |
| PIX | At most one PCIe switch | Good P2P over PCIe |
| PXB | Several PCIe switches, no host bridge | P2P works, a little slower |
| PHB | Through a PCIe host bridge (the CPU) | P2P may be slow or disabled |
| NODE | Between host bridges inside one NUMA node | Traffic crosses the CPU |
| SYS | Across the socket interconnect | Slowest path; often staged through host memory |
| X | The device itself | Diagonal of the matrix |
On HGX-style H100 boards every GPU reaches every other through NVSwitch, so the GPU-to-GPU matrix is uniformly NV18 and placement inside the box matters mostly for CPU and NIC affinity. On PCIe servers, on cloud VMs that expose part of a host, and on older or mixed systems, the GPU-to-GPU paths differ, and placement matters a lot. Bandwidth orders of magnitude to remember: NVLink 4 on H100 offers 900 GB/s per GPU in total, while a PCIe Gen5 x16 link is about 64 GB/s each way, before any socket-crossing penalty. The NVLink and PCIe articles cover the links themselves.
Reading the topology
The first command to run on any new machine is nvidia-smi topo -m. For the server in the diagram the output would look like this (illustrative, trimmed to four GPUs):
GPU0 GPU1 GPU2 GPU4 NIC0 NIC1 CPU Affinity NUMA Affinity
GPU0 X PIX NODE SYS PIX SYS 0-31 0
GPU1 PIX X NODE SYS PIX SYS 0-31 0
GPU2 NODE NODE X SYS NODE SYS 0-31 0
GPU4 SYS SYS SYS X SYS NODE 32-63 1
NIC0 PIX PIX NODE SYS X SYS
NIC1 SYS SYS SYS NODE SYS XRead it in three passes. Which GPU pairs are closest? Which CPU cores and NUMA node belong to each GPU? Which NIC is closest to each GPU? Those three answers become your rank layout, your CPU binding and your NIC selection. For scripts, query the same facts through NVML:
import pynvml as nv
nv.nvmlInit()
n = nv.nvmlDeviceGetCount()
h = [nv.nvmlDeviceGetHandleByIndex(i) for i in range(n)]
LEVEL = {nv.NVML_TOPOLOGY_INTERNAL: "X", nv.NVML_TOPOLOGY_SINGLE: "PIX",
nv.NVML_TOPOLOGY_MULTIPLE: "PXB", nv.NVML_TOPOLOGY_HOSTBRIDGE: "PHB",
nv.NVML_TOPOLOGY_NODE: "NODE", nv.NVML_TOPOLOGY_SYSTEM: "SYS"}
for i in range(n):
bus = nv.nvmlDeviceGetPciInfo(h[i]).busId
row = [LEVEL.get(nv.nvmlDeviceGetTopologyCommonAncestor(h[i], h[j]), "?")
if i != j else "X" for j in range(n)]
mask = nv.nvmlDeviceGetCpuAffinity(h[i], 2) # two 64-bit words covers 128 CPUs
cpus = [w * 64 + b for w, word in enumerate(mask) for b in range(64) if word >> b & 1]
print(i, bus, " ".join(row), f"cpus {cpus[0]}-{cpus[-1]}")
nv.nvmlShutdown()This call describes the PCIe path only. NVLink is a separate fabric: query it with nvmlDeviceGetNvLinkState or check link state with nvidia-smi nvlink -s. Save the output with every benchmark result: a number without the topology it ran on cannot be compared with anything.
How NCCL reads the topology
NCCL builds its own model of the node at communicator creation. It walks the PCIe tree, finds NVLinks and NICs, assigns each path a bandwidth, and searches for rings and trees that maximise the slowest link. Then it picks, per GPU, the NIC with the best path for network traffic. You can see and steer each step:
| Variable | What it does | When to use it |
|---|---|---|
NCCL_DEBUG=INFO with NCCL_DEBUG_SUBSYS=INIT,GRAPH | Logs the detected topology and chosen rings and trees | Every new cluster, every regression |
NCCL_TOPO_DUMP_FILE=/tmp/topo.xml | Writes the detected topology as XML | To diff nodes or file a bug |
NCCL_TOPO_FILE=path.xml | Loads a topology description before detection | VMs whose virtual PCIe tree hides the real one |
NCCL_P2P_LEVEL | Farthest path allowed for GPU P2P: LOC, NVL, PIX, PXB, PHB, SYS | Disable P2P across bad paths |
NCCL_NET_GDR_LEVEL | Farthest GPU-NIC path for GPUDirect RDMA | Stop RDMA across sockets |
NCCL_CROSS_NIC | 0 same NIC, 1 allow different, 2 default prefer same | Rail-aligned fabrics |
Cloud providers often ship a topology XML for their GPU VM images because the hypervisor flattens the PCIe tree; if NCCL's log shows every GPU at the same distance from every NIC on a VM that you know has structure, find that file. NCCL loads /var/run/nvidia-topologyd/virtualTopology.xml by default if it exists. The algorithms NCCL then runs are described in NCCL all-reduce.
Binding processes, cores and memory
Topology awareness inside a job comes down to three bindings per process: which GPU, which CPU cores, and which memory node. Frameworks set the first from the local rank; the other two are usually left to chance, which means data loader threads on the wrong socket and pinned host buffers in remote memory. The framework keeps choosing the GPU from LOCAL_RANK; a small launch wrapper adds the CPU and memory binding for that GPU without changing which devices are visible:
#!/usr/bin/env bash
# bind.sh: run as `torchrun --nproc-per-node 8 --no-python ./bind.sh python train.py`
# (--no-python lets torchrun exec this script directly; it exports LOCAL_RANK)
export CUDA_DEVICE_ORDER=PCI_BUS_ID # match nvidia-smi numbering
IFS=, read -ra VIS <<< "${CUDA_VISIBLE_DEVICES:-}" # respect the scheduler's allocation
GPU=${VIS[$LOCAL_RANK]:-$LOCAL_RANK} # physical index or UUID for this rank
BUS=$(nvidia-smi --query-gpu=pci.bus_id --format=csv,noheader -i "$GPU") # 00000000:3B:00.0
BUS=$(echo "${BUS#0000}" | tr 'A-F' 'a-f') # sysfs: 0000:3b:00.0
NODE=$(cat "/sys/bus/pci/devices/$BUS/numa_node")
[ "$NODE" -lt 0 ] && NODE=0 # -1 means unknown; fall back to node 0
exec numactl --cpunodebind="$NODE" --membind="$NODE" "$@"Two details matter. CUDA_DEVICE_ORDER=PCI_BUS_ID makes CUDA's device numbering match nvidia-smi; the default orders by estimated speed, which on mixed systems silently shuffles indices. And --membind is strict: if the node runs out of memory the process fails rather than spilling, which is usually what you want on a training box but can surprise you; --preferred is the soft variant.
Making the scheduler topology-aware
Bindings inside a job only help if the scheduler handed you a good set of devices in the first place. Both common schedulers can be told about topology.
Slurm. In gres.conf, list each GPU with the CPU cores local to it (the Cores= field), so the scheduler knows the affinity. Users then request --gpu-bind=closest so each task gets the GPU nearest its CPUs. For multi-node jobs, ask for whole nodes or for GPU counts that match switch boundaries; see Slurm GPU scheduling.
Kubernetes. The kubelet's Topology Manager coordinates the CPU manager and device plugins. Its policy decides what happens when resources cannot be aligned: none ignores topology, best-effort prefers alignment, restricted rejects pods whose request cannot be aligned, and single-numa-node requires every resource from one NUMA node. CPU alignment only applies to Guaranteed-QoS pods with integer CPU requests and the static CPU manager policy. Use pod scope so all containers in a training pod land together.
Worked example: two GPUs, two placements
Worked example: a team fine-tunes a 1.3-billion-parameter model with plain data parallelism on two GPUs of the PCIe server above. Each step all-reduces bf16 gradients: 1.3 billion times 2 bytes, 2.6 GB. A two-GPU ring sends and receives about the full buffer per GPU, so the step moves roughly 2.6 GB over the GPU-GPU path.
Measure rather than guess. With nccl-tests built, compare the two placements directly:
# same switch (PIX)
CUDA_VISIBLE_DEVICES=0,1 ./build/all_reduce_perf -b 256M -e 2G -f 2 -g 2
# opposite sockets (SYS)
CUDA_VISIBLE_DEVICES=0,4 ./build/all_reduce_perf -b 256M -e 2G -f 2 -g 2Read the bus bandwidth column at 2 GB. Suppose the PIX pair reaches 40 GB/s and the SYS pair 12 GB/s (illustrative; your numbers depend on the platform). The all-reduce costs about 65 ms versus 217 ms. If compute per step is 400 ms and communication is not overlapped, the step goes from 465 to 617 ms: the badly placed job is about a third slower for no reason anyone will see in their training code. Overlapping communication with backward computation hides part of the gap, never all of it. The fix costs nothing: request GPUs 0 and 1, and bind the processes to socket 0.
Catching topology drift
Topology drifts. A firmware update re-enables ACS, a technician reseats a NIC in a different slot, a BIOS change flips NUMA settings, or one NVLink starts failing. None of these stop a node from passing a basic GPU health check, and all of them slow every job that lands there. The cheap defence is a golden file: capture nvidia-smi topo -m once per node type when the hardware is known good, and compare every node against it at boot and before each job.
import json, re, subprocess, sys, pathlib
def topo_matrix():
out = subprocess.run(["nvidia-smi", "topo", "-m"], capture_output=True,
text=True, check=True).stdout
rows = [l.split() for l in out.splitlines() if re.match(r"(GPU|NIC)\d+\s", l)]
# keep only the link labels; affinity columns differ by CPU numbering
return {r[0]: r[1:1 + len(rows)] for r in rows}
# once, on a known-good node: json.dump(topo_matrix(), open("golden.json", "w"))
golden = json.loads(pathlib.Path(sys.argv[1]).read_text())
live = topo_matrix()
diffs = [(dev, g, l) for dev in golden
for g, l in zip(golden[dev], live.get(dev, [])) if g != l]
if diffs or live.keys() != golden.keys():
print("TOPOLOGY DRIFT", diffs[:10])
sys.exit(1) # drain the node, do not run jobsWire the exit code into whatever drains nodes: a Slurm health check program, a Kubernetes node-problem detector, or a pre-job hook. A drained node that is investigated costs one machine; an undrained one silently slows every multi-node job that includes it, because collectives run at the speed of their slowest member.
Failure modes
Topology problems rarely announce themselves. These are the usual ones:
- ACS on PCIe switches. Access Control Services redirect peer-to-peer traffic up to the root complex, so PIX pairs perform like PHB. NCCL's guidance is to disable ACS on bare metal; inside VMs the hypervisor may require it.
- Flattened VM topology. The guest sees every GPU behind one bridge, NCCL picks poor rings and NICs. Use the provider's topology file.
- Index mismatch. Without
CUDA_DEVICE_ORDER=PCI_BUS_ID, GPU 3 in your code may not be GPU 3 innvidia-smi, and you pin CPUs to the wrong socket. - A dead NVLink. One link down drops a pair to PCIe, and a whole ring runs at the slowest edge. Check
nvidia-smi nvlink -sin node health checks. - Fabric manager not running. On NVSwitch systems, GPUs cannot use the switch fabric until
nvidia-fabricmanageris up; jobs fail or fall back. - Container cpusets that ignore NUMA. The container runtime hands out cores from both sockets, undoing the binding you set inside.
Trade-offs
Strict placement improves performance but reduces how many jobs a cluster can pack. A scheduler that insists on same-switch pairs will leave odd GPUs idle; one that packs freely will run some jobs slowly. A sensible default is strict alignment for multi-GPU training and relaxed alignment for single-GPU inference, which has no GPU-to-GPU traffic. Likewise, hard memory binding avoids slow remote access but turns memory pressure into failures. Choose per workload, and record the choice so slow runs can be traced back to it.
What to do next
- Run
nvidia-smi topo -mon every node type you own and save it next to your benchmark results. - Set
CUDA_DEVICE_ORDER=PCI_BUS_IDin every launcher. - Turn on
NCCL_DEBUG=INFOandNCCL_DEBUG_SUBSYS=INIT,GRAPHfor one run per cluster and confirm the rings and NICs match the topology. - Run
all_reduce_perfon a good and a bad GPU pair to learn the size of your placement penalty. - Add CPU and memory binding to your launch wrapper.
- Configure
gres.confCores or the Kubernetes Topology Manager so the scheduler stops handing out bad sets. - Add NVLink state and fabric manager status to node health checks.