Two jobs ask for two GPUs on the same eight-GPU server. One gets a pair under the same PCIe switch; the other gets a pair on opposite CPU sockets. Both see torch.cuda.device_count() == 2, both run the same code, and the second one's gradient all-reduce can take several times longer. Nothing in the framework warns you. That gap is what topology awareness is about.

A server is not a flat bag of GPUs. It is a graph of links with very different bandwidth and latency: NVLink between some GPUs, PCIe switches, CPU root complexes, the socket-to-socket link, and NICs hanging off particular switches. This article shows how to read that graph, how NCCL reads it, how to bind processes and memory to match it, and how schedulers can enforce it, with a worked example and the failure modes that hide it from you. For a specific system's wiring see the DGX H100 architecture; for cross-node placement see rail-aligned topology.

The node is a graph

NVIDIA tools describe the path between two devices with a short label. Learn these seven and most topology output becomes readable:

LabelPath between the two devicesTypical consequence
NV#Direct NVLink, # links bondedFastest path; peer-to-peer copies at NVLink speed
PIXAt most one PCIe switchGood P2P over PCIe
PXBSeveral PCIe switches, no host bridgeP2P works, a little slower
PHBThrough a PCIe host bridge (the CPU)P2P may be slow or disabled
NODEBetween host bridges inside one NUMA nodeTraffic crosses the CPU
SYSAcross the socket interconnectSlowest path; often staged through host memory
XThe device itselfDiagonal of the matrix

On HGX-style H100 boards every GPU reaches every other through NVSwitch, so the GPU-to-GPU matrix is uniformly NV18 and placement inside the box matters mostly for CPU and NIC affinity. On PCIe servers, on cloud VMs that expose part of a host, and on older or mixed systems, the GPU-to-GPU paths differ, and placement matters a lot. Bandwidth orders of magnitude to remember: NVLink 4 on H100 offers 900 GB/s per GPU in total, while a PCIe Gen5 x16 link is about 64 GB/s each way, before any socket-crossing penalty. The NVLink and PCIe articles cover the links themselves.

A two-socket PCIe server as NCCL and nvidia-smi see itCPU socket 0 (NUMA 0)cores 0-31, local DRAMCPU socket 1 (NUMA 1)cores 32-63, local DRAMSYS: socket linkPCIe switch APCIe switch BPCIe switch CPCIe switch DGPU0GPU1GPU2GPU3GPU4GPU5GPU6GPU7NIC0NIC1GPU0-GPU1: PIXGPU0-GPU2: NODEGPU0-GPU4: SYSPlacement rule: put ranks that talk most under the same switch, the same socket, and the NIC on that switch
A two-socket PCIe server. GPU0 and GPU1 share a switch (PIX); GPU0 and GPU2 meet only at socket 0 (NODE); GPU0 and GPU4 cross the socket link (SYS). Each NIC is close to only a quarter of the GPUs.

Reading the topology

The first command to run on any new machine is nvidia-smi topo -m. For the server in the diagram the output would look like this (illustrative, trimmed to four GPUs):

        GPU0  GPU1  GPU2  GPU4  NIC0  NIC1  CPU Affinity  NUMA Affinity
GPU0     X    PIX   NODE  SYS   PIX   SYS   0-31          0
GPU1    PIX    X    NODE  SYS   PIX   SYS   0-31          0
GPU2    NODE  NODE   X    SYS   NODE  SYS   0-31          0
GPU4    SYS   SYS   SYS    X    SYS   NODE  32-63         1
NIC0    PIX   PIX   NODE  SYS    X    SYS
NIC1    SYS   SYS   SYS   NODE  SYS    X

Read it in three passes. Which GPU pairs are closest? Which CPU cores and NUMA node belong to each GPU? Which NIC is closest to each GPU? Those three answers become your rank layout, your CPU binding and your NIC selection. For scripts, query the same facts through NVML:

import pynvml as nv

nv.nvmlInit()
n = nv.nvmlDeviceGetCount()
h = [nv.nvmlDeviceGetHandleByIndex(i) for i in range(n)]
LEVEL = {nv.NVML_TOPOLOGY_INTERNAL: "X", nv.NVML_TOPOLOGY_SINGLE: "PIX",
         nv.NVML_TOPOLOGY_MULTIPLE: "PXB", nv.NVML_TOPOLOGY_HOSTBRIDGE: "PHB",
         nv.NVML_TOPOLOGY_NODE: "NODE", nv.NVML_TOPOLOGY_SYSTEM: "SYS"}

for i in range(n):
    bus = nv.nvmlDeviceGetPciInfo(h[i]).busId
    row = [LEVEL.get(nv.nvmlDeviceGetTopologyCommonAncestor(h[i], h[j]), "?")
           if i != j else "X" for j in range(n)]
    mask = nv.nvmlDeviceGetCpuAffinity(h[i], 2)   # two 64-bit words covers 128 CPUs
    cpus = [w * 64 + b for w, word in enumerate(mask) for b in range(64) if word >> b & 1]
    print(i, bus, " ".join(row), f"cpus {cpus[0]}-{cpus[-1]}")
nv.nvmlShutdown()

This call describes the PCIe path only. NVLink is a separate fabric: query it with nvmlDeviceGetNvLinkState or check link state with nvidia-smi nvlink -s. Save the output with every benchmark result: a number without the topology it ran on cannot be compared with anything.

How NCCL reads the topology

NCCL builds its own model of the node at communicator creation. It walks the PCIe tree, finds NVLinks and NICs, assigns each path a bandwidth, and searches for rings and trees that maximise the slowest link. Then it picks, per GPU, the NIC with the best path for network traffic. You can see and steer each step:

VariableWhat it doesWhen to use it
NCCL_DEBUG=INFO with NCCL_DEBUG_SUBSYS=INIT,GRAPHLogs the detected topology and chosen rings and treesEvery new cluster, every regression
NCCL_TOPO_DUMP_FILE=/tmp/topo.xmlWrites the detected topology as XMLTo diff nodes or file a bug
NCCL_TOPO_FILE=path.xmlLoads a topology description before detectionVMs whose virtual PCIe tree hides the real one
NCCL_P2P_LEVELFarthest path allowed for GPU P2P: LOC, NVL, PIX, PXB, PHB, SYSDisable P2P across bad paths
NCCL_NET_GDR_LEVELFarthest GPU-NIC path for GPUDirect RDMAStop RDMA across sockets
NCCL_CROSS_NIC0 same NIC, 1 allow different, 2 default prefer sameRail-aligned fabrics

Cloud providers often ship a topology XML for their GPU VM images because the hypervisor flattens the PCIe tree; if NCCL's log shows every GPU at the same distance from every NIC on a VM that you know has structure, find that file. NCCL loads /var/run/nvidia-topologyd/virtualTopology.xml by default if it exists. The algorithms NCCL then runs are described in NCCL all-reduce.

Binding processes, cores and memory

Topology awareness inside a job comes down to three bindings per process: which GPU, which CPU cores, and which memory node. Frameworks set the first from the local rank; the other two are usually left to chance, which means data loader threads on the wrong socket and pinned host buffers in remote memory. The framework keeps choosing the GPU from LOCAL_RANK; a small launch wrapper adds the CPU and memory binding for that GPU without changing which devices are visible:

#!/usr/bin/env bash
# bind.sh: run as `torchrun --nproc-per-node 8 --no-python ./bind.sh python train.py`
# (--no-python lets torchrun exec this script directly; it exports LOCAL_RANK)
export CUDA_DEVICE_ORDER=PCI_BUS_ID          # match nvidia-smi numbering
IFS=, read -ra VIS <<< "${CUDA_VISIBLE_DEVICES:-}"   # respect the scheduler's allocation
GPU=${VIS[$LOCAL_RANK]:-$LOCAL_RANK}                  # physical index or UUID for this rank
BUS=$(nvidia-smi --query-gpu=pci.bus_id --format=csv,noheader -i "$GPU")  # 00000000:3B:00.0
BUS=$(echo "${BUS#0000}" | tr 'A-F' 'a-f')                               # sysfs: 0000:3b:00.0
NODE=$(cat "/sys/bus/pci/devices/$BUS/numa_node")
[ "$NODE" -lt 0 ] && NODE=0                   # -1 means unknown; fall back to node 0
exec numactl --cpunodebind="$NODE" --membind="$NODE" "$@"

Two details matter. CUDA_DEVICE_ORDER=PCI_BUS_ID makes CUDA's device numbering match nvidia-smi; the default orders by estimated speed, which on mixed systems silently shuffles indices. And --membind is strict: if the node runs out of memory the process fails rather than spilling, which is usually what you want on a training box but can surprise you; --preferred is the soft variant.

Making the scheduler topology-aware

Bindings inside a job only help if the scheduler handed you a good set of devices in the first place. Both common schedulers can be told about topology.

Slurm. In gres.conf, list each GPU with the CPU cores local to it (the Cores= field), so the scheduler knows the affinity. Users then request --gpu-bind=closest so each task gets the GPU nearest its CPUs. For multi-node jobs, ask for whole nodes or for GPU counts that match switch boundaries; see Slurm GPU scheduling.

Kubernetes. The kubelet's Topology Manager coordinates the CPU manager and device plugins. Its policy decides what happens when resources cannot be aligned: none ignores topology, best-effort prefers alignment, restricted rejects pods whose request cannot be aligned, and single-numa-node requires every resource from one NUMA node. CPU alignment only applies to Guaranteed-QoS pods with integer CPU requests and the static CPU manager policy. Use pod scope so all containers in a training pod land together.

Worked example: two GPUs, two placements

Worked example: a team fine-tunes a 1.3-billion-parameter model with plain data parallelism on two GPUs of the PCIe server above. Each step all-reduces bf16 gradients: 1.3 billion times 2 bytes, 2.6 GB. A two-GPU ring sends and receives about the full buffer per GPU, so the step moves roughly 2.6 GB over the GPU-GPU path.

Measure rather than guess. With nccl-tests built, compare the two placements directly:

# same switch (PIX)
CUDA_VISIBLE_DEVICES=0,1 ./build/all_reduce_perf -b 256M -e 2G -f 2 -g 2
# opposite sockets (SYS)
CUDA_VISIBLE_DEVICES=0,4 ./build/all_reduce_perf -b 256M -e 2G -f 2 -g 2

Read the bus bandwidth column at 2 GB. Suppose the PIX pair reaches 40 GB/s and the SYS pair 12 GB/s (illustrative; your numbers depend on the platform). The all-reduce costs about 65 ms versus 217 ms. If compute per step is 400 ms and communication is not overlapped, the step goes from 465 to 617 ms: the badly placed job is about a third slower for no reason anyone will see in their training code. Overlapping communication with backward computation hides part of the gap, never all of it. The fix costs nothing: request GPUs 0 and 1, and bind the processes to socket 0.

Catching topology drift

Topology drifts. A firmware update re-enables ACS, a technician reseats a NIC in a different slot, a BIOS change flips NUMA settings, or one NVLink starts failing. None of these stop a node from passing a basic GPU health check, and all of them slow every job that lands there. The cheap defence is a golden file: capture nvidia-smi topo -m once per node type when the hardware is known good, and compare every node against it at boot and before each job.

import json, re, subprocess, sys, pathlib

def topo_matrix():
    out = subprocess.run(["nvidia-smi", "topo", "-m"], capture_output=True,
                         text=True, check=True).stdout
    rows = [l.split() for l in out.splitlines() if re.match(r"(GPU|NIC)\d+\s", l)]
    # keep only the link labels; affinity columns differ by CPU numbering
    return {r[0]: r[1:1 + len(rows)] for r in rows}

# once, on a known-good node: json.dump(topo_matrix(), open("golden.json", "w"))
golden = json.loads(pathlib.Path(sys.argv[1]).read_text())
live = topo_matrix()
diffs = [(dev, g, l) for dev in golden
         for g, l in zip(golden[dev], live.get(dev, [])) if g != l]
if diffs or live.keys() != golden.keys():
    print("TOPOLOGY DRIFT", diffs[:10])
    sys.exit(1)                                     # drain the node, do not run jobs

Wire the exit code into whatever drains nodes: a Slurm health check program, a Kubernetes node-problem detector, or a pre-job hook. A drained node that is investigated costs one machine; an undrained one silently slows every multi-node job that includes it, because collectives run at the speed of their slowest member.

Failure modes

Topology problems rarely announce themselves. These are the usual ones:

  • ACS on PCIe switches. Access Control Services redirect peer-to-peer traffic up to the root complex, so PIX pairs perform like PHB. NCCL's guidance is to disable ACS on bare metal; inside VMs the hypervisor may require it.
  • Flattened VM topology. The guest sees every GPU behind one bridge, NCCL picks poor rings and NICs. Use the provider's topology file.
  • Index mismatch. Without CUDA_DEVICE_ORDER=PCI_BUS_ID, GPU 3 in your code may not be GPU 3 in nvidia-smi, and you pin CPUs to the wrong socket.
  • A dead NVLink. One link down drops a pair to PCIe, and a whole ring runs at the slowest edge. Check nvidia-smi nvlink -s in node health checks.
  • Fabric manager not running. On NVSwitch systems, GPUs cannot use the switch fabric until nvidia-fabricmanager is up; jobs fail or fall back.
  • Container cpusets that ignore NUMA. The container runtime hands out cores from both sockets, undoing the binding you set inside.

Trade-offs

Strict placement improves performance but reduces how many jobs a cluster can pack. A scheduler that insists on same-switch pairs will leave odd GPUs idle; one that packs freely will run some jobs slowly. A sensible default is strict alignment for multi-GPU training and relaxed alignment for single-GPU inference, which has no GPU-to-GPU traffic. Likewise, hard memory binding avoids slow remote access but turns memory pressure into failures. Choose per workload, and record the choice so slow runs can be traced back to it.

What to do next

  1. Run nvidia-smi topo -m on every node type you own and save it next to your benchmark results.
  2. Set CUDA_DEVICE_ORDER=PCI_BUS_ID in every launcher.
  3. Turn on NCCL_DEBUG=INFO and NCCL_DEBUG_SUBSYS=INIT,GRAPH for one run per cluster and confirm the rings and NICs match the topology.
  4. Run all_reduce_perf on a good and a bad GPU pair to learn the size of your placement penalty.
  5. Add CPU and memory binding to your launch wrapper.
  6. Configure gres.conf Cores or the Kubernetes Topology Manager so the scheduler stops handing out bad sets.
  7. Add NVLink state and fabric manager status to node health checks.
Key takeaway: A GPU server is a graph, not a pool. Read it with nvidia-smi and NVML, let NCCL show you what it detected, bind each rank to its GPU, socket and memory node, and configure the scheduler so jobs get close devices. A bad placement can cost a third of your step time and nothing else will tell you.