A DGX H100 is eight H100 GPUs, two Xeon CPUs, ten ConnectX-7 network adapters, two NVMe arrays and six power supplies in an 8U chassis. To a training job, though, it is a set of bandwidth tiers and affinities. Most slow or failing jobs on these machines are not the GPUs' fault. They come from a process pinned to the wrong socket, an inter-node ring that uses the wrong adapter, a fabric service that did not start, or a dataset streamed over the management network instead of read from local flash.
This article walks the box from the GPU outward, the way software sees it: what each fabric is for, which NIC belongs to which GPU, where data should live, what a worked all-reduce estimate says about where time goes, and the commands and failure signatures an operator needs. Specifications were checked against NVIDIA's DGX H100/H200 user guide; where the guide does not state a detail, this article leaves it out rather than guess.
What is in the box
The table lists what is in the chassis and why software cares about each part. The H200 variant uses the same chassis and topology with 1,128 GB of total GPU memory instead of 640 GB, so almost everything below applies to both.
| Component | DGX H100 | Why software cares |
|---|---|---|
| GPUs | 8 x H100 SXM, 80 GB HBM3 each, 640 GB total | Model state, activations and KV cache must fit here or be sharded |
| GPU fabric | 4 NVSwitch chips, 900 GB/s GPU-to-GPU | Tensor parallel and intra-node collectives live here |
| CPUs | 2 x Xeon 8480C, 56 cores each | Data loading, tokenisation, launcher processes, NUMA domains |
| System memory | 2 TB in 32 DIMMs | Pinned host buffers, page cache, CPU offload |
| Compute network | 8 x single-port ConnectX-7, up to 400 Gb/s | One rail per GPU for inter-node collectives |
| Storage network | 2 x dual-port ConnectX-7 | Shared file system, checkpoints, in-band management |
| Local storage | OS RAID 1 on 2 x 1.92 TB; cache RAID 0 on 8 x 3.84 TB | Fast scratch and dataset cache, not durable storage |
| Power | 6 x 3.3 kW, 4+2 redundancy, 10.2 kW max | Rack budget, and what happens when a feed drops |
Three bandwidth tiers
Three fabrics carry three kinds of traffic, and their bandwidths differ by close to an order of magnitude. NVLink gives each GPU 900 GB/s of total bandwidth, which is 450 GB/s in each direction, to any other GPU in the box through the NVSwitch chips. Every GPU pair sees the same path, so placement inside the node does not matter for NVLink traffic. Each compute NIC runs at up to 400 Gb/s, which is 50 GB/s per direction. That is one ninth of a GPU's NVLink bandwidth, even before protocol overhead.
The rule that follows is simple: put the chattiest parallelism inside the box and the least chatty across boxes. Tensor parallelism exchanges activations in every layer, so a tensor-parallel group of up to 8 stays inside one DGX. Data parallelism exchanges gradients once per step, and pipeline parallelism sends activations only at stage boundaries, so they span nodes. A tensor-parallel group of 16 across two DGX systems forces per-layer traffic onto the 50 GB/s tier and is almost always a mistake.
The NVSwitch chips do more than switch. Third-generation NVSwitch can perform reductions inside the switch, and recent NCCL versions expose this as the NVLS algorithm for collectives within a node. Whether NCCL picks it depends on the NCCL version, the driver and the message size, so read the NCCL debug log rather than assuming. The switches are configured by a host service, NVIDIA Fabric Manager. If that service is not running, CUDA applications fail to initialise even though every GPU looks healthy in nvidia-smi.
GPU, NIC and NUMA affinity
Each GPU has a compute NIC that shares its PCIe path, and each half of the box hangs off one network module and CPU. NVIDIA's user guide publishes the mapping. The RDMA device names below are the defaults the guide lists. Names depend on the OS image and firmware, so confirm them on your own system.
| GPU | RDMA device | OSFP port | Network module / CPU |
|---|---|---|---|
| 0 | mlx5_0 | OSFP4 P2 | 0 |
| 1 | mlx5_3 | OSFP3 P2 | 0 |
| 2 | mlx5_4 | OSFP3 P1 | 0 |
| 3 | mlx5_5 | OSFP4 P1 | 0 |
| 4 | mlx5_6 | OSFP1 P2 | 1 |
| 5 | mlx5_9 | OSFP2 P2 | 1 |
| 6 | mlx5_10 | OSFP2 P1 | 1 |
| 7 | mlx5_11 | OSFP1 P1 | 1 |
Two practical consequences follow. First, GPU-to-NIC traffic should stay on the matching pair. GPUDirect RDMA lets the NIC read GPU memory directly, and that only stays fast when the NIC and GPU are close in the PCIe tree. NCCL detects the topology itself and normally gets this right, so only set NCCL_IB_HCA when you are deliberately excluding the storage adapters or debugging. Second, the CPU threads that feed a GPU should run on that GPU's socket, so pinned host buffers are allocated in local memory and the data loader does not cross the socket link on every batch.
The launcher wrapper below binds each local rank to its socket. It reads the NUMA node from sysfs instead of hard-coding it, which keeps it correct if the mapping differs on your image.
#!/usr/bin/env bash
# bind_rank.sh -- exec'd by torchrun/srun once per local rank
set -euo pipefail
rank=${LOCAL_RANK:?LOCAL_RANK not set}
bus=$(nvidia-smi --query-gpu=pci.bus_id --format=csv,noheader -i "$rank")
bus=$(echo "${bus#00000000:}" | tr 'A-F' 'a-f') # 00000000:1B:00.0 -> 1b:00.0
node=$(cat "/sys/bus/pci/devices/0000:${bus}/numa_node")
[ "$node" -lt 0 ] && node=0 # -1 means no affinity reported
exec numactl --cpunodebind="$node" --membind="$node" "$@"Run it as torchrun --nproc-per-node 8 --no-python bind_rank.sh python train.py or through your scheduler's equivalent. Check the binding afterwards: nvidia-smi topo -m prints a CPU affinity and NUMA column per GPU, and a rank whose process runs on the other socket shows up as a data loader that is consistently slower than its peers.
Local storage: cache, not home
The two NVMe arrays have opposite purposes. The OS sits on two M.2 drives in RAID 1, so it survives a drive failure. The cache array stripes eight 3.84 TB U.2 drives in RAID 0, about 30 TB raw. Striping gives the array the combined bandwidth of eight drives, which is what a data loader reading many shards wants. It also means one failed drive loses the whole array, so treat it as a cache that can be rebuilt from the shared file system.
A pattern that works: stage the training dataset from shared storage into /raid (the usual mount point on DGX OS) at job start or ahead of time, read shards locally during training, and write checkpoints to shared storage over the storage adapters. Never write the only copy of a checkpoint to the cache array. If your file system and libraries support GPUDirect Storage, reads can go from NVMe into GPU memory without a bounce through host memory, but the bigger win for most jobs is simply reading locally instead of across the in-band network that also carries everything else.
Worked example: where a step's communication goes
Take data-parallel training of a 7-billion-parameter model with bf16 gradients across two DGX H100 systems, 16 GPUs in all. The gradient buffer is 7e9 x 2 bytes = 14 GB per step. The bandwidth figures below are planning assumptions, not measurements. Replace them with your own nccl-tests results.
- Reduce-scatter inside each node. A ring reduce-scatter over 8 GPUs moves (8 - 1) / 8 x 14 GB = 12.25 GB per GPU. At an assumed 350 GB/s effective NVLink bandwidth that is about 35 ms. Each GPU now holds the node's partial sum for its 1.75 GB shard.
- All-reduce across nodes, one rail per GPU. GPU k on node A exchanges its 1.75 GB shard with GPU k on node B through its own NIC. A two-party all-reduce moves 2 x (2 - 1) / 2 x 1.75 GB = 1.75 GB per GPU. At an assumed 40 GB/s effective per NIC that is about 44 ms. All eight rails run in parallel.
- All-gather inside each node. This mirrors step 1: another 12.25 GB per GPU, about 35 ms.
The total is roughly 114 ms per step if nothing overlaps. Two lessons come out of the arithmetic. The inter-node phase moves the least data but takes the longest, because it runs on the 50 GB/s tier. And the inter-node phase only scales because there are eight rails. If NCCL used two NICs instead of eight, because of a bad NCCL_IB_HCA or a down link, step 2 alone would take about four times as long. In practice, frameworks overlap this communication with the backward pass in buckets, so the exposed cost is lower. The estimate tells you which link to watch. See the all-reduce deep dive for the bus bandwidth formulas, and rail-aligned topology for how the switches outside the box keep those eight rails separate.
Validating a node
Validate a node before it joins a cluster, and again after any hardware change. The checks below run in order from cheapest to most expensive, and each one rules out a different layer.
# 1. Fabric service and GPU visibility
systemctl is-active nvidia-fabricmanager # must print "active"
nvidia-smi -L # expect 8 GPUs
nvidia-smi topo -m # NV18 between every GPU pair; NUMA column per GPU
nvidia-smi nvlink --status -i 0 # every link up at full speed
# 2. Compute rails
ibstat | grep -E "CA '|State|Rate" # 8 compute ports Active at the expected rate
# 3. Health diagnostics (level 3 takes tens of minutes)
dcgmi diag -r 3
# 4. Collectives inside the node, then across nodes (nccl-tests)
./build/all_reduce_perf -b 8 -e 8G -f 2 -g 8
mpirun -np 16 -N 8 ./build/all_reduce_perf -b 8 -e 8G -f 2 -g 1Record the bus bandwidth of the largest message sizes as a baseline for each node. A node that passes every health check but posts noticeably lower bus bandwidth than its siblings usually has a degraded link, a NIC at reduced speed, or a cable problem. Remove it from the pool before it slows a 1,000-GPU job, because synchronous training runs at the speed of its slowest rank.
Failure modes
- Fabric Manager not running. CUDA calls fail at initialisation with "system not yet initialized" (error 802), while
nvidia-smilooks normal. Check the service and its log, and make sure its version matches the driver. - NVLink errors. Xid 74 in the kernel log reports an NVLink error. One bad link can make collectives slow or hang. Drain the node and run diagnostics rather than retrying the job.
- GPU memory errors. Xid 48 (double-bit ECC) and Xid 94/95 (contained and uncontained ECC errors) mean a GPU needs a reset or service. An uncontained error takes down the processes on that GPU.
- GPU fell off the bus. Xid 79 means the GPU is no longer reachable over PCIe. Only a reboot or hardware service fixes it.
- Rail flapping. A compute port that goes down and up again turns into NCCL timeouts or watchdog aborts on every job that uses it. Watch link state and symbol error counters, not just job failures.
- Cross-socket binding. No error at all, just one rank with a slower data loader and therefore a slower step for the whole job.
- Cache array loss. A failed U.2 drive takes out the RAID 0 array. Jobs that relied on it fail on read. Rebuild the array and re-stage the data from shared storage.
- Power. With 4+2 redundancy the system rides through two PSU failures. If three PSUs lose power, for example because one feed drops, NVIDIA's guide says it keeps running at reduced performance. Below three it will not boot. A rack with both feeds on one PDU circuit turns a breaker trip into a performance incident across the rack.
Trade-offs
DGX versus HGX-based servers. OEM servers built on the HGX H100 board share the same eight-GPU NVSwitch baseboard, so the NVLink tier is identical. What differs is CPU choice, NIC count and placement, storage, cooling and the software stack. DGX trades flexibility for a validated configuration, a published topology and DGX OS. With an HGX server, re-derive the GPU-to-NIC table from your own nvidia-smi topo -m rather than reusing the one above.
H100 versus H200 in the same chassis. The H200 model raises total GPU memory from 640 GB to 1,128 GB with the same topology. Memory-bound inference and long context benefit most. For training that already fits, the bottleneck in the worked example above does not change. For the chip itself, see the H100 architecture article, and for the switch fabric, NVLink and NVSwitch.
What to do next
- Run
nvidia-smi topo -mon one node and write down the GPU, NIC and NUMA mapping. Compare it with the table above. - Add the binding wrapper to your launcher and check that each rank's threads stay on one socket.
- Build nccl-tests and record intra-node and two-node bus bandwidth baselines for every node.
- Redo the worked estimate with your own model size and measured bandwidths, and find your exposed communication time.
- Move dataset reads to the local cache array, and make sure checkpoints only go to shared storage.
- Alert on Fabric Manager state, Xid 48/74/79/94/95 and compute-port state changes.
- Check that each rack's PSUs are split across independent power feeds.