An H100 server is easy to describe and hard to debug. Eight GPUs, an NVLink switch fabric, eight 400 Gb/s network adapters and some storage ports: the parts list fits on one line. But a training job that runs at 60 percent of its expected speed rarely tells you which part is to blame: a PCIe link that trained at half width, a miswired rail, a layout that pushes per-layer traffic onto the network, or a checkpoint that saturates storage.
This article breaks the network of an H100 system into hops. For each hop it gives the bandwidth per direction, the traffic that should use it, the tool that measures it in isolation and what it looks like when it fails. A worked example then budgets one training step across those hops, and a measurement ladder tests them bottom-up so you can localise a slowdown in an hour instead of a week. The chassis inventory, the GPU-to-NIC table and NUMA binding are covered in the DGX H100 architecture article; rail wiring and PXN have their own page on rail-aligned topology.
Following one byte out of the GPU
Follow one byte of a gradient from the memory of GPU 3 to GPU 3 on another node. It starts in HBM3, which the H100 SXM reads at about 3.35 TB/s. If its destination is another GPU in the same box, it leaves over NVLink 4: eighteen links giving 900 GB/s in total, which is 450 GB/s in each direction, through the NVSwitch chips to any peer. If its destination is another node, it leaves through PCIe instead. The GPU sits behind a PCIe Gen5 switch that it shares with its own ConnectX-7 adapter, and with GPUDirect RDMA the adapter reads the byte straight from GPU memory without touching the CPU or system RAM. The adapter sends it at 400 Gb/s, which is 50 GB/s per direction, to the leaf switch for that rail, and from there to the matching GPU on the destination node, through a spine only if the two nodes hang off different leaves.
A thin stub on this site once gave PCIe Gen5 as 128 GB/s per GPU. That is the bidirectional figure for an x16 link. Per direction it is about 64 GB/s raw and somewhat less after protocol overhead, which is only a little above the 50 GB/s the adapter needs. That margin is small, and it explains one of the most common silent failures below: a link that trains at Gen4 or at x8 halves the rail and nothing crashes.
The hop table
| Hop | Per-direction bandwidth | Traffic that belongs here | Isolating tool | Failure signature |
|---|---|---|---|---|
| HBM3 | about 3.35 TB/s | every kernel | a bandwidth microbenchmark | rare; ECC or thermal throttling shows in nvidia-smi |
| NVLink 4 via NVSwitch | 450 GB/s | tensor, sequence and expert parallel; intra-node all-reduce | nvbandwidth, nvidia-smi nvlink -s / -e | inactive links, rising error counters, Fabric Manager down |
| PCIe Gen5 x16 to the switch | about 64 GB/s raw | GPUDirect RDMA to the NIC, host copies | nvidia-smi PCIe query under load, lspci -vv | link below Gen5 or x8: one rank slow, rail capped near half |
| ConnectX-7 port | 50 GB/s | data parallel, FSDP, pipeline sends | ib_write_bw --use_cuda | port down, wrong rate in ibstat, symbol errors |
| Leaf (one per rail) | non-blocking within a scalable unit in reference designs | same-rail traffic between nodes | nccl-tests across nodes in one unit | miswired cable: one pair of nodes slow |
| Spine | depends on oversubscription | cross-unit and cross-rail traffic | nccl-tests across units | busbw drops when the job spans units |
| Storage and front-end ports | separate adapters, separate fabric | datasets, checkpoints, logging, SSH | fio against the shared file system | step-time spikes at checkpoint intervals |
Which parallelism rides which hop
The design rule is that bandwidth falls at every hop, by about nine times from NVLink to a rail, so each type of parallelism should run on the fastest hop that can contain it. Tensor parallelism and sequence parallelism exchange activations inside every transformer layer, twice in the forward pass and twice in the backward pass, so they belong on NVLink and the group size should not exceed eight. Expert parallelism exchanges tokens with an all-to-all in every MoE layer; it also prefers NVLink, and when it must cross nodes it becomes the most network-sensitive traffic in the job. Data parallelism, and the reduce-scatter and all-gather of FSDP or ZeRO, happen once per step on gradients and parameters, which is large but easy to overlap with the backward pass, so they ride the rails. Pipeline parallelism sends activations only at stage boundaries and tolerates the network well.
Datasets and checkpoints do not belong on the compute fabric at all; they go over the storage adapters, or they compete with gradient traffic exactly when the step waits on it.
Worked example: one step of a 70B model on 64 GPUs
Take a 70-billion-parameter dense model with hidden size 8,192 and 80 layers, trained on 64 H100s in eight nodes. The layout is tensor parallel 8 inside each node and data parallel 8 across nodes, with a global batch of 2,097,152 tokens per step. Each data-parallel replica, one node, therefore processes 262,144 tokens per step. Every number below is a planning estimate to be replaced by your own measurements.
NVLink: tensor parallel. One activation tensor for the replica's tokens is 262,144 × 8,192 × 2 bytes, about 4.3 GB summed across micro-batches. Four all-reduces per layer over 80 layers make about 1,374 GB of payload. A ring all-reduce over eight GPUs moves 2 × 7/8 of the payload through each GPU, about 2,400 GB. At an assumed 400 GB/s of effective bus bandwidth that is about 6 seconds per step.
Rails: data parallel. Each GPU holds one eighth of the model, 8.75 billion parameters, and reduces 17.5 GB of bf16 gradients with its seven same-rank peers on the other nodes. A ring moves 2 × 7/8 × 17.5 GB, about 30.6 GB, through each adapter. At an assumed 45 GB/s that is about 0.7 seconds, and most of it can overlap with the backward pass.
Compute, for scale. Training costs about 6 × parameters × tokens FLOPs, so each GPU does 6 × 70e9 × 262,144 / 8, about 1.4e16 FLOPs per step. At 400 TFLOPS achieved, about 40 percent of the dense bf16 peak, that is about 34 seconds.
Storage: the checkpoint. With bf16 weights, fp32 master weights and two Adam moments, a full checkpoint is about 14 bytes per parameter, roughly 980 GB, or about 122 GB per node. The node's storage adapters could move that in seconds; the shared file system's aggregate write bandwidth almost always sets the time instead. At 20 GB/s for the cluster it takes about 49 seconds, which is longer than a step.
The budget tells you where to look. Tensor-parallel traffic is about 18 percent of the compute time and is the hop to protect. The rails are lightly loaded at this scale until you add expert parallelism or shrink the step. The checkpoint is the largest single burst, so asynchronous checkpointing pays for itself.
The budget as code
The arithmetic fits in a few lines. Run it for your own model and layout before a large job; the ratio between the communication times and the compute time is what matters, not the absolute numbers.
# per_hop_budget.py -- rough bytes and seconds per training step, per hop
def budget(params, layers, hidden, tokens_per_step, tp, dp,
nvlink_busbw=400e9, rail_busbw=45e9, achieved_flops=400e12,
storage_bw=20e9, ckpt_bytes_per_param=14):
tok_rep = tokens_per_step / dp # tokens per DP replica
act = tok_rep * hidden * 2 # one bf16 activation tensor
tp_payload = 4 * layers * act # 2 fwd + 2 bwd all-reduces
tp_bytes = 2 * (tp - 1) / tp * tp_payload # ring traffic per GPU
grad = params / tp * 2 # bf16 gradient shard per GPU
dp_bytes = 2 * (dp - 1) / dp * grad
flops = 6 * params * tok_rep / tp
return {
"nvlink_s": tp_bytes / nvlink_busbw,
"rail_s": dp_bytes / rail_busbw,
"compute_s": flops / achieved_flops,
"checkpoint_s": params * ckpt_bytes_per_param / storage_bw,
}
print(budget(70e9, 80, 8192, 2_097_152, tp=8, dp=8))
# {'nvlink_s': ~6.0, 'rail_s': ~0.68, 'compute_s': ~34.4, 'checkpoint_s': 49.0}Try changing the layout. Setting tp=16 and dp=4 makes every tensor-parallel ring leave the box, so its speed is set by the rails, about one ninth of NVLink. Pass nvlink_busbw=45e9 to model that and the tensor-parallel term grows from 6 seconds to over 100. That is why a tensor-parallel group should never span two boxes.
A measurement ladder, bottom-up
Test the hops in the order a byte meets them, and do not move up a rung until the one below is clean. Each rung confirms a hop or names the component to replace.
# 1. Topology and links, per node
nvidia-smi topo -m # GPU/NIC affinity matrix; NIC beside its GPU = PIX/PXB
nvidia-smi nvlink -s # every link active, at the expected rate
nvidia-smi nvlink -e # error counters: note them, re-check after load
# PCIe speed drops when a GPU idles, so read it while ./nvbandwidth runs in another shell
nvidia-smi --query-gpu=index,pci.bus_id,pcie.link.gen.current,pcie.link.gen.max,\
pcie.link.width.current,pcie.link.width.max --format=csv
for d in $(lspci -d 15b3: | awk '{print $1}'); do # the ConnectX-7 end of each path
echo "$d $(sudo lspci -vv -s $d | grep -E -m2 'LnkCap:|LnkSta:' | tr -s ' ' | tr '\n' ' ')"
done # LnkSta should equal LnkCap for speed and width
ibstat | grep -E "CA '|State|Rate" # compute ports Active at 400
# 2. NVLink in isolation
./nvbandwidth # device-to-device rows should be uniform
# 3. One rail in isolation, GPU memory to GPU memory (perftest built with CUDA)
ib_write_bw -d mlx5_0 --use_cuda=0 -s 1048576 -F --report_gbits # on node B
ib_write_bw -d mlx5_0 --use_cuda=0 -s 1048576 -F --report_gbits nodeB # on node A
# 4. Collectives: one node, then two, then the whole allocation
NCCL_DEBUG=INFO NCCL_DEBUG_SUBSYS=INIT,NET \
./build/all_reduce_perf -b 8M -e 8G -f 2 -g 8
mpirun -np 16 -H nodeA:8,nodeB:8 ./build/all_reduce_perf -b 8M -e 8G -f 2 -g 1Read the results relative to each other, not against a number from a slide. The eight PCIe links in a node should all report the same speed and width. The device-to-device rows in nvbandwidth should be uniform; one low row is a degraded link or GPU. Each ib_write_bw run should land within a few percent of line rate, and all eight rails should match. In all_reduce_perf, compare the bus bandwidth column for large messages: the single-node figure is bounded by NVLink, the two-node figure by the rails, and a full-allocation figure that falls well below the two-node one points at the spine or at a specific pair of nodes. Set NCCL_TOPO_DUMP_FILE to save the topology NCCL detected, and keep the healthy output from your first good node as the baseline for every later one.
Failure modes
- A link trained low. A GPU or its adapter reports a lower PCIe generation or x8 under load after a reseat or a firmware update. The rail it serves runs at about half speed; in a ring every other rank waits for it, so the whole job slows. Found in rung 1, by comparing current against maximum link speed and width at both ends, not by NCCL.
- Fabric Manager not running.
nvidia-smilists eight healthy GPUs but CUDA applications fail to initialise or NVLink peers are missing. Check the service before anything else after a driver change. - Tensor parallel across nodes. A launcher change or an odd GPU count places one tensor-parallel group across two boxes. Step time multiplies and nothing errors. Print the process-group membership at start-up and assert it maps onto whole nodes.
- Storage on the compute rails. The file system mounts over a compute adapter, and checkpoint bursts collide with gradient traffic. Step time spikes on a regular interval that matches checkpointing.
- Error counters that climb. NVLink or InfiniBand symbol errors that rise under load are a cable or transceiver on its way out. Retransmission hides them as lost bandwidth until the link drops. Record counters before and after every burn-in.
Trade-offs
The rail-per-GPU design gives every GPU its own 50 GB/s and keeps same-rank traffic on one leaf, which is ideal for data-parallel rings, at the cost of eight compute adapters, eight cables and eight leaf ports per node. Cheaper designs share adapters between GPUs or oversubscribe the spine; they cost little for pipeline-heavy or small data-parallel jobs and a lot for expert parallelism across nodes. A separate storage fabric adds adapters and switches but keeps checkpoint bursts from stalling collectives. Finally, more NVLink domain beats more network: newer rack-scale NVLink systems, covered in the NVLink Switch article, exist because the hop between NVLink and the network is the steepest step in the whole budget. For the generations of InfiniBand behind the rails, see InfiniBand NDR and XDR, and for how ring and tree all-reduce turn these bandwidths into step time, see all-reduce in depth.
What to do next
- Draw your own system's hop table: HBM, NVLink, PCIe, adapter, leaf, spine, storage, with per-direction bandwidth for each.
- Run the budget script for your model and layout, and check that tensor and expert parallelism stay inside a node.
- Run rung 1 of the ladder on every node and record PCIe speed, width, NVLink state and error counters as a baseline.
- Run
ib_write_bwwith--use_cudaon every rail of one node pair, then single-node and two-nodeall_reduce_perf. - Confirm that file system mounts use the storage adapters, and time a full checkpoint against the step time.
- Re-run the ladder after every driver, firmware or cabling change, and compare against the baseline.