A GPU pod is the unit you schedule large training jobs onto: a few hundred to a few thousand GPUs that can run one job at full speed. People call its interconnect "the pod network", but it is really four networks with different jobs, and most performance and reliability problems come from traffic on the wrong one or from a mismatch between the job's parallelism and the network's shape.

This page explains the pod as a system: what each network carries, why the scale-out fabric is usually rail-optimised, how to size it, how to map tensor, pipeline and data parallelism onto it, and how to prove a new pod is healthy before a job finds out the hard way. Link speeds are quoted for H100 and GB200 systems and were checked against vendor material on the date above; other generations differ, so read your own platform's numbers.

Architecture at a glance

A rail-optimised pod: NVLink inside the node, one leaf switch per railNode A (8 GPUs)NVSwitch domain900 GB/s per GPUGPU0GPU1GPU7NIC0NIC1NIC7Node B (8 GPUs)NVSwitch domain900 GB/s per GPUGPU0GPU1GPU7NIC0NIC1NIC7Rail 0 leafRail 1 leafRail 7 leafSpine switchescross-rail and pod-to-podFront-end / storage networkseparate NICs: data loading, checkpoints, SSHManagement / out-of-bandBMC, provisioning, telemetry
Each GPU has its own NIC wired to the leaf for its rail, so same-rank collectives stay on one switch; NVLink carries traffic inside the node, spines carry cross-rail traffic, and storage and management run on separate networks.

From first principles: a pod is four networks

Start from what a training step needs to move. Within a step, tensor-parallel layers exchange activations several times per layer; pipeline stages pass activations forward and gradients back once per micro-batch; data-parallel or sharded replicas reduce gradients once per step; mixture-of-experts layers do all-to-all exchanges per layer. Separately, the job loads data, writes checkpoints of many gigabytes, and is managed by a scheduler and monitoring stack. These flows differ by orders of magnitude in volume and latency sensitivity, so a pod separates them:

NetworkCarriesTypical technology
Scale-upTensor-parallel and expert traffic inside the NVLink domainNVLink and NVSwitch
Scale-out (back-end)Data-parallel, pipeline and cross-node expert trafficInfiniBand or RoCE, one 400 Gb/s-class NIC per GPU on H100-class nodes
Front-end / storageData loading, checkpoints, container pulls, user accessEthernet, separate NICs
ManagementBMC, provisioning, telemetryLow-speed Ethernet, out-of-band

The separation matters operationally. A checkpoint write that shares links with gradient all-reduce stalls every rank in the job, because collectives proceed at the speed of the slowest participant. Keeping the back-end fabric for collectives only makes step time predictable.

The scale-up domain

The scale-up domain is the set of GPUs that can read and write each other's memory over NVLink. On an H100 HGX node that is eight GPUs, each with 900 GB/s of total NVLink bandwidth, counting both directions, through the node's NVSwitch chips. GB200 NVL72 extends the domain to a rack: 72 Blackwell GPUs, each with 1.8 TB/s of fifth-generation NVLink (900 GB/s each way), about 130 TB/s across the rack. Note the units: NVLink figures are gigabytes per second summed over both directions, while NICs are rated in gigabits per direction. A 400 Gb/s NIC moves about 50 GB/s each way and H100 NVLink about 450 GB/s each way, so on an H100 node the scale-up path per GPU is roughly nine times faster than the scale-out path.

That ratio is the most important design fact in the pod. Traffic that crosses the scale-out network is many times more expensive than traffic inside the NVLink domain, so the parallelism layout should keep the chattiest communication inside it. A bigger NVLink domain, as in NVL72, lets you put larger tensor-parallel or expert-parallel groups inside the fast domain, which is why rack-scale systems matter for very large and MoE models. See NVLink in depth for the link itself.

The scale-out fabric: rails, radix and oversubscription

The scale-out fabric connects NVLink domains. Most large pods use a rail-optimised fat tree. GPU i on every node connects through its own NIC to leaf switch i, so a node with eight GPUs has eight rails. Collectives such as data-parallel all-reduce run between GPUs with the same local rank on different nodes, and on a rail-optimised fabric that traffic stays on one leaf switch: one hop, no spine, no contention with the other seven rails.

Sizing is arithmetic on switch radix. With radix k switches in a non-blocking two-tier fat tree, each leaf uses half its ports down and half up, so the fabric supports k*k/2 endpoints. A 64-port switch (NVIDIA's Quantum-2 QM9700 InfiniBand switch has 64 ports of 400 Gb/s) gives 2,048 endpoints in two tiers; a third tier raises the ceiling to k*k*k/4, 65,536 at radix 64. Every tier adds latency, cables and optics, and most of the cost of a fabric is optics, so designers push hard to fit a pod in two tiers.

Oversubscription, having more down-link than up-link capacity at a leaf, cuts cost but slows any traffic that must leave the leaf. On a rail-optimised fabric, same-rail traffic never leaves the leaf, so moderate oversubscription to the spine hurts less than on a plain fat tree, provided the job's placement keeps collectives on rails. Cross-rail traffic, such as all-to-all between different local ranks, is the case that pays. NCCL's PXN feature helps: a GPU can send through NVLink to the GPU on the destination rail, which then uses its own NIC, turning a cross-rail hop into an NVLink hop plus a same-rail hop. It can be turned off with NCCL_PXN_DISABLE=1 when debugging.

InfiniBand and RoCE both work at scale. Meta has described two clusters of 24,576 H100 GPUs, one on RoCE and one on Quantum-2 InfiniBand, both with 400 Gbps endpoints, and reported training large models on both. The choice is about operational skill, tooling and supply, not a fixed performance verdict. See InfiniBand for GPU clusters and RoCE v2 for congestion control and lossless configuration.

Mapping parallelism onto the topology

Parallelism should follow bandwidth. A standard layout for a dense model on H100 nodes is:

DimensionCommunicationWhere to place it
Tensor parallelAll-reduce or all-gather of activations, several per layerInside the NVLink domain, never across nodes on H100-class systems
Pipeline parallelPoint-to-point activations per micro-batchAcross nodes; low volume, latency matters
Data parallel / shardedGradient reduce-scatter and all-gather once per stepAcross nodes on the same rail; overlap with backward
Expert parallelAll-to-all per MoE layerInside the NVLink domain when it fits; otherwise the most network-hungry dimension

To estimate a collective's cost, use the bus-bandwidth model from nccl-tests. A ring all-reduce of S bytes over n ranks moves 2(n-1)/n * S bytes through each rank's link, so its time is roughly that divided by achieved bus bandwidth. For all-gather and reduce-scatter the factor is (n-1)/n. The code below turns that into a planning tool:

def collective_seconds(bytes_per_rank, ranks, link_gbit, efficiency=0.85, kind="allreduce"):
    # link_gbit: NIC line rate in Gb/s; efficiency: achieved busbw / line rate,
    # measured on your fabric with nccl-tests, not assumed.
    factor = {"allreduce": 2 * (ranks - 1) / ranks,
              "allgather": (ranks - 1) / ranks,
              "reducescatter": (ranks - 1) / ranks}[kind]
    busbw_bytes = link_gbit * 1e9 / 8 * efficiency
    return factor * bytes_per_rank / busbw_bytes

# 70B dense model, TP=8 inside each node, DP=32 across nodes, bf16 gradients
params_per_gpu = 70e9 / 8
grad_bytes = params_per_gpu * 2
print(round(collective_seconds(grad_bytes, 32, 400), 2), "s per step for gradient all-reduce")

If backward computation takes longer than the result, the reduction can hide behind it with bucketed overlap; if not, the network is on the critical path and you should look at larger batches, gradient sharding or a faster fabric. NCCL collectives and tensor parallelism cover the algorithms.

Worked example: a 256-GPU pod and a 70B job

Plan a 256-GPU pod: 32 H100 nodes, eight GPUs each, eight 400 Gb/s back-end NICs per node. With rail-optimisation each rail has 32 endpoints, so one 64-port leaf per rail is enough, with 32 ports down and up to 32 up. Eight leaves cover the pod. Spines carry only cross-rail traffic and links to other pods; four 64-port spines give each leaf eight uplinks to each, a full 32 uplinks per leaf and a non-blocking fabric. Cutting to two spines halves the spine optics and makes the fabric 2:1 oversubscribed for traffic leaving a leaf.

Now run a 70B dense model on it with TP=8 inside each node and DP=32 across nodes. Each GPU holds one eighth of the parameters, about 8.75 billion, so its bf16 gradients are about 17.5 GB. The all-reduce runs among the 32 GPUs of each rail, entirely on one leaf. At 400 Gb/s and 85% efficiency, the planner gives about 0.8 s per step. If the backward pass takes 3 s, that hides well behind compute. With ZeRO-style sharding the data volume is similar, split into a reduce-scatter and an all-gather, but memory per GPU falls sharply.

The design decision is the spine count. This job puts nothing on the spine, so two spines lose nothing for it. An MoE model with expert parallelism across nodes would put all-to-all traffic across rails, where PXN converts much of it to same-rail traffic but not all. If MoE is on the roadmap, buy four spines now; recabling a live pod is slower and riskier than over-provisioning a little.

Bring-up and validation

A new pod is not ready when it powers on. Validate from the bottom up, and keep the results as a baseline:

# 1. Per node: NVLink and PCIe topology, GPU-to-NIC affinity
nvidia-smi topo -m            # each GPU should share a PCIe switch with its own NIC
nvidia-smi nvlink -s          # every link up at expected speed
ibstat                        # InfiniBand ports Active, expected rate (RoCE: check link state and rate)

# 2. Per node: NVLink collective bandwidth
NCCL_TOPO_DUMP_FILE=/tmp/topo.xml ./build/all_reduce_perf -b 8 -e 8G -f 2 -g 8

# 3. Across nodes, one rail at a time and then all rails (MPI build of nccl-tests)
NCCL_DEBUG=INFO NCCL_DEBUG_SUBSYS=INIT,NET \
mpirun -np 256 -N 8 --hostfile pod.hosts ./build/all_reduce_perf -b 1M -e 8G -f 2 -g 1

# 4. Compare busbw at large sizes against the per-rail baseline; bisect any outlier node

Read the NCCL_DEBUG=INFO output for the transport chosen (it should report the InfiniBand or RoCE transport with GPUDirect RDMA, not sockets) and which NICs each rank uses. NCCL_IB_HCA restricts NCCL to named adapters, NCCL_SOCKET_IFNAME picks the interface for bootstrap traffic, and on RoCE NCCL_IB_GID_INDEX must match the GID your fabric uses. Run the cross-node test on every node, not a sample: one bad cable makes the whole collective slow, and only a full sweep finds it. Then teach the scheduler the topology. Slurm's tree topology plugin reads switch-to-node mappings so jobs land on as few leaves as possible; on Kubernetes, use a scheduler that understands node topology labels.

Failure modes

  • One slow link, whole job slow. A degraded cable or a port that negotiated a lower speed caps every collective it touches. Detect with per-node sweeps and port counters; drain the node rather than tolerating it.
  • Wrong GPU-NIC pairing. A GPU using a NIC on the other CPU socket crosses the inter-socket link and loses bandwidth. Check nvidia-smi topo -m and NCCL's NIC choice in the logs.
  • Fallback to sockets. A missing driver module or wrong NCCL_IB_HCA sends collectives over TCP. The job runs, but many times slower. Alert on transport in the INFO logs.
  • Miswired rails. A cable plugged into the wrong leaf turns same-rail traffic into cross-rail traffic. Verify cabling against the plan with link-layer neighbour discovery before acceptance.
  • Lossless Ethernet misconfigured. On RoCE, PFC and ECN settings that disagree across switches cause pause storms or drops. Treat them as one versioned configuration.
  • Storage on the back-end fabric. Checkpoints or data loading routed over compute NICs cause periodic step-time spikes that line up with checkpoint intervals.
  • Fragmented placement. A job scattered across many leaves pushes rail traffic onto spines. Use topology-aware scheduling and watch spine utilisation per job.

Trade-offs

Rail-optimised designs give one-hop collectives and contain failures to a rail, but need eight times as many leaf ports per node as single-NIC designs and depend on placement discipline. A larger NVLink domain shrinks network traffic for TP and MoE, at the price of rack-scale power, liquid cooling and a larger failure blast radius per rack. InfiniBand brings mature congestion control and in-network reduction (see SHARP); RoCE uses Ethernet tooling and suppliers but needs careful lossless tuning. Oversubscribing the spine saves optics and money and is cheap for data-parallel dense training but expensive for MoE. Measure your own workloads on a small slice before fixing the design.

What to do next

  1. Write down the four networks for your pod and what traffic each may carry; enforce it with interface names and NCCL variables.
  2. Record your scale-up and scale-out bandwidth per GPU in the same units and the same direction, and compute the ratio.
  3. For each production model, list its TP, PP, DP and EP sizes and check that TP and EP fit inside the NVLink domain.
  4. Estimate collective time per step with the planner above, using measured busbw, and compare it to backward time.
  5. Run the bring-up sweep on every node, save the topology dumps and busbw baselines, and re-run after any hardware change.
  6. Configure topology-aware placement and track how many leaves each job spans.
  7. Alert on socket-transport fallback, link speed downgrades and per-node busbw outliers.
  8. Decide spine count from your roadmap, especially MoE, not only from today's jobs.
Key takeaway: A GPU pod is four networks: NVLink for scale-up, a rail-optimised InfiniBand or RoCE fabric for scale-out collectives, and separate front-end and management networks. Scale-up bandwidth per GPU is many times the scale-out rate, so keep tensor and expert parallelism inside the NVLink domain and run data-parallel reductions on rails. Size the fabric from switch radix, estimate collective time with the bus-bandwidth model, validate every node with nccl-tests, and schedule jobs with topology awareness.