A DGX SuperPOD is NVIDIA's reference architecture for building a large AI cluster out of DGX systems. It specifies the nodes, four separate networks, storage, management servers and the software that runs on them, all designed to be built in identical blocks. The topology is not just a wiring diagram. It determines which jobs fit, which GPUs talk in one switch hop, what fails together, and what your scheduler should be told.

Other pages cover the parts. Rail-aligned topology explains rails, fat-tree topology explains radix and oversubscription, and DGX H100 architecture explains the node. This page treats the pod as a system. It covers the scalable unit, the four fabrics, the switch arithmetic of a full pod, the node slot the pod gives up for UFM, the Grace Blackwell variant, and how Slurm and NCCL use the topology. Figures come from NVIDIA's DGX SuperPOD reference architecture for DGX B200, checked on 7 October 2026. The architecture is revised for each generation, so check the edition that matches your hardware.

The scalable unit

The scalable unit, or SU, is the building block. In the DGX B200 design, an SU has room for 32 DGX nodes with eight GPUs each, plus the leaf switches they connect to. A pod is built by adding SUs and a spine layer above them. The reference build has four SUs, and the architecture says it can scale much further.

Use the SU as the unit in planning, in software and when things fail. It is the unit you buy, cable and burn in. It is the largest group of nodes in which every rail is one switch hop away. It is also the unit that is most convenient to take out of service for a firmware change. Jobs that fit inside one SU get the best network. Jobs that span SUs depend on the spine. Give your scheduler the SU as a topology level.

SUsNodesGPUsLeaf switchesSpine switchesCompute and UFM cablesSpine-leaf cables
13124884252256
263504168508512
3957602416764768
41271,01632161,0201,024

The node counts are one short of a multiple of 32. The reference architecture explains why: the SU is a 32-node design, but one DGX system is removed to make room for UFM connectivity. UFM, NVIDIA's Unified Fabric Manager, manages the InfiniBand fabric. It runs the subnet manager, collects telemetry and runs diagnostics, and it needs its own ports on the compute fabric. Plan for 127 nodes, not 128. If your job sizes assume a power-of-two node count, the largest job is 64 nodes, or 512 GPUs, unless you accept odd sizes.

Four fabrics, kept apart

Compute fabric of a 4-SU DGX B200 SuperPOD (rail-aligned, two tier)Spine layer: 16 QM9700 switches1,024 leaf-spine cablesSU 1: 8 leavesone leaf per rail32 node slots8 x 400G NICs per nodeSU 2: 8 leavesone leaf per rail32 node slots8 x 400G NICs per nodeSU 3: 8 leavesone leaf per rail32 node slots8 x 400G NICs per nodeSU 4: 8 leavesone leaf per rail32 node slots8 x 400G NICs per nodeRail r of every node in an SU lands on leaf r of that SU: one hop between nodes on the same railCross-rail or cross-SU traffic climbs to the spine: three switch hopsStorage fabricIB or Ethernet, near 4:3 at the DGXIn-band managementEthernet: services, registry, SSHOut-of-band managementEthernet: BMCs, switch consoles
The rail-aligned compute fabric of a four-SU pod, with the three other networks every node also joins.

Each node connects to four networks. They are kept separate so that one kind of traffic cannot slow down another.

FabricWhat travels on itBuilt from (B200 RA)What you notice when it fails
ComputeGPU-to-GPU collectives: all-reduce, all-gather, all-to-allInfiniBand NDR, QM9700 switches, eight 400 Gb/s ports per node, rail-alignedslow or hung NCCL jobs
Storagedataset reads, checkpoint writesInfiniBand (MQM9700-NS2F) or Ethernet (SN5600); DGX side near 4:3, storage side 1:1dataloader stalls, slow checkpoints
In-band managementcluster services, SSH, container registry, Slurm control trafficEthernet, SN5600 and SN2201jobs fail to launch, image pulls time out
Out-of-band managementBMC, switch and PDU management portsEthernet, SN2201, kept logically separateyou lose remote power and console control

The storage fabric is the one teams most often get wrong. The reference architecture requires more than 40 GB/s of I/O per node, and it oversubscribes the DGX side slightly, at about 4:3. Storage traffic is bursty: checkpoints arrive from every node at the same moment. Keeping it off the compute fabric is what stops a checkpoint from slowing an all-reduce.

The switch arithmetic

The compute fabric can be checked with arithmetic. A QM9700 switch has 64 ports of 400 Gb/s NDR. A DGX B200 has eight compute-fabric ports, one per GPU, and port r of every node is rail r. In an SU, rail r of all 32 node slots goes to leaf r. Each leaf therefore uses 32 ports facing down, toward nodes and UFM, and has 32 left to face up, toward the spine.

  • Leaves in four SUs: 4 x 8 = 32. Uplinks: 32 x 32 = 1,024, which matches the spine-leaf cable count.
  • Spine ports: 16 switches x 64 ports = 1,024. Every uplink has a spine port, and each leaf reaches each spine over two cables.
  • Down equals up at every leaf, so the compute fabric is non-blocking. Any traffic pattern that leaves the SU still has a full path, as long as routing spreads it evenly.
  • Per node, the scale-out bandwidth is 8 x 400 Gb/s = 3.2 Tb/s, or about 400 GB/s. Inside the node, NVLink gives each B200 far more bandwidth. That gap explains the parallelism layout below.

The table lists 16 spines for three SUs as well as for four. A reasonable reading is that a pod is cabled for its final size, so the spine layer is built once. Whatever the reason, size the spine for the largest pod you expect to build. Adding spines later means recabling every leaf.

The Grace Blackwell SU

The Grace Blackwell version changes the building block. In the DGX SuperPOD with DGX GB200, an SU is eight DGX GB200 rack systems, 576 GPUs in total. Each rack is an NVL72: 72 GPUs in a single NVLink domain, connected through NVLink switch trays. The NVLink domain grows from 8 GPUs to 72, and the rack, not the node, becomes the unit of fast communication and the unit of failure. Racks are connected to each other through the InfiniBand scale-out fabric. NVIDIA's launch material named Quantum-X800 InfiniBand for it. The switch counts and rail layout are in the GB200 edition of the reference architecture. They are not repeated here, so read that edition before planning cabling.

For software, the rule is the same at a different scale. Keep the bandwidth-hungry parallelism inside the NVLink domain and the lighter traffic on InfiniBand. The NVL72 rack article covers the rack as a failure domain, and NVLink Switch in depth covers the switches.

Telling software about the topology

The scheduler cannot place jobs well unless it knows the topology. For Slurm's tree plugin, describe each SU as one switch, even though physically it contains eight rail leaves. Every node in an SU reaches every other in one hop on its own rail. That is the property the scheduler needs, and listing the eight leaves would only confuse it. This generator writes topology.conf from an inventory of SUs:

def topology_conf(sus, prefix="dgx", width=3):
    """sus: list of (first, last) node numbers per SU, for example [(1, 31), (32, 63)]."""
    lines = []
    for k, (a, b) in enumerate(sus, start=1):
        rng = f"{prefix}[{a:0{width}d}-{b:0{width}d}]"
        lines.append(f"SwitchName=su{k} Nodes={rng}")
    lines.append(f"SwitchName=spine Switches=su[1-{len(sus)}]")
    return "\n".join(lines) + "\n"

# 127 nodes: SU 1 has 31 (its 32nd slot went to UFM), the rest have 32
print(topology_conf([(1, 31), (32, 63), (64, 95), (96, 127)]))
# SwitchName=su1 Nodes=dgx[001-031]
# ...
# SwitchName=spine Switches=su[1-4]

Set TopologyPlugin=topology/tree in slurm.conf. A job can then ask to stay inside one SU with sbatch --switches=1@30:00. That means at most one leaf-level switch, and it is willing to wait up to 30 minutes for one. Check which node really gave up its slot in your build, because that depends on how the pod was cabled. Kubernetes clusters can do the same with node labels for SU and rack, plus a scheduler that understands topology.

NCCL finds the rails by itself. It reads the PCIe and NIC topology and pairs each GPU with its own NIC. Leave NCCL_IB_HCA unset unless there is a reason to set it, and confirm the pairing once with NCCL_DEBUG=INFO.

Worked example: placing a job

Suppose a team trains a model with tensor parallelism of 8, pipeline parallelism of 4 and data parallelism for the rest. Map each to the fabric that suits its traffic.

  • Tensor parallel (8) inside each node, on NVLink. It has the heaviest and most frequent traffic.
  • Pipeline parallel (4) across four nodes in the same SU. Point-to-point activations travel one leaf hop on matching rails.
  • Data parallel across the rest. One replica is 4 nodes, and the first SU has 31. Seven replicas use 28 nodes, so a 224-GPU job fits inside one SU and leaves three nodes free as hot spares. Eight replicas would need 32 nodes and spill onto the spine.
  • A 1,000-GPU job needs all four SUs. Its data-parallel all-reduce crosses the spine. Because the fabric is non-blocking, that costs latency, not bandwidth. Overlapping the gradient all-reduce with the backward pass hides most of it.

The general rule: fit the job to the SU first, then pick parallel degrees. A job that spills one node into a second SU pays spine latency for all of its traffic in order to gain a little extra capacity.

Rank order matters as much as node choice. Launchers number ranks node by node, so ranks 0 to 7 share a node and form one tensor-parallel group. Build the pipeline groups from consecutive nodes in the same SU, and the data-parallel groups from matching positions across replicas. Then every data-parallel peer of a GPU sits on the same rail, and its gradient traffic stays on one leaf when the job fits in an SU. Frameworks differ in the order in which they build these groups from the parallel degrees, and some make it configurable. Check the group membership the framework prints at start-up against the node list Slurm gave you, because a hostfile in the wrong order silently undoes the placement.

Failure modes

  • Miscabled rails. A node with rails 3 and 4 swapped still passes ping tests, but its traffic crosses the spine and slows every ring it joins. Check cabling with UFM or ibnetdiscover against the plan.
  • Topology the scheduler does not know. Without topology.conf, Slurm scatters jobs across SUs, and step time varies with placement.
  • Storage on the compute fabric. When checkpoint traffic is routed over the compute fabric, every checkpoint makes the training step slower.
  • A lost management network. If the in-band network fails, jobs cannot launch even though the GPUs are healthy. If the OOB network fails, you cannot power-cycle a hung node.
  • Power-of-two assumptions. Job sizes and launch scripts that expect 128 nodes fail on a 127-node pod.
  • One slow link. One flapping cable or link with a high bit error rate slows every collective that uses it. Track error counters for each port, not just link state.

Trade-offs

The SuperPOD design buys predictability. It is non-blocking, rail-aligned, uses identical SUs and keeps its networks separate. The cost is optics, switches and the cabling work for four networks. Some operators choose an oversubscribed spine, rail-only designs or Ethernet fabrics to save money. Those can work well for jobs that fit in one SU, or for inference. The penalty falls on large data-parallel jobs, which is exactly the workload a SuperPOD is bought for. InfiniBand NDR and XDR in depth covers the link-level trade-offs if you are pricing a next-generation fabric.

What to do next

  1. Get the reference architecture edition for your exact DGX generation, and copy its component table into your capacity plan.
  2. Record which node slot the pod gave up for UFM. Generate topology.conf from that inventory, never by hand.
  3. Burn in each SU on its own with nccl-tests, then test across SUs, and keep both baselines.
  4. Pick job sizes that fit inside one SU, and keep a few spare nodes per SU.
  5. Check that storage traffic stays on the storage fabric, by watching port counters during a checkpoint.
  6. Alert on per-port error counters from UFM, and run a cabling audit after every hardware change.
Key takeaway: A DGX SuperPOD is built from 32-node scalable units, and the pod gives one node slot to UFM, so four SUs make 127 nodes. The rail-aligned compute fabric is non-blocking; storage and two management networks stay separate. Describe each SU to Slurm as one switch, fit jobs inside an SU, and audit cabling and port errors continuously.