When people compare GPU clusters they quote a bandwidth number per GPU, and it is frequently the wrong one. A GPU in a modern training node has at least three different bandwidths attached to it: the scale-up link to its neighbours (NVLink), the scale-out link into the network (its NIC), and its share of the network's bisection bandwidth, which is what remains when every GPU talks across the middle of the cluster at once. Collective communication libraries such as NCCL hit different ones depending on the operation, and a training job slows down whenever the one it hits is smaller than you assumed.

This article defines bisection bandwidth precisely, shows how to turn a topology diagram into a per-GPU number, explains which collectives are bound by it and which are not, and gives you a small model you can run against your own fabric. It closes with how to measure the real figure, the failure modes that make the measured figure lower than the spec sheet, and the decisions that depend on it: oversubscription, job placement and how you map tensor, expert and data parallelism onto the hardware.

Defining bisection bandwidth

Take a network of N endpoints. Split the endpoints into two halves of N/2 each, and add up the capacity of every link that connects one half to the other. Do this for every possible balanced split; the smallest total is the bisection bandwidth. It is a worst case by construction: it tells you how much traffic can cross the narrowest middle of the network if the traffic pattern is as unkind as possible.

To turn it into a per-GPU figure, this article uses one convention throughout: per-GPU bisection bandwidth = bisection bandwidth / (N/2), measured in one direction. With that convention a network has full bisection when the per-GPU figure equals the per-GPU injection bandwidth: every GPU in one half can stream to a partner in the other half at its NIC's full line rate simultaneously. If the per-GPU figure is half the injection rate, the network is 2:1 oversubscribed at its narrowest point.

Two unit traps cause most of the confusion. Network gear is rated in gigabits per second (Gb/s) per direction, while GPU interconnects are usually quoted in gigabytes per second (GB/s) summed over both directions. A 400 Gb/s NIC moves 50 GB/s each way. An H100 SXM's NVLink figure of 900 GB/s is a bidirectional total, so roughly 450 GB/s in each direction. Convert everything to GB/s per direction before you compare anything.

Three bandwidths every GPU has

The table puts the three bandwidths side by side for two well-known systems. The NVLink figures are NVIDIA's published numbers; the scale-out column assumes the common one-400-Gb/s-NIC-per-GPU design, which is a cluster choice rather than a property of the GPU.

BandwidthWhat it connectsH100 SXM node (8 GPUs)GB200 NVL72 rack
Scale-up (NVLink)GPU to GPU inside the NVLink domain900 GB/s bidirectional per GPU, ~450 GB/s each way, domain of 81.8 TB/s bidirectional per GPU, 130 TB/s aggregate, domain of 72
Injection (NIC)One GPU into the scale-out fabric400 Gb/s = 50 GB/s each way if one NIC per GPUdepends on the NICs fitted; check the build
Per-GPU bisectionOne GPU's share of the narrowest cut50 GB/s if non-blocking; 25 GB/s at 2:1set by the scale-out fabric, not the rack

The ratio between the first row and the other two is the reason parallelism mapping matters. Inside the NVLink domain a GPU has roughly nine times the per-direction bandwidth of its NIC on the H100 design. Anything you can keep inside that domain is cheap; anything that must cross the fabric is paid for at the injection rate at best and at the bisection rate at worst. See NVLink Switch, in depth for how the scale-up domain is built and GPU Pod Network for how scale-up and scale-out domains fit together.

From topology to a per-GPU number

Most AI clusters use a two- or three-tier Clos (leaf-spine or fat-tree) network. A leaf switch has some ports facing down to NICs and some facing up to spines. If a 64-port switch uses 32 ports down and 32 up, every byte that enters the leaf from below can leave it upward at full rate: the leaf is non-blocking. If it uses 32 down and 16 up, the uplinks carry half the downlink capacity and the leaf is 2:1 oversubscribed. With a non-blocking spine layer, the per-GPU bisection bandwidth is simply the injection bandwidth divided by the leaf oversubscription ratio.

Rail-optimised fabrics add a twist. GPU r of every node connects to a leaf dedicated to rail r, so each rail is its own small Clos. Traffic between GPUs with the same local index stays on one rail. Traffic between different indices is first moved across NVLink inside the sending node to the GPU on the right rail, then sent over that rail; NCCL calls this PXN. The effect is that the fabric only ever carries same-rail traffic, which is why rail-aligned topology pairs so well with ring and tree collectives.

Where a balanced cut falls in a two-tier rail fabricSpine 1Spine 2Leaf 132 down / U upLeaf 232 down / U upLeaf 332 down / U upLeaf 432 down / U upNodes 1-328 GPUs + NVLink eachNICNodes 33-648 GPUs + NVLink eachNICNodes 65-968 GPUs + NVLink eachNICNodes 97-1288 GPUs + NVLink eachNICBisection cut: half the GPUs on each side; every byte crossing it rides a leaf uplink and the spineLeaf-uplink cut: the same question asked of one leaf; with oversubscription it binds first
Two cuts that matter. The bisection cut is the textbook definition; the leaf-uplink cut is what an oversubscribed leaf-spine actually runs out of first.

Which collectives are bisection-bound

Whether a collective cares about bisection depends on its traffic pattern, not on how much data it moves.

Ring all-reduce. Each GPU sends 2(N-1)/N times the buffer size, but only to its ring neighbour. If the ring is laid out so neighbours share a leaf or a rail, very little traffic crosses the middle of the network and the collective is bound by injection bandwidth (or NVLink, inside a node). nccl-tests reports this as bus bandwidth: busbw = algbw x 2(n-1)/n, where algbw is buffer size divided by time. Busbw is the number to compare against link speed. Tree and in-network reductions such as SHARP change the constant but not the conclusion: well-placed all-reduce is rarely bisection-bound.

All-to-all. Mixture-of-experts layers send each token to the GPU that holds its expert, so every GPU sends a slice of its buffer to every other GPU. For nccl-tests alltoall the bus factor is (n-1)/n. Here placement cannot help: whatever cut you draw, a fixed fraction of the traffic crosses it. Uniform all-to-all is the canonical bisection-bound workload, and in a leaf-spine the leaf-uplink cut binds before the bisection cut, because nearly all of a leaf's traffic is headed elsewhere.

Point-to-point. Pipeline parallel sends move activations between adjacent stages. They are injection-bound if stages sit on neighbouring nodes and can become bisection-bound if a scheduler scatters them across the cluster.

A detailed treatment of the collectives themselves is in NCCL collectives.

A model you can run

The model below gives a lower bound for a uniform all-to-all on a rail-optimised two-tier fabric. It ignores latency, NVLink staging time and congestion, all of which only make things slower, so treat its output as a floor rather than a prediction.

from dataclasses import dataclass


@dataclass
class Fabric:
    gpus: int              # N, total GPUs in the job
    gpus_per_node: int     # GPUs that share one node's NVLink domain (rail count)
    nic_gbps: float        # NIC line rate per GPU, Gb/s, one direction
    nodes_per_leaf: int    # nodes whose rail-r NIC lands on the same rail-r leaf
    oversub: float = 1.0   # leaf downlink:uplink capacity ratio, 1.0 = non-blocking

    @property
    def inj(self):         # injection bandwidth per GPU, GB/s, one direction
        return self.nic_gbps / 8

    def bisection_per_gpu(self):
        return self.inj / self.oversub


def alltoall_ms(f, bytes_per_gpu):
    """Lower bound for a uniform all-to-all on a rail-optimised two-tier fabric with PXN."""
    nodes = f.gpus // f.gpus_per_node
    off_node = (f.gpus - f.gpus_per_node) / f.gpus           # share that must use the NIC
    off_leaf = (nodes - f.nodes_per_leaf) / (nodes - 1)       # share of that leaving the leaf
    sent = bytes_per_gpu * off_node
    t_nic = sent / (f.inj * 1e9)
    t_up = sent * off_leaf / (f.bisection_per_gpu() * 1e9)
    return 1e3 * max(t_nic, t_up), "NIC injection" if t_nic >= t_up else "leaf uplinks"


for r in (1.0, 2.0):
    f = Fabric(gpus=1024, gpus_per_node=8, nic_gbps=400, nodes_per_leaf=32, oversub=r)
    ms, limit = alltoall_ms(f, 256e6)
    print(f"{r:.0f}:1  per-GPU bisection {f.bisection_per_gpu():5.1f} GB/s  "
          f"all-to-all {ms:5.2f} ms  bound by {limit}")

Running it prints:

1:1  per-GPU bisection  50.0 GB/s  all-to-all  5.08 ms  bound by NIC injection
2:1  per-GPU bisection  25.0 GB/s  all-to-all  7.68 ms  bound by leaf uplinks

Worked example: a 1,024-GPU MoE cluster

Walk through the numbers for that 1,024-GPU cluster: 128 nodes of eight GPUs, one 400 Gb/s NIC per GPU, eight rails, and 64-port leaf switches so each rail has four leaves of 32 nodes. Each MoE layer's dispatch sends 256 MB from every GPU.

Injection is 50 GB/s per direction. With PXN, 1/128 of each GPU's data stays inside its node and 127/128 goes out through a NIC. Of a NIC's traffic, the destinations on its own leaf are 31 of the 127 other nodes, so 96/127, about 75%, must climb to the spine.

At 1:1 the leaf uplinks offer 50 GB/s per GPU but only 75% of the traffic needs them, so the NIC is the bottleneck: 254 MB at 50 GB/s is 5.08 ms. At 2:1 the uplinks offer 25 GB/s per GPU, and 75% of 254 MB at 25 GB/s is 7.68 ms. Halving the spine saved switch and optics cost but made every expert-parallel dispatch at least 51% slower, and an MoE layer does a dispatch and a combine, forward and backward.

Run the same arithmetic for a dense model that only does all-reduce on well-placed rings and the 2:1 fabric costs almost nothing. That asymmetry is the whole design decision: the right per-GPU bisection figure is a property of your workload mix, not of the hardware.

Measuring the real figure

Spec sheets give line rates; you need delivered bandwidth. The standard tool is NVIDIA's open-source nccl-tests. Build it against your NCCL and MPI, then sweep message sizes:

# all-reduce and all-to-all across 16 nodes, one process per GPU
mpirun -np 128 -N 8 ./build/all_reduce_perf -b 8M -e 8G -f 2 -g 1
mpirun -np 128 -N 8 ./build/alltoall_perf   -b 8M -e 8G -f 2 -g 1

# read the busbw column at the largest sizes; compare with:
#   all-reduce  -> injection (or NVLink) rate per direction
#   all-to-all  -> min(injection, per-GPU bisection) per direction

For the bisection figure itself, pair every GPU in one half of the allocation with a partner in the other half and run point-to-point bandwidth tests (perftest's ib_write_bw, or NCCL send/receive) on all pairs at the same time. Sum the results and divide by N/2. Running one pair at a time measures injection bandwidth and will flatter an oversubscribed fabric every time. Repeat with several different pairings, because the worst one is the number that matters, and record per-pair results so a single bad cable shows up as an outlier instead of a lower average.

Failure modes

  • Unit and direction mix-ups. Comparing a bidirectional NVLink figure with a per-direction NIC figure overstates the gap by two; comparing Gb/s with GB/s understates it by eight.
  • ECMP hash collisions. Static flow hashing can put two large flows on the same uplink while another sits idle, so delivered bisection falls below nominal even at 1:1. Adaptive routing and NCCL's use of multiple queue pairs per connection reduce this; measure with all pairs active to see it.
  • Scattered placement. A scheduler that fills gaps across many leaves turns rail-local rings into cross-spine traffic. Topology-aware placement is a bandwidth feature.
  • Degraded links. One link that has fallen back to a lower speed, or a flapping optic, drags every ring that crosses it down to its rate. Collectives run at the speed of the slowest participant.
  • PCIe affinity. If a GPU's traffic reaches its NIC through the CPU interconnect instead of a shared PCIe switch, GPUDirect RDMA loses bandwidth before the network is even involved.
  • Shared fabrics. Another tenant's all-to-all consumes the same spine. Per-GPU bisection on a shared cluster is a statistical figure, not a guarantee.

Trade-offs

Full bisection is expensive: the spine layer and its optics can approach the cost of the leaf layer, and a three-tier network adds a further layer of both. Tapering to 2:1 or more is reasonable when the workload is dominated by data-parallel all-reduce and pipeline sends that placement can keep local. It is a poor bet for clusters that will run large mixture-of-experts models, many small jobs with unpredictable placement, or inference with disaggregated prefill and decode, all of which generate cross-cluster traffic.

The other lever is the scale-up domain. A 72-GPU NVLink domain lets tensor parallelism and much of expert parallelism stay off the scale-out fabric entirely, which lowers the per-GPU bisection you need from the network. The rule of thumb that follows: map the most bandwidth-hungry parallelism (tensor, then expert) inside the NVLink domain, pipeline across neighbouring nodes, and data parallel across the fabric where the all-reduce can be overlapped with compute.

What to do next

  1. Write down three numbers for your cluster in GB/s per direction: NVLink per GPU, NIC per GPU, and per-GPU bisection from the leaf oversubscription ratio.
  2. Run nccl-tests all_reduce_perf and alltoall_perf at your job's real scale and compare busbw with those numbers; a gap over about 20% is worth chasing.
  3. Measure bisection with all cross-half pairs active at once, in several pairings, and keep the per-pair results.
  4. Profile one training step and find what fraction of time is all-to-all versus all-reduce.
  5. Plug your fabric and message sizes into the model above to see which cut binds.
  6. Ask your scheduler team whether jobs are placed by rail and leaf, and check one real job.
  7. If MoE is on the roadmap, size the spine for it now; re-cabling later costs far more.
Key takeaway: Per-GPU bisection bandwidth is the share of the network's narrowest cut that each GPU gets when everyone talks across the middle at once. It equals the NIC rate on a non-blocking fabric and falls with oversubscription. Ring all-reduce on well-placed GPUs barely notices it; all-to-all for mixture-of-experts is bound by it. Convert every figure to GB/s per direction, measure with all pairs active, and size the fabric for the traffic pattern you will actually run.