NVIDIA Spectrum-X is NVIDIA's Ethernet platform for AI clusters. It is not a new protocol. GPUs still talk RoCE v2, the switches still forward Ethernet frames, and NCCL still sees RDMA verbs. What changes is that the switch and the NIC are designed as one system: Spectrum switches spread the packets of a single flow across every available path, and NVIDIA SuperNICs (BlueField-3, and ConnectX-8 in newer systems) put those packets back in order, run the congestion control, and keep tenants from slowing each other down.

This article explains that design from the problem it solves, through the mechanics, to what you configure, measure and debug. It assumes you know what RoCE is; the lossless-Ethernet plumbing underneath (PFC, ECN, DCQCN, GID selection) is covered in RoCE v2, in depth. Performance figures quoted here are NVIDIA's own claims, checked against NVIDIA's product page and developer blog in October 2026; treat them as the vendor's numbers until you have measured your own.

The problem: ECMP and elephant flows

Standard Ethernet data centres balance load with ECMP: the switch hashes each packet's five-tuple and sends every packet of a flow down the same uplink. That works for web traffic, which is millions of small, independent flows whose hashes average out. Training traffic is the opposite. A data-parallel all-reduce is a handful of very large, long-lived flows per GPU, all starting at the same moment, and the step finishes only when the slowest flow finishes.

Worked example. A leaf switch has 8 uplinks to 8 spines and 8 GPUs below it each send one elephant flow up. With a perfect assignment every flow gets a full link. With random hashing, the chance that all 8 land on different uplinks is 8!/8^8, about 0.24 percent. On average only 5.25 of the 8 uplinks carry anything, the busiest one carries 2.6 flows, and half the time it carries 3 or more. The flows on that link run at a third of line rate and the whole collective waits for them.

import random, statistics

def busiest_link(flows, links):
    load = [0] * links
    for _ in range(flows):
        load[random.randrange(links)] += 1      # ECMP: one hash, one path
    return max(load)

trials = [busiest_link(8, 8) for _ in range(200_000)]
print(statistics.mean(trials))                  # about 2.6 flows on the worst link
print(1 / statistics.mean(trials))              # about 38% of line rate for the collective

# More queue pairs per connection = more, smaller flows to hash
trials = [busiest_link(32, 8) for _ in range(200_000)]
print(4 / statistics.mean(trials))              # about 57%: better, still far from 100%

That simulation is the whole case for Spectrum-X in a dozen lines. Adding entropy, for example more queue pairs per connection, helps but leaves a large gap. Closing the gap requires deciding the path per packet, using the actual state of the queues, rather than per flow using a hash. How rails and fat-tree tiers shape the number of paths is in Fat-Tree Network Topology, in depth.

The parts of the platform

The platform has three layers, and the benefit depends on all three being present.

LayerComponentRole
SwitchSpectrum-4 ASIC, for example the SN5600: 51.2 Tb/s, 64 ports of 800GPer-packet adaptive routing, in-band telemetry, shared-buffer and QoS
Host NICBlueField-3 SuperNIC (400G RoCE); ConnectX-8 SuperNIC in newer systemsReordering and direct placement, congestion control, GPUDirect RDMA
SoftwareCumulus Linux or SONiC on switches, NVIDIA NetQ, DOCA and NIC firmware on hostsConfiguration, validation, telemetry and flow analysis

Newer Spectrum generations and co-packaged optics versions have been announced since the SN5600; check the current datasheets for port counts rather than relying on any article, including this one. The structure described below is the same across them.

Spectrum-X: per-packet spraying in the fabric, reordering at the receiving SuperNICGPU 0 + NCCLsends one flowSuperNIC (send)rate set by CCLeaf Apicks least-loadedSpine 1Spine 2Spine 3pkt 1,4pkt 2,5pkt 3,6Leaf BSuperNIC (recv)places by address2,1,3,5,4,6GPU 7 memorycomplete when all landtelemetry and congestion signals feed the sender's rate controlSwitches report queue state to each other; the NIC hides out-of-order arrival from NCCL.
One flow from GPU 0 to GPU 7 in a two-tier fabric. Leaf A sends consecutive packets to whichever spine uplink has the shortest queue; they arrive out of order and the receiving SuperNIC writes each into its final place in GPU memory. Telemetry returns to the sender's NIC, which sets the rate.

Per-packet adaptive routing

According to NVIDIA's description, a Spectrum-4 switch makes the routing decision for every packet: among the eligible egress ports toward the destination, it picks the one whose egress queue has the least load, and it also takes into account status notifications from neighbouring switches, so a leaf can avoid a spine whose downstream link is congested even when its own uplink queue to that spine looks fine. NVIDIA calls this RoCE adaptive routing.

Spreading at packet granularity makes the fabric behave much more like one big pipe. In the worked example, 8 flows across 8 uplinks become a stream of packets that fill all 8 links evenly, and the collective runs near line rate. NVIDIA's claim is up to 95 percent effective bandwidth at scale and 1.6 times the AI network performance of off-the-shelf Ethernet. The second number depends heavily on the baseline: the gap is largest for exactly the low-entropy, synchronous traffic in the example, and smallest for many-flow inference or storage traffic that ECMP already balances well.

In NVIDIA's design adaptive routing targets the RoCE traffic that the SuperNICs can reorder. Ordinary TCP traffic should stay on flow-consistent ECMP, because TCP stacks treat heavy reordering as loss and slow down.

Out-of-order arrival and direct placement

Per-packet spraying means packets of one message arrive out of order. Ordinary RoCE NICs handle that badly: the RoCE transport expects packet sequence numbers in order, and a gap triggers a NAK and go-back-N retransmission of everything after the gap, as described in RoCE (RDMA over Ethernet), in depth. Spraying onto such a NIC would turn every path difference into a retransmit storm.

The SuperNIC is built to accept it. NVIDIA states that the SuperNIC reorders arriving packets and places them in host or GPU memory so that the reordering is invisible to the application, and that completion is reported only once the data is in place. NVIDIA does not publish the per-packet mechanism; note that in standard RoCE only the first packet of an RDMA write carries the remote address, so this is a capability of the NIC, not of the protocol. What matters for software is the result: NCCL never sees an out-of-order packet, only a completed write.

The consequence for your fleet: every endpoint that receives sprayed traffic must be a NIC that supports this. Mix in a third-party or older NIC and either adaptive routing must be off for traffic to it, or that endpoint suffers retransmissions.

Congestion control on the SuperNIC

Adaptive routing fixes collisions inside the fabric. It cannot fix incast, where many senders target one receiver and the last-hop link is simply oversubscribed; there is only one path to the destination port. That needs senders to slow down, which is congestion control.

Classic RoCE uses DCQCN: switches mark packets with ECN when a queue passes a threshold, the receiver echoes congestion notification packets, and the sender's NIC cuts its rate. Spectrum-X replaces the signal with richer in-band telemetry from the switches, including queue information for estimating congestion and port utilisation for recovering rate quickly, and runs the control algorithm on the SuperNIC itself. NVIDIA says the NIC handles millions of congestion events per second with microsecond reaction latency. The practical effect is shallower queues during incast and faster ramp back to full rate afterwards, which is where tail latency and step-time jitter come from.

Lossless behaviour, through PFC on the RoCE traffic class, normally stays configured as a backstop in reference designs. Follow the reference architecture for your release on whether and where PFC is enabled; it is not something to improvise.

Performance isolation between tenants

A shared AI cloud runs many jobs on one fabric. Without isolation, one tenant's all-to-all can fill spine queues and slow every other job's all-reduce: the noisy neighbour problem. Spectrum-X addresses it with three mechanisms together: QoS separation between traffic classes, adaptive routing so one job's traffic is spread rather than concentrated on a few links, and per-flow congestion control so a congesting sender is slowed at its own NIC before it fills shared buffers. NVIDIA's stated goal is that no workload can create congestion that affects another's data movement.

Isolation still depends on placement. A job split across many leaves crosses the spine layer and competes there; a job packed under few leaves mostly does not. Mapping parallelism to rails and leaves is covered in GPU Pod Network, in depth.

What NCCL sees and what you set

From NCCL's point of view a Spectrum-X fabric is a RoCE network, and NCCL's IB verbs transport drives it. The variables you meet are the RoCE ones, all documented in the NCCL environment-variable reference.

# Typical things to pin on a RoCE fabric, Spectrum-X included.
export NCCL_IB_HCA=mlx5_0,mlx5_1,mlx5_2,mlx5_3   # which RDMA devices NCCL may use
export NCCL_IB_GID_INDEX=3       # RoCE v2 GID; recent NCCL defaults to automatic selection
export NCCL_IB_TC=106            # traffic class: must match the fabric's RoCE QoS policy
export NCCL_DEBUG=INFO           # log which devices, GIDs and transports were chosen

# Adaptive routing for the IB verbs transport: NCCL enables it by default on
# InfiniBand and disables it by default on RoCE. Whether to set it on your
# Spectrum-X release is a reference-architecture decision, not a guess.
# export NCCL_IB_ADAPTIVE_ROUTING=1

# Queue pairs per connection (default 1). Extra QPs add hash entropy on
# ECMP fabrics; with per-packet spraying the benefit is smaller.
# export NCCL_IB_QPS_PER_CONNECTION=4

The GID index and traffic-class values above are examples; the right values come from your NIC configuration and the fabric's QoS policy. Read the NCCL INFO log on the first run of every new cluster and confirm the devices and GID it chose. For communicator setup, topology search and NCCL's own failure handling, see NCCL, in depth.

Bringing a fabric into service

Bring a Spectrum-X cluster into service the same way as any RDMA fabric, in layers, and keep the results as a baseline.

  1. Firmware and OS: one approved NIC firmware, DOCA or driver version and switch OS version across the fleet. Mixed versions are the most common cause of features silently not engaging.
  2. Link layer: every port up at the expected speed with clean FEC and error counters; a single degraded link drags every collective that touches it.
  3. Pairwise RDMA: perftest ib_write_bw between hosts on different leaves should reach near line rate on every NIC.
  4. Collectives: nccl-tests all_reduce_perf and alltoall_perf at full scale; record bus bandwidth per message size. Run them with adaptive routing on and off once, so you know what the feature is buying on your topology.
  5. Contention: run two jobs that cross the spine at the same time and compare against each alone. This is the test of the isolation claim that matters to you.
  6. Telemetry: use NetQ or your switch telemetry pipeline to watch queue depth, drops, PFC pause frames and link utilisation, and alert on deviation from the baseline.

Failure modes

  • Collective bandwidth stuck near ECMP levels. Adaptive routing is not enabled for the traffic class NCCL actually uses. Check that NCCL_IB_TC and the switch QoS mapping agree.
  • Retransmission counters climbing on some hosts. Those endpoints cannot absorb out-of-order arrival: wrong NIC model, old firmware, or the feature is off on the NIC while on in the fabric.
  • PFC pause storms. Congestion control is not engaging, so the fabric falls back on pausing. Look for a misconfigured congestion-control profile or a host whose NIC settings drifted.
  • One slow rank. A flapping or high-error optic. Because collectives are synchronous, one bad link shows up as whole-job slowdown; per-port error rates find it faster than job profiling.
  • Storage or checkpoint traffic hurting training. Put storage on its own traffic class or its own rail, and check that it does not share the training class's buffers.
  • Good benchmarks, poor real jobs. nccl-tests run one job on an idle fabric; repeat the contention test with production placement.

Trade-offs

ChoiceGainCost
Spectrum-X vs standard RoCE fabricNear line-rate collectives without hand-tuning entropy; better incast behaviourNVIDIA switches and SuperNICs end to end to get the benefit
Spectrum-X vs InfiniBandEthernet operations, tooling and multi-tenancy modelInfiniBand has a longer record for AI fabrics and in-network reduction (SHARP)
Spectrum-X vs open Ultra Ethernet designsShipping, integrated product todayTies the fabric to one vendor's NIC and switch roadmap
Per-packet spraying vs flow ECMPBalanced links, higher effective bandwidthRequires reorder-capable receivers; TCP classes must stay on ECMP

Make the decision on measured job throughput for your model mix, not on peak link speed. For the InfiniBand side of that comparison see InfiniBand NDR and XDR, in depth.

What to do next

  1. Run the ECMP simulation above with your own leaf uplink count and flows per GPU to estimate what flow hashing is costing you today.
  2. Measure all_reduce_perf bus bandwidth on your current fabric as a baseline before any change.
  3. Inventory NICs, firmware and switch OS versions; list every endpoint that could not accept out-of-order delivery.
  4. Obtain NVIDIA's reference architecture for your generation and record its traffic class, PFC and congestion-control settings in configuration management.
  5. Pin NCCL_IB_HCA, the GID index and the traffic class explicitly, and read the NCCL INFO log on first run.
  6. Run the on/off adaptive routing and two-job contention tests and keep the numbers.
  7. Alert on queue depth, PFC pauses, retransmissions and per-port error rates against that baseline.
Key takeaway: Spectrum-X keeps RoCE but moves load balancing from per-flow hashing to per-packet decisions based on queue state, and makes the SuperNIC absorb the resulting reordering and run congestion control. The gain is largest for synchronous, low-entropy collectives, it requires reorder-capable NICs at every endpoint, and it should be proven with your own before-and-after measurements.