Large training jobs spend a measurable share of every step moving gradients and activations between nodes. Inside a node, NVLink does that; between nodes, the traffic leaves through a network adapter, and on a growing share of clusters that adapter speaks RoCE v2: RDMA over Converged Ethernet, version 2. The appeal is obvious: Ethernet switches, Ethernet operations teams and Ethernet routing, with the same verbs programming model and GPUDirect path that InfiniBand offers. The catch is that RDMA transports were designed for a network that almost never drops packets, and Ethernet is not that network unless you make it one.
This article assumes the RDMA and GPUDirect basics covered in InfiniBand for GPU clusters and the GPUDirect family, and focuses on what is specific to Ethernet: the packet format, why it routes, how congestion is controlled, how NCCL picks the right address, how to prove the fabric is healthy, and what goes wrong in production.
From verbs to a UDP packet
RDMA applications, including NCCL's network plugin, talk to the adapter through the verbs API: they register memory, create queue pairs (QPs) and post send, write and read work requests. The adapter's transport engine segments a message into packets, numbers them with a packet sequence number (PSN), and expects acknowledgements, all without the host CPU. That transport and its Base Transport Header (BTH) come from InfiniBand.
RoCE v1 carried that transport directly inside an Ethernet frame with its own EtherType, which confined it to one layer-2 domain. RoCE v2 puts the BTH inside UDP over IP, with UDP destination port 4791. Because the outer headers are ordinary IPv4 or IPv6, the traffic can cross routers, which is what makes a large leaf-spine fabric possible. The UDP source port is not a real port: the adapter derives it from the queue pair, so different QPs hash onto different equal-cost paths, while each QP's packets stay on one path and arrive in order.
The header overhead is small. With IPv4, each packet carries 14 bytes of Ethernet, 20 of IP, 8 of UDP, 12 of BTH, a 4-byte invariant CRC and a 4-byte frame check, 62 bytes in total, plus 20 bytes of preamble and inter-frame gap on the wire. With a 9000-byte Ethernet MTU the RDMA path MTU is usually set to 4096, so a full packet is about 98 percent payload.
GIDs: which address does the QP use?
Every RDMA port has a table of Global Identifiers. On Ethernet, the table holds one entry per IP address on the interface per RoCE version: an IPv6 link-local entry, the IPv4-mapped entry, and so on, each listed for both v1 and v2. A connection uses a GID index, and picking the wrong one is the classic RoCE misconfiguration: v1 entries do not route, and an entry for the wrong address family or VLAN interface sends traffic somewhere unexpected.
Current NCCL releases select the GID automatically: the documentation lists NCCL_IB_GID_INDEX as defaulting to AUTO since 2.32u1, choosing an entry whose RoCE version matches NCCL_IB_ROCE_VERSION_NUM (default 2) and whose address family matches NCCL_IB_ADDR_FAMILY (default AF_INET), optionally restricted by NCCL_IB_ADDR_RANGE. Older releases documented a default of -1, and many cluster recipes hard-code an index read from the GID table. Read the table on your own hosts rather than copying a number from someone else's recipe.
# GID table for one port; the type column says RoCE v1 or v2.
# show_gids ships with NVIDIA's OFED tools; the sysfs files below work everywhere.
for i in /sys/class/infiniband/mlx5_0/ports/1/gid_attrs/types/*; do
idx=$(basename "$i"); t=$(cat "$i" 2>/dev/null) || continue
gid=$(cat /sys/class/infiniband/mlx5_0/ports/1/gids/$idx)
[ "$gid" != "0000:0000:0000:0000:0000:0000:0000:0000" ] && echo "$idx $t $gid"
done
Lossless, lossy, and why it matters
The RDMA reliable-connection transport recovers lost packets, but historically not gracefully: many adapters used go-back-N, so one dropped packet causes everything after it on that QP to be resent, and a lost tail packet waits for a transport timeout. NCCL's NCCL_IB_TIMEOUT defaults to 20, and the verbs timeout is 4.096 microseconds times 2 to that power, about 4.3 seconds per retry, with NCCL_IB_RETRY_CNT defaulting to 7. A collective proceeds at the speed of its slowest participant, so a single stalled QP stalls the whole job.
That is why classic RoCE deployments make Ethernet lossless for the RDMA traffic class. Priority Flow Control (IEEE 802.1Qbb) lets a switch port tell its upstream neighbour to pause one of eight priorities when its buffer for that priority fills, so packets wait instead of being dropped. Newer adapters have better loss recovery, and some large operators run RoCE with PFC disabled and rely on congestion control plus fast retransmission. Which approach your hardware supports well is a question for your NIC and switch vendor, and the answer should come with test results on your topology.
ECN and DCQCN: the loop that should do most of the work
PFC is a blunt instrument. A pause stops every flow in that priority on the link, including flows that are not causing congestion, and pauses propagate hop by hop back towards the senders. A fabric that relies on PFC for routine congestion has head-of-line blocking, victim flows, and in the worst case deadlock when buffer dependencies form a cycle.
Congestion control is supposed to keep queues short enough that PFC rarely fires. The widely deployed scheme is DCQCN, built on Explicit Congestion Notification. Senders mark packets ECN-capable. When a switch queue grows past a configured threshold, the switch marks packets Congestion Experienced, with probability rising as the queue grows. The receiving adapter, the notification point, sees marked packets and sends a Congestion Notification Packet (CNP) back to the sender for that QP. The sending adapter, the reaction point, cuts that QP's rate multiplicatively, then recovers it in steps over time if no more CNPs arrive.
The tuning rule follows from the design: ECN marking thresholds must sit well below the PFC trigger, so that rate reduction starts long before buffers are full. If PFC fires before ECN, the fabric behaves like a lossless network with no congestion control. The CNPs themselves should travel in a high, strict priority so they are not stuck behind the congestion they report.
Classifying the traffic: DSCP, priorities and NCCL
For any of this to work, the host, the adapter and every switch must agree which packets are RDMA. Modern designs classify on the DSCP field in the IP header, which survives routing, rather than on VLAN priority bits, which do not. The adapter is set to trust DSCP, the RDMA traffic class maps to one priority with PFC enabled (if you run lossless), CNPs map to a higher one, and ordinary TCP traffic stays lossy in its own class.
NCCL sets the traffic class on its QPs through NCCL_IB_TC. With the current AUTO default it uses 0, which only works if the adapter applies a default DSCP for RoCE traffic. The DSCP is the upper six bits of the traffic class byte, so DSCP 26 corresponds to a traffic class of 104. The example below shows the shape of a host configuration on NVIDIA ConnectX adapters using mlnx_qos; tool names and options differ for other vendors, and the switch side must match.
# Host side, per RoCE interface (NVIDIA ConnectX example; verify against your vendor docs)
mlnx_qos -i ens1f0np0 --trust dscp
mlnx_qos -i ens1f0np0 --pfc 0,0,0,1,0,0,0,0 # lossless priority 3 only
# NCCL: mark RDMA traffic with DSCP 26 (traffic class = 26 << 2 = 104)
export NCCL_IB_TC=104
export NCCL_IB_HCA=mlx5_0,mlx5_1,mlx5_2,mlx5_3,mlx5_4,mlx5_5,mlx5_6,mlx5_7
export NCCL_SOCKET_IFNAME=eth0 # bootstrap over the management NIC
# Leave NCCL_IB_GID_INDEX unset on current NCCL; set it from the GID table on old releases.
Proving the fabric before training on it
Never let a training job be the first workload on a new RoCE fabric. Test in three layers. First, point to point with perftest: ib_write_bw between two hosts on every rail should come close to line rate for large messages, and ib_write_lat should show low single-digit microseconds. Second, collectives with nccl-tests across all nodes, reading the bus bandwidth column, which normalises for algorithm so that numbers are comparable across node counts (see NCCL collectives). Third, the same collectives with an incast or background load running, which is when congestion control is actually exercised.
# 1. per-rail bandwidth (server, then client)
GID_IDX=... # the RoCE v2 entry for this interface, from the GID-table loop above
ib_write_bw -d mlx5_0 -x $GID_IDX -q 4 --report_gbits -s 1048576 -D 10
ib_write_bw -d mlx5_0 -x $GID_IDX -q 4 --report_gbits -s 1048576 -D 10 <server-ip>
# -q runs several QPs so ECMP spreads them across paths
# 2. collectives across 16 nodes x 8 GPUs (nccl-tests)
mpirun -np 128 -N 8 --hostfile hosts ./build/all_reduce_perf -b 8M -e 4G -f 2 -g 1
# 3. watch congestion counters while 2 runs (names vary by driver and firmware)
watch -n1 'for f in np_cnp_sent rp_cnp_handled np_ecn_marked_roce_packets out_of_sequence; do
printf "%s %s\n" $f $(cat /sys/class/infiniband/mlx5_0/ports/1/hw_counters/$f); done;
ethtool -S ens1f0np0 | grep -E "prio3_pause|discard"'Counter names in the example are from the mlx5 driver and differ between driver and firmware versions; list /sys/class/infiniband/<dev>/ports/1/hw_counters and ethtool -S output on your hosts to find the equivalents. What you want to see under load: CNPs sent and handled rising, ECN marks rising, out-of-sequence and discard counters flat, and pause frames rare. Many pauses with few CNPs means ECN thresholds are too high or ECN is not enabled somewhere on the path.
Worked example: what a sick link costs a training step
Take 16 nodes with eight GPUs each and one 400 Gb/s RoCE adapter per GPU in a rail-optimised leaf-spine fabric. A data-parallel all-reduce of S bytes over n ranks moves about 2(n-1)/n times S per rank, so with nccl-tests reporting 40 GB/s of bus bandwidth across nodes, a 2 GB gradient all-reduce takes roughly 2 x 2 GB / 40 GB/s, about 100 milliseconds. If the framework overlaps it with backward compute, as described in collective overlap, much of that is hidden.
Now one spine link starts flapping its optics and dropping a fraction of a percent of packets. Every ring or tree that crosses it slows to the pace of retransmissions on that link. Bus bandwidth halves, the all-reduce takes 200 milliseconds, it no longer fits behind the backward pass, and step time rises by perhaps 10 percent across all 128 GPUs. If a tail packet is lost and waits for a transport timeout, one step takes over four seconds longer. In a lossless design, the failure looks different: a host that stops draining its receive buffers sends continuous pause frames, the pause propagates, and unrelated jobs on the same leaves slow down too. That is a pause storm, and the switch PFC watchdog, which disables PFC on a port that stays paused too long, exists to contain it.
In both cases the training framework sees only slower steps or an NCCL timeout. The fix is to correlate step time with fabric counters per link, which is why the counters belong on the same dashboard as the training metrics.
Trade-offs against InfiniBand
| Question | RoCE v2 | InfiniBand |
|---|---|---|
| Who runs it | Ethernet network team, familiar tools | Dedicated fabric skills, subnet manager |
| Congestion handling | PFC, ECN and DCQCN tuned by you | Credit-based link flow control built in |
| Routing | IP routing and ECMP hashing on UDP source port | Subnet-manager routing, adaptive routing on recent switches |
| In-network reduction | Depends on vendor features | SHARP on supported switches |
| Failure signature | Pause storms, ECMP hash collisions, GID mistakes | Subnet manager issues, cabling, credits |
Neither is free. RoCE moves complexity from buying a specialised fabric to configuring a general one correctly on every device, every time it is changed. In-switch reductions, covered in SHARP in-network reductions, are a further point to compare when choosing.
What to do next
- Dump the GID table on one host per node type and confirm which index is RoCE v2 on the right address.
- Upgrade NCCL to a release with automatic GID selection, or pin the index per host from that table.
- Write down the end-to-end QoS design: DSCP for RDMA and CNP, priority mapping, PFC on or off, ECN thresholds.
- Verify the same settings on every NIC and switch with a script, not by inspection.
- Run ib_write_bw per rail, then nccl-tests at full scale, then nccl-tests with background load.
- Export CNP, ECN, pause, discard and out-of-sequence counters per link into your training dashboards.
- Enable the switch PFC watchdog and alert on it.
- Rehearse a link failure and record what step time and NCCL logs look like, so the next one is recognised.