RoCE, RDMA over Converged Ethernet, lets one machine read and write another machine's registered memory without either CPU copying data. In GPU clusters it carries the NCCL traffic between servers: gradient all-reduces, FSDP all-gathers and expert-parallel all-to-alls, usually with GPUDirect RDMA so the NIC reads and writes GPU memory directly. Configuring the Ethernet fabric for it, with PFC, ECN, DCQCN and DSCP, is covered in RoCE v2 for GPU training.
This article looks one layer up, at the transport as software sees it. Which objects does NCCL create, how does a queue pair come to life, how does the NIC guarantee delivery over a network that can drop packets, and what are the numbers behind the dreaded 'transport retry counter exceeded' error? Knowing this turns a hung training job from a mystery into a short investigation.
The verbs objects behind every transfer
RDMA is programmed through the verbs API, provided on Linux by rdma-core's libibverbs. The same API drives InfiniBand and RoCE; only the addressing differs. Six object types matter:
| Object | What it is | Why you care |
|---|---|---|
| Protection domain (PD) | A container that ties memory regions to queue pairs | Isolation between users of one NIC |
| Memory region (MR) | Pinned, registered memory with a local key (lkey) and remote key (rkey) | The peer needs the rkey and address to write into your buffer |
| Completion queue (CQ) | Where the NIC reports finished work as completion entries | Errors surface here, as a status code |
| Queue pair (QP) | A send queue and receive queue bound to one remote QP | The unit of reliability, ordering and ECMP hashing |
| Work request (WR) | One operation posted to a queue: send, RDMA write, RDMA read | NCCL posts these from its proxy thread |
| Address handle / GID | The remote address; for RoCE v2, an IP-based GID | A wrong GID index means RoCE v1 or the wrong subnet |
Memory registration pins pages and gives the NIC a translation table, so it is expensive and done once per buffer, not per message. For GPU memory the NIC must be able to address the GPU's BAR: older stacks use the nvidia-peermem kernel module, newer ones register a dma-buf exported by the GPU driver through ibv_reg_dmabuf_mr. If neither path works, NCCL silently falls back to staging through host memory, which costs bandwidth; the mechanism is described in GPUDirect.
Bringing up a queue pair
A reliable-connected (RC) queue pair moves through states: RESET, INIT, RTR (ready to receive) and RTS (ready to send). The two sides exchange QP numbers, starting packet sequence numbers (PSNs), GIDs and MR keys out of band, usually over a TCP socket during NCCL bootstrap, and then each side moves its QP forward with ibv_modify_qp. The attributes set here are the transport's whole personality:
struct ibv_qp_attr a = {0};
a.qp_state = IBV_QPS_RTR;
a.path_mtu = IBV_MTU_4096; /* must fit the Ethernet MTU */
a.dest_qp_num = remote.qpn;
a.rq_psn = remote.psn;
a.max_dest_rd_atomic = 1;
a.min_rnr_timer = 12;
a.ah_attr.is_global = 1; /* RoCE always uses a GRH */
a.ah_attr.grh.dgid = remote.gid;
a.ah_attr.grh.sgid_index = gid_index; /* selects RoCE v2 and the IP */
a.ah_attr.grh.hop_limit = 64;
a.ah_attr.grh.traffic_class = tclass; /* DSCP and ECN bits */
a.ah_attr.port_num = 1;
ibv_modify_qp(qp, &a, IBV_QP_STATE | IBV_QP_AV | IBV_QP_PATH_MTU |
IBV_QP_DEST_QPN | IBV_QP_RQ_PSN |
IBV_QP_MAX_DEST_RD_ATOMIC | IBV_QP_MIN_RNR_TIMER);
a.qp_state = IBV_QPS_RTS;
a.timeout = 20; /* 4.096 us * 2^20 per attempt */
a.retry_cnt = 7; /* transport retries */
a.rnr_retry = 7; /* 7 means retry forever */
a.sq_psn = my.psn;
a.max_rd_atomic = 1;
ibv_modify_qp(qp, &a, IBV_QP_STATE | IBV_QP_TIMEOUT | IBV_QP_RETRY_CNT |
IBV_QP_RNR_RETRY | IBV_QP_SQ_PSN | IBV_QP_MAX_QP_RD_ATOMIC);NCCL fills these from environment variables: NCCL_IB_GID_INDEX picks the GID, NCCL_IB_TC the traffic class, NCCL_IB_TIMEOUT and NCCL_IB_RETRY_CNT the timer and retries. Once a QP enters the error state, it stays there; every outstanding work request is flushed with an error and the connection must be torn down and rebuilt, which NCCL does not do mid-job. That single fact explains why one bad link can kill a thousand-GPU job.
How RC delivers reliably, and why loss hurts
Every RC packet carries a 24-bit PSN. The receiver expects PSNs in order. When one arrives on time it writes the payload and periodically returns an ACK. When a gap appears, because a packet was dropped, the receiver discards the out-of-order packets and sends a sequence-error NAK naming the PSN it expected. The sender then goes back to that PSN and resends everything from there: go-back-N. If the lost packet was the last in a burst, nothing arrives to reveal the gap, and the sender waits for its local ACK timeout instead.
Go-back-N is cheap in NIC memory but brutal on bandwidth, because one loss throws away a whole window of in-flight data. At 400 Gb/s with a 10 microsecond round trip and 4096-byte packets, about 122 packets are in flight. A simple model counts transmissions per delivered packet:
import random
def window_packets(gbps, rtt_us, mtu=4096):
return int(gbps * 1e9 * rtt_us * 1e-6 / 8 // mtu)
def goodput(p, W, n=1_000_000, mode="gbn", seed=0):
"""Fraction of transmissions that deliver new data (NAK-detected losses)."""
rng = random.Random(seed)
sent = delivered = 0
while delivered < n:
sent += 1
if rng.random() < p:
if mode == "gbn":
sent += W - 1 # in-flight packets behind the hole are discarded
continue # the lost packet is resent
delivered += 1
return delivered / sent| Packet loss rate | Go-back-N goodput | Selective repeat goodput |
|---|---|---|
| 1e-5 | 0.9985 | 1.0000 |
| 1e-4 | 0.9871 | 0.9999 |
| 1e-3 | 0.8914 | 0.9990 |
| 1e-2 | 0.4465 | 0.9899 |
At one loss per thousand packets, go-back-N wastes about 11 percent of the link while selective repeat would waste 0.1 percent; at one percent, go-back-N halves throughput. This model ignores timeouts, which make tail losses much worse. It is the quantitative reason classic RoCE deployments demand a lossless fabric with PFC, and why newer NICs and the Ultra Ethernet effort add selective retransmission so that lossy operation becomes practical. Whether your adapter and firmware do this is a vendor question; check before you turn PFC off.
Timeouts, retries and what error 12 means
The local ACK timeout is 4.096 microseconds times 2 to the power of the timeout attribute, and the NIC retries up to retry_cnt times before declaring failure. NCCL's default has changed over time: 14 before NCCL 2.14, 18 from 2.14 and 20 from 2.23, with a retry count of 7.
| NCCL_IB_TIMEOUT | Per attempt | Worst case with 1 + 7 attempts |
|---|---|---|
| 14 | 0.067 s | 0.5 s |
| 18 | 1.074 s | 8.6 s |
| 20 | 4.295 s | 34.4 s |
Short timeouts fail fast on a dead link but misfire when a congested fabric pauses traffic for a few hundred milliseconds; long ones ride out congestion but turn a dead link into half a minute of stall per QP before anything is reported. When retries run out, the work request completes with IBV_WC_RETRY_EXC_ERR, status 12, which NCCL prints as a completion with error 12, and every other request on that QP follows as IBV_WC_WR_FLUSH_ERR, status 5. Two other codes are worth knowing: status 13, IBV_WC_RNR_RETRY_EXC_ERR, means the receiver had no posted receive buffer, and status 10, IBV_WC_REM_ACCESS_ERR, means the rkey or address was wrong, which is a software bug, not a network problem. Read the first error, not the flood of flush errors behind it.
Reading the NIC counters
The NIC counts everything the transport does. With the mlx5 driver, per-port counters live under /sys/class/infiniband/<dev>/ports/1/hw_counters/, and Ethernet-level counters, including per-priority pause frames such as rx_prio3_pause, are shown by ethtool -S <iface>.
| Counter | Meaning | Points to |
|---|---|---|
out_of_sequence | Packets received out of order | Drops or reordering on the path |
packet_seq_err | Sequence-error NAKs received | Loss detected by the peer |
implied_nak_seq_err | Gaps implied by responses, mainly for reads | Loss on read traffic |
local_ack_timeout_err | ACK timeouts on the sender | Tail loss or a dead path |
np_cnp_sent | Congestion notifications sent in response to ECN marks | Congestion at this receiver |
rp_cnp_handled | Congestion notifications that slowed this sender | DCQCN is acting |
#!/bin/sh
# snapshot the counters that matter, on every node, every 10 seconds
D=/sys/class/infiniband/mlx5_0/ports/1/hw_counters
for c in out_of_sequence packet_seq_err local_ack_timeout_err np_cnp_sent rp_cnp_handled; do
printf "%s %s %s\n" "$(hostname)" "$c" "$(cat $D/$c)"
done
ethtool -S eth0 | grep -E "prio3_pause"Rates matter more than totals. CNP counters rising during heavy collectives is normal congestion control. Sequence errors and timeouts rising on one node while its peers stay flat is a bad link, cable or transceiver. Pause counters climbing everywhere at once is a pause storm.
Worked example: one optic stalls 512 GPUs
Suppose a 512-GPU job hangs at step 18,204, with every rank stuck inside an all-reduce, and after the process-group timeout PyTorch reports an NCCL error; one rank log, earlier than the rest, shows a work completion with error 12. The investigation takes four steps:
- Find the first failing rank and its NIC from the NCCL log; ignore the later flush errors (status 5) on other ranks, which are consequences.
- Compare
local_ack_timeout_errandpacket_seq_errdeltas on that node with its neighbours. One NIC shows thousands of timeouts and sequence errors in the minutes before the hang; the others show none. - Check the port with
ethtool -Sfor symbol and FEC errors, and the switch port for the same; here the corrected-error count on one 400G optic is rising fast. - Drain the node, replace the optic, run an
ib_write_bwand nccl-tests pass between that node and a healthy one, and only then return it to the pool.
The arithmetic explains the symptoms: the link was dropping enough packets to push go-back-N into heavy retransmission, which slowed every ring through it, and finally a burst of loss exhausted seven 4.3-second retries. Because ring and tree all-reduce advance at the speed of the slowest link, as explained in all-reduce in depth, one optic stalled 512 GPUs.
Failure modes
- MTU mismatch. A 4096-byte RoCE path MTU needs an Ethernet MTU of at least about 4200 on every hop; one 1500-byte port drops large packets and every QP through it times out.
- Wrong GID index. Picking a RoCE v1 or link-local GID works on one switch and fails once traffic is routed.
- Pinned-memory limits. A low
ulimit -lin the container makes registration fail at start-up; set memlock to unlimited. - PCIe ACS on. Access control services force peer traffic through the root complex, breaking or slowing GPUDirect; disable ACS for the GPU and NIC switches on bare metal.
- Pause storms. A stuck receiver sends PFC pauses that propagate upstream and can deadlock a cyclic topology; enable the switch PFC watchdog.
- Too little entropy. One QP per connection hashes onto one ECMP path;
NCCL_IB_QPS_PER_CONNECTIONspreads traffic across several.
Trade-offs
Lossless RoCE with PFC gives the best goodput from go-back-N NICs but adds pause propagation, head-of-line blocking and deadlock risk. Lossy RoCE removes that complexity but needs selective retransmission and strong congestion control in the NIC. Compared with InfiniBand, described in InfiniBand and NVLink, RoCE runs on standard Ethernet switching, routing and tooling, at the price of more tuning and less mature congestion management out of the box. Timer settings trade fast failure against tolerance of congestion. Whichever you choose, size the fabric and its rails first, as in GPU pod network design.
What to do next
- Confirm GPUDirect RDMA is in use: check NCCL's debug output for GDRDMA rather than staged copies.
- Pin
NCCL_IB_GID_INDEX,NCCL_IB_TC,NCCL_IB_TIMEOUTandNCCL_IB_RETRY_CNTexplicitly so NCCL upgrades cannot change them silently. - Export the six hw_counters above and per-priority pause counters from every node, as rates, to your monitoring system.
- Alert on sequence errors or ACK timeouts on one node that its peers do not show.
- Teach on-call to read the first completion error: 12 is network, 13 is receiver buffers, 10 is a software bug.
- Add an
ib_write_bwand nccl-tests gate before any repaired node rejoins the pool. - Ask your NIC vendor whether your firmware supports selective retransmission before you consider lossy operation.