RoCE, RDMA over Converged Ethernet, lets one machine read and write another machine's registered memory without either CPU copying data. In GPU clusters it carries the NCCL traffic between servers: gradient all-reduces, FSDP all-gathers and expert-parallel all-to-alls, usually with GPUDirect RDMA so the NIC reads and writes GPU memory directly. Configuring the Ethernet fabric for it, with PFC, ECN, DCQCN and DSCP, is covered in RoCE v2 for GPU training.

This article looks one layer up, at the transport as software sees it. Which objects does NCCL create, how does a queue pair come to life, how does the NIC guarantee delivery over a network that can drop packets, and what are the numbers behind the dreaded 'transport retry counter exceeded' error? Knowing this turns a hung training job from a mystery into a short investigation.

The verbs objects behind every transfer

RDMA is programmed through the verbs API, provided on Linux by rdma-core's libibverbs. The same API drives InfiniBand and RoCE; only the addressing differs. Six object types matter:

ObjectWhat it isWhy you care
Protection domain (PD)A container that ties memory regions to queue pairsIsolation between users of one NIC
Memory region (MR)Pinned, registered memory with a local key (lkey) and remote key (rkey)The peer needs the rkey and address to write into your buffer
Completion queue (CQ)Where the NIC reports finished work as completion entriesErrors surface here, as a status code
Queue pair (QP)A send queue and receive queue bound to one remote QPThe unit of reliability, ordering and ECMP hashing
Work request (WR)One operation posted to a queue: send, RDMA write, RDMA readNCCL posts these from its proxy thread
Address handle / GIDThe remote address; for RoCE v2, an IP-based GIDA wrong GID index means RoCE v1 or the wrong subnet

Memory registration pins pages and gives the NIC a translation table, so it is expensive and done once per buffer, not per message. For GPU memory the NIC must be able to address the GPU's BAR: older stacks use the nvidia-peermem kernel module, newer ones register a dma-buf exported by the GPU driver through ibv_reg_dmabuf_mr. If neither path works, NCCL silently falls back to staging through host memory, which costs bandwidth; the mechanism is described in GPUDirect.

The RoCE RC data path for one NCCL sendNCCL proxyposts work requestsSend queueWQE: addr, lkey, rkeyNIC APSN, segment, retransmitGPU A memoryregistered MRpostdoorbellDMA readEthernet fabricUDP 4791, ECN, PFCpacketsNIC Bchecks PSN, ACK/NAKGPU B memoryrkey-checked writeDMA writeACK / NAKback to NIC ACompletion queueCQE: status, wr_idon ACKpollRetransmission and timeouts live in the NIC; software only sees a CQE with a status code
The NIC reads the source buffer from GPU memory, segments it into packets with sequence numbers, and only produces a completion when the peer acknowledges them.

Bringing up a queue pair

A reliable-connected (RC) queue pair moves through states: RESET, INIT, RTR (ready to receive) and RTS (ready to send). The two sides exchange QP numbers, starting packet sequence numbers (PSNs), GIDs and MR keys out of band, usually over a TCP socket during NCCL bootstrap, and then each side moves its QP forward with ibv_modify_qp. The attributes set here are the transport's whole personality:

struct ibv_qp_attr a = {0};
a.qp_state           = IBV_QPS_RTR;
a.path_mtu           = IBV_MTU_4096;      /* must fit the Ethernet MTU */
a.dest_qp_num        = remote.qpn;
a.rq_psn             = remote.psn;
a.max_dest_rd_atomic = 1;
a.min_rnr_timer      = 12;
a.ah_attr.is_global      = 1;             /* RoCE always uses a GRH */
a.ah_attr.grh.dgid       = remote.gid;
a.ah_attr.grh.sgid_index = gid_index;     /* selects RoCE v2 and the IP */
a.ah_attr.grh.hop_limit  = 64;
a.ah_attr.grh.traffic_class = tclass;     /* DSCP and ECN bits */
a.ah_attr.port_num       = 1;
ibv_modify_qp(qp, &a, IBV_QP_STATE | IBV_QP_AV | IBV_QP_PATH_MTU |
              IBV_QP_DEST_QPN | IBV_QP_RQ_PSN |
              IBV_QP_MAX_DEST_RD_ATOMIC | IBV_QP_MIN_RNR_TIMER);

a.qp_state      = IBV_QPS_RTS;
a.timeout       = 20;                     /* 4.096 us * 2^20 per attempt */
a.retry_cnt     = 7;                      /* transport retries */
a.rnr_retry     = 7;                      /* 7 means retry forever */
a.sq_psn        = my.psn;
a.max_rd_atomic = 1;
ibv_modify_qp(qp, &a, IBV_QP_STATE | IBV_QP_TIMEOUT | IBV_QP_RETRY_CNT |
              IBV_QP_RNR_RETRY | IBV_QP_SQ_PSN | IBV_QP_MAX_QP_RD_ATOMIC);

NCCL fills these from environment variables: NCCL_IB_GID_INDEX picks the GID, NCCL_IB_TC the traffic class, NCCL_IB_TIMEOUT and NCCL_IB_RETRY_CNT the timer and retries. Once a QP enters the error state, it stays there; every outstanding work request is flushed with an error and the connection must be torn down and rebuilt, which NCCL does not do mid-job. That single fact explains why one bad link can kill a thousand-GPU job.

How RC delivers reliably, and why loss hurts

Every RC packet carries a 24-bit PSN. The receiver expects PSNs in order. When one arrives on time it writes the payload and periodically returns an ACK. When a gap appears, because a packet was dropped, the receiver discards the out-of-order packets and sends a sequence-error NAK naming the PSN it expected. The sender then goes back to that PSN and resends everything from there: go-back-N. If the lost packet was the last in a burst, nothing arrives to reveal the gap, and the sender waits for its local ACK timeout instead.

Go-back-N is cheap in NIC memory but brutal on bandwidth, because one loss throws away a whole window of in-flight data. At 400 Gb/s with a 10 microsecond round trip and 4096-byte packets, about 122 packets are in flight. A simple model counts transmissions per delivered packet:

import random

def window_packets(gbps, rtt_us, mtu=4096):
    return int(gbps * 1e9 * rtt_us * 1e-6 / 8 // mtu)

def goodput(p, W, n=1_000_000, mode="gbn", seed=0):
    """Fraction of transmissions that deliver new data (NAK-detected losses)."""
    rng = random.Random(seed)
    sent = delivered = 0
    while delivered < n:
        sent += 1
        if rng.random() < p:
            if mode == "gbn":
                sent += W - 1    # in-flight packets behind the hole are discarded
            continue             # the lost packet is resent
        delivered += 1
    return delivered / sent
Packet loss rateGo-back-N goodputSelective repeat goodput
1e-50.99851.0000
1e-40.98710.9999
1e-30.89140.9990
1e-20.44650.9899

At one loss per thousand packets, go-back-N wastes about 11 percent of the link while selective repeat would waste 0.1 percent; at one percent, go-back-N halves throughput. This model ignores timeouts, which make tail losses much worse. It is the quantitative reason classic RoCE deployments demand a lossless fabric with PFC, and why newer NICs and the Ultra Ethernet effort add selective retransmission so that lossy operation becomes practical. Whether your adapter and firmware do this is a vendor question; check before you turn PFC off.

Timeouts, retries and what error 12 means

The local ACK timeout is 4.096 microseconds times 2 to the power of the timeout attribute, and the NIC retries up to retry_cnt times before declaring failure. NCCL's default has changed over time: 14 before NCCL 2.14, 18 from 2.14 and 20 from 2.23, with a retry count of 7.

NCCL_IB_TIMEOUTPer attemptWorst case with 1 + 7 attempts
140.067 s0.5 s
181.074 s8.6 s
204.295 s34.4 s

Short timeouts fail fast on a dead link but misfire when a congested fabric pauses traffic for a few hundred milliseconds; long ones ride out congestion but turn a dead link into half a minute of stall per QP before anything is reported. When retries run out, the work request completes with IBV_WC_RETRY_EXC_ERR, status 12, which NCCL prints as a completion with error 12, and every other request on that QP follows as IBV_WC_WR_FLUSH_ERR, status 5. Two other codes are worth knowing: status 13, IBV_WC_RNR_RETRY_EXC_ERR, means the receiver had no posted receive buffer, and status 10, IBV_WC_REM_ACCESS_ERR, means the rkey or address was wrong, which is a software bug, not a network problem. Read the first error, not the flood of flush errors behind it.

Reading the NIC counters

The NIC counts everything the transport does. With the mlx5 driver, per-port counters live under /sys/class/infiniband/<dev>/ports/1/hw_counters/, and Ethernet-level counters, including per-priority pause frames such as rx_prio3_pause, are shown by ethtool -S <iface>.

CounterMeaningPoints to
out_of_sequencePackets received out of orderDrops or reordering on the path
packet_seq_errSequence-error NAKs receivedLoss detected by the peer
implied_nak_seq_errGaps implied by responses, mainly for readsLoss on read traffic
local_ack_timeout_errACK timeouts on the senderTail loss or a dead path
np_cnp_sentCongestion notifications sent in response to ECN marksCongestion at this receiver
rp_cnp_handledCongestion notifications that slowed this senderDCQCN is acting
#!/bin/sh
# snapshot the counters that matter, on every node, every 10 seconds
D=/sys/class/infiniband/mlx5_0/ports/1/hw_counters
for c in out_of_sequence packet_seq_err local_ack_timeout_err np_cnp_sent rp_cnp_handled; do
  printf "%s %s %s\n" "$(hostname)" "$c" "$(cat $D/$c)"
done
ethtool -S eth0 | grep -E "prio3_pause"

Rates matter more than totals. CNP counters rising during heavy collectives is normal congestion control. Sequence errors and timeouts rising on one node while its peers stay flat is a bad link, cable or transceiver. Pause counters climbing everywhere at once is a pause storm.

Worked example: one optic stalls 512 GPUs

Suppose a 512-GPU job hangs at step 18,204, with every rank stuck inside an all-reduce, and after the process-group timeout PyTorch reports an NCCL error; one rank log, earlier than the rest, shows a work completion with error 12. The investigation takes four steps:

  1. Find the first failing rank and its NIC from the NCCL log; ignore the later flush errors (status 5) on other ranks, which are consequences.
  2. Compare local_ack_timeout_err and packet_seq_err deltas on that node with its neighbours. One NIC shows thousands of timeouts and sequence errors in the minutes before the hang; the others show none.
  3. Check the port with ethtool -S for symbol and FEC errors, and the switch port for the same; here the corrected-error count on one 400G optic is rising fast.
  4. Drain the node, replace the optic, run an ib_write_bw and nccl-tests pass between that node and a healthy one, and only then return it to the pool.

The arithmetic explains the symptoms: the link was dropping enough packets to push go-back-N into heavy retransmission, which slowed every ring through it, and finally a burst of loss exhausted seven 4.3-second retries. Because ring and tree all-reduce advance at the speed of the slowest link, as explained in all-reduce in depth, one optic stalled 512 GPUs.

Failure modes

  • MTU mismatch. A 4096-byte RoCE path MTU needs an Ethernet MTU of at least about 4200 on every hop; one 1500-byte port drops large packets and every QP through it times out.
  • Wrong GID index. Picking a RoCE v1 or link-local GID works on one switch and fails once traffic is routed.
  • Pinned-memory limits. A low ulimit -l in the container makes registration fail at start-up; set memlock to unlimited.
  • PCIe ACS on. Access control services force peer traffic through the root complex, breaking or slowing GPUDirect; disable ACS for the GPU and NIC switches on bare metal.
  • Pause storms. A stuck receiver sends PFC pauses that propagate upstream and can deadlock a cyclic topology; enable the switch PFC watchdog.
  • Too little entropy. One QP per connection hashes onto one ECMP path; NCCL_IB_QPS_PER_CONNECTION spreads traffic across several.

Trade-offs

Lossless RoCE with PFC gives the best goodput from go-back-N NICs but adds pause propagation, head-of-line blocking and deadlock risk. Lossy RoCE removes that complexity but needs selective retransmission and strong congestion control in the NIC. Compared with InfiniBand, described in InfiniBand and NVLink, RoCE runs on standard Ethernet switching, routing and tooling, at the price of more tuning and less mature congestion management out of the box. Timer settings trade fast failure against tolerance of congestion. Whichever you choose, size the fabric and its rails first, as in GPU pod network design.

What to do next

  1. Confirm GPUDirect RDMA is in use: check NCCL's debug output for GDRDMA rather than staged copies.
  2. Pin NCCL_IB_GID_INDEX, NCCL_IB_TC, NCCL_IB_TIMEOUT and NCCL_IB_RETRY_CNT explicitly so NCCL upgrades cannot change them silently.
  3. Export the six hw_counters above and per-priority pause counters from every node, as rates, to your monitoring system.
  4. Alert on sequence errors or ACK timeouts on one node that its peers do not show.
  5. Teach on-call to read the first completion error: 12 is network, 13 is receiver buffers, 10 is a software bug.
  6. Add an ib_write_bw and nccl-tests gate before any repaired node rejoins the pool.
  7. Ask your NIC vendor whether your firmware supports selective retransmission before you consider lossy operation.
Key takeaway: In RoCE, reliability lives in the NIC: PSNs, ACKs, go-back-N and a timer of 4.096 microseconds times 2 to the timeout. Loss costs a window per drop and an exhausted retry kills the queue pair, so watch sequence-error and timeout counters per node, read the first completion error, and pin the NCCL transport settings.