A large training step ends in collectives: all-reduce for gradients, all-gather and reduce-scatter for sharded weights, all-to-all for mixture-of-experts routing. Each one moves gigabytes between thousands of GPUs, and the step cannot finish until the slowest transfer does. On Ethernet fabrics those transfers have mostly used RoCE v2, the InfiniBand transport carried in UDP. The Ultra Ethernet Consortium (UEC), a Linux Foundation project, was formed to design a transport for this traffic from scratch rather than inherit one, and published UEC Specification 1.0 on 11 June 2025, with point revisions since.

This article explains what the specification defines, layer by layer, and why each choice matters to a training job. It assumes you know the basics of RDMA; if not, read RoCE v2, in depth first. Details come from the specification's public overview paper (Hoefler et al., arXiv 2508.08906); product availability is changing quickly, so check vendors' current conformance claims before planning a purchase.

Why a new transport

RoCE v2 works, and very large clusters run on it, but four of its properties fight AI traffic at scale.

  1. One path per connection. A queue pair's packets share one 5-tuple, so ECMP pins the flow to one path; few huge flows is exactly when hashing collides.
  2. Loss is expensive. The classic RoCE NIC recovers with go-back-N: one lost packet resends everything after it. That pushes operators to run the fabric lossless with PFC.
  3. PFC has fabric-wide side effects. Pause frames stop a whole priority on a link, which spreads congestion to innocent flows and, in bad cases, deadlocks. The details are in RoCE, in depth.
  4. Connection state grows with peers. Each queue pair needs a handshake and state.

UEC's answer is to make the transport tolerate reordering and loss cheaply, so the network can spray packets over every path and drop or trim when queues overflow, rather than being kept lossless.

The stack at a glance

The UE transport (UET) sits between libfabric, the API that communication libraries call, and an IP network. It is divided into sublayers with separate jobs.

SublayerResponsibility
SES, semanticsSend/receive, RMA, atomics, tag matching, job addressing; maps to libfabric, inspired by Portals 4
PDS, packet deliveryReliability and ordering per delivery mode; packet sequence numbers, ACKs, selective ACKs and NACKs, retransmission
CMS, congestion managementByte-level send window per congestion control context; NSCC at the source, optional RCCC at the receiver; load balancing across paths
TSS, transport securityOptional authentication and encryption within a secure domain

Underneath, packets are routable IP. The IANA-assigned UDP port for UET is 4793; an IP-only mode replaces the 8-byte UDP header with a 4-byte entropy header.

Ultra Ethernet Transport (UET) on one NIC, and what it needs from the fabricCCL / MPI / applicationtalks to libfabricSES: semantics sublayersend/recv, RMA, tagged match, JobIDPDS: packet deliveryRUD / ROD / RUDI / UUDCMSNSCC, RCCCTSS: transport security (optional)UDP/IP (port 4793) or IP-only modeEthernet link layeroptional LLR and CBFCSwitchesECMP on EV, ECN marking, optional trimmingRemote NIC (target FEP)JobID, PIDonFEP, RI lookup, ACK/SACK/NACKpackets, EV varies per packetACKs: RTT, ECN, trim NACKsMinimum switch requirement: ECMP plus ECN.Trimming is optional but speeds loss detection.
The UET sublayers on a NIC and the signals exchanged with switches and the target NIC. Only ECMP and ECN are required of switches.

Addressing jobs, not connections

RDMA addresses a queue pair. UET addresses a job. A Fabric Address (an IP address) selects a Fabric Endpoint (FEP) on a NIC. Inside, a 24-bit JobID identifies the parallel job, a 12-bit PIDonFEP identifies a process within that job on that endpoint, and a 12-bit Resource Index selects a receive context such as a queue. In relative addressing, used for distributed jobs, the target NIC looks up the JobID in its job table, then the process in that job's table. In absolute addressing, used for client-server services, PIDonFEP acts like a UDP port and the JobID serves as an authorisation token.

So a rank addresses peers by node, job and local index, with no per-peer connection to create first. JobIDs must be allocated and installed in NIC tables outside the application, typically by the job launcher, so integrating UET is partly a cluster-manager task.

Packet delivery: four modes and connectionless contexts

The PDS offers four delivery modes, chosen per operation by the layers above.

ModeGuaranteeTypical use
RUD, reliable unorderedevery packet arrives once, in any orderthe default for bulk transfers; enables spraying
ROD, reliable orderedevery packet, in order, on one pathtraffic that needs ordering, such as wildcard matching in the HPC profile
RUDI, reliable unordered for idempotent operationsreliable, but without receiver-side duplicate filteringoperations safe to apply twice, such as plain RMA writes; the most scalable
UUD, unreliable unordereddatagram semanticscontrol traffic and protocols with their own recovery

Reliability state lives in a Packet Delivery Context (PDC). PDCs are connectionless in the sense that matters: the information needed to set one up travels in the headers of the first data packets, so a sender can transmit immediately without a round-trip handshake, even when those first packets arrive out of order. Packets carry sequence numbers; the receiver returns cumulative ACKs (CACK) and selective ACKs with a 64-bit bitmap, and a maximum PSN range (MP_RANGE) caps how many packets may be outstanding.

The consequence for training: loss of one packet costs one retransmission, not a replay of everything behind it, so the fabric no longer needs to be lossless to perform well.

Spraying, entropy and trimming

Every UET packet carries an entropy value (EV). In UDP mode it is the UDP source port, which is otherwise unused, so ordinary switches hash it into their ECMP decision without knowing anything about UET. Changing the EV per packet sprays a single message over many paths. Since RUD tolerates reordering, nothing downstream has to reassemble in order.

Schemes range from oblivious spraying (new EV every packet) to recycled-entropy spraying (REPS) and path-aware schemes that steer away from EVs that return ECN marks or trims.

When a switch queue overflows, a trimming switch cuts the payload off the packet and forwards just the headers, possibly at higher priority. The receiver sees a trimmed packet, knows exactly which PSN lost its payload, and NACKs it at once instead of waiting for a timeout. Trimming detects congestion drops only; corruption still needs timeouts or link retry. Without trimming, a sprayed transport infers loss from ACKs for later packets sent on the same EV, which is slower and less precise.

Worked example: hash collisions in a collective

Why does spraying matter so much? Take 8 large flows leaving a leaf switch over 8 uplinks, a common situation when an all-reduce spans leaves. With per-flow ECMP each flow picks an uplink by hash. The chance that all 8 land on different uplinks is 8!/88, about 0.24 percent. On average only 5.25 uplinks carry any traffic, and the busiest carries about 2.6 flows.

import random

def mean_worst_link(flows, links, trials=100_000, spray=False):
    """Average load on the busiest uplink, in units of one flow's demand."""
    total = 0.0
    for _ in range(trials):
        load = [0.0] * links
        for _ in range(flows):
            if spray:                        # packets spread evenly over all paths
                for l in range(links):
                    load[l] += 1 / links
            else:                            # ECMP: one hash, one path, whole flow
                load[random.randrange(links)] += 1
        total += max(load)
    return total / trials

print(mean_worst_link(8, 8))              # about 2.59
print(mean_worst_link(8, 8, spray=True))  # 1.0

A collective completes when its slowest flow does, so the flows on that busiest link run at roughly 1/2.6, under 40 percent, of the rate they would get with perfect balance, and the whole step waits for them. Spraying makes the worst link carry exactly one flow's worth. Tuned hashing and multiple connections per peer soften this in practice; they are workarounds for the problem spraying removes. All-Reduce, in depth shows how this bandwidth loss appears in busbw.

Congestion control: NSCC and RCCC

Spraying spreads load but cannot create capacity; incast, where many senders target one receiver, still overflows the last hop. UET's congestion management has two algorithms that can run together.

  • NSCC (network-signal congestion control) runs at the source on every UE NIC. It combines round-trip time with ECN marks echoed in ACKs. The two signals together distinguish a queue that is building, one that is standing, spare capacity, and a queue that is already draining, which a single signal cannot do.
  • RCCC (receiver credit congestion control) is optional. The receiver hands out credits, which suits incast; congestion inside the network is still NSCC's job.

The sketch shows the shape of a two-signal controller; it is not the specified NSCC algorithm.

# Illustrative only: the shape of a sender-side controller that reads two
# signals per ACK. UE's NSCC defines its own constants and update rules.
def on_ack(ctx, rtt, ecn_marked):
    queued = rtt > ctx.target_rtt
    if ecn_marked and queued:        # queue is standing: back off proportionally
        ctx.cwnd -= ctx.beta * ctx.cwnd * (rtt - ctx.target_rtt) / rtt
    elif ecn_marked and not queued:  # queue is building but still short: hold or ease off gently
        ctx.cwnd -= ctx.small_step
    elif not ecn_marked and not queued:
        ctx.cwnd += ctx.fast_increase   # spare capacity: grow quickly
    else:                            # high RTT without marks: congestion is draining
        pass                         # do not punish the flow for an old queue
    ctx.cwnd = max(ctx.min_cwnd, min(ctx.cwnd, ctx.max_cwnd))

The link layer: LLR and CBFC

Two optional link-layer extensions work hop by hop, below the transport.

  • Link Level Retry (LLR) keeps a replay buffer on the sending side of a link, numbers frames, and retransmits with go-back-N when the receiver reports a gap. A bit error on a flaky link then costs a local replay instead of an end-to-end retransmission.
  • Credit-Based Flow Control (CBFC) replaces PFC pause frames with per-virtual-channel credits: a sender transmits only when the receiver has buffer space, which needs less buffer than PFC.

Both need support at both ends of a link, so they arrive with new hardware, not software updates.

Profiles, software and header cost

Not every NIC must implement everything. The specification defines three profiles.

ProfileAddsAimed at
AI Basethe simplest implementation; tag matching can be done by the libfabric provider in softwarecollective libraries
AI Fullsuperset of AI Base, adding exact tag matchingcollective libraries wanting more offload
HPCthe richest feature set, including in-order wildcard tag matchingMPI and OpenSHMEM

Both AI profiles provide deferrable sends, designed for offloading collective libraries: if a message arrives before the receiver has posted its buffer, the target can defer it and later ask the sender to resume, instead of buffering unexpected data.

For software the contract is libfabric. A collective library reaches UET through a libfabric provider, the same pattern as the existing open-source aws-ofi-nccl plugin that lets NCCL use libfabric. In practice you will consume UET through your NIC vendor's provider and collective-library plugin; check which profile and which optional features (RCCC, trimming, LLR, CBFC) your NIC and switches actually implement.

Header cost is modest: Ethernet, IPv4, UDP, PDS (12 bytes) and standard SES (44) headers plus FCS total 102 bytes, so a 4,096-byte payload uses about 97.1 percent of the wire including preamble and gap, against about 98.0 percent for a RoCE v2 middle packet.

Bring-up and operations

If you are evaluating or bringing up a UET fabric, operate it as a new transport, not a firmware option.

  1. Confirm the feature matrix per device: profile, RCCC, trimming, LLR, CBFC. Optional features are optional; a UEC label does not imply them.
  2. Configure ECN everywhere. NSCC depends on marks; an egress queue without ECN marking is invisible to it.
  3. Check ECMP hashes the source port. If switches hash only addresses, spraying collapses back onto one path per pair.
  4. Measure with collectives. Compare all-reduce and all-to-all busbw against your RoCE baseline at the same scale.
  5. Watch the new counters. Trim counts, NACK and retransmission rates, out-of-order depth and congestion-window statistics replace PFC pause counts as the health signals.
  6. Plan JobID management. The scheduler must allocate JobIDs and program NICs at job start, and revoke them at job end.

Mixed fabrics, where RoCE and UET share switches, need queue and ECN policies that keep a lossless RoCE class and a lossy UET class from harming each other. The topology itself, rails and oversubscription, is unchanged; see GPU Pod Network, in depth.

Failure modes

  • Spraying without EV hashing. All packets of a pair still take one path; throughput looks like RoCE with more reordering.
  • ROD where RUD would do. Ordered mode uses a single path; a library that requests it for bulk data gives up the main benefit.
  • No trimming and long timeouts. Without trimming, losses under incast wait for timeouts or inference, and tail latency of collectives grows.
  • ECN thresholds copied from RoCE. Thresholds tuned for DCQCN on a lossless class may be wrong for NSCC on a lossy one; tune them with the vendor's guidance.
  • Assuming RDMA verbs code ports directly. UET is programmed through libfabric; software written against ibverbs needs a different path.

Trade-offs against RoCE and InfiniBand

UETRoCE v2InfiniBand
Multipathper-packet spraying, designed inper-flow ECMP; vendor extensionsadaptive routing in switches
Loss handlingselective retransmit, trimminggo-back-N on many NICs; lossless fabric preferredlossless, credit-based link layer
Connection setupconnectionless PDCsQP handshake per peerQP handshake per peer
APIlibfabricverbsverbs
Maturitynew; ecosystem formingmature, widely deployedmature, single dominant vendor

UET's strengths are multipath, cheap loss recovery and an open multi-vendor standard on commodity Ethernet. Its weakness today is maturity: fewer production deployments, tooling and tuning experience. InfiniBand remains the established choice where its ecosystem fits; see InfiniBand NDR and XDR for that side of the comparison.

What to do next

  1. Measure your current collective efficiency: busbw for all-reduce and all-to-all at your real scale, and how much step time is exposed communication.
  2. Check whether ECMP collisions are a measurable problem today by comparing per-uplink utilisation during a collective.
  3. Ask NIC and switch vendors for a written UEC feature matrix: profile, NSCC, RCCC, trimming, LLR, CBFC.
  4. Confirm the software path: libfabric provider, collective-library plugin, and supported framework versions.
  5. Run a pilot at a few racks with identical benchmarks against RoCE, including incast and failure tests.
  6. Plan scheduler integration for JobIDs and secure domains before production.
Key takeaway: Ultra Ethernet's transport makes reordering and loss cheap so the network can spray every packet across every path: connectionless delivery contexts, selective acknowledgement, trimming for fast loss detection, NSCC at the sender and optional RCCC at the receiver. For training, the gain is collectives that are no longer limited by ECMP collisions or PFC side effects. Treat optional features as optional, verify what your devices implement, configure ECN and source-port hashing, and judge it with collective benchmarks against your current fabric.