Elastic Fabric Adapter is how GPU instances on AWS talk to each other fast enough for distributed training. It is not InfiniBand and not RoCE, and the differences explain most of the surprises people hit on it: a transport that delivers packets reliably but out of order, a fabric whose traffic is not routable and cannot cross Availability Zones or VPCs, and a software stack where one missing plugin quietly drops NCCL back to TCP.
This article explains EFA as software sees it: the interface types, the SRD transport and why it was designed that way, the layers between NCCL and the wire, the rules of the network, the counters to watch, running it on EKS, and a worked triage of a slow all-reduce. For per-instance launch recipes on P5 and P4, see AWS P5 and P4 instances.
What an EFA is
An EFA is a network device attached to an instance. AWS offers it in two shapes, plus the ordinary ENA interface every instance has:
| Interface | IP networking | EFA device | Can be primary | API name |
|---|---|---|---|---|
| ENA | Yes | No | Yes | interface |
| EFA (EFA with ENA) | Yes | Yes | Yes | efa |
| EFA-only | No | Yes | No | efa-only |
EFA-only interfaces exist because large GPU instances need many EFA devices but only one IP address. The usual layout is one primary interface for IP traffic and an EFA-only interface on each additional network card. Every interface, EFA-only included, counts towards the instance's network-interface limit, and supported types allow one EFA per network card.
The device provides OS bypass: after the kernel driver sets up queues and registers memory, the application posts work directly to the device from user space and polls completions, with no system call or kernel TCP stack on the data path.
SRD: reliable, unordered, multipath
The transport under EFA is the Scalable Reliable Datagram protocol, implemented on the AWS Nitro card rather than in the host. Its design, published by AWS in IEEE Micro, makes three choices that differ from InfiniBand's reliable connections.
- Reliable but unordered. SRD guarantees each packet arrives, but not in order. Strict ordering would make one slow path block everything behind it, the head-of-line problem. Order, where needed, is restored above the device.
- Multipath spraying. Packets of a single flow are spread across many network paths instead of being pinned to one by a flow hash. A large all-reduce between two hosts therefore does not collide on one congested link, the common failure of single-path RDMA on Ethernet.
- Retransmission in hardware. Loss detection and resend run on the Nitro card with timeouts far shorter than TCP's, so a dropped packet costs little and the host never sees it. When a path keeps timing out, SRD moves traffic off it.
For software, unordered delivery is the consequence that matters. Libfabric's efa provider, not the device, handles message matching and ordering semantics for the RDM endpoints NCCL uses. You rarely touch this, but it explains why the libfabric version matters so much to performance: protocol logic that a NIC would do lives in user-space code shipped with the EFA installer.
From NCCL to the wire
Five pieces have to line up for NCCL traffic to take EFA. NCCL calls a network plugin interface; aws-ofi-nccl implements it on top of libfabric; libfabric's efa provider drives the device through rdma-core and the efa kernel driver. On GPU instances that support GPUDirect RDMA, the NIC reads and writes GPU memory directly, skipping a bounce through host memory; see GPUDirect for the mechanism.
Capabilities differ by Nitro generation, and AWS publishes a per-instance table. A few rows that matter for training:
| Instance | Nitro / EFA | RDMA read | RDMA write | GPUDirect RDMA |
|---|---|---|---|---|
| p5.48xlarge, p5e.48xlarge | v4 / EFA v2 | Yes | Yes | Yes |
| p5en.48xlarge | v5 / EFA v3 | Yes | Yes | Yes |
| p6-b200.48xlarge | v6 / EFA v4 | Yes | Yes | Yes |
The table comes from the EFA user guide as of October 2026; check it for your instance before relying on a feature, because rows are added and revised.
The rules of the network
EFA traffic follows rules that IP traffic on the same interface does not.
- Not routable. EFA packets cannot cross a router. They cannot cross Availability Zones or VPCs. All instances in a job belong in one AZ, and in practice in a cluster placement group.
- Security groups apply. The group must allow all inbound and outbound traffic to and from itself. A rule allowing only TCP leaves EFA dead while SSH and the rendezvous port work, which looks like a hang inside NCCL init.
- Generation islands. AWS states that EFA traffic between P4d, P4de or DL1 instances and other instance types is not supported. Do not mix P4d and P5 nodes in one job.
- Shared bandwidth. On P5 the IP and EFA traffic share the network cards, so heavy S3 or checkpoint traffic over ENA competes with collectives.
Counters that tell you what is wrong
The EFA driver publishes cumulative counters per device, readable with rdma -p statistic show or from sysfs under /sys/class/infiniband/<device>/ports/1/hw_counters/. Totals since boot are hard to read; deltas over an interval are what you want. This script prints per-device deltas for the counters that indicate trouble:
#!/usr/bin/env python3
# efa_deltas.py: per-device EFA counter deltas over an interval
import glob, os, sys, time
WATCH = ["tx_bytes", "rx_bytes", "rx_drops", "retrans_pkts",
"retrans_timeout_events", "impaired_remote_conn_events",
"unresponsive_remote_events", "rdma_read_wr_err", "rdma_write_wr_err"]
def snap():
out = {}
for d in glob.glob("/sys/class/infiniband/*/ports/1/hw_counters"):
dev = d.split("/")[4]
out[dev] = {}
for name in WATCH:
path = os.path.join(d, name)
if os.path.exists(path): # retrans_* need Nitro v4 or later
out[dev][name] = int(open(path).read())
return out
interval = float(sys.argv[1]) if len(sys.argv) > 1 else 10
a = snap(); time.sleep(interval); b = snap()
for dev in sorted(b):
d = {k: b[dev][k] - a[dev].get(k, 0) for k in b[dev]}
gbps = (d.get("tx_bytes", 0) + d.get("rx_bytes", 0)) * 8 / interval / 1e9
bad = {k: v for k, v in d.items() if v and k not in ("tx_bytes", "rx_bytes")}
print(f"{dev:14s} {gbps:7.1f} Gb/s {bad if bad else 'clean'}")| Counter | Meaning | Action |
|---|---|---|
retrans_pkts | SRD packets resent | Some is normal under load; compare across nodes |
retrans_timeout_events | Timeouts that moved traffic to another path | Rising on one node only: suspect that host |
impaired_remote_conn_events | Connections rate-limited as impaired | Find the remote peer; drain if persistent |
unresponsive_remote_events | Peer stopped responding | Peer crashed, hung or lost its network |
rx_drops | Received then dropped | Often receive-side resource exhaustion |
One gap to know: CloudWatch Container Insights on EKS collects the EFA metrics except the five SRD health counters (retrans_bytes, retrans_pkts, retrans_timeout_events, unresponsive_remote_events, impaired_remote_conn_events). If those matter to you, run your own collector.
EFA on EKS
On Kubernetes, the AWS EFA device plugin advertises each node's EFA devices as the extended resource vpc.amazonaws.com/efa. A pod requests them like GPUs, and for multi-node training it should request all of them on the node, or NCCL sees only part of the bandwidth:
apiVersion: v1
kind: Pod
metadata:
name: trainer-0
spec:
containers:
- name: train
image: <registry>/train:efa # libfabric + aws-ofi-nccl + NCCL, matched versions
resources:
limits:
nvidia.com/gpu: 8
vpc.amazonaws.com/efa: 32 # p5.48xlarge advertises 32
hugepages-2Mi: 5120Mi # example sizing, not an AWS requirement
memory: 1000Gi # example sizing
env:
- name: FI_PROVIDER
value: efa
- name: NCCL_DEBUG
value: INFOCheck the node first with kubectl describe node: a zero or missing EFA count means the plugin is not running or the node launched without EFA interfaces. The container must carry libfabric and the plugin itself; the host's copies are not visible inside it. Use an image built from the AWS deep learning containers, or install the EFA software into your image and record the versions.
Worked example: a slow 16-node all-reduce
A 16-node P5 job reports step times 40 percent slower than last week. Work outward from NCCL.
- Is EFA in use at all? In the
NCCL_DEBUG=INFOlog, look for the NET/OFI plugin lines naming the efa provider. If NCCL reports its socket transport instead, the image lost the plugin orFI_PROVIDERpoints elsewhere. Nothing else will fail. - Is the fabric slow everywhere? Run
all_reduce_perffrom nccl-tests across all 16 nodes and compare bus bandwidth with the baseline recorded for this image. If two-node pairs are all healthy but 16 nodes are slow, suspect one node. - Which node? Run
efa_deltas.py 30on every node during the benchmark. Fifteen show only modestretrans_pkts; one showsretrans_timeout_eventsclimbing and a third of the throughput of its peers. - Act. Cordon and drain that node, restart from the last checkpoint on a replacement, and open a case with the instance ID and counter output.
Because every rank waits for the slowest in a collective, one impaired host sets the pace of all 128 GPUs; see all-reduce in depth for why bus bandwidth is the number to compare.
A preflight check before every job
The cheapest time to find a sick node is before a job starts. A preflight check runs on the allocated nodes, in the same container image as the training job, and refuses to start training if any node falls below its baseline. Pairwise tests matter because a full-cluster benchmark only tells you that something is slow, not which node.
#!/usr/bin/env bash
# preflight.sh hosts.txt BASELINE_GBPS: pairwise nccl-tests, flag slow pairs
set -euo pipefail
mapfile -t H < "$1"; BASE=$2; FAIL=0
for ((i = 0; i + 1 < ${#H[@]}; i += 2)); do
a=${H[$i]}; b=${H[$((i + 1))]}
bw=$(mpirun -np 16 -N 8 -H "$a:8,$b:8" -x FI_PROVIDER=efa \
./all_reduce_perf -b 1G -e 1G -g 1 | awk '/Avg bus bandwidth/ {print $NF}')
ok=$(awk -v x="$bw" -v b="$BASE" 'BEGIN {print (x >= 0.85 * b) ? 1 : 0}')
echo "$a $b busbw=$bw ok=$ok"
[ "$ok" = 1 ] || FAIL=1
done
exit $FAILPair nodes in a different order on the next run, so a slow result can be pinned to one host by intersection rather than guessed. The 85 percent threshold is a starting point; set it from the spread you see across healthy pairs on your own fleet. Run the counter script alongside the test and keep both outputs with the job record, so a later regression can be compared against what healthy looked like on the same image.
Failure modes
- Silent TCP fallback from a missing plugin or wrong provider: the job runs at a fraction of expected speed.
- Init hang from a security group that does not allow all traffic within itself, or nodes split across AZs.
- Partial bandwidth from launching with fewer EFA interfaces than network cards, or pods requesting fewer EFA devices than the node has.
- Version skew between libfabric, aws-ofi-nccl and NCCL in a custom image: crashes at init or poor bandwidth. Pin all three together.
- Mixed generations in one job, such as P4d with P5, which EFA does not support.
- One sick host showing growing timeout or impaired-connection counters and slowing every collective.
Trade-offs
EFA's design trades some things away. Multipath spraying and hardware retransmission make tail latency steadier on a shared, oversubscribed cloud network than single-path RDMA, without lossless Ethernet configuration. In exchange you get a proprietary transport that works only inside AWS, a non-routable fabric confined to one zone, and a performance profile that depends on user-space software versions. Compared with dedicated InfiniBand clusters, small-message latency is typically higher, which matters more for tightly coupled HPC codes and small-message collectives than for large gradient all-reduces.
The practical trade is operational: you cannot tune switches, only measure from the hosts. That makes per-node counters and a recorded bandwidth baseline the core of running EFA well.
What to do next
- Confirm your instance's row in the EFA supported-types table: RDMA read, write and GPUDirect.
- Put every node in one AZ and a cluster placement group, with a self-referencing all-traffic security group.
- Launch with an EFA interface on every network card; on EKS, request every
vpc.amazonaws.com/efadevice. - Pin libfabric, aws-ofi-nccl and NCCL as one set and record the versions in the image.
- On every new image, read the
NCCL_DEBUG=INFOlog once and confirm the efa provider is selected. - Record nccl-tests bus bandwidth at 2 and N nodes as a baseline.
- Run a counter-delta collector, including the SRD health counters Container Insights omits, and alert on per-node outliers.
- Write a drain-and-replace runbook for a node with rising timeout counters.