Elastic Fabric Adapter is how GPU instances on AWS talk to each other fast enough for distributed training. It is not InfiniBand and not RoCE, and the differences explain most of the surprises people hit on it: a transport that delivers packets reliably but out of order, a fabric whose traffic is not routable and cannot cross Availability Zones or VPCs, and a software stack where one missing plugin quietly drops NCCL back to TCP.

This article explains EFA as software sees it: the interface types, the SRD transport and why it was designed that way, the layers between NCCL and the wire, the rules of the network, the counters to watch, running it on EKS, and a worked triage of a slow all-reduce. For per-instance launch recipes on P5 and P4, see AWS P5 and P4 instances.

What an EFA is

An EFA is a network device attached to an instance. AWS offers it in two shapes, plus the ordinary ENA interface every instance has:

InterfaceIP networkingEFA deviceCan be primaryAPI name
ENAYesNoYesinterface
EFA (EFA with ENA)YesYesYesefa
EFA-onlyNoYesNoefa-only

EFA-only interfaces exist because large GPU instances need many EFA devices but only one IP address. The usual layout is one primary interface for IP traffic and an EFA-only interface on each additional network card. Every interface, EFA-only included, counts towards the instance's network-interface limit, and supported types allow one EFA per network card.

The device provides OS bypass: after the kernel driver sets up queues and registers memory, the application posts work directly to the device from user space and polls completions, with no system call or kernel TCP stack on the data path.

Where EFA sits: NCCL bypasses the kernel and hands packets to the Nitro cardTraining process (PyTorch)NCCLcollectives, rings and treesaws-ofi-nccl pluginNCCL net API to libfabriclibfabric, efa providerRDM endpoints, protocolsrdma-core + efa kernel driversetup only, not data pathNitro card: EFA deviceSRD: reliability, multipathOS bypassdata path skips kernelENA deviceTCP/IP, routable, S3AWS network fabricmany paths between hostsSRD packets sprayednot routable, one AZ, one VPCsame security group rules apply
The EFA software stack. The kernel driver and rdma-core set up queues and memory registration; data moves from libfabric straight to the Nitro card, which runs SRD across many fabric paths.

SRD: reliable, unordered, multipath

The transport under EFA is the Scalable Reliable Datagram protocol, implemented on the AWS Nitro card rather than in the host. Its design, published by AWS in IEEE Micro, makes three choices that differ from InfiniBand's reliable connections.

  • Reliable but unordered. SRD guarantees each packet arrives, but not in order. Strict ordering would make one slow path block everything behind it, the head-of-line problem. Order, where needed, is restored above the device.
  • Multipath spraying. Packets of a single flow are spread across many network paths instead of being pinned to one by a flow hash. A large all-reduce between two hosts therefore does not collide on one congested link, the common failure of single-path RDMA on Ethernet.
  • Retransmission in hardware. Loss detection and resend run on the Nitro card with timeouts far shorter than TCP's, so a dropped packet costs little and the host never sees it. When a path keeps timing out, SRD moves traffic off it.

For software, unordered delivery is the consequence that matters. Libfabric's efa provider, not the device, handles message matching and ordering semantics for the RDM endpoints NCCL uses. You rarely touch this, but it explains why the libfabric version matters so much to performance: protocol logic that a NIC would do lives in user-space code shipped with the EFA installer.

From NCCL to the wire

Five pieces have to line up for NCCL traffic to take EFA. NCCL calls a network plugin interface; aws-ofi-nccl implements it on top of libfabric; libfabric's efa provider drives the device through rdma-core and the efa kernel driver. On GPU instances that support GPUDirect RDMA, the NIC reads and writes GPU memory directly, skipping a bounce through host memory; see GPUDirect for the mechanism.

Capabilities differ by Nitro generation, and AWS publishes a per-instance table. A few rows that matter for training:

InstanceNitro / EFARDMA readRDMA writeGPUDirect RDMA
p5.48xlarge, p5e.48xlargev4 / EFA v2YesYesYes
p5en.48xlargev5 / EFA v3YesYesYes
p6-b200.48xlargev6 / EFA v4YesYesYes

The table comes from the EFA user guide as of October 2026; check it for your instance before relying on a feature, because rows are added and revised.

The rules of the network

EFA traffic follows rules that IP traffic on the same interface does not.

  • Not routable. EFA packets cannot cross a router. They cannot cross Availability Zones or VPCs. All instances in a job belong in one AZ, and in practice in a cluster placement group.
  • Security groups apply. The group must allow all inbound and outbound traffic to and from itself. A rule allowing only TCP leaves EFA dead while SSH and the rendezvous port work, which looks like a hang inside NCCL init.
  • Generation islands. AWS states that EFA traffic between P4d, P4de or DL1 instances and other instance types is not supported. Do not mix P4d and P5 nodes in one job.
  • Shared bandwidth. On P5 the IP and EFA traffic share the network cards, so heavy S3 or checkpoint traffic over ENA competes with collectives.

Counters that tell you what is wrong

The EFA driver publishes cumulative counters per device, readable with rdma -p statistic show or from sysfs under /sys/class/infiniband/<device>/ports/1/hw_counters/. Totals since boot are hard to read; deltas over an interval are what you want. This script prints per-device deltas for the counters that indicate trouble:

#!/usr/bin/env python3
# efa_deltas.py: per-device EFA counter deltas over an interval
import glob, os, sys, time

WATCH = ["tx_bytes", "rx_bytes", "rx_drops", "retrans_pkts",
         "retrans_timeout_events", "impaired_remote_conn_events",
         "unresponsive_remote_events", "rdma_read_wr_err", "rdma_write_wr_err"]

def snap():
    out = {}
    for d in glob.glob("/sys/class/infiniband/*/ports/1/hw_counters"):
        dev = d.split("/")[4]
        out[dev] = {}
        for name in WATCH:
            path = os.path.join(d, name)
            if os.path.exists(path):          # retrans_* need Nitro v4 or later
                out[dev][name] = int(open(path).read())
    return out

interval = float(sys.argv[1]) if len(sys.argv) > 1 else 10
a = snap(); time.sleep(interval); b = snap()
for dev in sorted(b):
    d = {k: b[dev][k] - a[dev].get(k, 0) for k in b[dev]}
    gbps = (d.get("tx_bytes", 0) + d.get("rx_bytes", 0)) * 8 / interval / 1e9
    bad = {k: v for k, v in d.items() if v and k not in ("tx_bytes", "rx_bytes")}
    print(f"{dev:14s} {gbps:7.1f} Gb/s  {bad if bad else 'clean'}")

CounterMeaningAction
retrans_pktsSRD packets resentSome is normal under load; compare across nodes
retrans_timeout_eventsTimeouts that moved traffic to another pathRising on one node only: suspect that host
impaired_remote_conn_eventsConnections rate-limited as impairedFind the remote peer; drain if persistent
unresponsive_remote_eventsPeer stopped respondingPeer crashed, hung or lost its network
rx_dropsReceived then droppedOften receive-side resource exhaustion

One gap to know: CloudWatch Container Insights on EKS collects the EFA metrics except the five SRD health counters (retrans_bytes, retrans_pkts, retrans_timeout_events, unresponsive_remote_events, impaired_remote_conn_events). If those matter to you, run your own collector.

EFA on EKS

On Kubernetes, the AWS EFA device plugin advertises each node's EFA devices as the extended resource vpc.amazonaws.com/efa. A pod requests them like GPUs, and for multi-node training it should request all of them on the node, or NCCL sees only part of the bandwidth:

apiVersion: v1
kind: Pod
metadata:
  name: trainer-0
spec:
  containers:
  - name: train
    image: <registry>/train:efa   # libfabric + aws-ofi-nccl + NCCL, matched versions
    resources:
      limits:
        nvidia.com/gpu: 8
        vpc.amazonaws.com/efa: 32      # p5.48xlarge advertises 32
        hugepages-2Mi: 5120Mi          # example sizing, not an AWS requirement
        memory: 1000Gi                 # example sizing
    env:
    - name: FI_PROVIDER
      value: efa
    - name: NCCL_DEBUG
      value: INFO

Check the node first with kubectl describe node: a zero or missing EFA count means the plugin is not running or the node launched without EFA interfaces. The container must carry libfabric and the plugin itself; the host's copies are not visible inside it. Use an image built from the AWS deep learning containers, or install the EFA software into your image and record the versions.

Worked example: a slow 16-node all-reduce

A 16-node P5 job reports step times 40 percent slower than last week. Work outward from NCCL.

  1. Is EFA in use at all? In the NCCL_DEBUG=INFO log, look for the NET/OFI plugin lines naming the efa provider. If NCCL reports its socket transport instead, the image lost the plugin or FI_PROVIDER points elsewhere. Nothing else will fail.
  2. Is the fabric slow everywhere? Run all_reduce_perf from nccl-tests across all 16 nodes and compare bus bandwidth with the baseline recorded for this image. If two-node pairs are all healthy but 16 nodes are slow, suspect one node.
  3. Which node? Run efa_deltas.py 30 on every node during the benchmark. Fifteen show only modest retrans_pkts; one shows retrans_timeout_events climbing and a third of the throughput of its peers.
  4. Act. Cordon and drain that node, restart from the last checkpoint on a replacement, and open a case with the instance ID and counter output.

Because every rank waits for the slowest in a collective, one impaired host sets the pace of all 128 GPUs; see all-reduce in depth for why bus bandwidth is the number to compare.

A preflight check before every job

The cheapest time to find a sick node is before a job starts. A preflight check runs on the allocated nodes, in the same container image as the training job, and refuses to start training if any node falls below its baseline. Pairwise tests matter because a full-cluster benchmark only tells you that something is slow, not which node.

#!/usr/bin/env bash
# preflight.sh hosts.txt BASELINE_GBPS: pairwise nccl-tests, flag slow pairs
set -euo pipefail
mapfile -t H < "$1"; BASE=$2; FAIL=0
for ((i = 0; i + 1 < ${#H[@]}; i += 2)); do
  a=${H[$i]}; b=${H[$((i + 1))]}
  bw=$(mpirun -np 16 -N 8 -H "$a:8,$b:8" -x FI_PROVIDER=efa \
        ./all_reduce_perf -b 1G -e 1G -g 1 | awk '/Avg bus bandwidth/ {print $NF}')
  ok=$(awk -v x="$bw" -v b="$BASE" 'BEGIN {print (x >= 0.85 * b) ? 1 : 0}')
  echo "$a $b busbw=$bw ok=$ok"
  [ "$ok" = 1 ] || FAIL=1
done
exit $FAIL

Pair nodes in a different order on the next run, so a slow result can be pinned to one host by intersection rather than guessed. The 85 percent threshold is a starting point; set it from the spread you see across healthy pairs on your own fleet. Run the counter script alongside the test and keep both outputs with the job record, so a later regression can be compared against what healthy looked like on the same image.

Failure modes

  • Silent TCP fallback from a missing plugin or wrong provider: the job runs at a fraction of expected speed.
  • Init hang from a security group that does not allow all traffic within itself, or nodes split across AZs.
  • Partial bandwidth from launching with fewer EFA interfaces than network cards, or pods requesting fewer EFA devices than the node has.
  • Version skew between libfabric, aws-ofi-nccl and NCCL in a custom image: crashes at init or poor bandwidth. Pin all three together.
  • Mixed generations in one job, such as P4d with P5, which EFA does not support.
  • One sick host showing growing timeout or impaired-connection counters and slowing every collective.

Trade-offs

EFA's design trades some things away. Multipath spraying and hardware retransmission make tail latency steadier on a shared, oversubscribed cloud network than single-path RDMA, without lossless Ethernet configuration. In exchange you get a proprietary transport that works only inside AWS, a non-routable fabric confined to one zone, and a performance profile that depends on user-space software versions. Compared with dedicated InfiniBand clusters, small-message latency is typically higher, which matters more for tightly coupled HPC codes and small-message collectives than for large gradient all-reduces.

The practical trade is operational: you cannot tune switches, only measure from the hosts. That makes per-node counters and a recorded bandwidth baseline the core of running EFA well.

What to do next

  1. Confirm your instance's row in the EFA supported-types table: RDMA read, write and GPUDirect.
  2. Put every node in one AZ and a cluster placement group, with a self-referencing all-traffic security group.
  3. Launch with an EFA interface on every network card; on EKS, request every vpc.amazonaws.com/efa device.
  4. Pin libfabric, aws-ofi-nccl and NCCL as one set and record the versions in the image.
  5. On every new image, read the NCCL_DEBUG=INFO log once and confirm the efa provider is selected.
  6. Record nccl-tests bus bandwidth at 2 and N nodes as a baseline.
  7. Run a counter-delta collector, including the SRD health counters Container Insights omits, and alert on per-node outliers.
  8. Write a drain-and-replace runbook for a node with rising timeout counters.
Key takeaway: EFA gives AWS instances OS-bypass networking over SRD, a transport that is reliable but unordered and sprays packets across many paths, with recovery on the Nitro card. It is confined to one AZ and VPC, and its speed depends on matched libfabric, aws-ofi-nccl and NCCL versions. Verify the provider, baseline bandwidth, and watch per-node SRD counters.