AWS sells its large NVIDIA training instances as two families. P4 is the A100 generation: p4d with 40 GB GPUs and p4de with 80 GB. P5 is the Hopper generation: p5 with H100, p5e and p5en with H200, plus a single-GPU p5.4xlarge. On a spec sheet they look like the same shape with bigger numbers. To software they are quite different machines, because the thing that changed most between them is not the GPU but the network per GPU, which went up eightfold.
This article explains the six instance types as a training or serving engineer meets them: what is inside, how bandwidth divides per GPU, how to launch with EFA correctly, which software layers must line up for NCCL to use the fast path, where to put data and checkpoints, and how to choose. Specifications are from AWS's instance pages and EC2 documentation as of October 2026; check them again before you size a purchase. For the GPU itself see the H100 architecture and the H200.
The six instance types
| Type | GPUs | GPU memory | vCPU / RAM | Network | Cards | Local NVMe | EBS |
|---|---|---|---|---|---|---|---|
| p4d.24xlarge | 8 x A100 | 8 x 40 GB | 96 / 1152 GiB | 400 Gbps EFA | 4 | 8 x 1 TB | 19 Gbps |
| p4de.24xlarge | 8 x A100 | 8 x 80 GB | 96 / 1152 GiB | 400 Gbps EFA | 4 | 8 x 1 TB | 19 Gbps |
| p5.4xlarge | 1 x H100 | 80 GB | 16 / 256 GiB | 100 Gbps EFA | 1 | 1 x 3.84 TB | 10 Gbps |
| p5.48xlarge | 8 x H100 | 8 x 80 GB | 192 / 2 TiB | 3,200 Gbps EFA | 32 | 8 x 3.84 TB | 80 Gbps |
| p5e.48xlarge | 8 x H200 | 8 x 141 GB | 192 / 2 TiB | 3,200 Gbps EFA | 32 | 8 x 3.84 TB | 80 Gbps |
| p5en.48xlarge | 8 x H200 | 8 x 141 GB | 192 / 2 TiB | 3,200 Gbps EFAv3 | 16 | 8 x 3.84 TB | 100 Gbps |
The P4 types use Intel Xeon Platinum 8275CL CPUs on Nitro v3 and NVSwitch at 600 GB/s per GPU. p5 and p5e use AMD EPYC 7R13 on Nitro v4 with NVSwitch at 900 GB/s. p5en moves to Intel Sapphire Rapids with PCIe Gen5 between CPU and GPU, third-generation EFA on Nitro v5, and the same 3,200 Gbps over 16 network cards instead of 32; AWS says EFAv3 gives up to 35 percent lower latency than P5. The p5.4xlarge supports EFA but not GPUDirect RDMA and has no NVSwitch peer, so it is a development and single-GPU serving shape, not a building block for multi-node training.
Inside an instance: two fabrics
Every 8-GPU instance has two fabrics. Inside, NVLink through NVSwitch connects every GPU to every other at full bandwidth, so tensor parallelism and the intra-node half of every collective stay inside the box. Outside, EFA, AWS's RDMA-capable network interface, connects instances using SRD, a reliable datagram transport that sprays packets across many paths and tolerates reordering. With GPUDirect RDMA, the NIC reads and writes GPU memory directly, so inter-node traffic never bounces through host RAM. The layout is close to NVIDIA's own HGX design; DGX H100 system architecture walks through the same GPU-to-NIC and NUMA affinity questions.
Per-GPU network math
Divide the network by the GPUs and the generational gap becomes obvious. p4d gives each GPU 400 / 8 = 50 Gbps, about 6.25 GB/s. p5 gives each GPU 3,200 / 8 = 400 Gbps, about 50 GB/s. NVSwitch inside the instance is 600 or 900 GB/s per GPU, so on p4d the inter-node link is roughly 100 times slower than the intra-node one, and on p5 about 18 times. That ratio decides which parallelism strategies are viable across instances.
def ring_allreduce_seconds(bytes_per_gpu, gpus, link_gbytes_per_s):
"""Bandwidth term of a ring all-reduce: each GPU sends 2(N-1)/N of the buffer."""
return 2 * (gpus - 1) / gpus * bytes_per_gpu / link_gbytes_per_s
grads = 7e9 * 2 # 7B parameters, bf16 gradients = 14 GB
for name, per_gpu in [("p4d", 6.25), ("p5", 50.0)]:
t = ring_allreduce_seconds(grads, gpus=128, link_gbytes_per_s=per_gpu)
print(f"{name}: {t:.2f} s per full-gradient all-reduce (bandwidth bound only)")
# p4d: 4.45 s p5: 0.56 sThis is a lower bound on a hierarchical all-reduce whose slowest hop is the network; real NCCL bus bandwidth lands below the line rate, and overlap with the backward pass hides part of it. The point is the ratio. On p4d, pure data parallelism of a 7B model at 128 GPUs spends seconds per step on gradient exchange, so you shard (ZeRO or FSDP), accumulate gradients or compress. On p5, the same exchange fits under the backward pass of a moderately sized batch. See NCCL collectives for the algorithm details.
Launching with EFA
Bandwidth is only available if the instance launches with its EFA interfaces attached. On p5.48xlarge and p5e.48xlarge that means 32 network cards: an ENA interface for IP traffic as the primary interface, an EFA-only interface on card 0 at device index 1, and an EFA-only interface on each of cards 1 to 31. AWS documents that the 3,200 Gbps is shared between EFA and IP traffic, with IP capped at 800 Gbps, so heavy IP traffic (for example S3 reads) reduces collective bandwidth. p5en has 16 cards and p4d 4.
# p5.48xlarge: primary ENA plus 32 EFA-only interfaces in one placement group.
NI="NetworkCardIndex=0,DeviceIndex=0,Groups=$SG,SubnetId=$SUBNET,InterfaceType=interface"
NI="$NI NetworkCardIndex=0,DeviceIndex=1,Groups=$SG,SubnetId=$SUBNET,InterfaceType=efa-only"
for i in $(seq 1 31); do
NI="$NI NetworkCardIndex=$i,DeviceIndex=0,Groups=$SG,SubnetId=$SUBNET,InterfaceType=efa-only"
done
aws ec2 run-instances --instance-type p5.48xlarge --count 2 \
--image-id "$AMI" --key-name "$KEY" \
--placement "GroupName=$CLUSTER_PG" \
--network-interfaces $NIThree details cause most launch-time problems. The security group must allow all traffic to and from itself, because EFA traffic is filtered like any other. All instances in a job should share a cluster placement group in one Availability Zone. And capacity is the real constraint: large P5 fleets are usually bought as reservations or EC2 Capacity Blocks for a fixed future window rather than launched on demand, so the launch template must match the reserved type and zone exactly.
The software path from NCCL to the wire
For NCCL to use EFA, five layers must agree: the NVIDIA driver, NVIDIA Fabric Manager (which configures NVSwitch, and without which CUDA reports the system as not initialised), the EFA kernel driver and libfabric from the AWS EFA installer, the aws-ofi-nccl plugin that lets NCCL talk to libfabric, and NCCL itself. The AWS Deep Learning AMIs and containers ship matched versions; if you build your own image, install them as a set and record their versions.
fi_info -p efa -t FI_EP_RDM | grep -c 'provider: efa' # non-zero: libfabric sees EFA
systemctl is-active nvidia-fabricmanager # must be active on 8-GPU types
nvidia-smi topo -m # GPU, NIC and NUMA affinity
export FI_PROVIDER=efa
export FI_EFA_USE_DEVICE_RDMA=1 # p4d with older libfabric; newer releases detect it
export NCCL_DEBUG=INFO # confirm the OFI plugin and efa provider are selected
mpirun -np 16 -N 8 --hostfile hosts ./build/all_reduce_perf -b 8 -e 8G -f 2 -g 1Read the NCCL_DEBUG=INFO output from the first run of every new image. If it shows NCCL falling back to its socket transport instead of the OFI plugin, the job will run, slowly, and nothing else will tell you. nccl-tests across two instances should reach a bus bandwidth that is a large fraction of the per-GPU figure above; record it as the baseline for that image.
Storage: instance store, FSx and S3
Each 8-GPU instance has local NVMe instance store: 8 x 1 TB on P4, 8 x 3.84 TB on P5. It is fast and free, and it disappears when the instance stops, so use it for dataset caches, compiled kernels and the first write of a checkpoint, never as the only copy. A RAID 0 across the drives gives one large scratch volume. For shared training data, FSx for Lustre linked to S3 is the common choice; for durable checkpoints, copy to S3 asynchronously after writing locally. EBS bandwidth (19 Gbps on P4, 80 on p5 and p5e, 100 on p5en) is enough for boot and small volumes but is the wrong path for terabyte checkpoints.
Worked example: sizing serving and fine-tuning
Suppose you need to serve a 70B-parameter model in bf16 and also fine-tune a 13B model. The 70B weights are about 140 GB before KV cache. On p4d (320 GB total) they fit across eight GPUs with about 180 GB left, roughly 22 GB per GPU after weights, which caps batch size and context length relative to the larger types. On p4de or p5 (640 GB) there is room for a useful cache at tensor parallelism of 8. On p5e or p5en (1,128 GB, 141 GB per GPU) the weights fit on two GPUs with cache to spare, so one instance can host several replicas at tensor parallelism 2 or 4, which usually serves more requests per dollar than one TP=8 replica.
For the 13B fine-tune on 32 GPUs, gradients are 26 GB in bf16. By the formula above, a bandwidth-bound all-reduce takes about 8 s on p4d and 1 s on p5. On p4d you would use FSDP with gradient accumulation to amortise that cost; on p5 plain data parallelism with overlap is viable. If the job fits in 8 GPUs, it never touches EFA, and a single p4de can be the cheaper answer. Bare metal versus cloud GPU cost covers the utilisation side of that decision.
Failure modes
- Launched without EFA interfaces. The job runs over TCP at a fraction of the speed. Check
fi_info -p efain the job prologue and fail fast. - Fabric Manager not running. CUDA initialisation fails with a system-not-ready error on 8-GPU types. Enable the service in the image.
- Mismatched plugin and NCCL. NCCL silently falls back to sockets. Pin the versions together and grep the debug log.
- Placement spread. Instances outside one cluster placement group see higher latency and lower bisection bandwidth; collectives slow without errors.
- Security group blocks self-traffic. The first collective hangs.
- Instance store treated as durable. A stop or a hardware retirement loses the only checkpoint.
- A bad GPU or NIC. One slow rank sets the pace for all. Run a short nccl-tests and GPU burn-in on every new instance and replace outliers before the job.
Trade-offs
P4 instances are older, more often available and cheaper per hour, and they are good at work that stays inside one instance: single-node fine-tuning, serving models that fit in 320 or 640 GB, and evaluation. They are weak at multi-node data parallelism because of the 50 Gbps per GPU. P5 is the default for multi-node training on Hopper; p5e and p5en add memory per GPU for long contexts and larger serving replicas, and p5en's newer EFA and PCIe Gen5 help latency-bound collectives and host-to-device transfers. Newer P6 Blackwell instances exist above them; the same analysis applies. Spot capacity can work for checkpoint-tolerant jobs, see spot GPU instances.
What to do next
- Compute per-GPU network bandwidth and gradient size for your job, and decide whether it needs multi-node training at all.
- Build or pick an image with matched driver, Fabric Manager, EFA installer, aws-ofi-nccl and NCCL, and record the versions.
- Create a launch template with the full EFA interface layout for your type and a cluster placement group.
- On first launch, run
fi_info -p efa,nvidia-smi topo -mand a two-instance nccl-tests run; save the bus bandwidth as a baseline. - Stage data on local NVMe or FSx and write checkpoints locally first, then copy to S3.
- Plan capacity with reservations or Capacity Blocks before committing to a schedule.