A Trn2 UltraServer is four Trainium2 instances joined into one 64-chip NeuronLink domain. Together they hold 6,144 GiB of HBM and deliver 42.8 PFLOPS of dense BF16, or 83.2 PFLOPS of FP8. The headline numbers are easy to quote. Getting a training job to use them is harder, and it comes down to one thing: the bandwidth between two chips depends on where they sit. Inside an instance, inside the UltraServer, and across EFA are three different networks, and each is several times slower than the one before it.
This page treats the UltraServer as a system. It explains the topology and reconciles the published bandwidth figures. It estimates what collectives cost at each level, maps tensor, pipeline and data parallelism onto the 4×4×4 shape, and works through the memory plan for a 405-billion-parameter model. It also covers what a single failed instance does to a 64-chip job and how Trn3's switched fabric changes the picture. The NeuronCore itself, numerics and the generation table are on AWS Trainium in depth, and this page assumes them.
The shape of the domain
The Neuron architecture documentation describes the topology precisely. In a trn2.48xlarge, or its UltraServer variant trn2u.48xlarge, 16 Trainium2 chips form a 4×4 two-dimensional torus. In an UltraServer, chips with the same coordinates in each of the four instances are connected in a ring. Think of it as a 4×4×4 grid. Two dimensions are fast and local, and the third crosses instance boundaries.
| Level | Chips | Bandwidth per chip | HBM in the domain |
|---|---|---|---|
| One Trainium2 chip | 1 | HBM 2.9 TB/s | 96 GiB |
| trn2u.48xlarge, 4x4 torus | 16 | NeuronLink-v3: 1,024 GB/s | 1,536 GiB |
| UltraServer, rings across 4 instances | 64 | NeuronLink-v3: 256 GB/s | 6,144 GiB |
| Beyond the UltraServer | any | EFA: 3,200 Gbps per instance, about 25 GB/s | - |
The 1,024 and 256 GB/s figures add up to the 1.28 TB/s per chip that the Trainium2 chip page quotes. That headline is the sum of two different networks, not bandwidth you get toward any one peer. Traffic that crosses instances gets a quarter of what traffic inside the torus gets. EFA is a 400 GB/s pipe per instance, shared by 16 chips, which leaves about 25 GB/s each, roughly a tenth of the cross-instance ring.
What collectives cost at each level
A useful first-order model for an all-reduce of S bytes over n participants with per-participant bandwidth B is the ring formula. It ignores latency, so it is optimistic for small messages, but it orders the options correctly:
def allreduce_seconds(size_bytes, n, gbytes_per_s):
"""Ring all-reduce: each participant sends and receives 2(n-1)/n of the data."""
return 2 * (n - 1) / n * size_bytes / (gbytes_per_s * 1e9)
act = 2 * 4096 * 16384 # one BF16 activation block: seq 4096 x hidden 16384
grads = 2 * 405e9 / 64 # BF16 gradient shard per chip for a 405B model on 64 chips
print(f"TP all-reduce, 16 chips in torus : {allreduce_seconds(act, 16, 1024)*1e3:.2f} ms")
print(f"TP all-reduce, ring of 4 : {allreduce_seconds(act, 4, 256)*1e3:.2f} ms")
print(f"Same block over EFA : {allreduce_seconds(act, 16, 25)*1e3:.2f} ms")
print(f"Gradient shard, DP=2 over EFA : {allreduce_seconds(grads, 2, 25):.2f} s")The output is about 0.25 ms, 0.79 ms, 10.1 ms and 0.51 s. A transformer layer runs several tensor-parallel all-reduces in every forward and backward pass. At a quarter of a millisecond they overlap with compute. At 10 ms they would dominate the step. That is the whole argument for keeping tensor parallelism on NeuronLink. The gradient all-reduce for data parallelism happens once per step and can overlap with the backward pass, so EFA is acceptable for it. Before you commit, measure real numbers with a collective micro-benchmark on the actual instances, because the torus layout and per-message latency change the constants. Where the 2(n-1)/n factor comes from is derived in ring all-reduce in depth.
Mapping parallelism onto 4x4x4
Mapping parallelism onto an UltraServer follows the bandwidth table from fastest to slowest:
| Dimension | Communication | Where to place it |
|---|---|---|
| Tensor parallel (TP) | All-reduce or reduce-scatter inside every layer | Inside the 16-chip torus: across 8 or 16 chips of one instance |
| Sequence or context parallel | Exchanges K/V blocks per attention layer | Inside the torus, or along the 4-instance ring if TP already fills it |
| Pipeline parallel (PP) | Point-to-point activations per micro-batch | Along the ring: one stage per instance, PP=4 per UltraServer |
| Expert parallel (EP) | All-to-all token routing | Inside the NeuronLink domain. All-to-all is what the torus does worst |
| Data parallel or sharded optimizer | Gradient reduce once per step | Across UltraServers over EFA |
Pipeline stages fit the ring well because point-to-point traffic between neighbours is exactly what a ring provides, and per-micro-batch activation volume is modest. Tensor parallelism wider than 16 crosses the ring at a quarter of the bandwidth, so do it only for layers too large to shard 16 ways. Expert parallelism is the awkward case. All-to-all on a torus takes multiple hops, so keep the expert group inside one instance where you can. The general patterns are in tensor parallelism and pipeline parallelism.
One more layer of mapping sits below the chip. Trn2 defaults to a Logical NeuronCore configuration of 2 (LNC=2), so each chip's eight physical cores appear as four logical cores. The runtime variable NEURON_LOGICAL_NC_CONFIG and the compiler option --logical-nc-config must agree. Count ranks in logical cores when you build the process grid: 64 per instance at LNC=2, 256 per UltraServer. Tensor parallelism across all 16 chips of an instance is therefore a TP degree of 64 in logical cores. This page quotes layouts in chips; convert before you configure the job.
Worked example: a 405B model
Does a 405B dense model train on one UltraServer? Count the bytes per parameter for mixed-precision training with Adam: 2 for BF16 weights, 2 for BF16 gradients, 4 for an FP32 master copy and 8 for the two Adam moments, 16 in total.
GiB = 2**30
params = 405e9
state = params * 16 / GiB # weights, grads, master, Adam moments
print(f"training state : {state:,.0f} GiB") # ~6,035 GiB
print(f"one UltraServer HBM: {64 * 96:,} GiB") # 6,144 GiB
for n_us in (1, 2, 4):
chips = 64 * n_us
per_chip_state = state / chips # fully sharded (ZeRO-3 / FSDP style)
headroom = 96 - per_chip_state
print(f"{n_us} UltraServer(s): {per_chip_state:5.1f} GiB state/chip, {headroom:5.1f} GiB left")
# A concrete layout: TP over 16 chips x PP 4 = 64 chips per replica, DP = 2,
# optimizer state (master + Adam, 12 B/param) sharded over both replicas (ZeRO-1)
replica_chips, dp = 64, 2
zero1 = params * 4 / replica_chips / GiB + params * 12 / (replica_chips * dp) / GiB
print(f"TP16 x PP4 x DP2, ZeRO-1: {zero1:.1f} GiB/chip, {96 - zero1:.1f} GiB left")The state alone is about 6,035 GiB, which is 98 percent of one UltraServer's 6,144 GiB. That leaves around 1.7 GiB per chip for activations, the runtime and fragmentation, which is not a workable plan. Fully sharded across two UltraServers, state falls to about 47 GiB per chip. That is the lower bound. A practical layout for two UltraServers is tensor parallelism across the 16 chips of each instance, PP=4 along the ring and DP=2 across EFA, with only the optimizer state sharded across the two replicas. Weights and gradients are then split 64 ways and master weights plus Adam moments 128 ways, about 59 GiB per chip, leaving about 37 GiB for activations. The one EFA collective per step is the 0.51 s gradient reduction estimated above. It must overlap with the backward pass, or it adds directly to step time. Activation checkpointing and the micro-batch size then decide how much of the remaining headroom you use.
The 64-chip failure domain
A 64-chip NeuronLink domain is also a 64-chip failure domain. When one of the four instances fails, its 16 chips are gone and the rings that cross it are broken. With TP and PP spread over the domain, every rank stalls on its next collective. The job does not degrade gracefully: it stops, and it restarts from the last checkpoint once the domain is whole again. Across many UltraServers, the job fails whenever any one of them fails, so the job-level mean time between failures shrinks as you scale out.
Choose the checkpoint interval deliberately. Young's approximation gives the interval that minimises expected lost work: the square root of 2 × checkpoint cost × MTBF.
import math
def checkpoint_interval_s(ckpt_seconds, mtbf_hours):
return math.sqrt(2 * ckpt_seconds * mtbf_hours * 3600)
# e.g. 90 s to write a sharded checkpoint, job-level MTBF of 24 h
print(checkpoint_interval_s(90, 24) / 60) # ~ 66 minutes between checkpointsThe inputs are yours to measure: the time to write a sharded checkpoint at your model size, and the failure rate you actually observe. Asynchronous checkpointing, where you snapshot to host memory and then upload, cuts the cost term and allows shorter intervals. Restarting needs a replacement for the whole UltraServer, not just one instance, so automate the restart path and rehearse it before the first long run. Sharded checkpoint formats are covered in training checkpointing.
Trn3 UltraServers: a switched fabric
Trainium3 changes the topology question. The Neuron Trn3 architecture page describes two UltraServer configurations. Gen1 has 64 Trainium3 chips in four servers of 16. Gen2 has 144 chips in 36 servers of four. Both connect chips through NeuronSwitch-v1, an all-to-all fabric, with NeuronLink-v4 at 2,048 GiB/s per device. HBM is 9,216 GiB for Gen1 and 20,736 GiB for Gen2, which is 144 GiB per chip.
For software, the switch removes most of the placement puzzle above. Inside a switched domain, any two chips have the same path, so the difference between the torus and the ring disappears. All-to-all traffic for expert parallelism stops being the worst case. The domain also grows to 144 chips, so a 405B-class model's training state fits inside one domain with real headroom. The failure-domain lesson still applies, and with 36 servers per Gen2 domain it matters more. For a team on Trn2 today, the practical consequence is to keep the parallel layout in configuration rather than hard-coding it, so that moving to Trn3 means changing TP, EP and PP degrees, not rewriting the job.
Failure modes
- TP across the ring by accident. A process grid that orders ranks by instance first puts tensor-parallel groups across instances at a quarter of the bandwidth. Step time grows with no error. Print the rank-to-chip map at start-up and check it.
- LNC mismatch. A graph compiled for one logical-core configuration and run with another fails to load or uses half the cores. Set both values explicitly in the job spec.
- Memory plan with no headroom. Fitting the optimizer state with a few GiB to spare fails at the first long sequence or larger micro-batch. Plan for activations first.
- Partial UltraServer. Scheduling the four instances as independent nodes, or replacing one with an instance outside the UltraServer, silently drops the ring links. Schedule and replace the four as a unit.
- EFA saturation. Data-parallel gradient traffic that does not overlap with compute shows up as idle NeuronCores at the end of every step. Profile, then bucket and overlap the reductions.
- Untested restart. The first failure is the wrong time to find that checkpoints cannot be resharded onto a new domain.
Trade-offs
An UltraServer buys a scale-up domain four times bigger than one instance. That makes wide tensor and pipeline layouts possible without EFA on the critical path. The costs are a 4:1 bandwidth step inside the domain that your layout must respect, a 64-chip unit of failure and scheduling, and a capacity unit that is harder to obtain than single instances. Check current purchase options and Region availability before planning around one. For models whose state fits comfortably in 16 chips, plain trn2.48xlarge instances with data parallelism over EFA are simpler and easier to get. UltraServers pay off for models that need more than 16 chips of tensor and pipeline parallelism, and Trn3's switched domains raise that threshold further.
What to do next
- Compute training state at 16 bytes per parameter and divide by 96 GiB per chip to see whether you need one instance, one UltraServer or several.
- Lay out the process grid: TP inside the torus, PP along the ring, DP over EFA, counted in logical NeuronCores at your LNC setting.
- Print and verify the rank-to-chip map at the start of every job.
- Run collective benchmarks at each level and replace the model's constants with your measurements.
- Measure checkpoint write time, set the interval with Young's formula and rehearse a full-domain restart.
- Keep TP, PP and EP degrees in configuration so that the same job can move to Trn3.