AWS Trainium is Amazon's own accelerator for training deep learning models. It is the training counterpart to Inferentia. Like a GPU it pairs dense matrix units with high-bandwidth memory and a fast chip-to-chip fabric, but it is programmed differently. Your model is compiled ahead of time into a program that schedules several specialised engines inside each core. That changes how you think about shapes, numerics, parallelism and sizing.
This article explains Trainium from the hardware up, in terms of what training code experiences: what is inside a NeuronCore, how the three chip generations compare, why stochastic rounding matters for BF16 training, what logical NeuronCores are, how the instance and UltraServer topology should shape your parallelism plan, and how to size a job before you reserve capacity. The software stack has its own page, AWS Neuron SDK in depth, and serving is covered in AWS Inferentia in depth. Specifications were checked against the Neuron documentation and AWS announcements on 2026-10-03.
The hardware in one picture
Three generations compared
Three generations matter in 2026. The figures below come from the Neuron architecture pages, except Trainium3, which is quoted as AWS announced it on 2 December 2025 when Trn3 UltraServers became generally available.
| Trainium (Trn1) | Trainium2 (Trn2) | Trainium3 (Trn3) | |
|---|---|---|---|
| NeuronCores per chip | 2 × NeuronCore-v2 | 8 × NeuronCore-v3 | not covered here |
| HBM per chip | 32 GiB, 820 GiB/s | 96 GiB, 2.9 TB/s | 144 GB HBM3e, 4.9 TB/s |
| Dense compute per chip | 190 TFLOPS BF16 | 667 TFLOPS BF16, 1,299 FP8 | 2.52 PFLOPS FP8 |
| Chips per instance or server | 16 (trn1.32xlarge) | 16 (trn2.48xlarge); 64 per UltraServer | up to 144 per UltraServer |
| Chip-to-chip interconnect | NeuronLink-v2, 2D torus | NeuronLink-v3, 1.28 TB/s per chip | NeuronSwitch-v1 all-to-all fabric |
| Scale-out network | 800 Gbps EFA (1,600 on trn1n) | 3,200 Gbps EFA | not covered here |
A trn2.48xlarge therefore holds 1,536 GiB of HBM and about 10.7 PFLOPS of dense BF16 (16 × 667 TFLOPS). Trn1 remains useful for smaller models and for teams that already have it running. New large training jobs start on Trn2 or Trn3.
Inside a NeuronCore
A NeuronCore is not a cluster of thousands of small threads like a GPU streaming multiprocessor. It is a small number of large engines, each with its own instruction stream. On NeuronCore-v3:
- The tensor engine runs matrix multiplies, convolutions and transposes. It provides 158 dense FP8 TFLOPS per core and supports several structured-sparsity patterns. Nearly all of a transformer's FLOPs land here.
- The vector engine handles operations where each output depends on many inputs, such as layer normalisation, softmax reductions and pooling.
- The scalar engine handles element-wise work such as activations and scaling.
- The GpSimd engine is eight fully programmable 512-bit vector processors that can run custom C code with direct access to on-chip SRAM. This is the escape hatch for operations nothing else covers.
- 28 MB of software-managed SRAM per core sits between the engines and HBM. It is not a hardware cache, so the compiler decides what to load and when, and DMA engines move tiles in and out.
The practical consequences: the compiler must see the whole graph with fixed shapes to schedule four engines and the DMA queues. Matmul-heavy models with large, regular dimensions keep the tensor engine busy. Models dominated by gathers, scatters, odd reshapes or tiny matrices leave it idle while slower engines work. When profiling shows a hot operation the compiler handles badly, the answer is usually a custom kernel written with NKI, the Neuron Kernel Interface, which the SDK article covers.
Numerics and stochastic rounding
Trainium supports FP32, TF32, BF16, FP16 and configurable FP8 (cFP8) on the tensor engine, and its rounding mode is programmable: round-to-nearest-even or stochastic rounding. This matters most for training in pure BF16.
BF16 has 8 bits of mantissa precision. Add a tiny update, such as learning rate times gradient, to a weight of magnitude 1.0, and round-to-nearest returns the same weight whenever the update is below half a unit in the last place. Millions of small updates vanish, and training stalls. The usual fix is to keep FP32 master weights, which costs 4 extra bytes per parameter. Stochastic rounding rounds up with probability proportional to the remainder, so the expected value of each update is preserved and small updates accumulate on average. This simulation shows the difference:
import numpy as np
def to_bf16(x): # truncate an fp32 array to bf16 bit patterns
return (x.astype(np.float32).view(np.uint32) & 0xFFFF0000).view(np.float32)
def round_bf16(x, stochastic, rng):
bits = x.astype(np.float32).view(np.uint32).astype(np.uint64)
if stochastic:
bits = bits + rng.integers(0, 1 << 16, size=bits.shape, dtype=np.uint64)
else:
bits = bits + 0x7FFF + ((bits >> 16) & 1) # round to nearest even
return (bits & 0xFFFF0000).astype(np.uint32).view(np.float32)
rng = np.random.default_rng(0)
for stochastic in (False, True):
w = to_bf16(np.ones(10_000, dtype=np.float32))
for _ in range(1_000):
w = round_bf16(w + np.float32(1e-4), stochastic, rng) # 1e-4 << bf16 spacing at 1.0
print("stochastic" if stochastic else "nearest ", float(w.mean()))
# nearest stays at 1.0; stochastic lands near 1.1, the exact sumStochastic rounding lets some workloads drop FP32 master weights and save memory. Whether model quality holds is an empirical question for your model, so compare loss curves against a mixed-precision baseline before you depend on it. Mixed-precision training explains the baseline.
Logical NeuronCores
On Trainium2, the Neuron runtime can group physical cores into logical NeuronCores. The NEURON_LOGICAL_NC_CONFIG setting accepts only 1 or 2, and 2 is the default. With LNC=2, each chip's eight physical cores appear as four logical cores, so a trn2.48xlarge exposes 64 logical cores instead of 128 physical ones. Each logical core combines the compute and memory resources of two physical cores.
This matters in two places. First, process counts: if you launch one worker per NeuronCore, the count is per logical core, so 64 workers per instance at the default. Second, compilation: the compiler and runtime must agree on the LNC setting, so pin it in the environment of both the compile job and the training job, and include it in your compile cache key. LNC=2 suits large models, where bigger cores mean fewer, larger collectives. LNC=1 gives more, smaller cores, which helps workloads that do not scale to large per-core shapes.
Topology and parallelism
Training throughput at scale depends on bandwidth at each level. Inside a trn2.48xlarge, 16 chips form a 4x4 2D torus over NeuronLink-v3. A Trn2 UltraServer joins four trn2u.48xlarge instances into a 64-chip NeuronLink domain with 6,144 GiB of HBM. Between instances and UltraServers, traffic crosses EFA at 3,200 Gbps per instance, which is 0.4 TB/s, or about 25 GB/s per chip. Compare that with 1.28 TB/s of NeuronLink per chip: roughly a fifty-fold drop.
Map parallelism to that hierarchy the same way you would on GPUs:
- Tensor parallelism runs collectives inside every layer, so keep it inside the NeuronLink domain: within an instance, or within an UltraServer for very wide layers. See tensor parallelism.
- Pipeline parallelism sends activations between stages once per micro-batch. It tolerates EFA, at the cost of pipeline bubbles.
- Data parallelism and sharded optimisers (ZeRO or FSDP style) exchange gradients and parameters once per step, and they overlap well with compute across EFA. See ZeRO optimizer sharding.
A typical plan for a model that needs a few dozen chips: tensor parallel of 8 or 16 inside an instance or UltraServer, pipeline parallel only if one replica does not fit, and data parallel across the rest.
Sizing a training job
Size the job before you ask for capacity. Two numbers decide it: memory for model state and activations, and total FLOPs against sustained throughput. For mixed precision with Adam, a common budget is 16 bytes per parameter: BF16 weights (2) and gradients (2), FP32 master weights (4), and two FP32 Adam moments (8). Activations come on top and depend on batch, sequence length and recomputation.
# Back-of-envelope Trainium2 sizing. MFU is an assumption; measure yours.
CHIP_HBM_GIB, CHIP_BF16_TFLOPS, CHIPS_PER_INSTANCE = 96, 667, 16
def size_job(params_b, tokens_t, mfu=0.40, bytes_per_param=16, activation_share=0.35):
state_gib = params_b * 1e9 * bytes_per_param / 2**30
usable_gib = CHIP_HBM_GIB * (1 - activation_share) # leave room for activations
min_chips = -(-state_gib // usable_gib) # ceil: fully sharded state
flops = 6 * params_b * 1e9 * tokens_t * 1e12 # ~6 FLOPs per parameter per token
inst_flops = CHIPS_PER_INSTANCE * CHIP_BF16_TFLOPS * 1e12 * mfu
inst_days = flops / inst_flops / 86400
return state_gib, int(min_chips), inst_days
for p_b, t_t in [(8, 1.0), (70, 2.0)]:
st, chips, days = size_job(p_b, t_t)
print(f"{p_b}B params, {t_t}T tokens: state {st:,.0f} GiB, >= {chips} chips, "
f"{days:,.0f} instance-days at 40% MFU")For an 8B model on 1 trillion tokens, the model state is about 119 GiB, which fits one instance with plenty of room. The compute is about 4.8 × 1022 FLOPs: roughly 130 instance-days at an assumed 40 percent model FLOPs utilisation, or about 8 days on 16 instances. A 70B model has about 1,043 GiB of state, which needs most of a trn2.48xlarge's 1,536 GiB before any activations. That pushes it to at least two instances or an UltraServer, with optimizer sharding across chips. Replace the 40 percent with the utilisation you measure on a short run, because it is the least certain number here.
Running a job
The workflow differs from GPUs mainly at the start. Install a pinned Neuron SDK release (driver, runtime, compiler and framework packages together), then run neuron-ls to confirm the devices and core counts the instance exposes. Compile once on a short run and store the compile cache on shared storage. With the torch-neuronx integration that means setting NEURON_COMPILE_CACHE_URL, for example to an S3 prefix, and passing compiler options through NEURON_CC_FLAGS. Launch with the usual distributed launcher, one process per logical core:
# Example environment for a 2-instance Trn2 job (torch-neuronx integration).
export NEURON_LOGICAL_NC_CONFIG=2 # default, pinned explicitly
export NEURON_COMPILE_CACHE_URL=s3://ml-artifacts/neuron-cache/llm-8b/sdk-pinned/
torchrun --nnodes 2 --nproc_per_node 64 \
--rdzv_backend c10d --rdzv_endpoint $HEAD_NODE:29500 \
train.py --tp 16 --global-batch 1024 --seq-len 4096
neuron-top # live NeuronCore utilisation and HBM usage per coreThe training script itself depends on the integration. TorchNeuron is the native PyTorch path for Trn2 and Trn3, and torch-neuronx is the older XLA-based one. Follow the SDK page for the current API, and do not copy launch snippets across SDK releases. Checkpoint to Amazon S3 or FSx for Lustre often enough that losing a node costs minutes, not hours. At hundreds of chips, node failures are routine events, not rare ones.
Failure modes
- Recompilation every few steps: variable sequence lengths or a final partial batch create new shapes. Pad to fixed buckets and drop or pad the last batch.
- Long first step: compilation, not a hang. A warm shared cache fixes it for every later run, but only if SDK version, flags and LNC match.
- Out of device memory at load: state plus activations exceed HBM per core. Increase sharding, enable activation recomputation, or reduce micro-batch size.
- Collective hangs: ranks compiled different graphs, for example because one rank took a different code path. Keep control flow identical across ranks.
- Low utilisation despite a busy chip: time spent on vector, scalar or GpSimd engines, or waiting on DMA. Profile, then fuse operations or write a kernel for the hot one.
- Version mismatch: a runtime too old for the compiler that built a NEFF refuses to load it. Upgrade the whole SDK together.
Trade-offs
Price-performance versus ecosystem. Trainium can cost less per trained token than comparable GPU instances, and capacity is sometimes easier to get. In exchange you rely on the Neuron compiler's operator coverage and on a smaller community, and custom CUDA kernels must be rewritten.
Compile-time determinism versus flexibility. Ahead-of-time compilation gives predictable step times and makes optimisation explicit. It punishes dynamic shapes and data-dependent control flow that GPUs absorb without complaint.
UltraServer versus instances. A 64-chip NeuronLink domain makes large tensor-parallel groups practical. If your model fits comfortably in one instance, separate instances over EFA are simpler to schedule.
What to do next
- Run the sizing script with your parameter count, token budget and an honest utilisation guess, and decide between Trn2 instances, a Trn2 UltraServer or Trn3.
- Check that your model's operators compile under the current Neuron SDK on a small instance before reserving large capacity.
- Pin one SDK release, and pin
NEURON_LOGICAL_NC_CONFIGin both compile and training environments. - Make every input shape static: fixed sequence buckets, fixed micro-batch, padded final batch.
- Set up a shared compile cache keyed by SDK version, flags and LNC, and warm it with a short run.
- Choose parallelism by topology: tensor parallel inside NeuronLink, data parallel or sharding across EFA.
- Compare BF16 with stochastic rounding against FP32 master weights on a short run before dropping master weights.
- Profile one step, find the engine or DMA bottleneck, and fix the top operation before scaling out.