If you searched for Inferentia3, here is the short answer. As of 7 October 2026, AWS has not announced a third Inferentia chip. The Inferentia product page lists two generations: the original Inferentia behind EC2 Inf1, and Inferentia2 behind Inf2. Meanwhile the newest Neuron inference software targets Trainium. Since Neuron 2.29, NxD Inference, the library behind Neuron's large-model serving, supports models only on Trn2 and newer. The Trainium3 architecture documentation names autoregressive inference serving as a target for its switch fabric.
So the practical question is not when Inferentia3 will arrive. It is what to do with the Inferentia capacity you already run, or are about to buy. This page answers that question. It covers what changed between generations and which workloads still belong on Inf2. It walks through migrating an Inf1 fleet, gives a method for comparing platforms in cost per million requests, and lays out the decision for large language models. Chip internals, compilation and shape bucketing have their own page, AWS Inferentia in depth, and this page builds on it.
Two generations, two toolchains
Each Inferentia generation pairs a chip with its own toolchain, and the toolchain is what your code depends on. A model compiled for one generation does not run on the other. Migrating means recompiling, so treat a generation change as a software migration, not a hardware swap.
| Inferentia (Inf1) | Inferentia2 (Inf2) | Inferentia3 | |
|---|---|---|---|
| Core | NeuronCore-v1, four per chip | NeuronCore-v2, two per chip (the core Trn1 also uses) | not announced |
| Memory per chip | small on-chip memory plus DRAM | 32 GiB HBM | - |
| Chip-to-chip | limited; pipelining across cores | NeuronLink-v2, 192 GiB/s per chip on 24xl and 48xl | - |
| PyTorch package | torch-neuron (torch.neuron) | torch-neuronx (torch_neuronx) | - |
| Compiler | neuron-cc | neuronx-cc | - |
| SDK status in 2026 | legacy toolchain; check your release notes | current for traced models; LLM serving pinned at 2.28 | - |
Two rows decide most migrations. The compiler row means every Inf1 artifact must be rebuilt from the source model. The chip-to-chip row explains why Inf2 can serve models larger than one chip with tensor parallelism, while Inf1 could not do this in any useful way. Inf2 comes in four sizes: 1, 1, 6 and 12 chips, with 4, 32, 96 and 192 vCPUs. Only the 6- and 12-chip sizes have NeuronLink between chips.
Triage: what still belongs on Inf2
Sort every model you serve on Inferentia into one of three buckets before you change anything. The bucket depends on which software path the model uses, not on its size.
| Workload | Software path on Inf2 | Verdict |
|---|---|---|
| Text encoders, embedding models, rerankers, classifiers | torch_neuronx.trace with shape buckets | Good fit. Supported, cheap per request, latency is predictable |
| Vision backbones, detection, OCR, speech encoders | torch_neuronx.trace | Good fit if inputs can be resized or padded to a few fixed shapes |
| Diffusion pipelines | Trace each sub-network separately | Workable. Test image quality after casting to lower precision |
| Decoder LLMs served with NxD Inference | Needs Neuron 2.28 or earlier on Inf2 | Frozen dependency. Plan the exit |
| Anything on Inf1 | Legacy torch-neuron | Migrate now |
The first two rows are why Inf2 is still worth running in 2026. Encoders and vision models compile into a static graph once. They run at fixed shapes, so the chip's low-precision compute stays busy. They also fit in one chip's 32 GiB, so you can pack many of them onto one instance. A retrieval stack that embeds and reranks millions of passages a day is the classic Inf2 customer.
Leaving Inf1: a migration runbook
An Inf1 fleet runs on a legacy toolchain that current Neuron releases have moved away from; check the release notes for your version to see exactly what is still shipped. Every security patch to an old host image is your problem. The migration is mechanical, but it has four steps, and teams that skip the last one learn about numeric drift from their users.
- Find the source model. You need the original PyTorch weights and code, not the Inf1 artifact. A compiled Inf1 model cannot be converted.
- Recompile with the Inf2 toolchain, using the same example shapes you served on Inf1 so that you can compare like with like.
- Compare outputs against the CPU or GPU reference on a held-out set of real requests.
- Re-measure throughput and latency per NeuronCore, because the core count, batch sizes and best number of workers all change.
# Inf1 (legacy): torch-neuron + neuron-cc
import torch, torch_neuron
inf1 = torch.neuron.trace(model, example_inputs=[ids, mask])
# Inf2: torch-neuronx + neuronx-cc. Same model object, new compiler.
import torch, torch_neuronx
model.eval()
inf2 = torch_neuronx.trace(model, (ids, mask))
torch.jit.save(inf2, "encoder_b8_s128.pt") # build in CI, ship the artifact
# Agreement gate: run before any traffic moves
import torch.nn.functional as F
def agreement(ref_model, neuron_model, batches, min_cos=0.999):
worst = 1.0
with torch.no_grad():
for ids, mask in batches:
a = ref_model(ids, mask).float()
b = neuron_model(ids, mask).float()
worst = min(worst, F.cosine_similarity(a, b, dim=-1).min().item())
assert worst >= min_cos, f"worst cosine {worst:.5f} below {min_cos}"
return worstThe threshold depends on the task. For embeddings, cosine similarity against the reference is the right check, together with a retrieval recall test on a labelled set. For classifiers, compare the top-1 label and alert on any change in agreement above a small fraction. Run the gate in CI on every SDK upgrade, not just once.
Cost per million requests, measured
Platform choices fail when teams compare peak benchmark numbers. The number that matters is cost per million requests at your latency target, measured with your own model and traffic shape. The method is the same for Inf2, Trainium and GPUs:
- Fix the latency objective, for example p99 under 50 ms for a reranker call.
- On one instance, raise offered load step by step and record throughput and p99 at each step.
- Take the highest throughput whose p99 still meets the objective. That is the usable capacity.
- Divide the hourly price by the requests served per hour.
def cost_per_million(price_per_hour, usable_rps):
"""price in your currency per instance-hour; usable_rps at the latency target."""
return price_per_hour / (usable_rps * 3600) * 1_000_000
def usable_rps(sweep, p99_target_ms):
"""sweep: list of (offered_rps, achieved_rps, p99_ms) from a load test."""
ok = [achieved for offered, achieved, p99 in sweep if p99 <= p99_target_ms]
return max(ok) if ok else 0.0Worked example (illustrative numbers, not a benchmark). Suppose a reranker sweep on an inf2.xlarge meets a 50 ms p99 up to 900 requests per second, and a GPU instance meets it up to 2,400. Suppose too that the GPU instance costs 3.5 times as much per hour. With the Inf2 price written as P, Inf2 costs P / (900 × 3600) per request and the GPU costs 3.5P / (2,400 × 3600). The ratio is 3.5 × 900 / 2,400 = 1.31, so the GPU costs 31 percent more per request, even though it is 2.7 times faster. Now replace the assumptions with your own sweep and current on-demand or reserved prices. The answer often flips for bigger models, which spill across chips on Inf2 but fit on one GPU.
Two corrections keep the comparison honest. First, include idle capacity. If autoscaling keeps you at 40 percent average utilisation on one platform and 70 percent on another, divide by the average utilisation, not the peak. Second, include engineering cost. A platform that needs a separate compile pipeline and a pinned SDK has a running cost that the hourly price does not show. For queueing behaviour under load, see serving capacity planning.
The LLM decision after Neuron 2.29
Large language models are where the Inferentia story changed in 2026. The Neuron 2.29.0 release notes (9 April 2026) say NxD Inference models are supported only on Trn2 and newer. They tell customers who need NxD Inference kernel support on Inf2 or Trn1 to pin to release 2.28. The vLLM integration on Neuron, which adds an OpenAI-compatible server with continuous batching, is aimed at Trainium too. So an Inf2 LLM deployment has three options:
| Option | What you get | What it costs you |
|---|---|---|
| Stay on Inf2, pinned at 2.28 | No migration work now; existing compiled artifacts keep working | Frozen drivers, AMIs, containers and model support; no new architectures; patching is your problem |
| Move to Trn2 or Trn3 | The supported NxD Inference and vLLM paths, larger HBM per chip, newer kernels | A recompile and re-validation; new capacity to reserve; a different price point |
| Move to GPUs | The broadest serving ecosystem and model support on day one | Usually a higher hourly price; a second hardware platform to operate |
Pinning is a reasonable choice for a model you will retire within a few quarters. It is a poor one for a model family you expect to keep upgrading, because every new architecture lands on the supported path first. Whichever option you pick, write it down with a review date. Continuous batching changes the capacity numbers so much that it must be in the comparison; continuous batching in depth explains why. The training-side view of the same chips is in AWS Trainium in depth.
Packing small models onto one instance
The workloads that stay on Inf2 are mostly small models, and small models waste a big instance unless you pack them. An inf2.48xlarge has 24 NeuronCores. A traced encoder normally occupies one core per worker process. That gives you up to 24 independent workers on one host, which can serve one model or several:
# one worker per NeuronCore, each pinned explicitly (inf2.48xlarge: cores 0-23)
for core in $(seq 0 15); do
NEURON_RT_VISIBLE_CORES=$core python serve.py --model embedder --port $((9000+core)) &
done
for core in $(seq 16 23); do
NEURON_RT_VISIBLE_CORES=$core python serve.py --model reranker --port $((9000+core)) &
doneSize the split from the cost sweep, not from guesswork. If the embedder takes two thirds of the traffic and has similar per-request cost, two thirds of the cores is right. Host CPU is the second constraint, because tokenisation and pre-processing run on the vCPUs. That is the reason inf2.8xlarge exists alongside inf2.xlarge, with the same chip and eight times the vCPUs. Watch per-core utilisation with neuron-top and host CPU together.
Failure modes
- Artifact and runtime mismatch. A model compiled with one SDK release and loaded by another fails to load or behaves differently. Version the artifact with the compiler version and refuse to load on mismatch.
- Drift after migration. Recompiling for Inf2 can change numerics, for example when you cast to BF16 or FP8. The agreement gate catches it; skipping the gate does not make the drift go away.
- Pinned stack decay. A 2.28 LLM deployment slowly falls behind on kernels, model support and OS patches. Track its age as a risk item, not as a stable fact.
- Regional capacity. Inf2 is not in every Region, and capacity for the largest size can be tight. Check availability before you design around it, and keep a fallback.
- Planning for a chip that does not exist. Roadmaps that assume Inferentia3 pricing or features are built on nothing that AWS has published. Plan with Inf2 and Trainium as they are documented today.
Trade-offs
Inf2 is cheap, predictable capacity for static-graph models, and it is still supported for that use. In return you accept a compile step, fixed shapes and a smaller ecosystem than GPUs. For LLMs in 2026, the trade has moved: the supported Neuron path runs on Trainium. A team with one model family on Inferentia should usually ask whether the saving pays for a second platform. A team running a large embedding and reranking fleet usually finds that it does. Software details for both chip families are in the Neuron SDK deep dive.
What to do next
- List every model on Inferentia, with its instance type, SDK version and software path.
- Put each one in a bucket: stay on Inf2, pinned LLM, or Inf1 to migrate.
- For Inf1 models, find the source weights, recompile with
torch_neuronxand pass an agreement gate before moving traffic. - Run a latency-bounded load sweep and compute cost per million requests on Inf2 and on one alternative.
- For pinned LLMs, record the decision and a review date, and price the move to Trn2 or Trn3.
- Pack small models one worker per NeuronCore and split cores in proportion to traffic.
- Remove any plan that depends on an Inferentia3 until AWS announces one.