If you searched for Inferentia3, here is the short answer. As of 7 October 2026, AWS has not announced a third Inferentia chip. The Inferentia product page lists two generations: the original Inferentia behind EC2 Inf1, and Inferentia2 behind Inf2. Meanwhile the newest Neuron inference software targets Trainium. Since Neuron 2.29, NxD Inference, the library behind Neuron's large-model serving, supports models only on Trn2 and newer. The Trainium3 architecture documentation names autoregressive inference serving as a target for its switch fabric.

So the practical question is not when Inferentia3 will arrive. It is what to do with the Inferentia capacity you already run, or are about to buy. This page answers that question. It covers what changed between generations and which workloads still belong on Inf2. It walks through migrating an Inf1 fleet, gives a method for comparing platforms in cost per million requests, and lays out the decision for large language models. Chip internals, compilation and shape bucketing have their own page, AWS Inferentia in depth, and this page builds on it.

Two generations, two toolchains

Each Inferentia generation pairs a chip with its own toolchain, and the toolchain is what your code depends on. A model compiled for one generation does not run on the other. Migrating means recompiling, so treat a generation change as a software migration, not a hardware swap.

Inferentia (Inf1)Inferentia2 (Inf2)Inferentia3
CoreNeuronCore-v1, four per chipNeuronCore-v2, two per chip (the core Trn1 also uses)not announced
Memory per chipsmall on-chip memory plus DRAM32 GiB HBM-
Chip-to-chiplimited; pipelining across coresNeuronLink-v2, 192 GiB/s per chip on 24xl and 48xl-
PyTorch packagetorch-neuron (torch.neuron)torch-neuronx (torch_neuronx)-
Compilerneuron-ccneuronx-cc-
SDK status in 2026legacy toolchain; check your release notescurrent for traced models; LLM serving pinned at 2.28-

Two rows decide most migrations. The compiler row means every Inf1 artifact must be rebuilt from the source model. The chip-to-chip row explains why Inf2 can serve models larger than one chip with tensor parallelism, while Inf1 could not do this in any useful way. Inf2 comes in four sizes: 1, 1, 6 and 12 chips, with 4, 32, 96 and 192 vCPUs. Only the 6- and 12-chip sizes have NeuronLink between chips.

Triage: what still belongs on Inf2

The Inferentia line in 2026 and where each workload goesInferentia (Inf1)torch-neuron + neuron-ccrecompileInferentia2 (Inf2)torch-neuronx + neuronx-ccInferentia3not announcedTriage of an existing Inferentia fleetTraced graphsencoders, embeddings, vision, speechLLMs via NxD Inferencedecoder serving with KV cacheAnything still on Inf1old SDK, old AMIsStay on Inf2supported torch-neuronx pathPin 2.28 on Inf2or move to Trn2 / Trn3 / GPUMigrate to Inf2recompile, re-validate outputsMeasure every option in cost per million requests at your latency target
Inferentia generations, the absent third chip, and the three buckets an existing fleet falls into.

Sort every model you serve on Inferentia into one of three buckets before you change anything. The bucket depends on which software path the model uses, not on its size.

WorkloadSoftware path on Inf2Verdict
Text encoders, embedding models, rerankers, classifierstorch_neuronx.trace with shape bucketsGood fit. Supported, cheap per request, latency is predictable
Vision backbones, detection, OCR, speech encoderstorch_neuronx.traceGood fit if inputs can be resized or padded to a few fixed shapes
Diffusion pipelinesTrace each sub-network separatelyWorkable. Test image quality after casting to lower precision
Decoder LLMs served with NxD InferenceNeeds Neuron 2.28 or earlier on Inf2Frozen dependency. Plan the exit
Anything on Inf1Legacy torch-neuronMigrate now

The first two rows are why Inf2 is still worth running in 2026. Encoders and vision models compile into a static graph once. They run at fixed shapes, so the chip's low-precision compute stays busy. They also fit in one chip's 32 GiB, so you can pack many of them onto one instance. A retrieval stack that embeds and reranks millions of passages a day is the classic Inf2 customer.

Leaving Inf1: a migration runbook

An Inf1 fleet runs on a legacy toolchain that current Neuron releases have moved away from; check the release notes for your version to see exactly what is still shipped. Every security patch to an old host image is your problem. The migration is mechanical, but it has four steps, and teams that skip the last one learn about numeric drift from their users.

  1. Find the source model. You need the original PyTorch weights and code, not the Inf1 artifact. A compiled Inf1 model cannot be converted.
  2. Recompile with the Inf2 toolchain, using the same example shapes you served on Inf1 so that you can compare like with like.
  3. Compare outputs against the CPU or GPU reference on a held-out set of real requests.
  4. Re-measure throughput and latency per NeuronCore, because the core count, batch sizes and best number of workers all change.
# Inf1 (legacy): torch-neuron + neuron-cc
import torch, torch_neuron
inf1 = torch.neuron.trace(model, example_inputs=[ids, mask])

# Inf2: torch-neuronx + neuronx-cc. Same model object, new compiler.
import torch, torch_neuronx
model.eval()
inf2 = torch_neuronx.trace(model, (ids, mask))
torch.jit.save(inf2, "encoder_b8_s128.pt")      # build in CI, ship the artifact

# Agreement gate: run before any traffic moves
import torch.nn.functional as F
def agreement(ref_model, neuron_model, batches, min_cos=0.999):
    worst = 1.0
    with torch.no_grad():
        for ids, mask in batches:
            a = ref_model(ids, mask).float()
            b = neuron_model(ids, mask).float()
            worst = min(worst, F.cosine_similarity(a, b, dim=-1).min().item())
    assert worst >= min_cos, f"worst cosine {worst:.5f} below {min_cos}"
    return worst

The threshold depends on the task. For embeddings, cosine similarity against the reference is the right check, together with a retrieval recall test on a labelled set. For classifiers, compare the top-1 label and alert on any change in agreement above a small fraction. Run the gate in CI on every SDK upgrade, not just once.

Cost per million requests, measured

Platform choices fail when teams compare peak benchmark numbers. The number that matters is cost per million requests at your latency target, measured with your own model and traffic shape. The method is the same for Inf2, Trainium and GPUs:

  1. Fix the latency objective, for example p99 under 50 ms for a reranker call.
  2. On one instance, raise offered load step by step and record throughput and p99 at each step.
  3. Take the highest throughput whose p99 still meets the objective. That is the usable capacity.
  4. Divide the hourly price by the requests served per hour.
def cost_per_million(price_per_hour, usable_rps):
    """price in your currency per instance-hour; usable_rps at the latency target."""
    return price_per_hour / (usable_rps * 3600) * 1_000_000

def usable_rps(sweep, p99_target_ms):
    """sweep: list of (offered_rps, achieved_rps, p99_ms) from a load test."""
    ok = [achieved for offered, achieved, p99 in sweep if p99 <= p99_target_ms]
    return max(ok) if ok else 0.0

Worked example (illustrative numbers, not a benchmark). Suppose a reranker sweep on an inf2.xlarge meets a 50 ms p99 up to 900 requests per second, and a GPU instance meets it up to 2,400. Suppose too that the GPU instance costs 3.5 times as much per hour. With the Inf2 price written as P, Inf2 costs P / (900 × 3600) per request and the GPU costs 3.5P / (2,400 × 3600). The ratio is 3.5 × 900 / 2,400 = 1.31, so the GPU costs 31 percent more per request, even though it is 2.7 times faster. Now replace the assumptions with your own sweep and current on-demand or reserved prices. The answer often flips for bigger models, which spill across chips on Inf2 but fit on one GPU.

Two corrections keep the comparison honest. First, include idle capacity. If autoscaling keeps you at 40 percent average utilisation on one platform and 70 percent on another, divide by the average utilisation, not the peak. Second, include engineering cost. A platform that needs a separate compile pipeline and a pinned SDK has a running cost that the hourly price does not show. For queueing behaviour under load, see serving capacity planning.

The LLM decision after Neuron 2.29

Large language models are where the Inferentia story changed in 2026. The Neuron 2.29.0 release notes (9 April 2026) say NxD Inference models are supported only on Trn2 and newer. They tell customers who need NxD Inference kernel support on Inf2 or Trn1 to pin to release 2.28. The vLLM integration on Neuron, which adds an OpenAI-compatible server with continuous batching, is aimed at Trainium too. So an Inf2 LLM deployment has three options:

OptionWhat you getWhat it costs you
Stay on Inf2, pinned at 2.28No migration work now; existing compiled artifacts keep workingFrozen drivers, AMIs, containers and model support; no new architectures; patching is your problem
Move to Trn2 or Trn3The supported NxD Inference and vLLM paths, larger HBM per chip, newer kernelsA recompile and re-validation; new capacity to reserve; a different price point
Move to GPUsThe broadest serving ecosystem and model support on day oneUsually a higher hourly price; a second hardware platform to operate

Pinning is a reasonable choice for a model you will retire within a few quarters. It is a poor one for a model family you expect to keep upgrading, because every new architecture lands on the supported path first. Whichever option you pick, write it down with a review date. Continuous batching changes the capacity numbers so much that it must be in the comparison; continuous batching in depth explains why. The training-side view of the same chips is in AWS Trainium in depth.

Packing small models onto one instance

The workloads that stay on Inf2 are mostly small models, and small models waste a big instance unless you pack them. An inf2.48xlarge has 24 NeuronCores. A traced encoder normally occupies one core per worker process. That gives you up to 24 independent workers on one host, which can serve one model or several:

# one worker per NeuronCore, each pinned explicitly (inf2.48xlarge: cores 0-23)
for core in $(seq 0 15); do
  NEURON_RT_VISIBLE_CORES=$core python serve.py --model embedder --port $((9000+core)) &
done
for core in $(seq 16 23); do
  NEURON_RT_VISIBLE_CORES=$core python serve.py --model reranker --port $((9000+core)) &
done

Size the split from the cost sweep, not from guesswork. If the embedder takes two thirds of the traffic and has similar per-request cost, two thirds of the cores is right. Host CPU is the second constraint, because tokenisation and pre-processing run on the vCPUs. That is the reason inf2.8xlarge exists alongside inf2.xlarge, with the same chip and eight times the vCPUs. Watch per-core utilisation with neuron-top and host CPU together.

Failure modes

  • Artifact and runtime mismatch. A model compiled with one SDK release and loaded by another fails to load or behaves differently. Version the artifact with the compiler version and refuse to load on mismatch.
  • Drift after migration. Recompiling for Inf2 can change numerics, for example when you cast to BF16 or FP8. The agreement gate catches it; skipping the gate does not make the drift go away.
  • Pinned stack decay. A 2.28 LLM deployment slowly falls behind on kernels, model support and OS patches. Track its age as a risk item, not as a stable fact.
  • Regional capacity. Inf2 is not in every Region, and capacity for the largest size can be tight. Check availability before you design around it, and keep a fallback.
  • Planning for a chip that does not exist. Roadmaps that assume Inferentia3 pricing or features are built on nothing that AWS has published. Plan with Inf2 and Trainium as they are documented today.

Trade-offs

Inf2 is cheap, predictable capacity for static-graph models, and it is still supported for that use. In return you accept a compile step, fixed shapes and a smaller ecosystem than GPUs. For LLMs in 2026, the trade has moved: the supported Neuron path runs on Trainium. A team with one model family on Inferentia should usually ask whether the saving pays for a second platform. A team running a large embedding and reranking fleet usually finds that it does. Software details for both chip families are in the Neuron SDK deep dive.

What to do next

  1. List every model on Inferentia, with its instance type, SDK version and software path.
  2. Put each one in a bucket: stay on Inf2, pinned LLM, or Inf1 to migrate.
  3. For Inf1 models, find the source weights, recompile with torch_neuronx and pass an agreement gate before moving traffic.
  4. Run a latency-bounded load sweep and compute cost per million requests on Inf2 and on one alternative.
  5. For pinned LLMs, record the decision and a review date, and price the move to Trn2 or Trn3.
  6. Pack small models one worker per NeuronCore and split cores in proportion to traffic.
  7. Remove any plan that depends on an Inferentia3 until AWS announces one.
Key takeaway: There is no Inferentia3. Inf2 remains a supported, cost-effective home for traced encoders, embedding, vision and speech models, while Neuron's LLM serving has moved to Trainium. Migrate anything left on Inf1, gate every recompile on output agreement, compare platforms by cost per million requests at your latency target, and treat a pinned LLM stack as a dated decision.