AWS Inferentia is Amazon's own accelerator for running trained models, sold as EC2 instance families: Inf1 on the first-generation chip and Inf2 on Inferentia2. The pitch is cost per inference, but you cannot capture that saving by treating the chip as a drop-in GPU. Inferentia runs graphs that were compiled ahead of time for fixed tensor shapes, through a software stack (the AWS Neuron SDK) with its own compiler, runtime, support matrix and release cadence. Teams that ignore those facts discover them in production as multi-minute compiles, out-of-memory loads and an SDK upgrade that dropped their framework.
This article explains Inferentia2 from first principles: the chip, compilation and static shapes, LLM memory sizing, core pinning, the 2025-2026 SDK support changes, and when Inferentia belongs in your stack. Hardware figures come from the Neuron architecture documentation and the EC2 Inf2 page; SDK facts are as of Neuron 2.32 (August 2026) and should be rechecked against the release notes for whatever version you deploy.
What is on the chip, and why software cares
An Inferentia2 chip holds two NeuronCore-v2 cores, 32 GiB of high-bandwidth memory (HBM) with 820 GiB/s of bandwidth, DMA engines rated at 1 TB/s with inline compression, and NeuronLink-v2 links to neighbouring chips. AWS rates it at 190 TFLOPS for FP16, BF16, cFP8 and TF32, 380 TOPS for INT8 and 47.5 TFLOPS for FP32. Those numbers tell you two things straight away. Low-precision maths is four times faster than FP32, so precision is a performance decision. And memory bandwidth, not its FLOPS, will limit you whenever each step has to read a lot of weights for little compute, which is exactly the shape of LLM decoding.
Each NeuronCore-v2 has four engines. The TensorEngine is a systolic array for matrix multiplication, convolution and transposes, rated at over 90 TFLOPS of BF16/FP16 per core. The VectorEngine handles operations where each output depends on several inputs, such as layer-norm reductions and softmax. The ScalarEngine handles element-wise work such as activations. The GPSIMD engine is eight general-purpose 512-bit vector processors that run C code against on-chip SRAM; this is where custom operators and unusual control flow land. The compiler decides which engine runs each operation and schedules data movement between HBM and the software-managed on-chip SRAM. You do not write kernels for the common path, but you do pay when a model leans on operations the compiler maps poorly.
| Instance | Chips | NeuronCores | Accelerator memory | vCPU / host memory | NeuronLink |
|---|---|---|---|---|---|
| inf2.xlarge | 1 | 2 | 32 GB | 4 / 16 GiB | no |
| inf2.8xlarge | 1 | 2 | 32 GB | 32 / 128 GiB | no |
| inf2.24xlarge | 6 | 12 | 192 GB | 96 / 384 GiB | yes |
| inf2.48xlarge | 12 | 24 | 384 GB | 192 / 768 GiB | yes |
The two single-chip sizes differ only in host, so pick inf2.8xlarge when pre-processing needs the CPU. Only the multi-chip sizes have NeuronLink, so a model larger than one chip's 32 GiB forces you onto inf2.24xlarge or inf2.48xlarge.
The software path: compile once, run many times
On a GPU, PyTorch dispatches kernels eagerly and a new input shape just launches the same kernels with different sizes. Inferentia works differently. You hand the Neuron compiler (neuronx-cc) a captured graph together with example inputs, and it produces a NEFF: a compiled program for the NeuronCores that fixes every tensor shape, the placement of every operation on an engine, and the schedule for moving data. The Neuron Runtime loads the NEFF onto cores and executes it. For classic PyTorch inference the entry point is torch_neuronx.trace(func, example_inputs, *, compiler_workdir=None, compiler_args=None, ...), which returns a TorchScript module with the compiled graph embedded; you save it with torch.jit.save and load it on the serving host with torch.jit.load.
Three consequences follow. First, compilation takes minutes, so it belongs in your build pipeline, not container start-up. Store compiled artifacts next to the model weights, keyed by SDK version, model revision, shapes and compiler flags. Second, the artifact is tied to the SDK that produced it, so an SDK upgrade means recompiling and re-validating. Third, the compiler only knows the shapes it saw. Feed a traced model a sequence length it was not compiled for and you get an error, not a slower answer.
Worked example: an encoder classifier with shape buckets
Take a sentiment classifier (DistilBERT fine-tuned on SST-2) that must serve requests from a few words up to 256 tokens. Padding everything to 256 would work, but a ten-token request would then pay for 256 tokens of attention. The standard answer is buckets: compile one artifact per sequence length you will serve, and route each request to the smallest bucket that fits. Batch size is fixed at trace time too; torch_neuronx.dynamic_batch lets the compiled model accept other batch sizes by splitting and padding them into chunks of the traced size.
import torch, torch_neuronx
from transformers import AutoTokenizer, AutoModelForSequenceClassification
name = "distilbert-base-uncased-finetuned-sst-2-english"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForSequenceClassification.from_pretrained(name, torchscript=True).eval()
SEQ_BUCKETS = (64, 128, 256) # every length the service will ever run
BATCH = 4
def example(seq_len):
enc = tok(["warmup"] * BATCH, padding="max_length", max_length=seq_len,
truncation=True, return_tensors="pt")
return (enc["input_ids"], enc["attention_mask"])
for seq in SEQ_BUCKETS:
neuron_model = torch_neuronx.trace(
model, example(seq),
compiler_workdir=f"/tmp/neuron-work-{seq}",
# Set precision explicitly: the default has differed between releases.
compiler_args=["--auto-cast", "matmult", "--auto-cast-type", "bf16"],
)
torch.jit.save(neuron_model, f"sst2_b{BATCH}_s{seq}.pt")import bisect, torch, torch_neuronx
BUCKETS = [64, 128, 256]
models = {s: torch_neuronx.dynamic_batch(torch.jit.load(f"sst2_b4_s{s}.pt")) for s in BUCKETS}
def classify(texts):
enc = tok(texts, return_tensors="pt", padding=True)
n = enc["input_ids"].shape[1]
i = bisect.bisect_left(BUCKETS, n)
if i == len(BUCKETS):
raise ValueError(f"input of {n} tokens exceeds largest bucket {BUCKETS[-1]}")
seq = BUCKETS[i]
ids = torch.nn.functional.pad(enc["input_ids"], (0, seq - n))
mask = torch.nn.functional.pad(enc["attention_mask"], (0, seq - n))
# dynamic_batch splits/pads any batch size into chunks of the traced batch (4)
logits = models[seq](ids, mask)[0]
return logits.argmax(-1).tolist()Pick buckets from your real request-length histogram; each costs a compile and HBM for a loaded copy, so three to five is usually enough. The --auto-cast matmult --auto-cast-type bf16 flags tell the compiler to run FP32 matrix multiplications in BF16 and leave other operations alone. The compiler's default for --auto-cast has not been the same in every release, so state it explicitly and compare outputs against the CPU or GPU model on a held-out set before you ship. Compare top-1 agreement for classifiers and cosine similarity for embeddings, against a written tolerance.
Cores, processes and pinning
The runtime gives each process exclusive ownership of the NeuronCores it loads onto. By default a process claims all visible cores, which is why a second worker on the same host often fails at load time. Two environment variables control this: NEURON_RT_VISIBLE_CORES names exact cores or a range for the process, and NEURON_RT_NUM_CORES asks the runtime to reserve that many free cores. If both are set, the visible-cores setting wins. On Kubernetes, the Neuron device plugin advertises aws.amazon.com/neuroncore (individual cores) and aws.amazon.com/neuron (whole devices) as extended resources; EKS also documents a Neuron DRA driver for newer clusters.
# Two worker processes on one inf2.xlarge (1 chip = 2 NeuronCores), one core each.
# A NeuronCore is owned by one process; a third worker would fail to load.
NEURON_RT_VISIBLE_CORES=0 python server.py --port 8001 &
NEURON_RT_VISIBLE_CORES=1 python server.py --port 8002 &
# Kubernetes: request cores, not whole devices, so the scheduler can pack pods.
# resources:
# limits:
# aws.amazon.com/neuroncore: 1For small models, run one model copy and one worker per core: data parallelism with no cross-core traffic. torch_neuronx.DataParallel does the same thing inside one process if you prefer that. Models larger than a core need tensor parallelism from the Neuron distributed libraries, not hand-written tracing.
Sizing a large language model against HBM
Generative models make memory the first constraint. Weights in BF16 take two bytes per parameter. The key-value (KV) cache takes, per token, two (K and V) times layers times KV heads times head dimension times bytes per value. Add headroom for activations, compiler scratch and runtime buffers, and you have the budget. The function below does the arithmetic; the 15 percent headroom is a planning assumption, not a Neuron figure.
def fits(params_b, layers, kv_heads, head_dim, batch, ctx, chips, bytes_per=2,
hbm_gib_per_chip=32, headroom=0.85):
weights = params_b * 1e9 * bytes_per
kv_per_token = 2 * layers * kv_heads * head_dim * bytes_per # K and V
kv = kv_per_token * batch * ctx
budget = chips * hbm_gib_per_chip * 2**30 * headroom # leave room for scratch
return (weights + kv) / 2**30, budget / 2**30
# 8B model, 32 layers, 8 KV heads, head_dim 128, batch 4, 4096 context, BF16
print(fits(8, 32, 8, 128, 4, 4096, chips=1)) # (~16.9 GiB needed, ~27.2 GiB budget)
# 70B model, 80 layers, 8 KV heads, head_dim 128, batch 4, 4096 context, BF16
print(fits(70, 80, 8, 128, 4, 4096, chips=6)) # (~135 GiB needed, ~163 GiB budget)
print(fits(70, 80, 8, 128, 4, 4096, chips=12)) # (~135 GiB needed, ~326 GiB budget)For an 8-billion-parameter model with 32 layers, 8 KV heads and head dimension 128, the KV cache is 128 KiB per token, so a batch of 4 at 4,096 tokens adds 2 GiB to roughly 15 GiB of weights. That fits one chip, split across its two cores with tensor parallelism. A 70-billion-parameter model needs about 130 GiB for weights alone: on paper it fits inf2.24xlarge's six chips (about 163 GiB of budget), while inf2.48xlarge leaves room for larger batches or contexts. Either way it is sharded across NeuronLink.
Then check speed. Each decode step reads every weight once. At 820 GiB/s, reading 15 GiB takes about 18 ms, so a single sequence on one chip cannot exceed roughly 55 tokens per second however fast the TensorEngine is. Batching is how you buy throughput, because one weight read serves every sequence in the batch. Batching also grows the KV cache, so the batch-size ceiling comes from the memory sum above.
The SDK support matrix is an operational risk
The Neuron SDK has moved fast, and its support changes land directly on Inferentia fleets. As of the 2.32 documentation: Inf1 virtual environments and AMIs stopped being supported from Neuron 2.27, and Inf1 users are told to stay on 2.26 or earlier. TensorFlow support for Inf2 ended in Neuron 2.29, and Inf2 users are pointed at PyTorch. The same 2.29 release notes say NxD Inference, the library behind Neuron's LLM serving and its vLLM integration, now supports models only on Trn2 and newer, and that customers who need its kernels on Inf2 should pin to release 2.28. The torch-neuronx tracing path still lists Inf2 as supported. Neuron has also announced a move from PyTorch/XLA to a native PyTorch backend, with PyTorch 2.9 as the last XLA-based version.
In practice: an Inf2 fleet serving encoders, vision models or other traced graphs has a supported path today. An Inf2 fleet serving LLMs through NxD Inference is on a pinned SDK, which means pinned drivers, AMIs, containers and model support. Treat that like any other frozen dependency: record the decision, re-evaluate it every quarter, and price the move to Trainium-based instances, where the LLM stack is now aimed. Read the release notes for your target version before every upgrade, because the answers above will change.
Failure modes and how they show up
- Compile at start-up. Containers that trace on boot take minutes to become ready, autoscaling lags, and a crash loop recompiles forever. Compile in CI and ship the artifact.
- Shape miss. A request longer than every bucket, or a batch dimension baked in without
dynamic_batch, errors out. Enforce limits at the API edge and return a clear 413 or 400, not a 500. - Core contention. A second process cannot load because the first claimed every core. Set
NEURON_RT_VISIBLE_CORESper worker and requestaws.amazon.com/neuroncorein pods. - Silent precision drift. Casting to BF16 or INT8 changes outputs slightly, and occasionally a lot for a model with outlier activations. Gate releases on an agreement test against the reference model.
- CPU fallback. Operations the compiler cannot place may be partitioned to the CPU, so the graph ping-pongs across PCIe. Watch host CPU and per-request latency together, and read the compiler log for unsupported operators.
- Tooling blind spots. Use
neuron-lsto list devices and the processes using them,neuron-topfor live utilisation, andneuron-monitorto export metrics to your monitoring stack.
When Inferentia is the right choice
Inferentia2 fits steady, high-volume inference on models whose shapes you control: embedding services, rerankers, classifiers, speech and vision encoders, and mid-sized generative models where you have already confirmed SDK support. It fits poorly when shapes are wildly dynamic, you depend on custom CUDA kernels, models change weekly, or you need the newest LLM serving features, which in 2026 land on Trainium first.
Decide with a bake-off, not a datasheet: same model, same request-length distribution, cost per million requests at your promised p99 latency on Inf2 and on your current GPU instance, plus the engineering cost of new tooling and a pinned SDK. If you want a managed model rather than your own, compare against Bedrock as well, since the cheapest accelerator is sometimes the one you do not operate.
What to do next
- Write down your request-length and batch-size histograms and your p99 latency target; they decide buckets and instance size.
- Check the Neuron release notes for your model family on Inf2, including whether you would need NxD Inference and therefore a pin to 2.28.
- Trace your model on an inf2.xlarge (or larger, per the sizing arithmetic) with explicit
--auto-castflags and run an agreement test against the reference model. - Move compilation into CI, version artifacts by SDK release and model revision, and bake them into the serving image. Read EC2 fundamentals for AMI and launch-template hygiene.
- Deploy one worker per NeuronCore with
NEURON_RT_VISIBLE_CORESor on EKS withaws.amazon.com/neuroncorerequests, and scale with Auto Scaling groups using warm pools so new nodes do not start cold. - Export
neuron-monitormetrics, alert on core utilisation, host CPU and errors from requests outside every bucket, and re-run the cost bake-off whenever the SDK or your model changes.