Arm Neoverse is not a chip. It is a family of CPU core designs that Arm licenses to companies that build server processors: AWS Graviton, NVIDIA Grace, Google Axion and Microsoft Cobalt all start from Neoverse cores. When people ask whether Neoverse is good for AI they usually mean two different things. The first is the CPU that sits next to an accelerator, feeding it data, tokenising text, running the Python control loop and sometimes holding optimizer state. The second is the CPU doing inference itself, for embedding models, rerankers, classic ML and small quantized language models where a GPU would sit mostly idle.
This article explains both roles from the instruction set upward. You will learn which core generation ships in which cloud CPU, which vector and matrix instructions ML kernels actually use, how a library detects them at run time, why a 96-core CPU running a quantized 8B model is limited by memory bandwidth rather than arithmetic, and the operational mistakes that quietly cost a factor of two. Neoverse is not a GPU replacement for training; the goal is to know exactly where it is the cheaper correct answer.
The Neoverse families and where they ship
Arm splits Neoverse into three lines. The V series maximises per-core performance and vector throughput. The N series targets performance per watt and per square millimetre, so a vendor can fit more cores. The E series is aimed at throughput-oriented data-plane and edge systems. Each generation also moves the architecture version forward, which is what decides the instructions available to software.
| Core | Architecture | Vector units | ML-relevant features | Shipping examples |
|---|---|---|---|---|
| N1 | Armv8.2-A | 2 x 128-bit Neon | Neon, dotprod (SDOT/UDOT), fp16 | Graviton2, Ampere Altra |
| V1 | Armv8.4-A plus later features | 2 x 256-bit SVE | SVE, BF16, I8MM | Graviton3 |
| N2 | Armv9.0-A | 2 x 128-bit SVE2 | SVE2, BF16, I8MM | Microsoft Cobalt 100 |
| V2 | Armv9.0-A | 4 x 128-bit SVE2 | SVE2, BF16, I8MM | Graviton4, NVIDIA Grace, Google Axion |
| N3, V3 | Armv9.2-A | SVE2 | SVE2, BF16, I8MM; check newer bits at run time | check vendor documentation |
Two corrections are worth making explicitly because they circulate widely. First, none of the shipping N1, V1, N2 or V2 cores has a matrix engine in the sense of a dedicated tile unit. They have matrix-multiply instructions such as BFMMLA and SMMLA that operate inside ordinary vector registers. The Scalable Matrix Extension (SME), which does add tile storage, is a separate feature: treat its presence on any given part as something to check, never to assume. Second, Ampere Altra uses Neoverse N1, while AmpereOne uses Ampere's own Arm cores, so it is not a Neoverse part at all.
Note the vector widths. V1 has two 256-bit pipes; V2 has four 128-bit pipes. The total width per cycle is the same, but V2 runs four independent smaller operations. Code that assumed a 256-bit vector length on Graviton3 gets a 128-bit vector length on Graviton4, which is exactly why SVE code must be written to be vector-length agnostic.
The instructions ML kernels use
Four instruction groups carry nearly all ML work on Neoverse. Understanding what each computes lets you read a kernel library's dispatch table and know what you are getting.
- Dot product (SDOT, UDOT). Each 32-bit lane accumulates the dot product of four 8-bit values. This is the oldest int8 path and the only one on N1.
- SVE and SVE2. Scalable vectors with predication. A loop compiled once runs at whatever vector length the hardware implements, and predicates handle the loop tail without scalar cleanup code. SVE2 adds integer and DSP-style operations that quantized kernels use for unpacking and widening.
- BF16 (BFDOT, BFMMLA). BFMMLA treats each 128-bit segment as a small block: a 2 x 4 bfloat16 matrix times a 4 x 2 bfloat16 matrix, accumulated into a 2 x 2 fp32 block. That is 16 multiply-accumulates per segment per instruction, four times the multiply-accumulates of a 128-bit fp32 FMA.
- I8MM (SMMLA, UMMLA, USMMLA). The int8 equivalent: a 2 x 8 by 8 x 2 product into a 2 x 2 int32 block, 32 multiply-accumulates per 128-bit segment. Quantized LLM kernels lean on this for prefill and batched work.
The matrix-multiply forms change how weights must be laid out. A GEMM kernel that uses SMMLA wants two rows of A and two columns of B interleaved in eight-byte chunks, so libraries repack weights once at load time. When you see a model take longer to load on Arm than on x86 with the same file, repacking is often why; when you see two libraries with identical instruction support differ by 2x in throughput, the layout and blocking are usually why.
For a sense of scale, fp32 peak on a V2 core is four 128-bit pipes times four lanes times two operations for a fused multiply-add, or 32 FLOP per cycle. A 96-core Graviton4 at roughly 2.8 GHz therefore peaks near 8.6 fp32 TFLOP/s, with the bf16 and int8 matrix instructions raising the ceiling further. That is a respectable number for preprocessing and small models and a tiny one next to a data-centre GPU, which is the correct mental model.
Detecting features at run time
Because the same aarch64 binary may run on N1, V1 and V2 machines in one fleet, robust code detects features at run time and dispatches, rather than compiling for the newest core and hoping. On Linux the kernel publishes features in the auxiliary vector (AT_HWCAP and AT_HWCAP2) and mirrors them as names in /proc/cpuinfo. The C version:
#include <stdio.h>
#include <sys/auxv.h>
#include <asm/hwcap.h>
#include <sys/prctl.h> /* PR_SVE_GET_VL: no SVE code generation needed */
int main(void) {
unsigned long hw = getauxval(AT_HWCAP);
unsigned long hw2 = getauxval(AT_HWCAP2);
int dotprod = !!(hw & HWCAP_ASIMDDP);
int sve = !!(hw & HWCAP_SVE);
int sve2 = !!(hw2 & HWCAP2_SVE2);
int bf16 = !!(hw2 & HWCAP2_BF16);
int i8mm = !!(hw2 & HWCAP2_I8MM);
printf("dotprod=%d sve=%d sve2=%d bf16=%d i8mm=%d\n", dotprod, sve, sve2, bf16, i8mm);
if (sve) { /* only ask for the vector length when SVE is present */
int vl = prctl(PR_SVE_GET_VL) & PR_SVE_VL_LEN_MASK; /* bytes */
printf("SVE vector length: %d bits\n", vl * 8);
}
return 0;
}Fleet tooling usually wants the same answer from Python, for example to label nodes or choose a container image. Reading the flags line is enough:
def arm_features(path="/proc/cpuinfo"):
want = {"asimddp", "sve", "sve2", "bf16", "i8mm", "svebf16", "svei8mm", "sme"}
with open(path) as f:
for line in f:
if line.startswith("Features"):
return sorted(want & set(line.split(":", 1)[1].split()))
return []
feats = arm_features()
tier = ("v2-class" if {"sve2", "i8mm", "bf16"} <= set(feats)
else "v1-class" if {"sve", "i8mm"} <= set(feats)
else "n1-class" if "asimddp" in feats else "baseline")
print(feats, tier)Compiler flags are where most production incidents start. -mcpu=neoverse-v2 tells GCC or Clang to both tune for and emit instructions available on V2, so the resulting binary can die with SIGILL on a Graviton2 node. Build the portable baseline for the oldest core you run, and compile hot kernels separately per feature tier behind a dispatch check like the one above. That is what the major libraries do internally.
Framework paths and an honest benchmark
You rarely write SMMLA by hand. The practical question is which framework path picks up which instructions.
- PyTorch on aarch64 routes many fp32 operators through oneDNN with the Arm Compute Library backend. AWS's Graviton performance guide documents
DNNL_DEFAULT_FPMATH_MODE=BF16, which lets fp32 models use bf16 matrix instructions internally. It changes numerics, so validate accuracy before enabling it in production. - KleidiAI is Arm's library of matmul micro-kernels for int4, int8 and bf16, integrated into XNNPACK, ExecuTorch, ONNX Runtime and llama.cpp. Arm reports large uplifts for quantized LLM token generation on Graviton4 and Axion; treat those as vendor figures and measure on your own model.
- llama.cpp builds a ggml-cpu backend with aarch64 paths for Neon, dotprod, SVE and I8MM, and repacks supported quantized weights into an interleaved layout when the instructions are present.
A minimal, honest benchmark for a CPU inference service pins threads, fixes the input length and separates prefill from decode:
# Graviton has no SMT: one vCPU is one physical core; set threads to the instance vCPU count, e.g. 64.
export OMP_NUM_THREADS=64
export DNNL_DEFAULT_FPMATH_MODE=BF16 # only after an accuracy check
numactl --cpunodebind=0 --membind=0 \
python bench.py --model bge-large --batch 32 --seq 256 --warmup 20 --iters 200
# Quantized LLM: report prompt processing (pp) and token generation (tg) separately.
./llama-bench -m model-q4_0.gguf -t 64 -p 512 -n 128
How software reaches the hardware
Worked example: an 8B model on Graviton4
Suppose a team wants to serve an 8B-parameter chat model on a single-socket Graviton4 instance for an internal tool with modest traffic. Is that viable, and what throughput should they expect?
Start with decode, the phase that produces one token at a time. Every generated token must read every weight once. At 4-bit quantization with block scales, the weights occupy roughly 4.5 bits per parameter, so about 4.5 GB. Graviton4 has twelve DDR5-5600 channels: 12 x 5600 MT/s x 8 bytes = 537.6 GB/s theoretical. The ceiling for a single stream is therefore 537.6 / 4.5, about 119 tokens per second, before KV-cache reads. Sustained bandwidth is lower than theoretical, so plan against what a STREAM-style test actually measures on the instance, commonly well below the peak. If you measure 400 GB/s, the ceiling becomes about 89 tokens per second, and real kernels will land below that.
Prefill is different. Processing a 512-token prompt costs about 2 x 8e9 x 512, or 8.2 TFLOP. Here the arithmetic units matter, which is where I8MM and well-blocked kernels earn their keep, and where a GPU pulls far ahead. A prompt-heavy workload such as summarising long documents is a poor fit; a chat tool with short prompts and a few concurrent users can be a good one.
Batching helps decode on a CPU for the same reason it helps on a GPU: one pass over the weights serves several sequences. The bandwidth math also tells you what does not help. Adding more cores beyond the point where memory is saturated buys nothing, and a two-socket instance only helps if each socket serves its own replica from local memory.
The CPU in a training node
In training clusters the Neoverse CPU is a supporting actor, but a mis-sized one stalls expensive accelerators. Its jobs are data loading and decoding, tokenisation, augmentation, launching kernels, collective-communication bookkeeping and checkpoint I/O. On Grace-based systems it also serves as a large memory tier reachable over a coherent link; that design, and the porting and NUMA work it requires, is covered in the GH200 deep dive and the GB200 compute-tray guide.
One training technique puts real arithmetic on the CPU: running the optimizer step there while the GPU computes the next backward pass. Adam is element-wise and bandwidth-heavy, which suits wide memory systems and SVE loops; the trade-offs are worked through in optimizer offload to CPU.
A quick sizing rule: measure samples per second for your input pipeline on one CPU core with the accelerator removed, multiply by the cores you can dedicate, and keep the result at least 1.5 times above what the accelerators consume. If you cannot, move decoding offline into a pre-tokenised format rather than adding loader workers.
Failure modes
- SIGILL in a mixed fleet. A wheel or binary built with
-mcpu=neoverse-v2lands on an N1 node. Build for the oldest target and dispatch per feature. - Thread oversubscription. OpenMP threads times intra-op threads times data-loader workers exceeds the core count, and throughput collapses. Set every pool explicitly.
- Counting vCPUs like x86. On x86 a vCPU is usually a hyperthread; on Graviton it is a full core. Copying an x86 thread setting halves usable parallelism or doubles contention.
- Remote memory on two sockets. A process spanning both sockets reads half its weights over the inter-socket link. Run one replica per socket with
numactl. - Silent precision change. bf16 fast-math alters fp32 results. Gate it on an accuracy test, not on a throughput chart.
- Fixed-vector-length assumptions. Hand-written SVE code that assumes 256-bit vectors breaks or slows on 128-bit V2. Use
svcntb()and predicates.
Trade-offs
Choose Neoverse CPU inference when models are small or heavily quantized, traffic is moderate or spiky, latency targets are in hundreds of milliseconds, and you value the operational simplicity of ordinary instances. Choose a GPU when prompts are long, batch sizes are large, or you need tens of tokens per second per user at high concurrency. For the arithmetic-intensity reasoning behind that line, see GPU versus CPU for deep learning, and for how vector instructions map onto transformer kernels, see SIMD instructions for transformer math.
What to do next
- Run the feature probe on every instance type in your fleet and record the tier per node label.
- Rebuild native extensions for the oldest core you run, with per-tier hot kernels behind dispatch.
- Measure sustained memory bandwidth on your target instance and compute the decode ceiling for your model size before benchmarking anything else.
- Benchmark prefill and decode separately with pinned threads and one replica per socket.
- Try the KleidiAI-enabled path in your framework and keep it only if it wins on your model.
- Enable bf16 fast-math only after an accuracy comparison on a held-out set.
- For training nodes, verify the input pipeline delivers at least 1.5 times accelerator demand.