Every M-series Mac, and every recent iPhone and iPad, contains a Neural Engine (ANE): a fixed-function accelerator for neural network inference that sits beside the CPU and GPU on the same chip and shares their memory. It is often the most power-efficient way to run a model on the device, and it is also the hardest engine to target, because Apple publishes no instruction set, no low-level API and no detailed architecture. You reach it only through Core ML, which compiles your model and decides, operation by operation, whether the ANE will run it.
This page is about working with that reality. It covers what is publicly known, how a model gets from training to the ANE, how to write layers the ANE accepts, how to check where each operation actually ran, how precision and quantization behave, and how training decisions made on a GPU determine whether the exported model runs on the ANE at all. For the chip as a whole, start with Apple Silicon in depth.
What is publicly known
Apple states core counts and peak operations per second in marketing material, and little else. The base, Pro and Max chips of each M-series generation have a 16-core Neural Engine; the Ultra chips, built from two fused dies, list 32 cores. The headline base-chip figures are about 11 trillion operations per second for M1, 15.8 for M2, about 18 for M3 and 38 for M4. Treat these with care: the precision behind each number is not stated consistently, and the M4 figure is widely understood to be an int8 rate while earlier figures were not, so the jump overstates the real change for fp16 models. Apple has not been consistent in quoting a comparable figure for M5, so none is given here.
What matters more for software is the shape of the engine, which is consistent across generations and documented indirectly by Apple's own optimisation guidance. It runs fixed computation graphs compiled ahead of time. It computes in float16, with int8 compute available on newer chips. It prefers static shapes. It shares system memory bandwidth with everything else, so large weights streamed for each inference make it bandwidth-bound quickly. It is used for inference; Apple offers no public way to train on it.
| Engine | Reached through | Strength | Weakness |
|---|---|---|---|
| Neural Engine | Core ML only | Lowest energy per inference for supported graphs | Opaque; op and shape restrictions; fp16 or int8 |
| GPU | Metal, MPS, MLX, Core ML | Flexible, high throughput, float32 available | Higher power draw |
| CPU | Accelerate, Core ML, anything | Runs every op | Slowest for large tensors |
Architecture: how a model reaches the ANE
The flow has four stages. You train in PyTorch or MLX on a GPU. coremltools traces the model and converts it to an ML Program, a typed intermediate representation stored in an .mlpackage, optionally compressing weights. On first load the device specialises the program for its own hardware and caches the result, which is why first launch is slower than every later one and why an OS update can make the next launch slow again. At prediction time the runtime executes segments on the engines allowed by MLComputeUnits.
The partition is the key fact. If one operation in the middle of a block is not supported on the ANE, the graph is cut there, the intermediate tensor is handed to the GPU or CPU, and then handed back. Each hand-off costs synchronisation and memory traffic, so a model that is 95 percent ANE-compatible can run slower than the same model on the GPU alone.
Writing layers the ANE accepts
Apple's research note on deploying transformers to the Neural Engine, with its reference implementation ml-ane-transformers, gives the clearest public rules. Four matter most. Represent activations as a 4D tensor shaped (batch, channels, 1, sequence) rather than the usual (batch, sequence, channels). Replace nn.Linear with a 1x1 nn.Conv2d on that layout. Split multi-head attention into per-head computations so no single intermediate tensor is huge. Avoid reshapes and transposes, which turn into memory copies. The note also explains that the last axis of an ANE buffer is padded to 64 bytes, so a tiny last dimension wastes memory and bandwidth; putting the sequence last is what makes the layout efficient.
import torch
import torch.nn as nn
class ANEAttention(nn.Module):
"""Self-attention on (B, C, 1, S) tensors, split per head (after Apple's ml-ane-transformers)."""
def __init__(self, dim: int, heads: int):
super().__init__()
self.heads, self.dh = heads, dim // heads
self.q = nn.Conv2d(dim, dim, 1)
self.k = nn.Conv2d(dim, dim, 1)
self.v = nn.Conv2d(dim, dim, 1)
self.o = nn.Conv2d(dim, dim, 1)
def forward(self, x, mask): # x: (B, C, 1, S); mask: (B, S, 1, S), 0 or -1e4
scale = self.dh ** -0.5
qs = (self.q(x) * scale).split(self.dh, dim=1) # scale early: keeps fp16 logits small
ks = self.k(x).split(self.dh, dim=1)
vs = self.v(x).split(self.dh, dim=1)
outs = []
for qh, kh, vh in zip(qs, ks, vs):
w = torch.einsum("bchq,bchk->bkhq", qh, kh) + mask # (B, S_k, 1, S_q)
w = w.softmax(dim=1)
outs.append(torch.einsum("bkhq,bchk->bchq", w, vh))
return self.o(torch.cat(outs, dim=1))
def linear_to_conv(lin: nn.Linear) -> nn.Conv2d:
"""Reuse pretrained Linear weights in the 1x1 convolution."""
conv = nn.Conv2d(lin.in_features, lin.out_features, 1, bias=lin.bias is not None)
conv.weight.data = lin.weight.data[:, :, None, None].clone()
if lin.bias is not None:
conv.bias.data = lin.bias.data.clone()
return convLayer normalisation must then normalise over dimension 1 (channels) instead of the last dimension; the reference repository includes such a module. Note the mask value of -1e4 rather than negative infinity or -1e9: the latter is outside float16 range and becomes infinity.
Converting for the ANE
Convert with fp16 compute, a deployment target new enough for the features you use, and enumerated shapes instead of open ranges. Fully flexible ranges often keep operations off the ANE; a small set of fixed sequence lengths lets Core ML prepare each one.
import coremltools as ct
import numpy as np
model = Encoder().eval() # built from ANE-style blocks
example = (torch.rand(1, 768, 1, 128), torch.zeros(1, 128, 1, 128))
traced = torch.jit.trace(model, example)
lengths = [64, 128, 256]
mlmodel = ct.convert(
traced,
inputs=[
ct.TensorType(name="x", dtype=np.float16,
shape=ct.EnumeratedShapes(shapes=[[1, 768, 1, s] for s in lengths])),
ct.TensorType(name="mask", dtype=np.float16,
shape=ct.EnumeratedShapes(shapes=[[1, s, 1, s] for s in lengths])),
],
convert_to="mlprogram",
compute_precision=ct.precision.FLOAT16,
compute_units=ct.ComputeUnit.CPU_AND_NE,
minimum_deployment_target=ct.target.macOS14,
)
mlmodel.save("Encoder.mlpackage")Setting CPU_AND_NE during testing is deliberate: it removes the GPU as a silent fallback, so operations the ANE rejects land on the CPU and show up clearly in timings and in the compute plan. In production you may prefer .all and let Core ML choose. If two inputs share a dimension, check after conversion that enumerated shapes were accepted as you intended; some combinations are not supported, and the conversion error is the place to find out.
Checking where every op runs
Do not guess where operations ran. Since macOS 14.4 and iOS 17.4, MLComputePlan reports, for every operation in an ML Program, the preferred compute device, the devices that could run it, and an estimated share of the total cost. Xcode's Core ML performance report shows the same information interactively.
import CoreML
func reportPlacement(modelURL: URL) async throws {
let config = MLModelConfiguration()
config.computeUnits = .cpuAndNeuralEngine
let plan = try await MLComputePlan.load(contentsOf: modelURL, configuration: config)
guard case let .program(program) = plan.modelStructure,
let main = program.functions["main"] else { return }
var offANE: [(String, Double)] = []
for op in main.block.operations {
guard let usage = plan.deviceUsage(for: op) else { continue }
let cost = plan.estimatedCost(of: op)?.weight ?? 0
if case .neuralEngine = usage.preferred { continue }
offANE.append((op.operatorName, cost))
}
for (name, cost) in offANE.sorted(by: { $0.1 > $1.1 }) {
print(name, String(format: "%.3f", cost))
}
}The output is a ranked list of operations not placed on the ANE, weighted by estimated cost. Fix from the top. If you start from an .mlpackage, compile it first with MLModel.compileModel(at:) and pass the resulting .mlmodelc URL.
Worked example: moving an encoder onto the ANE
Worked example, as a workflow rather than a benchmark. A team exports a six-layer sentence encoder straight from a Hugging Face model with nn.Linear layers, rank-3 activations and a sequence dimension given as an open range. The compute plan shows most of the estimated cost preferred on the CPU, and timing with CPU_AND_NE is worse than the GPU run. Three changes follow:
- Swap the attention and feed-forward blocks for ANE-style modules, porting weights with
linear_to_conv, and verify the outputs match the original within fp16 tolerance on a test set. - Replace the open sequence range with enumerated lengths of 64, 128 and 256, padding inputs up to the next length.
- Change the attention mask constant from
-1e9to-1e4after an fp16 comparison shows infinities in the softmax inputs.
After re-conversion the plan shows the transformer blocks preferred on the ANE. The token embedding lookup still prefers the CPU; it is cheap, at the start of the graph, and costs one hand-off, so it stays. The team then measures, rather than assumes, the gain on each target device, because the balance between engines differs by chip.
Measuring latency and energy
Measure on the device and configuration you ship. Load once outside the timed region, discard warm-up runs, and compare compute-unit settings on identical inputs. On macOS, coremltools can run predictions directly:
import time
def bench(path, units, feed, runs=200, warmup=20):
m = ct.models.MLModel(path, compute_units=units) # load and specialise first
for _ in range(warmup):
m.predict(feed)
t = []
for _ in range(runs):
t0 = time.perf_counter(); m.predict(feed); t.append(time.perf_counter() - t0)
t.sort()
return t[len(t) // 2] * 1e3, t[int(len(t) * 0.99)] * 1e3 # p50, p99 in ms
feed = {"x": np.random.rand(1, 768, 1, 128).astype(np.float16),
"mask": np.zeros((1, 128, 1, 128), dtype=np.float16)}
for units in (ct.ComputeUnit.CPU_AND_NE, ct.ComputeUnit.CPU_AND_GPU, ct.ComputeUnit.ALL):
print(units, bench("Encoder.mlpackage", units, feed))Latency is half the story. The ANE's advantage is usually energy, which matters for battery devices and for keeping the GPU free for rendering. Use the Core ML and power profiling templates in Instruments to see which engine was busy and for how long.
Precision and quantisation
The ANE computes in float16, whose largest finite value is 65504. Models trained in float32 or bfloat16 can produce intermediate values beyond that: attention logits, residual streams in large models, and the sum of squares inside a normalisation. The symptom is infinities or NaN that appear only on the ANE. Scale queries before the dot product, keep mask constants inside range, and compare fp16 outputs against float32 on real inputs during conversion; FP16 in depth explains the range limits.
Weight compression in coremltools (linear int8 or int4 quantisation, palettisation, pruning) shrinks the model and cuts bandwidth, which is often the real ANE bottleneck. Quantising activations as well, the W8A8 mode, uses int8-int8 compute; Apple's coremltools documentation says that path is faster from iPhone 15 Pro onwards, and developers report the same for M4-class Macs. It needs calibration data:
import coremltools.optimize as cto
act_cfg = cto.coreml.OptimizationConfig(
global_config=cto.coreml.experimental.OpActivationLinearQuantizerConfig(mode="linear_symmetric"))
a8 = cto.coreml.experimental.linear_quantize_activations(mlmodel, act_cfg, calibration_samples)
w_cfg = cto.coreml.OptimizationConfig(
global_config=cto.coreml.OpLinearQuantizerConfig(mode="linear_symmetric", weight_threshold=512))
w8a8 = cto.coreml.linear_quantize_weights(a8, config=w_cfg)Re-run accuracy and the compute plan after every compression step; a quantised model is a new model.
Where training fits
Training never touches the ANE, but it decides what the ANE can do. Choose architectures whose operations convert cleanly, and test the export early rather than after training. Keep activations fp16-safe: watch activation maxima during training and prefer normalisation placements that bound them. If you plan int8 deployment, use quantisation-aware training or at least keep a representative calibration set. Apple's own on-device foundation model is customised by training adapters off the device with Apple's toolkit, which follows the same principle: the device runs fixed graphs, and learning happens elsewhere. For workflows where you do need to train or fine-tune locally, the GPU through MLX is the tool.
Failure modes
- Silent GPU fallback. With
.alla model that never touches the ANE still works, just with more power. Check the compute plan in CI for each release. - Fragmented graphs. A few unsupported operations in every block cause repeated hand-offs. Fix or move them to the edges of the graph.
- fp16 overflow. NaN only on device. Compare fp16 and float32 outputs at conversion time on real data.
- Open shape ranges. Flexible ranges keep operations off the ANE. Enumerate the lengths you need.
- First-load stalls. Specialisation on first load or after an OS update delays the first prediction. Load the model in the background before the user needs it.
- Comparing TOPS across generations. Headline numbers use different precisions; benchmark your own model.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Target the ANE | Low energy, GPU stays free | Layout rewrites, op limits, opaque tooling |
| Target the GPU (Metal, MLX) | Flexibility, float32, custom kernels | Higher power; competes with graphics |
| Enumerated shapes | ANE-eligible and pre-specialised | Padding waste between lengths |
| W8A8 quantisation | int8 compute on newer chips | Calibration and accuracy risk |
| Core ML decides placement | Less code | Placement changes between OS versions |
What to do next
- Export one model you already ship and run the compute plan script against it; record the share of cost not on the ANE.
- Rewrite the heaviest blocks into the (B, C, 1, S) layout with 1x1 convolutions and confirm outputs match.
- Replace open shape ranges with enumerated shapes and add an fp16-versus-float32 output check to conversion.
- Benchmark p50 and p99 with each compute-unit setting on every target device, and profile energy in Instruments.
- Try weight compression, then W8A8 where the hardware supports it, re-checking accuracy each time.
- Keep learning: Core ML for LLMs and edge NPUs in general.