Every M-series Mac, and every recent iPhone and iPad, contains a Neural Engine (ANE): a fixed-function accelerator for neural network inference that sits beside the CPU and GPU on the same chip and shares their memory. It is often the most power-efficient way to run a model on the device, and it is also the hardest engine to target, because Apple publishes no instruction set, no low-level API and no detailed architecture. You reach it only through Core ML, which compiles your model and decides, operation by operation, whether the ANE will run it.

This page is about working with that reality. It covers what is publicly known, how a model gets from training to the ANE, how to write layers the ANE accepts, how to check where each operation actually ran, how precision and quantization behave, and how training decisions made on a GPU determine whether the exported model runs on the ANE at all. For the chip as a whole, start with Apple Silicon in depth.

What is publicly known

Apple states core counts and peak operations per second in marketing material, and little else. The base, Pro and Max chips of each M-series generation have a 16-core Neural Engine; the Ultra chips, built from two fused dies, list 32 cores. The headline base-chip figures are about 11 trillion operations per second for M1, 15.8 for M2, about 18 for M3 and 38 for M4. Treat these with care: the precision behind each number is not stated consistently, and the M4 figure is widely understood to be an int8 rate while earlier figures were not, so the jump overstates the real change for fp16 models. Apple has not been consistent in quoting a comparable figure for M5, so none is given here.

What matters more for software is the shape of the engine, which is consistent across generations and documented indirectly by Apple's own optimisation guidance. It runs fixed computation graphs compiled ahead of time. It computes in float16, with int8 compute available on newer chips. It prefers static shapes. It shares system memory bandwidth with everything else, so large weights streamed for each inference make it bandwidth-bound quickly. It is used for inference; Apple offers no public way to train on it.

EngineReached throughStrengthWeakness
Neural EngineCore ML onlyLowest energy per inference for supported graphsOpaque; op and shape restrictions; fp16 or int8
GPUMetal, MPS, MLX, Core MLFlexible, high throughput, float32 availableHigher power draw
CPUAccelerate, Core ML, anythingRuns every opSlowest for large tensors

Architecture: how a model reaches the ANE

From a PyTorch module to Neural Engine executionPyTorch / MLXtrain on GPUcoremltoolstrace, convert, compress.mlpackageML Program, fp16on-device loadspecialise, cacheCore ML runtimepartitions ops by MLComputeUnitsNeural Enginefp16, int8 on newer chipsGPUMetal, fallback for odd opsCPUunsupported ops, gluesupported segmentfallbackunified memoryevery switch between engines is a hand-off and a sync pointMLComputePlan shows the preferred device and estimated cost for every operation before you ship.
Training happens elsewhere; Core ML compiles the exported program on the device and splits it into segments for each engine.

The flow has four stages. You train in PyTorch or MLX on a GPU. coremltools traces the model and converts it to an ML Program, a typed intermediate representation stored in an .mlpackage, optionally compressing weights. On first load the device specialises the program for its own hardware and caches the result, which is why first launch is slower than every later one and why an OS update can make the next launch slow again. At prediction time the runtime executes segments on the engines allowed by MLComputeUnits.

The partition is the key fact. If one operation in the middle of a block is not supported on the ANE, the graph is cut there, the intermediate tensor is handed to the GPU or CPU, and then handed back. Each hand-off costs synchronisation and memory traffic, so a model that is 95 percent ANE-compatible can run slower than the same model on the GPU alone.

Writing layers the ANE accepts

Apple's research note on deploying transformers to the Neural Engine, with its reference implementation ml-ane-transformers, gives the clearest public rules. Four matter most. Represent activations as a 4D tensor shaped (batch, channels, 1, sequence) rather than the usual (batch, sequence, channels). Replace nn.Linear with a 1x1 nn.Conv2d on that layout. Split multi-head attention into per-head computations so no single intermediate tensor is huge. Avoid reshapes and transposes, which turn into memory copies. The note also explains that the last axis of an ANE buffer is padded to 64 bytes, so a tiny last dimension wastes memory and bandwidth; putting the sequence last is what makes the layout efficient.

import torch
import torch.nn as nn

class ANEAttention(nn.Module):
    """Self-attention on (B, C, 1, S) tensors, split per head (after Apple's ml-ane-transformers)."""
    def __init__(self, dim: int, heads: int):
        super().__init__()
        self.heads, self.dh = heads, dim // heads
        self.q = nn.Conv2d(dim, dim, 1)
        self.k = nn.Conv2d(dim, dim, 1)
        self.v = nn.Conv2d(dim, dim, 1)
        self.o = nn.Conv2d(dim, dim, 1)

    def forward(self, x, mask):                    # x: (B, C, 1, S); mask: (B, S, 1, S), 0 or -1e4
        scale = self.dh ** -0.5
        qs = (self.q(x) * scale).split(self.dh, dim=1)   # scale early: keeps fp16 logits small
        ks = self.k(x).split(self.dh, dim=1)
        vs = self.v(x).split(self.dh, dim=1)
        outs = []
        for qh, kh, vh in zip(qs, ks, vs):
            w = torch.einsum("bchq,bchk->bkhq", qh, kh) + mask   # (B, S_k, 1, S_q)
            w = w.softmax(dim=1)
            outs.append(torch.einsum("bkhq,bchk->bchq", w, vh))
        return self.o(torch.cat(outs, dim=1))

def linear_to_conv(lin: nn.Linear) -> nn.Conv2d:
    """Reuse pretrained Linear weights in the 1x1 convolution."""
    conv = nn.Conv2d(lin.in_features, lin.out_features, 1, bias=lin.bias is not None)
    conv.weight.data = lin.weight.data[:, :, None, None].clone()
    if lin.bias is not None:
        conv.bias.data = lin.bias.data.clone()
    return conv

Layer normalisation must then normalise over dimension 1 (channels) instead of the last dimension; the reference repository includes such a module. Note the mask value of -1e4 rather than negative infinity or -1e9: the latter is outside float16 range and becomes infinity.

Converting for the ANE

Convert with fp16 compute, a deployment target new enough for the features you use, and enumerated shapes instead of open ranges. Fully flexible ranges often keep operations off the ANE; a small set of fixed sequence lengths lets Core ML prepare each one.

import coremltools as ct
import numpy as np

model = Encoder().eval()                      # built from ANE-style blocks
example = (torch.rand(1, 768, 1, 128), torch.zeros(1, 128, 1, 128))
traced = torch.jit.trace(model, example)

lengths = [64, 128, 256]
mlmodel = ct.convert(
    traced,
    inputs=[
        ct.TensorType(name="x", dtype=np.float16,
                      shape=ct.EnumeratedShapes(shapes=[[1, 768, 1, s] for s in lengths])),
        ct.TensorType(name="mask", dtype=np.float16,
                      shape=ct.EnumeratedShapes(shapes=[[1, s, 1, s] for s in lengths])),
    ],
    convert_to="mlprogram",
    compute_precision=ct.precision.FLOAT16,
    compute_units=ct.ComputeUnit.CPU_AND_NE,
    minimum_deployment_target=ct.target.macOS14,
)
mlmodel.save("Encoder.mlpackage")

Setting CPU_AND_NE during testing is deliberate: it removes the GPU as a silent fallback, so operations the ANE rejects land on the CPU and show up clearly in timings and in the compute plan. In production you may prefer .all and let Core ML choose. If two inputs share a dimension, check after conversion that enumerated shapes were accepted as you intended; some combinations are not supported, and the conversion error is the place to find out.

Checking where every op runs

Do not guess where operations ran. Since macOS 14.4 and iOS 17.4, MLComputePlan reports, for every operation in an ML Program, the preferred compute device, the devices that could run it, and an estimated share of the total cost. Xcode's Core ML performance report shows the same information interactively.

import CoreML

func reportPlacement(modelURL: URL) async throws {
    let config = MLModelConfiguration()
    config.computeUnits = .cpuAndNeuralEngine
    let plan = try await MLComputePlan.load(contentsOf: modelURL, configuration: config)
    guard case let .program(program) = plan.modelStructure,
          let main = program.functions["main"] else { return }

    var offANE: [(String, Double)] = []
    for op in main.block.operations {
        guard let usage = plan.deviceUsage(for: op) else { continue }
        let cost = plan.estimatedCost(of: op)?.weight ?? 0
        if case .neuralEngine = usage.preferred { continue }
        offANE.append((op.operatorName, cost))
    }
    for (name, cost) in offANE.sorted(by: { $0.1 > $1.1 }) {
        print(name, String(format: "%.3f", cost))
    }
}

The output is a ranked list of operations not placed on the ANE, weighted by estimated cost. Fix from the top. If you start from an .mlpackage, compile it first with MLModel.compileModel(at:) and pass the resulting .mlmodelc URL.

Worked example: moving an encoder onto the ANE

Worked example, as a workflow rather than a benchmark. A team exports a six-layer sentence encoder straight from a Hugging Face model with nn.Linear layers, rank-3 activations and a sequence dimension given as an open range. The compute plan shows most of the estimated cost preferred on the CPU, and timing with CPU_AND_NE is worse than the GPU run. Three changes follow:

  1. Swap the attention and feed-forward blocks for ANE-style modules, porting weights with linear_to_conv, and verify the outputs match the original within fp16 tolerance on a test set.
  2. Replace the open sequence range with enumerated lengths of 64, 128 and 256, padding inputs up to the next length.
  3. Change the attention mask constant from -1e9 to -1e4 after an fp16 comparison shows infinities in the softmax inputs.

After re-conversion the plan shows the transformer blocks preferred on the ANE. The token embedding lookup still prefers the CPU; it is cheap, at the start of the graph, and costs one hand-off, so it stays. The team then measures, rather than assumes, the gain on each target device, because the balance between engines differs by chip.

Measuring latency and energy

Measure on the device and configuration you ship. Load once outside the timed region, discard warm-up runs, and compare compute-unit settings on identical inputs. On macOS, coremltools can run predictions directly:

import time

def bench(path, units, feed, runs=200, warmup=20):
    m = ct.models.MLModel(path, compute_units=units)   # load and specialise first
    for _ in range(warmup):
        m.predict(feed)
    t = []
    for _ in range(runs):
        t0 = time.perf_counter(); m.predict(feed); t.append(time.perf_counter() - t0)
    t.sort()
    return t[len(t) // 2] * 1e3, t[int(len(t) * 0.99)] * 1e3   # p50, p99 in ms

feed = {"x": np.random.rand(1, 768, 1, 128).astype(np.float16),
        "mask": np.zeros((1, 128, 1, 128), dtype=np.float16)}
for units in (ct.ComputeUnit.CPU_AND_NE, ct.ComputeUnit.CPU_AND_GPU, ct.ComputeUnit.ALL):
    print(units, bench("Encoder.mlpackage", units, feed))

Latency is half the story. The ANE's advantage is usually energy, which matters for battery devices and for keeping the GPU free for rendering. Use the Core ML and power profiling templates in Instruments to see which engine was busy and for how long.

Precision and quantisation

The ANE computes in float16, whose largest finite value is 65504. Models trained in float32 or bfloat16 can produce intermediate values beyond that: attention logits, residual streams in large models, and the sum of squares inside a normalisation. The symptom is infinities or NaN that appear only on the ANE. Scale queries before the dot product, keep mask constants inside range, and compare fp16 outputs against float32 on real inputs during conversion; FP16 in depth explains the range limits.

Weight compression in coremltools (linear int8 or int4 quantisation, palettisation, pruning) shrinks the model and cuts bandwidth, which is often the real ANE bottleneck. Quantising activations as well, the W8A8 mode, uses int8-int8 compute; Apple's coremltools documentation says that path is faster from iPhone 15 Pro onwards, and developers report the same for M4-class Macs. It needs calibration data:

import coremltools.optimize as cto

act_cfg = cto.coreml.OptimizationConfig(
    global_config=cto.coreml.experimental.OpActivationLinearQuantizerConfig(mode="linear_symmetric"))
a8 = cto.coreml.experimental.linear_quantize_activations(mlmodel, act_cfg, calibration_samples)

w_cfg = cto.coreml.OptimizationConfig(
    global_config=cto.coreml.OpLinearQuantizerConfig(mode="linear_symmetric", weight_threshold=512))
w8a8 = cto.coreml.linear_quantize_weights(a8, config=w_cfg)

Re-run accuracy and the compute plan after every compression step; a quantised model is a new model.

Where training fits

Training never touches the ANE, but it decides what the ANE can do. Choose architectures whose operations convert cleanly, and test the export early rather than after training. Keep activations fp16-safe: watch activation maxima during training and prefer normalisation placements that bound them. If you plan int8 deployment, use quantisation-aware training or at least keep a representative calibration set. Apple's own on-device foundation model is customised by training adapters off the device with Apple's toolkit, which follows the same principle: the device runs fixed graphs, and learning happens elsewhere. For workflows where you do need to train or fine-tune locally, the GPU through MLX is the tool.

Failure modes

  • Silent GPU fallback. With .all a model that never touches the ANE still works, just with more power. Check the compute plan in CI for each release.
  • Fragmented graphs. A few unsupported operations in every block cause repeated hand-offs. Fix or move them to the edges of the graph.
  • fp16 overflow. NaN only on device. Compare fp16 and float32 outputs at conversion time on real data.
  • Open shape ranges. Flexible ranges keep operations off the ANE. Enumerate the lengths you need.
  • First-load stalls. Specialisation on first load or after an OS update delays the first prediction. Load the model in the background before the user needs it.
  • Comparing TOPS across generations. Headline numbers use different precisions; benchmark your own model.

Trade-offs

ChoiceGainCost
Target the ANELow energy, GPU stays freeLayout rewrites, op limits, opaque tooling
Target the GPU (Metal, MLX)Flexibility, float32, custom kernelsHigher power; competes with graphics
Enumerated shapesANE-eligible and pre-specialisedPadding waste between lengths
W8A8 quantisationint8 compute on newer chipsCalibration and accuracy risk
Core ML decides placementLess codePlacement changes between OS versions

What to do next

  1. Export one model you already ship and run the compute plan script against it; record the share of cost not on the ANE.
  2. Rewrite the heaviest blocks into the (B, C, 1, S) layout with 1x1 convolutions and confirm outputs match.
  3. Replace open shape ranges with enumerated shapes and add an fp16-versus-float32 output check to conversion.
  4. Benchmark p50 and p99 with each compute-unit setting on every target device, and profile energy in Instruments.
  5. Try weight compression, then W8A8 where the hardware supports it, re-checking accuracy each time.
  6. Keep learning: Core ML for LLMs and edge NPUs in general.
Key takeaway: The Neural Engine is reachable only through Core ML, runs fixed fp16 or int8 graphs, and is used only for operations it supports, with costly hand-offs around the rest. Write layers in the channels-first layout with 1x1 convolutions, enumerate shapes, keep activations in fp16 range, and verify placement with MLComputePlan and real measurements on every target device.