A neural processing unit (NPU) is a block of silicon on a phone, laptop, camera or microcontroller that does one thing well: multiply and accumulate large tensors of low-precision numbers at a fraction of the energy a CPU or GPU would spend on the same work. Apple calls its version the Neural Engine, Qualcomm ships the Hexagon NPU, Intel and AMD put NPUs in their laptop chips, and Arm licenses the Ethos-U family for microcontrollers. Microsoft's Copilot+ PC label requires an NPU rated at more than 40 TOPS.

For a developer, the NPU is not programmed like a GPU. You do not write kernels for it. You hand a model graph to a vendor compiler, usually through a runtime such as Core ML, LiteRT, ONNX Runtime or OpenVINO, and the compiler decides which operators run on the NPU and which fall back to the CPU. This article explains what happens inside the NPU, why quantization is required, how the software stacks partition a model, how to estimate performance before you buy or build, and what breaks in production. Data-center accelerators are covered separately in TPUs and tensor cores.

Advertisement

What an edge NPU is, and what it is not

An edge NPU is a fixed-function or lightly programmable accelerator designed around the operations that dominate neural-network inference: convolutions, matrix multiplications, element-wise activations, pooling and normalisation. It is optimised for energy per inference rather than peak flexibility.

It is not a small GPU. A GPU runs thousands of general threads, each executing arbitrary code, and relies on massive parallelism to hide memory latency. An NPU typically has a small controller that issues a precompiled command stream to a large multiply-accumulate (MAC) array, a vector unit, and DMA engines that move tiles between DRAM and on-chip SRAM. That constraint explains almost everything about working with NPUs: operator coverage, static shapes, quantization and fallback all follow from it.

An edge NPU inside a system-on-chip: the compiler decides what runs whereApp + runtime (CPU)LiteRT / Core ML / ORT / OpenVINONPU compilerpartition, tile, scheduleCompiled blobcommand stream + weightsShared LPDDRweights, activations, KVDMA enginestile in / tile outOn-chip SRAMscratchpad, 1s-10s of MBMAC arrayint8/int4/fp16 multiply-addVector unitactivations, requant, poolingCPU / GPU fallbacksubgraphs the NPU cannot runmodel graphloadEvery NPU-to-CPU boundary costs a synchronisation and often a layout or precision conversion.Decode-time LLM speed is set by DRAM bandwidth, not by the TOPS number on the box.
The runtime hands the graph to a vendor compiler, which emits a command stream for the NPU and leaves unsupported subgraphs on the CPU or GPU. At run time DMA engines stream tiles of weights and activations between shared DRAM and on-chip SRAM, where the MAC array and vector unit work on them.

Inside the NPU: MAC arrays, scratchpad SRAM and the DMA schedule

The MAC array is a grid of multiply-accumulate units, often organised as a systolic array or a set of dot-product engines. Each cycle it multiplies a block of weights by a block of activations and adds the products into wide accumulators, typically 32-bit for int8 inputs so that long dot products do not overflow. Arm's Ethos-U55, for example, is offered in configurations from 32 to 256 MACs per cycle; a laptop NPU has more than ten thousand.

The MAC array is fed from on-chip SRAM, a software-managed scratchpad rather than a cache. The compiler decides exactly which tile of which tensor lives in SRAM at each moment and programs the DMA engines to fetch the next tile while the array works on the current one. This double buffering is the heart of NPU performance: if the compiler cannot hide DRAM transfers behind compute, the array sits idle. Layer fusion matters for the same reason. When a convolution, its bias, its activation and the requantization step run back to back on a tile that stays in SRAM, intermediate tensors never touch DRAM.

The vector unit handles everything that is not a large matrix product: activation functions, often via lookup tables or piecewise approximations, softmax, layer normalisation, pooling, and rescaling int32 accumulators back to int8.

Advertisement

Why integers: quantization as the price of admission

Most edge NPUs are fastest, and some only work, with 8-bit integers, and many now support 4-bit weights. Integer MACs are smaller and use far less energy than floating-point ones, and 8-bit tensors halve memory traffic compared with fp16. The standard scheme is affine quantization: a real value r is represented as r = scale * (q - zero_point), where q is an int8 value. Weights usually get per-channel scales, one per output channel, and activations get a per-tensor scale found by calibration: running a few hundred representative inputs through the float model and recording the range of each activation.

Post-training quantization (PTQ) needs only that calibration data and works well for most convolutional networks. Quantization-aware training (QAT) inserts fake-quantization operations during fine-tuning so the model learns to tolerate rounding, and is worth it when PTQ loses more accuracy than your product can accept. Small language models on NPUs usually combine 4-bit weights, with a scale per group of 32 to 128 weights, and 8- or 16-bit activations, because activations in transformers have outlier channels that per-tensor int8 handles poorly.

The following LiteRT (TensorFlow Lite) conversion produces a fully int8 model, including int8 inputs and outputs, which is what microcontroller NPUs such as Ethos-U require:

import tensorflow as tf

def representative_data():
    # 100-500 real, preprocessed samples from the deployment distribution
    for batch in calibration_images.take(300):
        yield [batch]

converter = tf.lite.TFLiteConverter.from_saved_model("person_detect_saved")
converter.optimizations = [tf.lite.Optimize.DEFAULT]
converter.representative_dataset = representative_data
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
converter.inference_input_type = tf.int8
converter.inference_output_type = tf.int8
open("person_detect_int8.tflite", "wb").write(converter.convert())

Calibrate with real data. Calibrating on random noise or on a handful of images produces activation ranges that clip real inputs, and the accuracy loss shows up only in the field.

The software stack: compilers, delegates and graph partitioning

Each platform has its own path from a trained model to the NPU. The table summarises the common ones; the pattern is the same everywhere.

PlatformRuntime entry pointHow the NPU is selected
Apple Neural EngineCore ML (convert with coremltools)compute_units=ct.ComputeUnit.CPU_AND_NE or ALL; Core ML decides per layer
Qualcomm Hexagon NPUONNX Runtime QNN execution provider, or the Qualcomm AI stackQNN EP with the HTP backend library
Intel NPUOpenVINO, or ONNX Runtime with the OpenVINO EPdevice name "NPU"
Arm Ethos-U (microcontrollers)TensorFlow Lite Microoffline Vela compiler rewrites supported ops into one custom op
Android devices generallyLiteRT with vendor delegatesNNAPI was deprecated in Android 15; use LiteRT's GPU delegate or vendor-specific paths

Every stack performs graph partitioning. The compiler walks the model, marks each operator as supported or not for the NPU given its data type, shape and attributes, and cuts the graph into NPU subgraphs and fallback subgraphs. A single unsupported operator in the middle of a network splits it in two, and each boundary costs a synchronisation, a copy and often a conversion between int8 and float or between memory layouts. A model that is 95% supported can be slower on the NPU than on the CPU if the 5% is scattered.

Static shapes are the other constraint. Most NPU compilers plan tiling and SRAM allocation for fixed tensor shapes at compile time. Dynamic batch sizes, variable sequence lengths or data-dependent control flow either force recompilation, fall back, or are rejected. For transformers this means compiling a few fixed prompt-chunk lengths and a fixed maximum KV-cache size rather than a fully dynamic graph.

Worked example: a vision model and a 3B language model on a 45 TOPS laptop

Consider a laptop whose NPU is advertised at 45 TOPS and whose LPDDR5X memory delivers, say, 100 GB/s shared between CPU, GPU and NPU. Vendor TOPS figures are peak int8 operations per second, counting a multiply-add as two operations, often with every MAC busy on ideal shapes. Plan with 20 to 40% of peak for real graphs.

Vision model. A MobileNetV2-class image classifier at 224 by 224 pixels needs roughly 0.6 billion operations per frame; detectors need several times more. At an effective 10 TOPS that is 60 microseconds of compute. In practice you will measure one to a few milliseconds, because the per-inference cost is dominated by dispatch, DMA of activations, pre- and post-processing on the CPU and any fallback subgraphs. For small models the lesson is to minimise boundaries and CPU work, not to chase TOPS.

Language model, prefill. A 3-billion-parameter model needs about 2 x 3e9 = 6 GFLOP per token. A 1,000-token prompt is 6e12 operations; at 12 effective TOPS that is about half a second. Prefill is compute-bound, so this is where the NPU earns its keep.

Language model, decode. Each generated token must read every weight once. With 4-bit weights plus group scales the model is about 1.65 GB, so at 100 GB/s the ceiling is roughly 100 / 1.65, about 60 tokens per second, before the KV cache, the CPU and the display compete for the same bus. The NPU's TOPS are irrelevant here; memory bandwidth and weight size decide. That is why weight quantization matters more than activation precision for on-device chat, and why the same model often decodes at similar speed on the laptop's GPU and NPU but at very different power.

def npu_estimate(params_b, weight_bits, prompt_tokens, eff_tops, mem_gbps, overhead=1.1):
    weight_gb = params_b * weight_bits / 8 * overhead   # overhead: scales, zero points
    prefill_s = 2 * params_b * 1e9 * prompt_tokens / (eff_tops * 1e12)
    decode_tok_s = mem_gbps / weight_gb                  # upper bound, bandwidth-bound
    return weight_gb, prefill_s, decode_tok_s

print(npu_estimate(3, 4, 1000, 12, 100))   # (~1.65 GB, ~0.5 s, ~60 tok/s ceiling)

Targeting the NPU from three runtimes

The code below shows the minimum needed to put a model on the NPU in three ecosystems. In each case, verify where operators actually ran rather than trusting the request.

# 1. Apple: Core ML, requesting CPU + Neural Engine only (no GPU)
import coremltools as ct
mlmodel = ct.convert(traced_model,
                     inputs=[ct.TensorType(shape=(1, 3, 224, 224))],
                     compute_units=ct.ComputeUnit.CPU_AND_NE)
mlmodel.save("classifier.mlpackage")

# 2. Windows on Snapdragon: ONNX Runtime with the QNN execution provider (HTP = NPU)
import onnxruntime as ort
so = ort.SessionOptions()
so.add_session_config_entry("session.disable_cpu_ep_fallback", "1")  # fail loudly in CI
sess = ort.InferenceSession("classifier.qdq.onnx", sess_options=so,
                            providers=[("QNNExecutionProvider", {"backend_path": "QnnHtp.dll"})])

# 3. Intel NPU: OpenVINO
import openvino as ov
core = ov.Core()
print(core.available_devices)            # e.g. ['CPU', 'GPU', 'NPU'] when the driver is present
compiled = core.compile_model("classifier.xml", "NPU")

For microcontrollers the flow is offline: run vela person_detect_int8.tflite --accelerator-config ethos-u55-128 and deploy the rewritten model with TensorFlow Lite Micro. Vela reports which operators it mapped to the Ethos-U and which remain on the Cortex-M CPU, and it reports SRAM and flash usage.

Failure modes

  • Silent fallback. A model update adds an operator the NPU compiler does not support, the runtime quietly moves that subgraph to the CPU, and latency triples without any error. Track per-backend operator counts in CI.
  • Accuracy loss after quantization. Aggregate accuracy looks fine but a minority class, low-light images or accented speech degrade badly. Evaluate the quantized model on sliced test sets, not only overall metrics.
  • Driver and OS drift. The same compiled model behaves differently after an OS or NPU driver update, because the vendor compiler changed. Pin versions in your test fleet and re-run your benchmark suite on every update.
  • First-run compilation stalls. Several runtimes compile or specialise the graph on first load, which can take seconds. Cache compiled artifacts where the runtime supports it and warm up models off the critical path.
  • Thermal and power throttling. A benchmark run for ten seconds hits peak clocks; a video call running for an hour does not. Measure sustained performance with the device in its real enclosure.

Operational guidance

Treat the NPU as one backend among several and keep a working CPU path for every model. Ship a capability check at startup that records the device, driver version, which backend each model landed on and a short warm-up latency, and report it through telemetry. Keep the float reference model and a golden set of inputs and outputs so you can compare quantized output against the reference on real devices and flag drift.

Profile with the vendor tools rather than wall-clock timing alone: Xcode's Core ML performance reports show the compute unit per layer, Qualcomm and Intel ship profilers that break down time by operator and by NPU versus fallback, and Vela's summary shows cycles per operator. Budget energy as well as latency, since the NPU's main advantage over the GPU on a laptop is often joules per inference, not milliseconds.

Trade-offs: NPU, GPU or CPU on the device

BackendStrengthsWeaknesses
NPUbest energy per inference, frees GPU and CPU, strong int8/int4 throughputvendor compilers, limited operators, static shapes, per-vendor toolchains
Integrated GPUflexible, fp16 friendly, mature runtimes, handles dynamic shapes bettermore power, contends with graphics
CPUruns everything, easiest to debug, good for tiny or control-heavy modelsslowest for large tensors, worst energy

Always-on and background features belong on the NPU. Bursty, latency-critical work with unusual operators may be better on the GPU. For small language models, compare both on tokens per second and on joules per token; see Apple Silicon for ML for how unified memory changes the picture, and small language model edge deployment for the serving architecture around the model.

What to do next

  1. List the devices you ship to and record, for each, the NPU, the runtime entry point and the driver or OS version you will support.
  2. Convert one production model with post-training int8 quantization using a few hundred real calibration samples, and measure accuracy on sliced test sets against the float model.
  3. Run the vendor compiler or runtime report and count operators on the NPU versus fallback; eliminate or replace operators that create boundaries.
  4. Add a CI job that creates the session with CPU fallback disabled so any new unsupported operator fails the build.
  5. Estimate LLM decode speed from weight size and memory bandwidth before promising a tokens-per-second figure, then measure sustained performance for ten minutes on a real device.
  6. Ship startup telemetry recording which backend each model landed on, and keep a CPU path as a tested fallback.
Key takeaway: An edge NPU is a MAC array fed by a compiler-managed SRAM scratchpad and DMA engines, reached only through vendor compilers and runtimes. Quantized integer models, native operators and static shapes are what let the compiler keep a whole model on the NPU, and every fallback boundary costs time. TOPS figures describe peak prefill-style compute; small-LLM decode speed is set by memory bandwidth and weight size. Verify where every operator runs, test on real devices under sustained load, and keep a CPU path for when the NPU is unavailable.