Core ML is the system framework that runs machine learning models inside apps on iPhone, iPad, Mac and Vision Pro. It takes a model compiled ahead of time, chooses where each part runs among the CPU, GPU and Neural Engine, and gives the app a typed prediction API. That design is ideal for a classifier that sees one fixed-size image at a time. A large language model is a different animal: it runs hundreds of forward passes per answer, each one depending on a growing key-value cache, with an input length that changes between the prompt and every decoding step.
This article shows how to make that work. You will build a wrapper that turns the KV cache into Core ML state, convert with flexible sequence lengths, compress the weights to 4 bits, and write the generation loop on both the Python and Swift sides. A worked example budgets memory and bandwidth for an 8-billion-parameter model so that you can predict speed before you convert anything. For the hardware underneath, read Apple Silicon, in depth first.
Why an LLM is an awkward fit
Three properties of Core ML collide with autoregressive decoding. First, a Core ML model is a static program: you trace it once, it is compiled into an mlprogram, and the runtime plans memory and placement for the shapes it knows about. Second, inputs and outputs cross a boundary as tensors that the framework may copy. Third, before iOS 18 and macOS 15 a model had no memory of its own between calls.
Decoding needs the opposite. The prompt is processed in one pass, after which every step processes exactly one new token while reading the keys and values of every earlier token. If that cache is an ordinary input and output, the app copies the whole cache in and out on every token, hundreds of megabytes per step for a large model. That is the main reason early on-device ports were slow.
The fix arrived with stateful models: a tensor declared as state lives inside the model across predictions and is updated in place, never copied through the API.
The architecture
Read the diagram as two phases. The offline phase turns a PyTorch checkpoint into an .mlpackage: a wrapper module exposes the cache as registered buffers, torch.jit.trace captures the graph, coremltools converts it with the cache declared as state, and a compression pass shrinks the weights. The online phase belongs to the app: the package is compiled on the device the first time it is loaded, a state object is created per conversation, and a loop feeds token ids and a causal mask in and reads logits out.
Keep the tokenizer and sampler outside the model: they change more often than the weights, and in Swift you can implement temperature, top-p and stop sequences without reconverting.
Step 1: make the KV cache a buffer
Core ML recognises a state from a PyTorch buffer registered with register_buffer. The wrapper below owns two preallocated tensors, one for keys and one for values, shaped as layers by batch by KV heads by maximum context by head dimension. Each forward call writes the new keys and values into a slice and attends over everything written so far. This mirrors the approach in Apple's Llama 3.1 write-up; the class names are illustrative.
import torch
from transformers import AutoModelForCausalLM
class SliceCache:
"""Fixed-size cache updated by slice assignment, which converts to an in-place state write.
In practice, subclass the Cache class of your Transformers version so attention calls update()."""
def __init__(self, shape, dtype=torch.float16):
self.k = torch.zeros(shape, dtype=dtype)
self.v = torch.zeros(shape, dtype=dtype)
def update(self, layer, k_new, v_new):
b, e = self.begin, self.end # set by forward() before each call
self.k[layer, :, :, b:e, :] = k_new
self.v[layer, :, :, b:e, :] = v_new
return self.k[layer, :, :, :e, :], self.v[layer, :, :, :e, :]
class StatefulLM(torch.nn.Module):
def __init__(self, path, context=2048):
super().__init__()
self.model = AutoModelForCausalLM.from_pretrained(path, torch_dtype=torch.float16)
cfg = self.model.config
head_dim = cfg.hidden_size // cfg.num_attention_heads
self.shape = (cfg.num_hidden_layers, 1, cfg.num_key_value_heads, context, head_dim)
self.cache = SliceCache(self.shape)
self.register_buffer("keyCache", self.cache.k) # becomes Core ML state
self.register_buffer("valueCache", self.cache.v)
@torch.no_grad()
def forward(self, input_ids, causal_mask):
# causal_mask is (1, 1, query_len, end); end - query_len tokens are already cached
self.cache.end = causal_mask.shape[-1]
self.cache.begin = self.cache.end - input_ids.shape[-1]
return self.model(input_ids=input_ids, attention_mask=causal_mask,
past_key_values=self.cache, use_cache=True).logitsThe write position comes from tensor shapes, not a Python constant, so the traced graph stays valid for any length. The cache is allocated for the maximum context up front, so its memory is paid in full even for a short chat. Check that the traced graph contains the slice writes before converting.
Step 2: convert with state and flexible shapes
Conversion declares three things: the inputs, with a ct.RangeDim for the query length and the total length so one model serves both the prompt and single-token steps; the outputs; and the states, each wrapping a tensor type with the exact buffer shape and the buffer name. States require the mlprogram format and a minimum deployment target of iOS 18 or macOS 15.
import numpy as np
import coremltools as ct
wrapper = StatefulLM("meta-llama/Llama-3.1-8B-Instruct", context=2048).eval()
example = (torch.zeros((1, 2), dtype=torch.int32), torch.zeros((1, 1, 2, 5), dtype=torch.float16))
traced = torch.jit.trace(wrapper, example)
q = ct.RangeDim(lower_bound=1, upper_bound=2048, default=1)
k = ct.RangeDim(lower_bound=1, upper_bound=2048, default=1)
mlmodel = ct.convert(
traced,
inputs=[ct.TensorType(shape=(1, q), dtype=np.int32, name="inputIds"),
ct.TensorType(shape=(1, 1, q, k), dtype=np.float16, name="causalMask")],
outputs=[ct.TensorType(dtype=np.float16, name="logits")],
states=[ct.StateType(wrapped_type=ct.TensorType(shape=wrapper.shape, dtype=np.float16), name="keyCache"),
ct.StateType(wrapped_type=ct.TensorType(shape=wrapper.shape, dtype=np.float16), name="valueCache")],
minimum_deployment_target=ct.target.macOS15,
)
mlmodel.save("llm_fp16.mlpackage")Targeting macOS 15 or iOS 18 also lets the converter emit a fused scaled dot-product attention operation instead of separate matrix multiplies and a softmax. Before compressing anything, verify the float16 model: run the same prompt through PyTorch and Core ML and compare the logits of the first few steps. A mismatch here is a conversion bug, and every later step would hide it.
Step 3: compress the weights
An 8B model in float16 does not fit comfortably on most phones and saturates memory bandwidth on a Mac. Core ML Tools offers post-training weight compression on the converted model. Linear quantization stores integers plus a scale; with granularity="per_block" each block of consecutive weights in a row gets its own scale, which keeps outliers from ruining a whole channel. Palettization instead replaces weights by indices into a small lookup table of learned centroids.
import coremltools.optimize as cto
# int4, symmetric, one scale per 32 weights (the setting Apple used for Llama 3.1 8B)
lin = cto.coreml.OpLinearQuantizerConfig(mode="linear_symmetric", dtype="int4",
granularity="per_block", block_size=32)
int4_model = cto.coreml.linear_quantize_weights(
mlmodel, config=cto.coreml.OptimizationConfig(global_config=lin))
int4_model.save("llm_int4.mlpackage")
# alternative: 4-bit palettization, one 16-entry table per group of 16 output channels
pal = cto.coreml.OpPalettizerConfig(nbits=4, mode="kmeans",
granularity="per_grouped_channel", group_size=16)
lut_model = cto.coreml.palettize_weights(
mlmodel, config=cto.coreml.OptimizationConfig(global_config=pal))Both grouped forms need the iOS 18 / macOS 15 target. Compression here is weights only: activations and the KV cache stay in float16. Evaluate held-out prompts against the float16 model after compressing. For the theory behind the two schemes, see INT4 quantization and weight clustering.
Step 4: the generation loop
Generation has two phases. Prefill sends the whole prompt in one call, with a causal mask of shape prompt length by prompt length. Decode sends one token per call, with a mask of one row by the total length so far. The state object carries the cache between calls, so the app never touches it.
import numpy as np
model = ct.models.MLModel("llm_int4.mlpackage", compute_units=ct.ComputeUnit.CPU_AND_GPU)
def mask(q_len, total):
m = np.full((1, 1, q_len, total), -np.inf, dtype=np.float16)
for i in range(q_len): # row i may see positions up to total - q_len + i
m[0, 0, i, : total - q_len + i + 1] = 0
return m
def generate(ids, max_new=128, eos=128009):
state = model.make_state() # fresh cache for this conversation
out = model.predict({"inputIds": np.array([ids], dtype=np.int32),
"causalMask": mask(len(ids), len(ids))}, state=state)
total = len(ids)
for _ in range(max_new):
nxt = int(out["logits"][0, -1].argmax()) # greedy; replace with your sampler
if nxt == eos:
break
yield nxt
total += 1
out = model.predict({"inputIds": np.array([[nxt]], dtype=np.int32),
"causalMask": mask(1, total)}, state=state)On iOS 18 and macOS 15 the Swift side has the same shape: model.makeState() creates the state, and model.prediction(from: input, using: state) runs one step. Predictions that share a state must be serialised; give each concurrent conversation its own state. Stop before the total reaches the context size you converted with, because the cache cannot grow past its preallocated shape.
let state = model.makeState()
var output = try model.prediction(from: prefillInput, using: state)
while generated.count < maxNew {
let next = sampler.pick(from: output.featureValue(for: "logits")!.multiArrayValue!)
if next == eosId { break }
generated.append(next)
output = try model.prediction(from: decodeInput(next, total: promptCount + generated.count),
using: state)
}
Worked example: budgeting an 8B model
Llama 3.1 8B has about 8.03 billion parameters, 32 layers, 8 key-value heads and a head dimension of 128. Work out the three numbers that decide whether it runs well.
| Quantity | Arithmetic | Result |
|---|---|---|
| float16 weights | 8.03e9 x 2 bytes | about 16 GB |
| int4 per-block weights | 4 bits + one 16-bit scale per 32 weights = 4.5 bits | about 4.5 GB; Apple reports 4.2 GB |
| KV cache per token | 2 (K,V) x 32 layers x 8 heads x 128 x 2 bytes | 128 KiB |
| KV state at 2,048 context | 2,048 x 128 KiB | 256 MiB, allocated up front |
| Decode ceiling on 400 GB/s | 400 / 4.2 GB read per token | about 95 tokens/s |
Decoding one token reads every weight once, so memory bandwidth sets the ceiling. Apple's published measurements for this model on an M1 Max make the point concrete: about 0.19 tokens per second for the naive conversion, 16.26 after making the KV cache state, and 33.67 after int4 compression. The final figure is roughly a third of the bandwidth ceiling, which is a typical efficiency once attention, the cache reads and per-call overhead are added. Apple targeted the GPU for this run because the model is bandwidth-bound and the GPU offered the best mix of bandwidth and compute on that chip.
Apply the same arithmetic to a phone: a few gigabytes for your app and a fraction of a Mac's bandwidth point to a 1B to 3B model at 4 bits, not an 8B one.
Choosing compute units
The computeUnits setting on the model configuration takes .all, .cpuAndGPU, .cpuAndNeuralEngine or .cpuOnly. Core ML decides placement per operation within the units you allow, and an operation that a unit cannot run falls back to another.
- GPU is the safe default for decode on Macs: it handles flexible shapes and large matrix-vector products well and gets the full memory bandwidth.
- Neural Engine is the most power-efficient engine, which matters on a phone, but it favours fixed shapes and float16 layouts it supports natively. Teams that target it usually convert separate fixed-shape models for prefill and decode, or split the network into chunks, and then measure. Do not assume it is faster for a bandwidth-bound decoder.
- Measure placement with the Core ML performance report in Xcode; unexpected CPU fallbacks are a common cause of slow models.
How these engines compare with phone and laptop NPUs from other vendors is covered in edge NPUs.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| Conversion error about states | Deployment target below iOS 18 / macOS 15, or not mlprogram | Set the target and use the default mlprogram format |
| Output degrades after the first token | Cache slice indices or mask offsets are off by one | Compare per-step logits with PyTorch for the first ten steps |
| Every token is as slow as the prompt | Cache passed as input and output, not state | Declare StateType and verify the model spec lists states |
| Crash or garbage near the context limit | Total length exceeded the preallocated cache | Stop or truncate before the converted context size |
| Quality drop after compression | Per-tensor scales, or sensitive layers compressed | Use per_block or grouped palettization; leave embeddings or the head in float16 |
| Slower than expected on the Neural Engine | Flexible shapes or unsupported ops fall back | Inspect the performance report; try fixed-shape models or the GPU |
Trade-offs against other runtimes
Core ML is a system framework, so it adds nothing to your app's binary and handles placement across all three engines, including the Neural Engine. The cost is the conversion pipeline: every new architecture needs a wrapper, a trace and validation, and the context size is frozen. MLX and llama.cpp load new checkpoints in minutes and suit experimentation. A common pattern is to prototype with MLX on a Mac, then convert the chosen model to Core ML for the shipping app. For packaging, updates and fallbacks once a model ships, see deploying small models to the edge.
At WWDC 2026 Apple also introduced Core AI, a newer framework aimed at running local generative models on Apple silicon. Existing Core ML models keep working, but before starting a new project, check Core AI's documentation for its OS requirements and conversion tooling, which this article does not cover. The budgeting above, the stateful cache and the compressed weights apply either way.
What to do next
- Pick a model size from the budget table for your weakest target device, not your development Mac.
- Build the stateful wrapper and confirm with a trace that the cache writes are slice assignments.
- Convert in float16 with states and flexible shapes, and match PyTorch logits for the first ten tokens.
- Compress with int4 per-block quantization, then try grouped palettization, and evaluate both on held-out prompts.
- Benchmark tokens per second and time to first token with CPU and GPU and with the Neural Engine, using the performance report to find fallbacks.
- Ship with one state per conversation, a hard stop before the context limit, and the model loaded once at launch.