A Mac with an M-series chip has become a common place to run and fine-tune language models, and the reason is not raw compute. A workstation GPU still does far more arithmetic per second. The reason is memory: the CPU, the GPU and the Neural Engine share one pool of RAM, so a laptop with 128 GB can hold a model that would need several discrete GPUs, and it can do so without copying the weights anywhere.
This page explains the chip from the point of view of the software that runs on it: which part does what, which frameworks reach each part, how to predict tokens per second from the spec sheet, how to size a model and its KV cache against RAM, what training fits, and where the platform stops being the right tool. Figures come from Apple's published specifications and its MLX benchmark write-up.
The one idea: unified memory
On a discrete GPU system, the GPU has its own memory. A model is loaded into host RAM, then copied over PCIe into the GPU's memory, and anything the CPU needs from the GPU is copied back. The GPU's memory is fast but small and expensive, and it is the hard ceiling on model size.
Apple Silicon puts the CPU, GPU, Neural Engine and media engines on one system on a chip, attached to one pool of LPDDR memory in the same package. A buffer allocated by one engine can be read by another in place. For machine learning that has three consequences. First, capacity is the whole of RAM, less what the operating system and other applications use, so the question "does it fit" is answered by the memory configuration you bought. Second, there is no staging copy, so loading a model is mostly reading it from SSD, and CPU pre-processing can hand tensors to the GPU without a transfer. Third, every engine draws on the same bandwidth budget, so a busy CPU and a busy GPU slow each other down.
The hard limit: memory is soldered and chosen at purchase, so every decision below starts from the configuration you bought.
The compute units and who can program them
Four kinds of engine matter for machine learning, and they are not equally reachable.
| Engine | Good at | How software reaches it | Used by |
|---|---|---|---|
| CPU cores | control flow, tokenization, small ops, data loading | C, Swift, Python; Apple's Accelerate framework for vector and matrix routines | every framework, NumPy |
| GPU cores | large parallel matrix and elementwise work | Metal shaders and Metal Performance Shaders | MLX, PyTorch (mps device), llama.cpp, Core ML |
| Neural Accelerators (M5 family) | dense matrix multiplication inside each GPU core | Tensor APIs in Metal 4; MLX uses them on macOS 26.2 or later | MLX, frameworks built on Metal 4 |
| Neural Engine | low-power inference of fixed graphs | Core ML only; there is no public low-level API | Core ML models, Apple Intelligence |
The practical reading: for open-source LLM work, the GPU is the engine that matters, and on M5-family chips the Neural Accelerators inside it are what speed up prompt processing. The Neural Engine is efficient but reachable only through Core ML, which compiles a model ahead of time and decides for itself which layers run where. It suits a shipped app with a fixed model, not a research loop. All M5 variants include a 16-core Neural Engine; Apple says the M5 Pro and M5 Max give it a higher-bandwidth path to memory.
The numbers that decide performance
Three spec-sheet numbers predict almost everything: memory capacity, memory bandwidth and GPU compute. Apple publishes the first two clearly and describes the third mostly in relative terms, so plan with the first two.
| Chip | Max unified memory | Memory bandwidth |
|---|---|---|
| M4 | configuration dependent | 120 GB/s |
| M5 | 32 GB | 153 GB/s |
| M5 Pro | 64 GB | 307 GB/s |
| M4 Max | 128 GB | 546 GB/s |
| M5 Max | 128 GB | 614 GB/s |
| M3 Ultra | 512 GB | 819 GB/s |
The M5 Pro and M5 Max are built with what Apple calls the Fusion Architecture, two dies combined into one system on a chip. For comparison, a current data-centre GPU's HBM delivers several terabytes per second, so a Mac trades bandwidth for capacity per dollar. The HBM architecture page explains where that bandwidth comes from on the other side.
Worked example: predicting tokens per second
LLM inference has two phases with different bottlenecks, and the inference latency guide derives both. Prefill processes the whole prompt at once; it is a large matrix multiplication and is limited by compute. Decode produces one token at a time; each step must read every weight once, so it is limited by memory bandwidth. That gives a one-line ceiling for decode: tokens per second is at most bandwidth divided by the bytes read per token.
Take an 8-billion-parameter model quantized to 4 bits. Apple's MLX write-up measured Qwen3 8B at 4 bits as 5.61 GB. On an M5 at 153 GB/s the ceiling is 153 / 5.61, about 27 tokens per second. On an M5 Max at 614 GB/s it is about 109. Real engines reach a fraction of the ceiling, depending on the engine, the quantization format and the context length, because of KV-cache reads, kernel launch gaps and dequantization work. The same arithmetic shows why quantization matters so much here: the same model at bf16 (17.46 GB in Apple's table) has a ceiling one third as high.
Apple's own M5 versus M4 measurements match this model. Time to first token, which is prefill and compute-bound, improved by 3.33 to 4.06 times thanks to the Neural Accelerators. Generation speed, which is bandwidth-bound, improved by 1.19 to 1.27 times, close to the 28 percent bandwidth increase from 120 to 153 GB/s. If a benchmark claims a large decode speedup with no bandwidth change, look for a different model, quantization or batch size.
def decode_ceiling(bandwidth_gb_s, weight_gb, kv_gb_read_per_token=0.0):
# Upper bound: every weight (plus the live KV cache) is read once per token.
return bandwidth_gb_s / (weight_gb + kv_gb_read_per_token)
def kv_cache_gb(layers, kv_heads, head_dim, tokens, bytes_per_elem=2):
# K and V, per layer, per KV head, per token.
return 2 * layers * kv_heads * head_dim * tokens * bytes_per_elem / 1e9
# Llama-3-8B-like shape: 32 layers, 8 KV heads (grouped-query attention), head_dim 128.
kv = kv_cache_gb(32, 8, 128, tokens=32_768) # about 4.3 GB at fp16
print(decode_ceiling(153, 5.61)) # M5, short context: about 27 tok/s
print(decode_ceiling(153, 5.61, kv)) # M5, 32k context: about 15 tok/sThe KV cache is the part people forget. With the shape above, each token of context costs 128 KiB at fp16, so a 32k-token context adds about 4.3 GB of memory and, at full length, that much extra reading per generated token. On a 16 GB or 24 GB machine the model and its cache can fit at short context and fail at long context. Budget RAM as weights plus cache plus the operating system and your other applications, and leave headroom.
The software stack
MLX is Apple's open-source array framework built for this memory model. Arrays live in unified memory and are not bound to a device, so there are no host-to-device copies. Computation is lazy: operations build a graph that runs when you call mx.eval or read a value, and mx.compile fuses graphs into fewer kernels. The companion package mlx-lm loads, quantizes, runs and fine-tunes Hugging Face models from the command line.
pip install mlx mlx-lm
# Convert a Hugging Face checkpoint to MLX format with 4-bit quantization.
mlx_lm.convert --hf-path Qwen/Qwen3-8B -q
# Generate from the converted (or a pre-converted) model.
mlx_lm.generate --model mlx_model --prompt "Explain unified memory in one paragraph."
# Interactive chat.
mlx_lm.chat --model mlx_modelPyTorch reaches the GPU through the mps device, backed by Metal Performance Shaders. Most models run unchanged once moved to the device, but some operators are not implemented; setting PYTORCH_ENABLE_MPS_FALLBACK=1 runs those on the CPU instead of raising an error, which keeps code working and can quietly make it slow. float64 is not supported on the GPU.
llama.cpp and tools built on it run GGUF files with a Metal backend and are the most common way to serve quantized models locally; the GGUF runtime guide covers the format. Core ML is the route to the Neural Engine for apps that ship a fixed model.
A minimal MLX training loop shows the shape of the framework. Note the explicit mx.eval: without it the graph keeps growing and nothing is actually computed until later, which also makes naive timing meaningless.
import mlx.core as mx
import mlx.nn as nn
import mlx.optimizers as optim
model = nn.Sequential(nn.Linear(512, 1024), nn.GELU(), nn.Linear(1024, 10))
optimizer = optim.AdamW(learning_rate=3e-4)
def loss_fn(model, x, y):
return nn.losses.cross_entropy(model(x), y, reduction="mean")
step = nn.value_and_grad(model, loss_fn)
for x, y in batches(): # your data iterator yielding mx.array pairs
loss, grads = step(model, x, y)
optimizer.update(model, grads)
mx.eval(model.parameters(), optimizer.state, loss) # force the step to run
What training fits on a Mac
Inference needs roughly the weights plus the cache. Full training needs much more. With mixed-precision AdamW, a common budget is about 16 bytes per parameter: 2 for bf16 weights, 2 for gradients, 4 for an fp32 master copy and 8 for the two optimizer moments, before activations. An 8-billion-parameter model therefore needs about 128 GB for state alone, which does not fit on a 128 GB M5 Max once the operating system and activations are counted.
Parameter-efficient fine-tuning changes the picture. LoRA freezes the base model, which can itself be quantized, and trains small adapter matrices, so memory is close to the inference footprint plus the adapter's gradients and optimizer state plus activations. That is why LoRA and QLoRA fine-tuning of 7B to 14B models is a routine Mac workload, and mlx-lm includes a LoRA trainer for it. Full pre-training or full fine-tuning of anything large belongs on data-centre GPUs, where fast interconnects make scaling out practical.
Throughput is the other limit. Prefill-heavy work such as training on long sequences is compute-bound, and a Mac GPU has much less compute than a data-centre part; the M5 Neural Accelerators narrow the gap for matrix multiplication but do not close it. A good rule: use the Mac to iterate on data, prompts and small adapters, and move to rented GPUs when a run would take more than overnight.
Failure modes
- Swapping. When model plus cache plus everything else exceeds RAM, macOS compresses and swaps memory and throughput collapses by an order of magnitude rather than failing cleanly. Watch memory pressure in Activity Monitor during a run, not just free memory before it.
- The GPU working-set limit. macOS does not let the GPU wire all of RAM; Metal reports a recommended maximum working set per device (
recommendedMaxWorkingSetSize). A model that fits in RAM can still exceed it. Check the value on your machine before buying or sizing to the last gigabyte. - Silent CPU fallback. With the PyTorch MPS fallback enabled, an unsupported operator in the hot loop moves data to the CPU every step. Profile once with the fallback disabled to find them.
- Lazy evaluation and bad benchmarks. In MLX, timing code without
mx.evalmeasures graph construction. Measure prefill and decode separately and report the context length. - Thermal limits and shared bandwidth. A fanless laptop sustains less than it peaks, and a heavy CPU pipeline running beside inference takes bandwidth from the GPU.
- Quantization quality. Aggressive 2- or 3-bit models fit where 4-bit ones do not, but quality drops unevenly by task; the quantization guide explains how to measure it.
Trade-offs against a discrete GPU
The choice is capacity and simplicity against bandwidth, compute and ecosystem. A Mac with 64 to 128 GB holds models that would need several consumer GPUs, is silent, and needs no driver or CUDA management. It generates more slowly per dollar than a data-centre GPU at batch size one, processes prompts much more slowly, and cannot run CUDA-only code such as many custom attention kernels and training libraries. It also cannot be upgraded. Serving many concurrent users is a data-centre job, because batching turns decode from bandwidth-bound to compute-bound and that is exactly where discrete GPUs lead. For a single developer, a small team's private assistant, or offline evaluation, the Mac is often the cheapest capable option. The GPU memory hierarchy page is the companion for reasoning about the discrete side.
What to do next
- Decide the largest model and context you need, then compute weights plus KV cache with the functions above and add at least 8 GB for the system before choosing a memory configuration.
- Divide the chip's bandwidth by the model's size to set your decode expectation, and treat anything much above that ceiling as a measurement error.
- Install
mlxandmlx-lm, convert one model at 4 bits, and measure time to first token and tokens per second at short and long context. - If you use PyTorch, run once with
PYTORCH_ENABLE_MPS_FALLBACKunset to list unsupported operators in your model. - Fine-tune with LoRA locally on a small dataset; move full fine-tuning and long runs to rented GPUs.
- On M5-family machines, update to macOS 26.2 or later so MLX can use the Neural Accelerators, and re-measure prefill.