A Mac with an M-series chip has become a common place to run and fine-tune language models, and the reason is not raw compute. A workstation GPU still does far more arithmetic per second. The reason is memory: the CPU, the GPU and the Neural Engine share one pool of RAM, so a laptop with 128 GB can hold a model that would need several discrete GPUs, and it can do so without copying the weights anywhere.

This page explains the chip from the point of view of the software that runs on it: which part does what, which frameworks reach each part, how to predict tokens per second from the spec sheet, how to size a model and its KV cache against RAM, what training fits, and where the platform stops being the right tool. Figures come from Apple's published specifications and its MLX benchmark write-up.

Advertisement

The one idea: unified memory

On a discrete GPU system, the GPU has its own memory. A model is loaded into host RAM, then copied over PCIe into the GPU's memory, and anything the CPU needs from the GPU is copied back. The GPU's memory is fast but small and expensive, and it is the hard ceiling on model size.

Apple Silicon puts the CPU, GPU, Neural Engine and media engines on one system on a chip, attached to one pool of LPDDR memory in the same package. A buffer allocated by one engine can be read by another in place. For machine learning that has three consequences. First, capacity is the whole of RAM, less what the operating system and other applications use, so the question "does it fit" is answered by the memory configuration you bought. Second, there is no staging copy, so loading a model is mostly reading it from SSD, and CPU pre-processing can hand tensors to the GPU without a transfer. Third, every engine draws on the same bandwidth budget, so a busy CPU and a busy GPU slow each other down.

The hard limit: memory is soldered and chosen at purchase, so every decision below starts from the configuration you bought.

One package, one memory pool: every engine reads the same weights in placeSystem on a chipCPU clustersAccelerate, NumPy, tokenizersGPU coresMetal, MPS, MLX, llama.cppNeural EngineCore ML onlyMedia enginesencode, decodeNeural Accelerator per core (M5)System level cache and memory controllersshared by all engines; bandwidth is the budget every engine draws onUnified LPDDR memoryM5 153 GB/s, M5 Pro 307 GB/s, M5 Max 614 GB/sDiscrete GPU boxweights copied over PCIe into separate VRAMApple Siliconno copy: capacity is RAM, speed is bandwidthCapacity decides what fits; bandwidth decides decode speed; GPU compute decides prompt processing.
The engines on an Apple Silicon package and the single memory they share. The Neural Accelerators sit inside each GPU core on the M5 family.

The compute units and who can program them

Four kinds of engine matter for machine learning, and they are not equally reachable.

EngineGood atHow software reaches itUsed by
CPU corescontrol flow, tokenization, small ops, data loadingC, Swift, Python; Apple's Accelerate framework for vector and matrix routinesevery framework, NumPy
GPU coreslarge parallel matrix and elementwise workMetal shaders and Metal Performance ShadersMLX, PyTorch (mps device), llama.cpp, Core ML
Neural Accelerators (M5 family)dense matrix multiplication inside each GPU coreTensor APIs in Metal 4; MLX uses them on macOS 26.2 or laterMLX, frameworks built on Metal 4
Neural Enginelow-power inference of fixed graphsCore ML only; there is no public low-level APICore ML models, Apple Intelligence

The practical reading: for open-source LLM work, the GPU is the engine that matters, and on M5-family chips the Neural Accelerators inside it are what speed up prompt processing. The Neural Engine is efficient but reachable only through Core ML, which compiles a model ahead of time and decides for itself which layers run where. It suits a shipped app with a fixed model, not a research loop. All M5 variants include a 16-core Neural Engine; Apple says the M5 Pro and M5 Max give it a higher-bandwidth path to memory.

Advertisement

The numbers that decide performance

Three spec-sheet numbers predict almost everything: memory capacity, memory bandwidth and GPU compute. Apple publishes the first two clearly and describes the third mostly in relative terms, so plan with the first two.

ChipMax unified memoryMemory bandwidth
M4configuration dependent120 GB/s
M532 GB153 GB/s
M5 Pro64 GB307 GB/s
M4 Max128 GB546 GB/s
M5 Max128 GB614 GB/s
M3 Ultra512 GB819 GB/s

The M5 Pro and M5 Max are built with what Apple calls the Fusion Architecture, two dies combined into one system on a chip. For comparison, a current data-centre GPU's HBM delivers several terabytes per second, so a Mac trades bandwidth for capacity per dollar. The HBM architecture page explains where that bandwidth comes from on the other side.

Worked example: predicting tokens per second

LLM inference has two phases with different bottlenecks, and the inference latency guide derives both. Prefill processes the whole prompt at once; it is a large matrix multiplication and is limited by compute. Decode produces one token at a time; each step must read every weight once, so it is limited by memory bandwidth. That gives a one-line ceiling for decode: tokens per second is at most bandwidth divided by the bytes read per token.

Take an 8-billion-parameter model quantized to 4 bits. Apple's MLX write-up measured Qwen3 8B at 4 bits as 5.61 GB. On an M5 at 153 GB/s the ceiling is 153 / 5.61, about 27 tokens per second. On an M5 Max at 614 GB/s it is about 109. Real engines reach a fraction of the ceiling, depending on the engine, the quantization format and the context length, because of KV-cache reads, kernel launch gaps and dequantization work. The same arithmetic shows why quantization matters so much here: the same model at bf16 (17.46 GB in Apple's table) has a ceiling one third as high.

Apple's own M5 versus M4 measurements match this model. Time to first token, which is prefill and compute-bound, improved by 3.33 to 4.06 times thanks to the Neural Accelerators. Generation speed, which is bandwidth-bound, improved by 1.19 to 1.27 times, close to the 28 percent bandwidth increase from 120 to 153 GB/s. If a benchmark claims a large decode speedup with no bandwidth change, look for a different model, quantization or batch size.

def decode_ceiling(bandwidth_gb_s, weight_gb, kv_gb_read_per_token=0.0):
    # Upper bound: every weight (plus the live KV cache) is read once per token.
    return bandwidth_gb_s / (weight_gb + kv_gb_read_per_token)

def kv_cache_gb(layers, kv_heads, head_dim, tokens, bytes_per_elem=2):
    # K and V, per layer, per KV head, per token.
    return 2 * layers * kv_heads * head_dim * tokens * bytes_per_elem / 1e9

# Llama-3-8B-like shape: 32 layers, 8 KV heads (grouped-query attention), head_dim 128.
kv = kv_cache_gb(32, 8, 128, tokens=32_768)     # about 4.3 GB at fp16
print(decode_ceiling(153, 5.61))                # M5, short context: about 27 tok/s
print(decode_ceiling(153, 5.61, kv))            # M5, 32k context: about 15 tok/s

The KV cache is the part people forget. With the shape above, each token of context costs 128 KiB at fp16, so a 32k-token context adds about 4.3 GB of memory and, at full length, that much extra reading per generated token. On a 16 GB or 24 GB machine the model and its cache can fit at short context and fail at long context. Budget RAM as weights plus cache plus the operating system and your other applications, and leave headroom.

The software stack

MLX is Apple's open-source array framework built for this memory model. Arrays live in unified memory and are not bound to a device, so there are no host-to-device copies. Computation is lazy: operations build a graph that runs when you call mx.eval or read a value, and mx.compile fuses graphs into fewer kernels. The companion package mlx-lm loads, quantizes, runs and fine-tunes Hugging Face models from the command line.

pip install mlx mlx-lm
# Convert a Hugging Face checkpoint to MLX format with 4-bit quantization.
mlx_lm.convert --hf-path Qwen/Qwen3-8B -q
# Generate from the converted (or a pre-converted) model.
mlx_lm.generate --model mlx_model --prompt "Explain unified memory in one paragraph."
# Interactive chat.
mlx_lm.chat --model mlx_model

PyTorch reaches the GPU through the mps device, backed by Metal Performance Shaders. Most models run unchanged once moved to the device, but some operators are not implemented; setting PYTORCH_ENABLE_MPS_FALLBACK=1 runs those on the CPU instead of raising an error, which keeps code working and can quietly make it slow. float64 is not supported on the GPU.

llama.cpp and tools built on it run GGUF files with a Metal backend and are the most common way to serve quantized models locally; the GGUF runtime guide covers the format. Core ML is the route to the Neural Engine for apps that ship a fixed model.

A minimal MLX training loop shows the shape of the framework. Note the explicit mx.eval: without it the graph keeps growing and nothing is actually computed until later, which also makes naive timing meaningless.

import mlx.core as mx
import mlx.nn as nn
import mlx.optimizers as optim

model = nn.Sequential(nn.Linear(512, 1024), nn.GELU(), nn.Linear(1024, 10))
optimizer = optim.AdamW(learning_rate=3e-4)

def loss_fn(model, x, y):
    return nn.losses.cross_entropy(model(x), y, reduction="mean")

step = nn.value_and_grad(model, loss_fn)

for x, y in batches():                    # your data iterator yielding mx.array pairs
    loss, grads = step(model, x, y)
    optimizer.update(model, grads)
    mx.eval(model.parameters(), optimizer.state, loss)   # force the step to run

What training fits on a Mac

Inference needs roughly the weights plus the cache. Full training needs much more. With mixed-precision AdamW, a common budget is about 16 bytes per parameter: 2 for bf16 weights, 2 for gradients, 4 for an fp32 master copy and 8 for the two optimizer moments, before activations. An 8-billion-parameter model therefore needs about 128 GB for state alone, which does not fit on a 128 GB M5 Max once the operating system and activations are counted.

Parameter-efficient fine-tuning changes the picture. LoRA freezes the base model, which can itself be quantized, and trains small adapter matrices, so memory is close to the inference footprint plus the adapter's gradients and optimizer state plus activations. That is why LoRA and QLoRA fine-tuning of 7B to 14B models is a routine Mac workload, and mlx-lm includes a LoRA trainer for it. Full pre-training or full fine-tuning of anything large belongs on data-centre GPUs, where fast interconnects make scaling out practical.

Throughput is the other limit. Prefill-heavy work such as training on long sequences is compute-bound, and a Mac GPU has much less compute than a data-centre part; the M5 Neural Accelerators narrow the gap for matrix multiplication but do not close it. A good rule: use the Mac to iterate on data, prompts and small adapters, and move to rented GPUs when a run would take more than overnight.

Failure modes

  • Swapping. When model plus cache plus everything else exceeds RAM, macOS compresses and swaps memory and throughput collapses by an order of magnitude rather than failing cleanly. Watch memory pressure in Activity Monitor during a run, not just free memory before it.
  • The GPU working-set limit. macOS does not let the GPU wire all of RAM; Metal reports a recommended maximum working set per device (recommendedMaxWorkingSetSize). A model that fits in RAM can still exceed it. Check the value on your machine before buying or sizing to the last gigabyte.
  • Silent CPU fallback. With the PyTorch MPS fallback enabled, an unsupported operator in the hot loop moves data to the CPU every step. Profile once with the fallback disabled to find them.
  • Lazy evaluation and bad benchmarks. In MLX, timing code without mx.eval measures graph construction. Measure prefill and decode separately and report the context length.
  • Thermal limits and shared bandwidth. A fanless laptop sustains less than it peaks, and a heavy CPU pipeline running beside inference takes bandwidth from the GPU.
  • Quantization quality. Aggressive 2- or 3-bit models fit where 4-bit ones do not, but quality drops unevenly by task; the quantization guide explains how to measure it.

Trade-offs against a discrete GPU

The choice is capacity and simplicity against bandwidth, compute and ecosystem. A Mac with 64 to 128 GB holds models that would need several consumer GPUs, is silent, and needs no driver or CUDA management. It generates more slowly per dollar than a data-centre GPU at batch size one, processes prompts much more slowly, and cannot run CUDA-only code such as many custom attention kernels and training libraries. It also cannot be upgraded. Serving many concurrent users is a data-centre job, because batching turns decode from bandwidth-bound to compute-bound and that is exactly where discrete GPUs lead. For a single developer, a small team's private assistant, or offline evaluation, the Mac is often the cheapest capable option. The GPU memory hierarchy page is the companion for reasoning about the discrete side.

What to do next

  1. Decide the largest model and context you need, then compute weights plus KV cache with the functions above and add at least 8 GB for the system before choosing a memory configuration.
  2. Divide the chip's bandwidth by the model's size to set your decode expectation, and treat anything much above that ceiling as a measurement error.
  3. Install mlx and mlx-lm, convert one model at 4 bits, and measure time to first token and tokens per second at short and long context.
  4. If you use PyTorch, run once with PYTORCH_ENABLE_MPS_FALLBACK unset to list unsupported operators in your model.
  5. Fine-tune with LoRA locally on a small dataset; move full fine-tuning and long runs to rented GPUs.
  6. On M5-family machines, update to macOS 26.2 or later so MLX can use the Neural Accelerators, and re-measure prefill.
Key takeaway: Apple Silicon is a capacity machine. One pool of unified memory is shared by the CPU, the GPU and the Neural Engine, so a model fits if it fits in RAM and loads without copies, while decode speed is set by memory bandwidth: 153 GB/s on the M5, 307 on the M5 Pro and 614 on the M5 Max. Prefill is compute-bound, and the M5 Neural Accelerators inside each GPU core speed it up three to four times over the M4 in Apple's MLX tests. Use MLX or llama.cpp for local inference, LoRA for fine-tuning, budget RAM for the KV cache, and send full training and high-concurrency serving to data-centre GPUs.