Ollama makes running a local model a one-line command, which is exactly why its GPU behaviour surprises people. The same model can produce 80 tokens per second on Monday and 6 on Tuesday because someone raised the context length or a second model was loaded. None of that is visible from the chat window; it is visible from the GPU, which is the view this article takes.

The general serving story, including Modelfiles, the HTTP API and how to run Ollama for a team, is covered in serving models locally with Ollama. Here the subject is narrower and lower level: what Ollama puts in VRAM, how it decides how many layers go to the GPU, what happens across several GPUs, how KV cache precision changes the budget, and how to measure decode speed and predict it from memory bandwidth. Environment variables and API fields named here were checked against Ollama's FAQ, API reference and configuration source on 2026-10-04; defaults have changed between releases, so check your version.

Server, runners and layers

Ollama is two kinds of process. ollama serve is the long-running server: it exposes the API on port 11434, keeps a request queue, and runs a scheduler that decides which models are resident and on which devices. When a request needs a model that is not loaded, the scheduler starts a runner, a separate process that owns the model's memory and does the actual matrix maths using GGML backends for CUDA, ROCm, Metal or Vulkan, falling back to CPU code.

The model itself is a GGUF file: a set of named tensors, mostly quantized weight matrices, grouped into transformer layers (blocks). That grouping matters because offload is decided per layer. A layer is either on a GPU or on the CPU; its attention and feed-forward weights travel together, and so does its slice of the KV cache. Formats and quantization types are explained in the GGUF quantization guide.

So the unit of failure is not "the model does not fit" but "k layers do not fit": Ollama degrades quietly rather than refusing, and anything else holding VRAM changes how many layers fit at the next load.

The VRAM budget

Before loading, the scheduler estimates four things and compares them with free VRAM on each visible GPU:

  • Weights. Roughly the GGUF file size. An 8B model at Q4_K_M is about 4.9 GB; a 70B model at the same type is around 40 GB.
  • KV cache. Two tensors (K and V) per layer, per KV head, per head dimension, per token of context, per parallel slot.
  • Compute graph. Scratch buffers for intermediate activations during a forward pass; they grow with batch size and context.
  • Reserve. Whatever you hold back with OLLAMA_GPU_OVERHEAD, which takes a byte count per GPU.

The KV cache is the part people forget. Per token it is 2 x layers x kv_heads x head_dim x bytes_per_element. For a Llama 3.1 8B shape (32 layers, 8 KV heads with grouped-query attention, head dimension 128) at f16 that is 2 x 32 x 8 x 128 x 2 = 131,072 bytes, 128 KiB per token. An 8,192-token context costs 1 GiB; 32,768 tokens cost 4 GiB, close to the size of the weights. Parallelism multiplies it: Ollama serves OLLAMA_NUM_PARALLEL requests in one runner by giving each its own context slot, so four slots at 8K need the same KV memory as one 32K context. The general sizing method is worked through in sizing the KV cache.

The default context length is version-dependent: the FAQ documents 4096 tokens, while recent configuration code describes a default that scales with available VRAM. Do not rely on either; set num_ctx explicitly per model or request, so the memory estimate is something you chose.

Where an Ollama model lives: the VRAM budget and the layer splitollama servescheduler, API, queueRunner processone per loaded modelloadGPU 0 VRAMWeights: layers 0..N-kquantized GGUF tensorsKV cachenum_ctx x parallel slotsCompute graphscratch buffersSystem RAMWeights: last k layersrun on CPU coresTheir KV sliceactivationsollama ps PROCESSOR100% GPU, or e.g. 48%/52% CPU/GPUPer generated token, every layer's weights are read once: GPU layers at hundreds of GB/s,CPU layers at tens of GB/s. The slowest share sets the token rate.ReserveOLLAMA_GPU_OVERHEAD,desktop, other processes
Figure 1. Ollama's scheduler fits weights, KV cache and compute buffers into VRAM. Layers that do not fit run on the CPU from system RAM, and the bandwidth gap between the two makes even a small CPU share dominate decode time.

Layer offload and the CPU/GPU split

If the whole budget fits, the model loads 100% GPU. If it does not, Ollama computes how many layers do fit and offloads only those. The remainder run on the CPU from system RAM. ollama ps reports the result in its PROCESSOR column, for example 48%/52% CPU/GPU, and /api/ps returns both size and size_vram for each loaded model, so you can alert when they differ.

The num_gpu option (in a Modelfile or a request's options) overrides this; despite its name it is generally treated as the number of layers to offload, though the API reference does not spell that out, so verify on your version. Setting it higher than what fits trades a graceful partial offload for an out-of-memory failure; setting it to 0 forces a CPU-only baseline.

Why does a small CPU share hurt so much? Decode, generating one token at a time, is memory-bandwidth bound: for each token every weight is read once and used for very little arithmetic (see the maths of decode). Time per token is therefore roughly the bytes each device must read divided by that device's bandwidth, summed across devices because layers run in sequence:

# decode time per token, ignoring compute and transfer overhead
t_token ~= bytes_on_gpu / bw_gpu + bytes_on_cpu / bw_cpu
tokens_per_s ~= 1 / t_token

Take a 40 GB 70B model on a 24 GB card. Assume about 1,000 GB/s for a current high-end consumer GPU and about 80 GB/s for dual-channel DDR5; both are approximate, so measure your own. With 22 GB on the GPU and 18 GB on the CPU, t = 22/1000 + 18/80 = 0.022 + 0.225 = 0.247 s, about 4 tokens per second. The GPU portion takes 9% of the time. The CPU holds 45% of the weights and costs 91% of the time. That is the shape of every partial offload: the result is close to CPU speed, not halfway between.

Several GPUs: placement, spreading and residency

With more than one GPU the documented rule is simple: if the model fits entirely on a single GPU, Ollama loads it on that one; otherwise it spreads it across all available GPUs. Spreading is a layer split. GPU 0 holds the first block of layers, GPU 1 the next, and activations pass from one to the next for every token. This is pipeline-style placement, not tensor parallelism: the GPUs take turns on a single token, so two cards give you the capacity of both but per-token speed close to that of one card reading the same total bytes. Engines built for tensor parallelism, such as vLLM on the GPU, split each matrix instead and can use both cards at once.

The single-GPU preference is usually right because it avoids the hop between devices and leaves the other GPU free for a second model. Two settings change it. OLLAMA_SCHED_SPREAD=1 makes the scheduler always spread a model across all GPUs, useful when you want large contexts on a model that only barely fits on one card. Device visibility variables choose which GPUs Ollama sees at all: CUDA_VISIBLE_DEVICES for NVIDIA, ROCR_VISIBLE_DEVICES for AMD, GGML_VK_VISIBLE_DEVICES for Vulkan. The main_gpu request option exists for choosing a primary device, but placement is mostly the scheduler's job.

Residency is per GPU as well. OLLAMA_MAX_LOADED_MODELS caps how many models stay loaded; the FAQ gives the default as three per GPU, or three for CPU inference. When a new model needs room, the scheduler evicts idle ones. Models stay resident for OLLAMA_KEEP_ALIVE (5 minutes by default) after their last request, and requests beyond capacity wait in a queue of OLLAMA_MAX_QUEUE (512 by default) before being rejected.

KV cache precision and flash attention

Because KV memory grows linearly with context and slots, its precision is the cheapest lever after the weights. OLLAMA_KV_CACHE_TYPE accepts f16 (the default), q8_0 (about half the memory of f16) and q4_0 (about a quarter). In our 8B example, 32K of context drops from 4 GiB to roughly 2 GiB at q8_0, often the difference between a full GPU load and a split.

The costs are quality and scope. q8_0 is usually hard to distinguish from f16 on chat workloads; q4_0 can measurably hurt long-context recall, and models with few KV heads are more sensitive because each cached value carries more information. The setting is server-wide, not per model. Quantized KV cache works together with flash attention, OLLAMA_FLASH_ATTENTION=1; flash attention computes attention in tiles without materialising the full attention matrix, which also shrinks the compute graph for long prompts. Whether it is on by default has varied by version and by hardware, so set it explicitly and confirm in the server log.

Measuring prefill and decode

Every non-streamed response from /api/generate or /api/chat, and the final streamed chunk, carries timing fields in nanoseconds: load_duration, prompt_eval_count, prompt_eval_duration, eval_count and eval_duration. Prompt processing (prefill) is compute bound and runs many tokens at once; decode is bandwidth bound. Measure them separately or you will average two different regimes. This script does both and checks placement:

import json, urllib.request

BASE = "http://127.0.0.1:11434"

def post(path, body):
    req = urllib.request.Request(BASE + path, json.dumps(body).encode(),
                                 {"Content-Type": "application/json"})
    with urllib.request.urlopen(req) as r:
        return json.load(r)

def bench(model, ctx, prompt="Explain how a B-tree splits a node.", n=256):
    r = post("/api/generate", {"model": model, "prompt": prompt, "stream": False,
             "options": {"num_ctx": ctx, "num_predict": n, "temperature": 0}})
    prefill = r["prompt_eval_count"] / (r["prompt_eval_duration"] / 1e9)
    decode = r["eval_count"] / (r["eval_duration"] / 1e9)
    return r["load_duration"] / 1e9, prefill, decode

for ctx in (4096, 16384, 32768):
    load_s, pf, dec = bench("llama3.1:8b", ctx)
    with urllib.request.urlopen(BASE + "/api/ps") as r:
        m = next(x for x in json.load(r)["models"] if x["name"] == "llama3.1:8b")
    share = m["size_vram"] / m["size"]
    print(f"ctx={ctx:6d} load={load_s:5.1f}s prefill={pf:7.0f} tok/s "
          f"decode={dec:5.1f} tok/s vram_share={share:.0%}")

Each new context size reloads the model, so a non-zero load time is expected here; it only matters on repeated identical requests. Watch nvidia-smi or rocm-smi alongside. The pattern to look for is the context at which vram_share drops below 100%: decode speed falls off a cliff at exactly that point, and that is your real context ceiling for this GPU, model and slot count.

Worked example: one 24 GB GPU for a team

A team has one 24 GB GPU. They want Llama 3.1 8B at Q4_K_M for an internal assistant with four concurrent users and 16K of context each, and occasionally a 70B model for hard questions.

The 8B service. Weights 4.9 GB. KV at f16: 128 KiB x 16,384 tokens x 4 slots = 8 GiB. Add a compute graph of perhaps 1-2 GB and a 1 GB reserve for the desktop: about 15-16 GB. It fits, with OLLAMA_NUM_PARALLEL=4 and num_ctx 16384 set in the Modelfile. Switching the KV cache to q8_0 brings it near 11 GB, leaving room for an embedding model to stay resident beside it.

The 70B request. 40 GB of weights cannot fit, and while the 8B model is resident even fewer layers can go to the GPU. The scheduler will evict the idle 8B model or load the 70B with a large CPU share; by the arithmetic above, expect low single-digit tokens per second. The options are a second GPU (capacity, not speed) or routing hard questions to a remote endpoint. The team chose routing, kept the 8B model pinned with a long keep_alive, and set an alert when size_vram is less than size for the assistant model.

Containers and drivers

In containers the GPU must be passed through explicitly. The NVIDIA path is the official image with the NVIDIA Container Toolkit installed on the host:

docker run -d --gpus=all -v ollama:/root/.ollama -p 11434:11434 \
  -e OLLAMA_NUM_PARALLEL=4 -e OLLAMA_KV_CACHE_TYPE=q8_0 -e OLLAMA_FLASH_ATTENTION=1 \
  -e OLLAMA_GPU_OVERHEAD=1073741824 --name ollama ollama/ollama

AMD uses the ollama/ollama:rocm image with /dev/kfd and /dev/dri mapped into the container. In both cases, read the start-up log: it lists the GPUs discovered and, at each load, how many layers were offloaded. No compatible GPU means a driver or pass-through problem, and every request will silently run on the CPU.

Failure modes

  • Silent partial offload. Raising context or slots pushes layers to the CPU; decode drops several-fold with no error. Alert on size_vram < size.
  • VRAM stolen after load. Another process allocates memory after Ollama's estimate; the next long prompt fails with out-of-memory. Reserve headroom with OLLAMA_GPU_OVERHEAD or dedicate the GPU.
  • Thrashing. Two models that cannot co-reside alternate requests, and every switch pays a full load. Watch load_duration on warm traffic; it should be near zero.
  • Forced layers. A hand-set num_gpu that fitted at 4K context fails at 16K. Prefer letting the scheduler decide, and fix the context instead.
  • Expecting multi-GPU speed-ups. A layer split adds capacity; per-token latency does not halve.

Trade-offs

Ollama's scheduler optimises for convenience on one machine: automatic placement, graceful degradation and many models sharing a GPU. That suits a workstation or small team server, not predictable throughput under heavy load, where dedicated engines with continuous batching and tensor parallelism win. KV quantization buys context for a small, workload-dependent quality cost. Spreading across GPUs buys capacity at the cost of an inter-device hop per token. A partial CPU offload buys the ability to run something at all, at close to CPU speed.

What to do next

  1. Run ollama ps now and confirm every model you care about shows 100% GPU.
  2. Set num_ctx explicitly per model and compute its KV cache with the formula above, multiplied by your parallel slots.
  3. Run the benchmark script at three context lengths and record where vram_share drops below 100%.
  4. Decide on KV precision: try q8_0 with flash attention and compare answers on your own long-context tasks.
  5. Reserve VRAM for other processes with OLLAMA_GPU_OVERHEAD, or dedicate the GPU to Ollama.
  6. On multi-GPU hosts, decide between one model per GPU and OLLAMA_SCHED_SPREAD based on capacity needs, not speed hopes.
  7. Add an alert on size_vram versus size and on non-zero warm load_duration.
Key takeaway: Ollama's GPU behaviour is a memory budget: weights plus a KV cache that scales with context and parallel slots, plus compute buffers and a reserve. Whatever does not fit runs on the CPU, and because decode is bandwidth bound, a small CPU share sets the token rate. Fix the context, size the KV cache, verify 100% GPU, and treat extra GPUs as capacity rather than speed.