Ollama is the shortest path from a laptop to a running language model: one command pulls a quantized model and another starts a chat. That convenience hides real machinery. Behind the CLI is an HTTP server that schedules requests, loads models into separate runner processes, decides how many layers fit on the GPU, and reserves a KV cache whose size you probably did not choose on purpose. When Ollama is slow, truncates a long prompt or runs out of memory, the cause is almost always one of those hidden decisions.
This page explains Ollama from the request path up, so you can predict its behaviour instead of discovering it. It covers memory and context budgeting with a worked example, the concurrency settings and what they cost, Modelfiles, the API including structured output, and what changes when a team shares one Ollama server. The inference engine underneath shares its lineage with llama.cpp, which is covered in the llama.cpp deep dive; this page stays at the level Ollama exposes.
What Ollama is made of
Ollama has four parts. The client is the ollama CLI or any HTTP caller. The server, started by ollama serve or the desktop app, listens on 127.0.0.1 port 11434 by default and owns all state. The model store holds models as content-addressed blobs plus small manifests that name them, so two tags sharing a weights file store it once. The runners are child processes, one per loaded model, that hold the weights and KV cache and actually generate tokens.
Models are stored in GGUF, the single-file format that carries weights, tokenizer, chat template and metadata together; its layout is described in GGUF and llama.cpp. The inference code is built on the GGML tensor library that llama.cpp also uses. Ollama has been moving model architectures onto its own engine on top of GGML, so which code path a particular model takes depends on the Ollama version and the model, and that detail rarely matters for operating it.
The life of a request
A request to /api/chat names a model and may carry options such as num_ctx and keep_alive. The scheduler first looks for a runner already serving that model with compatible settings. If one exists and has a free slot, the request goes straight to it. If the model is loaded but every slot is busy, the request waits in a queue. If the model is not loaded, the scheduler estimates its memory need, evicts idle models if necessary, and starts a runner.
Loading reads the GGUF file, places as many transformer layers on the GPU as fit, keeps the rest in system RAM, and allocates the KV cache for the configured context times the number of parallel slots. After the request finishes, the runner stays resident for keep_alive, five minutes by default, so the next request skips the load.
Two consequences follow. A request that asks for a different num_ctx than the loaded runner has forces a reload, so mixed clients that each set their own context size will thrash. And the first request after an idle period is slow for a reason that has nothing to do with generation speed.
Budget memory before you pull
Memory reserved at load has three parts: weights, KV cache and compute buffers. Weights are roughly parameters times bits per weight divided by eight. A 4-bit K-quant averages a little under 5 bits per weight once scales are included, so an 8-billion-parameter model is about 4.5 to 5 GB; the exact size is the blob size Ollama shows on pull. How quantization levels trade quality for size is covered in SLM quantization.
The KV cache is where surprises come from, because it scales with context and with parallel slots, not with the model file. For each token the model stores a key and a value vector per layer per KV head. Models with grouped-query attention have far fewer KV heads than query heads, which keeps this manageable. A worked example for a Llama-3.1-8B-shaped model:
def kv_cache_bytes(layers, kv_heads, head_dim, ctx, parallel, bytes_per_elem=2):
# K and V, per layer, per KV head, per token, per parallel slot
return 2 * layers * kv_heads * head_dim * ctx * parallel * bytes_per_elem
# Llama-3.1-8B-shaped model: 32 layers, 8 KV heads (GQA), head_dim 128
gib = 1024 ** 3
print(kv_cache_bytes(32, 8, 128, 8192, 1) / gib) # 1.0 GiB at f16
print(kv_cache_bytes(32, 8, 128, 8192, 4) / gib) # 4.0 GiB with 4 slots
print(kv_cache_bytes(32, 8, 128, 8192, 4, 1) / gib) # ~2.0 GiB with q8_0 KVSo an 8B model at 4 bits with an 8K context needs about 5 GB of weights plus 1 GiB of cache plus compute buffers: comfortable on an 8 GB GPU. Raise the context to 32K and the cache alone is 4 GiB. Add four parallel slots at 32K and it is 16 GiB, and most of the model spills to the CPU. OLLAMA_KV_CACHE_TYPE accepts f16, the default, q8_0 and q4_0; q8_0 halves the cache with small quality loss and is the first lever to pull when context is what you are short of. KV cache quantization requires flash attention, which OLLAMA_FLASH_ATTENTION controls.
Verify the result rather than trusting the arithmetic. ollama ps shows each loaded model's total size and a PROCESSOR column such as 100% GPU or 40%/60% CPU/GPU. Any CPU share means some layers run on the CPU, and because every token must pass through every layer, generation speed drops sharply, often several-fold, even when only a few layers are off the GPU.
Context length: the silent truncation trap
Every runner has a fixed context window, and prompts that exceed it are truncated, not rejected. A retrieval pipeline that stuffs twelve chunks into a 4K window gets an answer built from whatever survived, and nothing in the HTTP status tells you. This is the single most common way an Ollama deployment produces confidently wrong answers.
The default window has changed across releases. The FAQ states a 4,096-token default, while the current context-length documentation describes defaults that scale with available VRAM, from 4K on small GPUs to much larger windows on large ones. Treat the default as version and hardware dependent and set it explicitly: OLLAMA_CONTEXT_LENGTH on the server for a global default, num_ctx in a Modelfile for a packaged model, or num_ctx in request options. Then confirm it in the CONTEXT column of ollama ps.
The reliable defence is on the client: count tokens with the model's tokenizer before sending, compare against num_ctx, trim retrieved context to fit and leave room for the answer. prompt_eval_count in the final chunk is only a rough secondary signal, because Ollama reuses cached prompt prefixes and then counts only the newly evaluated tokens.
Concurrency, queueing and residency
Four server settings shape throughput and latency.
| Setting | Documented default | What it trades |
|---|---|---|
OLLAMA_NUM_PARALLEL | 1 | Concurrent requests per model; each slot multiplies the context allocation, so 4 slots at 8K reserve a 32K cache. |
OLLAMA_MAX_LOADED_MODELS | 3 per GPU, or 3 on CPU | How many models stay resident at once; more models means less memory each. |
OLLAMA_MAX_QUEUE | 512 | Requests waiting before the server starts rejecting new ones. |
OLLAMA_KEEP_ALIVE | 5 minutes | How long an idle runner stays loaded; -1 keeps it forever, 0 unloads immediately. |
Parallel slots help when many short requests arrive together, because a GPU decoding one sequence is mostly waiting on memory bandwidth and can decode several for little extra time per token. They hurt when context is long, because the cache multiplies.
For a single-user laptop, the defaults are fine. For a shared box serving one model, pin it with keep_alive -1, set the context explicitly and size slots to the measured concurrency. Ollama is not a high-throughput batch server; if you need dozens of concurrent streams with continuous batching and paged KV memory, that is the job of an engine like vLLM.
Modelfiles: packaging a model as a product
A Modelfile turns a base model into a named, reproducible artifact with its own defaults. FROM names a base model or a local GGUF file, PARAMETER sets generation and runtime defaults, SYSTEM sets a system prompt, and TEMPLATE overrides the prompt template.
# Modelfile: a support assistant built on a small instruct model
FROM llama3.2:3b
PARAMETER num_ctx 8192
PARAMETER temperature 0.2
PARAMETER stop "<|eot_id|>"
SYSTEM """You answer questions about the Acme billing API.
If you are not sure, say so and point to the docs page."""ollama create acme-support -f Modelfile # build a new model from the Modelfile
ollama show acme-support --modelfile # print what was actually stored
ollama run acme-support "How do I rotate an API key?"
ollama ps # NAME ID SIZE PROCESSOR CONTEXT UNTILTwo rules keep Modelfiles honest. First, prefer the template that ships in the GGUF; write TEMPLATE only when importing a file whose template is missing or wrong, because a broken template degrades every answer without any error, as explained in chat templates. Second, baking num_ctx into the Modelfile means every client gets the same window, which avoids reload thrash from mixed clients.
The API: chat, structure and measurement
The native API has /api/chat for conversations, /api/generate for raw completion, /api/embed for embeddings, plus management endpoints such as /api/tags and /api/ps. Responses stream as newline-delimited JSON by default. Ollama also serves an OpenAI-compatible surface under /v1, so many existing SDKs work by changing the base URL; the native API exposes more Ollama-specific options.
The format field accepts a JSON schema and constrains decoding so the output parses. That is far more reliable than asking nicely in the prompt, and it is the main technique in structured output with small models. It guarantees shape, not truth: validate the values as you would any input.
import json, requests
OLLAMA = "http://127.0.0.1:11434"
def chat(messages, model="acme-support", schema=None):
body = {
"model": model,
"messages": messages,
"stream": True,
"keep_alive": "30m", # keep the runner warm between bursts
"options": {"num_ctx": 8192}, # same value every call, or the model reloads
}
if schema:
body["format"] = schema # JSON schema constrains the output
with requests.post(f"{OLLAMA}/api/chat", json=body, stream=True, timeout=300) as r:
r.raise_for_status()
text = []
for line in r.iter_lines():
chunk = json.loads(line)
text.append(chunk["message"]["content"])
if chunk.get("done"):
# durations are nanoseconds
# fields can be omitted when zero, e.g. on a warm, cached request
ev, dur = chunk.get("eval_count", 0), chunk.get("eval_duration", 0)
tps = ev / (dur / 1e9) if dur else 0.0
print(f"load {chunk.get('load_duration', 0)/1e9:.2f}s, "
f"prompt {chunk.get('prompt_eval_count', 0)} tok, {tps:.1f} tok/s")
return "".join(text)
ticket_schema = {
"type": "object",
"properties": {"category": {"type": "string",
"enum": ["billing", "auth", "bug", "other"]},
"summary": {"type": "string"}},
"required": ["category", "summary"],
}
print(chat([{"role": "user", "content": "My key stopped working after rotation"}],
schema=ticket_schema))The final streamed chunk carries timing fields in nanoseconds: total_duration, load_duration, prompt_eval_count and prompt_eval_duration, eval_count and eval_duration. These are your metrics. Log them per request: load_duration above zero tells you about evictions, prompt tokens per second measures prefill, and eval tokens per second measures decode, which is bound by memory bandwidth.
Running Ollama for a team
The server has no authentication. Setting OLLAMA_HOST=0.0.0.0 makes it reachable from the network, and then anyone who can reach port 11434 can run, pull and delete models. Keep it bound to localhost and put a reverse proxy in front that terminates TLS, checks a token or SSO identity and allows only the endpoints users need, typically /api/chat, /api/embed and /v1. Block the model management endpoints at the proxy. OLLAMA_ORIGINS controls which browser origins may call it and should stay narrow.
Pre-pull models during deployment rather than on first request, and set OLLAMA_MODELS if model storage should live on a larger volume. Pin the models you serve with keep_alive -1 so users never pay a load, and set MAX_LOADED_MODELS low enough that pinned models cannot be evicted by an ad-hoc pull. Health-check with a tiny generation, not just a TCP connect, because a server can be up with a runner that failed to load.
Failure modes
- Silent truncation: prompts longer than the window are cut without an error. Set the context explicitly and count tokens on the client.
- Partial offload cliff: ollama ps shows a CPU share and decode speed falls several-fold. Shrink context or slots, quantize the KV cache, or pick a smaller quantization.
- Reload thrash: clients send different num_ctx values and load_duration is non-zero on most requests. Bake the context into the Modelfile.
- Cold first request: keep_alive expired over lunch. Pin production models.
- Open port: 0.0.0.0 without a proxy exposes model deletion and pulls to the network.
- Template drift: a custom TEMPLATE that does not match the model's training format produces rambling or role confusion with no error.
- Moving tags: a pull of latest changes behaviour under you; record digests.
Trade-offs
| Choice | Ollama | Alternative |
|---|---|---|
| Getting started | One command, curated library | llama.cpp server: more flags, more control |
| Concurrency | Fixed slots per model | vLLM: continuous batching, paged KV cache |
| Best fit | Laptops, dev boxes, small teams, edge | High-concurrency serving |
What to do next
- Run ollama ps with your model loaded and confirm the PROCESSOR column reads 100% GPU and the CONTEXT column shows the window you expect.
- Compute your KV cache budget with the function above for your model's layer count, KV heads and head size, at your real context and slot count.
- Set the context explicitly in a Modelfile, commit the Modelfile and rebuild with ollama create.
- Add a client-side token count before every request and reject or trim prompts that would not fit in num_ctx.
- Log load_duration, prompt and eval tokens per second from the final chunk of every response.
- Use the format field with a JSON schema for any output a program parses, and validate the values.
- If others will use the server, keep it on localhost behind an authenticating proxy that blocks model management endpoints.