llama.cpp is a C and C++ inference engine for transformer models that runs on laptops, phones, servers and almost any GPU. Most people meet it indirectly, through a desktop app or a wrapper, and treat it as a box that turns a GGUF file into tokens. That works until a model runs out of memory, generates at a fraction of the expected speed or produces odd output, and then you need to know what the engine is doing.
This page covers the engine and its server: how the layers fit together, what happens to one request, how to compute memory before you load anything, which flags matter and how to measure the result. The file format itself, including conversion and quantization on disk, is covered in GGUF format, in depth. Flag names and defaults here were checked against the llama-server README in the ggml-org repository on 2 October 2026; the project moves quickly, so confirm them with llama-server --help on the build you run.
Three layers
ggml is the tensor library underneath. It represents a forward pass as a compute graph of operations such as matrix multiply, RMS norm and attention, and executes that graph on a backend: optimised CPU code using the SIMD instructions available, or a GPU backend such as CUDA, Metal or Vulkan. Its quantized types store weights in small blocks with shared scales, and its kernels multiply against those blocks directly, so a 4-bit model is never expanded to 16 bits in memory.
libllama knows about language models. It reads GGUF metadata to build the right architecture, maps the weights, owns the KV cache, splits work into batches and provides the sampler chain. The tools are thin programs over libllama: llama-server for an HTTP API, llama-bench for speed, llama-perplexity for quality, llama-imatrix and llama-quantize for producing quantized files.
One consequence of this design is that a feature arrives at different times in different places. A new architecture needs libllama support, a new kernel needs backend support, and a fast path on CUDA may not yet exist on Vulkan. When something is slow on one machine and fast on another, check which backend actually ran the graph before tuning anything else.
One request through llama-server
The server keeps a fixed set of slots, one per concurrent sequence, set by -np (automatic by default). A chat request arrives at the OpenAI-compatible /v1/chat/completions route, is assigned a slot and is turned into a prompt using the chat template stored in the GGUF, rendered with a Jinja engine. The prompt is tokenized and compared with what the slot already holds: if the new prompt starts with the same tokens, as it does in a multi-turn chat, those tokens are reused from the KV cache rather than processed again.
The remaining prompt tokens are processed in chunks: -b sets the logical batch size (default 2048) and -ub the physical micro-batch actually sent to the backend (default 512). This prefill phase is compute-bound, because many tokens share each weight read. Then generation starts. Each decode step runs one forward pass that advances every busy slot by one token, which is continuous batching: new requests join between steps instead of waiting for a batch to finish. Each slot's logits go through the sampler chain, the chosen token is streamed to the client as a server-sent event, and the loop continues until an end-of-sequence token, a stop string or the token limit.
If prompt plus output exceeds the slot's context, the default is to stop rather than silently discard history; context shifting is off unless you enable --context-shift. The response reports the cut-off, and your client should treat it as an error, not a short answer.
Compute memory before you load
Memory has three parts: the weights, the KV cache and compute buffers. Weights are roughly the GGUF file size. The KV cache stores one key and one value vector per layer, per KV head, per token:
kv_bytes_per_token = 2 * n_layers * n_kv_heads * head_dim * bytes_per_elementFor a Llama 3.1 8B model there are 32 layers, 8 KV heads because of grouped-query attention and a head dimension of 128. In f16, the default cache type, that is 2 × 32 × 8 × 128 × 2 = 131,072 bytes, 128 KiB per token. A 32,768-token context therefore needs 4 GiB of KV cache on top of roughly 4.9 GB of weights at Q4_K_M. Setting -ctk q8_0 -ctv q8_0 stores 32 elements in 34 bytes and cuts the cache to about 2.1 GiB, usually with a small quality cost you should measure rather than assume. Compute buffers add a few hundred megabytes and grow with -ub.
Context is a server-wide budget. -c sizes the KV pool (0 means the model's training length, which for many current models is far more than you can afford), and with several slots the pool is either shared or divided per slot. The current README enables the unified, shared pool by default only when the slot count is automatic, so setting -np 4 with -c 16384 may leave each slot 4,096 tokens unless you also pass -kvu. Read the effective n_ctx from /props and /slots after start-up instead of trusting arithmetic alone. Setting -c explicitly is the single most effective way to avoid surprise out-of-memory failures.
Offload, flash attention and MoE models
-ngl sets how many layers live in GPU memory; it accepts a number, auto or all, and the current default is auto. Layers left on the CPU run there, and every token then crosses between devices, so a model that is 90% offloaded can be much slower than the 90% suggests. With several GPUs, -sm layer (the default) places whole layers on each device and -ts sets the proportions.
Mixture-of-experts models change the calculation. Only a few experts run per token, so expert weights are large but rarely read. --n-cpu-moe N keeps the expert weights of the first N layers in system memory while attention and shared weights stay on the GPU, and -ot moves arbitrary tensors by name pattern. This is how a model whose total weights exceed GPU memory still decodes at usable speed. -fa controls flash attention (on, off or auto, default auto), which reduces attention memory traffic for long contexts.
# One 8B model, 4 slots sharing one unified 16k-token KV pool, metrics on, local only
llama-server -m models/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf \
-c 16384 -np 4 -kvu -ngl all -fa auto \
--host 127.0.0.1 --port 8080 --metrics --api-key-file keys.txt
# A large MoE model on a 24 GB GPU: experts of the first 20 layers stay in RAM
llama-server -m models/moe-model-Q4_K_M.gguf -c 8192 -ngl all --n-cpu-moe 20
Calling it
Because the server speaks the OpenAI chat format, any OpenAI client works by changing the base URL. The server also accepts llama.cpp sampling fields such as min_p and top_k in the same JSON body, and can constrain output with a JSON schema or a GBNF grammar, which is how small models produce reliable structured output.
import requests, time
BASE = "http://127.0.0.1:8080"
HEAD = {"Authorization": "Bearer " + open("keys.txt").readline().strip()}
while requests.get(BASE + "/health").status_code == 503: # 503 while the model loads
time.sleep(1)
r = requests.post(BASE + "/v1/chat/completions", headers=HEAD, json={
"messages": [{"role": "user", "content": "Name three uses of a KV cache."}],
"max_tokens": 200, "temperature": 0.3, "min_p": 0.05,
})
body = r.json()
print(body["choices"][0]["message"]["content"])
print(body["usage"]) # prompt and completion token counts
Measuring speed
llama-bench reports two numbers by default: pp512, prompt processing of 512 tokens, and tg128, generation of 128 tokens, both in tokens per second. They measure different bottlenecks. Prefill is compute-bound and scales with GPU throughput. Generation is memory-bandwidth-bound, because each new token reads every active weight once.
That gives a ceiling you can compute. Generation tokens per second cannot exceed memory bandwidth divided by bytes read per token. A 4.9 GB model on a laptop with 100 GB/s of memory bandwidth tops out near 20 tokens per second; on a GPU with 1,000 GB/s it tops out near 200. Real numbers land below the ceiling. If yours is far below, look for layers left on the CPU, a backend that is not the one you expected or thread oversubscription. The same arithmetic explains why unified-memory machines do well with large models; see Apple Silicon, in depth.
# Compare full offload with partial offload, three repetitions each
llama-bench -m model-Q4_K_M.gguf -ngl 99,20 -p 512 -n 128 -r 3
Measuring quality after quantization
Speed is easy to see; quality loss from quantization is not. llama-perplexity can save the logits of a reference model on a text file and then compare a quantized model against them with --kl-divergence-base and --kl-divergence. KL divergence and the rate at which the top token changes are more sensitive than perplexity alone. The saved logits file is large, tens of gigabytes for a full test set with a modern vocabulary, so use a modest text file. An importance matrix from llama-imatrix on representative text, passed to llama-quantize --imatrix, usually reduces the loss at low bit widths. How the k-quant blocks work is in GGUF quantization architecture, and how to judge a result statistically is in quantization evaluation methodology.
llama-perplexity -m model-BF16.gguf -f eval.txt --kl-divergence-base ref.kld
llama-perplexity -m model-Q4_K_M.gguf --kl-divergence-base ref.kld --kl-divergence
Running it in production
- Health:
/healthreturns 503 while the model loads and 200 when ready, and skips the API-key check, so it is safe for load-balancer probes. - Metrics: start with
--metricsand scrape/metrics. Watchllamacpp:requests_deferred(requests waiting for a slot),llamacpp:predicted_tokens_secondsandllamacpp:n_busy_slots_per_decode. - Access: the default bind address is 127.0.0.1. If you bind a public interface, set
--api-keyor--api-key-fileand put TLS in front. - Prompt reuse: prompt caching is on by default; keep stable system prompts at the start of every request so the shared prefix is reused.
- Pin versions: flags and defaults change between releases. Pin a release tag or container digest and read the changelog before upgrading.
Failure modes
- Out of memory at load or under load:
-c 0on a long-context model, or KV cache growth as slots fill. Set-cexplicitly and size from the formula above. - Wrong chat template: fine-tunes sometimes ship a template that does not match their training format, giving rambling answers or broken tool calls. Inspect the rendered prompt with
/apply-templateand override with--chat-template-file. - Silent CPU fallback: a build without the intended backend runs everything on the CPU. Check
--list-devicesand the start-up log. - Thread oversubscription: more CPU threads than physical cores, or several servers on one host, makes generation slower. Benchmark thread counts.
- Truncated answers: context exhausted with context shift off. On the OpenAI-compatible route check for a
finish_reasonoflength(the native/completionroute setstruncated) and size the context to the workload.
Trade-offs
llama.cpp wins on portability, a small dependency footprint, CPU and Apple performance, aggressive quantization and mixed CPU-GPU placement. A dedicated datacenter engine such as vLLM is usually the better choice for many concurrent users on large NVIDIA or AMD GPUs, where paged KV management and higher-precision kernels give more throughput per GPU. Wrappers built on llama.cpp trade control for convenience; when you need specific flags, run llama-server directly.
What to do next
- Pick one model and compute its KV bytes per token from the GGUF metadata, then choose
-cand-npthat fit your memory with headroom. - Run
llama-benchwith full and partial offload and compare generation speed with the bandwidth ceiling. - Start llama-server with an explicit context, metrics and an API key, and confirm the effective context in
/props. - Render one conversation through
/apply-templateand check it matches the model card's format. - Measure KL divergence of your chosen quantization against a higher-precision reference on text from your own domain.
- Pin the release you tested and add the health and metrics endpoints to your monitoring.