When a transformer generates text, each layer keeps the key and value vectors of every token it has already seen, so the next token can attend to them without recomputing them. That stored state is the KV cache. It grows by a fixed number of bytes for every token of every live request, and on a serving GPU it is usually the thing you run out of first. Weights are a fixed cost. The KV cache is the variable cost that decides how many requests you can serve at once and how long they can be.
This article is about sizing it for a deployment. Starting from a model's config file and a GPU, you will compute the bytes each token costs, the memory left for the cache after weights and activations, how many tokens that pool holds, and what concurrency that buys for your traffic. The mechanics of how a server divides the pool into blocks are covered in paged KV cache architecture; here we stay at the level of capacity planning: which GPU, how many, which settings.
Bytes per token, from first principles
For standard attention, every layer stores one key vector and one value vector per KV head per token. Each vector has head_dim elements. So the cache costs, per token:
bytes_per_token = 2 * num_layers * num_kv_heads * head_dim * bytes_per_element
# ^ K and VThe term that matters most is num_kv_heads. In original multi-head attention (MHA) it equals the number of query heads. Grouped-query attention (GQA) lets several query heads share one KV head, so the cache shrinks by the group factor. Llama 3.1 70B has 64 query heads but only 8 KV heads, an 8x saving compared with MHA. Read the value from the model's config.json (num_key_value_heads), never from the number of attention heads.
| Model (published config) | Layers | KV heads x head_dim | Bytes per token at BF16 |
|---|---|---|---|
| Llama 3.1 8B | 32 | 8 x 128 | 2 x 32 x 8 x 128 x 2 = 131,072 (128 KiB) |
| Llama 3.1 70B | 80 | 8 x 128 | 2 x 80 x 8 x 128 x 2 = 327,680 (320 KiB) |
| DeepSeek-V3 (MLA) | 61 | one 512-dim latent + 64-dim RoPE key | 61 x 576 x 2 = 70,272 (about 69 KiB) |
Multi-head latent attention (MLA), used by DeepSeek-V2 and V3, does not store per-head K and V at all. It stores one compressed latent vector per layer per token, plus a small decoupled positional key, and reconstructs keys and values from it. That is why a 671B-parameter model can cost less cache per token than a 70B dense model. The general formula above does not apply; compute MLA from kv_lora_rank + qk_rope_head_dim per layer, and check that your serving engine actually stores the compressed form rather than expanded heads.
Two more adjustments. Models that interleave sliding-window layers with global layers only keep the last window of tokens in the local layers, so their cost per token falls for long contexts; count the global and local layers separately. And the element size depends on the KV dtype, not the weight dtype: weights quantised to 4 bits with a BF16 cache still pay 2 bytes per element. An FP8 KV cache pays 1.
The per-GPU budget is whatever is left
The cache gets the memory left after everything else, so compute everything else first.
- Usable fraction. Serving engines reserve a share of memory for other users, fragmentation and safety. vLLM's
--gpu-memory-utilization(default 0.9) sets the fraction of each GPU the engine may use in total. - Weights. Parameter count times bytes per parameter, divided by the tensor-parallel degree. A 70B model in BF16 is about 141 GB, or about 35 GB per GPU at TP=4.
- Activations and workspace. The temporary tensors of one forward pass at the largest batch the engine will schedule. vLLM measures this with a profiling forward pass at startup, sized by
--max-num-batched-tokens; a larger value means a larger peak and therefore a smaller cache. - Runtime overheads. CUDA context, communication buffers for tensor parallelism, and memory captured by CUDA graphs. Budget one to a few GB per GPU and confirm by measurement.
kv_pool_bytes_per_gpu = gpu_mem * utilization - weights_per_gpu - peak_activations - overheadsThis is why the startup log is your ground truth. Your arithmetic tells you whether a configuration is plausible; the engine's own measurement tells you what it got. vLLM prints lines such as GPU KV cache size: 161,344 tokens and Maximum concurrency for 16,384 tokens per request: 9.85x at startup. Record them for every configuration you deploy.
Tensor parallelism divides the cache too
With tensor parallelism, attention heads are split across GPUs, and each GPU stores the cache only for its own KV heads. At TP=4, Llama 3.1 70B's 8 KV heads become 2 per GPU, so each GPU stores 80 KiB per token instead of 320 KiB. The pool in tokens is the per-GPU pool divided by the per-GPU cost, and every token occupies a slice on all four GPUs at once. See tensor parallelism for how the split works.
There is a limit. Once the TP degree exceeds the number of KV heads, heads must be replicated rather than split, and adding GPUs stops reducing the per-GPU cache cost. Llama 3.1 70B at TP=16 would need each of its 8 KV heads on two GPUs. Pipeline parallelism, by contrast, splits layers, so each stage stores the cache only for its own layers. Both are worth knowing when you are choosing between more, smaller GPUs and fewer, larger ones.
From pool size to concurrency
A pool of N tokens can hold N tokens of live requests, in any mix. The question is what mix your traffic produces. A request occupies cache for its prompt plus every token generated so far, from the moment it is admitted until it finishes. Its average footprint over its life is roughly the prompt length plus half the output length, and its peak footprint is prompt plus full output.
Little's law turns that into capacity. The average number of requests in the system equals the arrival rate times the average time each spends there. Multiply by the average footprint and you get the average tokens the pool must hold:
def required_kv_tokens(arrival_rps, avg_prompt, avg_output, decode_tps_per_req, prefill_s=0.2):
"""Average live KV tokens for a steady arrival rate (Little's law)."""
time_in_system = prefill_s + avg_output / decode_tps_per_req # seconds
live_requests = arrival_rps * time_in_system
avg_footprint = avg_prompt + avg_output / 2 # tokens
return live_requests * avg_footprint
# 3 req/s, 2,000-token prompts, 400-token answers, 40 tok/s per stream
need = required_kv_tokens(3, 2000, 400, 40) # 3 * 10.2 s * 2,200 = about 67,000 tokensSize the pool for the average with headroom, typically 1.5x to 2x, because arrivals are bursty and long requests cluster. Then check the tail separately: the longest request you accept, set by --max-model-len, must fit in the pool on its own, and vLLM refuses to start if the configured maximum length cannot fit. Remember that decode speed per stream falls as concurrency rises, which increases time in system, so iterate the calculation with a measured decode rate from a load test rather than a single-request benchmark.
Worked example 1: Llama 3.1 70B on four 80 GB GPUs
Deploy Llama 3.1 70B in BF16 on four 80 GB GPUs with TP=4 and the default utilisation of 0.9.
- Usable per GPU: 80 x 0.9 = 72 GB.
- Weights per GPU: about 141 GB / 4 = 35.3 GB.
- Activations, CUDA graphs and communication buffers: assume 4 GB per GPU, to be confirmed from the log.
- KV pool per GPU: 72 - 35.3 - 4 = 32.7 GB.
- Bytes per token per GPU: 327,680 / 4 = 81,920.
- Pool: 32.7e9 / 81,920 = about 399,000 tokens.
- At an average live footprint of 8,000 tokens, about 50 concurrent requests; with the full 128K context for one request, about 3 such requests fit at once.
Now apply the levers. An FP8 KV cache halves bytes per token and roughly doubles the pool to about 800,000 tokens, at a small accuracy cost you must measure on your own evaluations. Capping --max-model-len at 32K does not enlarge the pool but stops a few huge requests from crowding out many normal ones. Moving to eight GPUs at TP=8 halves both weights and per-token cost per GPU, which more than doubles the pool, but also doubles cost; compare cost per served token, not pool size.
Worked example 2: Llama 3.1 8B on a 24 GB GPU
Small GPUs make the arithmetic unforgiving. Llama 3.1 8B in BF16 is about 16 GB of weights. On a 24 GB card at 0.9 utilisation the engine may use 21.6 GB. Subtract 16 GB of weights and around 1.5 GB of activations and overheads, and about 4 GB is left: 4e9 / 131,072 = about 30,500 tokens. That is not enough for even one request at the model's full 128K context; it is three or four requests at 8K.
The fixes, in the order to try them: cap the maximum context to what your traffic actually needs, use an FP8 KV cache (about 61,000 tokens), and quantise the weights (an 8-bit weight format frees about 8 GB, adding roughly another 60,000 tokens at BF16 cache). If you still need long contexts at concurrency, the answer is a bigger GPU, not more tuning.
Operating it: what to watch
- Startup log lines. Keep the reported pool size and maximum concurrency for each deployed configuration. A change after an engine upgrade or flag change is a capacity change, even when the model is the same.
- Cache utilisation. vLLM exposes
vllm:kv_cache_usage_percon its Prometheus endpoint, the fraction of KV blocks in use from 0 to 1. Sustained values near 1 mean admissions are limited by cache, not compute. - Waiting and preempted requests. When the cache fills, new requests wait in the queue and running ones may be preempted and recomputed or swapped. Watch the engine's queue-length and preemption metrics alongside time to first token; rising preemptions mean you are paying for wasted computation.
- Prefix-cache hit rate. With prefix caching enabled, shared system prompts occupy the pool once rather than per request, so a high hit rate stretches the effective pool considerably. Measure it before counting on it.
Scale on cache pressure, not only on GPU utilisation. A server can show modest compute use while rejecting work because its cache is full, which is the usual state of decode-heavy traffic with long contexts. Continuous batching makes good use of the pool you have, but it cannot create a bigger one.
Failure modes
- Wrong head count. Using the query-head count instead of
num_key_value_headsoverstates GQA models' cache by the group factor and leads you to buy GPUs you do not need. - Weight dtype mistaken for cache dtype. A 4-bit-weight model still pays 2 bytes per cache element unless the KV dtype is changed.
- Sizing on averages only. Average footprints fit, but a cluster of maximum-length requests fills the pool and preempts everyone else.
- Single-request decode speed. Little's law with an optimistic decode rate understates time in system and therefore required tokens.
- TP beyond the KV heads. Adding GPUs past the KV-head count replicates the cache instead of splitting it.
- Untracked engine changes. A new engine version changes activation profiling or graph capture and silently shrinks the pool; only the startup log shows it.
- Raising utilisation to 0.98. It grows the pool on paper and fails with out-of-memory errors when another process or a fragmentation spike needs the headroom.
What to do next
- Read
num_hidden_layers,num_key_value_headsandhead_dim(orhidden_size / num_attention_headswhen absent, or the MLA fields) from your model's config and compute bytes per token for your KV dtype. - Compute the per-GPU budget for each candidate GPU and TP degree, and the pool in tokens.
- Measure prompt and output length distributions from real traffic and apply Little's law with a load-tested decode rate.
- Choose
--max-model-lenfrom your real long tail, not the model's maximum. - Start the engine and compare the logged KV cache size and maximum concurrency with your estimate; investigate any gap above about 10%.
- Evaluate an FP8 KV cache on your quality benchmarks; adopt it if the drop is acceptable.
- Alert on
vllm:kv_cache_usage_perc, queue length and preemptions, and scale on cache pressure.