vLLM is an open-source inference engine for large language models. Point it at a model and it gives you an OpenAI-compatible HTTP server whose throughput on one GPU is usually several times that of a naive loop over model.generate(). The gain comes from two ideas: store the key-value (KV) cache in small fixed-size blocks instead of one contiguous slab per request, and re-plan the batch at every decoding step instead of once per batch.
Those mechanisms have their own articles on this site. This one is about running the engine: where each gigabyte of GPU memory goes, how that memory becomes a hard limit on concurrent tokens, which four settings control the trade between latency and throughput, how to spread a model across GPUs, and how to read the engine's behaviour from its metrics. The worked examples use Llama-3.1-8B and 70B on 80 GB GPUs, but the arithmetic transfers to any decoder model.
What runs where
A vllm serve deployment is several cooperating processes rather than one. The API server accepts HTTP requests, applies the chat template, tokenises prompts and detokenises output as it streams back. The engine core owns the scheduler and the KV block manager: it decides, step by step, which requests run and which cache blocks each one holds. The GPU workers, one per GPU, hold the model weights and a slice of the KV pool and execute the forward pass. Splitting them keeps CPU-heavy work such as tokenisation off the loop that feeds the GPU.
The loop is simple to state. Every step, the scheduler picks a set of requests, the workers run one forward pass that processes all of their scheduled tokens together, one new token is sampled for every request that is decoding, and the results stream back. A request that finishes frees its blocks immediately and a waiting request can take its place on the very next step. That is continuous batching; see continuous batching for why iteration-level scheduling beats fixed batches.
The startup memory profile
When the engine starts, it loads the weights, then runs a dummy forward pass at the largest batch the configuration allows, to measure peak activation memory. It then computes how much memory it may use in total: --gpu-memory-utilization times the GPU's total memory. Whatever remains under that cap after weights, peak activations and other runtime overheads becomes the KV cache pool, which is allocated up front and carved into blocks. The engine logs the resulting capacity in tokens at startup; it is the most important number about your server.
Two consequences follow. First, the engine takes nearly the whole GPU on purpose. nvidia-smi showing 90 percent memory in use on an idle server is correct behaviour, not a leak. Second, the fraction is of total memory, so anything else using the same GPU, such as a second vLLM instance or a stray notebook, can make startup fail or leave less room than expected. The default has changed across releases (0.9 for a long time; the current docs show 0.92), so check vllm serve --help for your version and set it explicitly.
Worked example: turning gigabytes into a token budget
The KV cache stores one key vector and one value vector per layer per KV head for every token held. With grouped-query attention the number of KV heads is smaller than the number of query heads, which is why modern models are far cheaper to cache than older ones. The arithmetic is short enough to do in a script:
def kv_bytes_per_token(layers, kv_heads, head_dim, dtype_bytes, tp=1):
# K and V, every layer, every KV head this GPU holds
return 2 * layers * (kv_heads // tp) * head_dim * dtype_bytes
def kv_token_budget(gpu_gib, util, weights_gib, reserve_gib, per_token):
usable = gpu_gib * util - weights_gib - reserve_gib
return int(usable * 2**30 // per_token)
# Llama-3.1-8B: 32 layers, 8 KV heads (GQA), head_dim 128, bf16
t8 = kv_bytes_per_token(32, 8, 128, 2) # 131072 B = 128 KiB
n8 = kv_token_budget(80, 0.90, 15, 3, t8) # 442368 tokens
# Llama-3.1-70B on 4 GPUs: 80 layers, 8 KV heads -> 2 per GPU
t70 = kv_bytes_per_token(80, 8, 128, 2, tp=4) # 81920 B = 80 KiB
n70 = kv_token_budget(80, 0.90, 33, 4, t70) # 458752 tokens for the replica
print(t8, n8, n8 // 16384) # full-length 16k requests that fit at once: 27Llama-3.1-8B in bf16 needs 128 KiB of cache per token. On an 80 GB GPU (treated as 80 GiB for round numbers) at a 0.90 cap, with about 15 GiB of weights and an assumed 3 GiB reserve for activations and runtime overheads (your profile will report the real figure), roughly 54 GiB remains, which holds about 442,000 tokens. That is the real limit on concurrency: 27 requests at a full 16,384-token context, or about 200 requests averaging 2,000 tokens of prompt plus output.
For Llama-3.1-70B across four GPUs with tensor parallelism, each GPU holds a quarter of the weights (about 33 GiB) and two of the eight KV heads, so each GPU needs 80 KiB per token. With an assumed 4 GiB reserve, about 35 GiB remains per GPU, giving roughly 459,000 tokens. Every GPU holds its slice of every token, so this is the budget for the whole four-GPU replica, not per GPU.
Setting --kv-cache-dtype fp8 halves bytes per token and so roughly doubles the budget, at some accuracy cost that you must measure on your own evaluation set. Quantising weights (for example an FP8 or 4-bit checkpoint) frees weight memory that flows straight into the pool.
The four settings that shape behaviour
| Setting | What it bounds | Raise it when | Lower it when |
|---|---|---|---|
--max-model-len | Longest prompt plus output one request may use | Users need longer context | Startup fails because one full-length request will not fit in the pool |
--gpu-memory-utilization | Share of total GPU memory the engine may claim | GPU is dedicated and you want more KV blocks | Out-of-memory at startup or during CUDA graph capture |
--max-num-seqs | Requests in the running batch at once | KV usage is low and requests queue | Preemptions climb or per-token latency is too high |
--max-num-batched-tokens | Tokens processed per step, prefill and decode together | Time to first token matters, or you want throughput | Inter-token latency spikes when long prompts arrive |
They interact through the step. vLLM's scheduler first gives each running decode its one token, then fills the rest of the token budget with chunks of waiting prompts; long prompts are split across several steps (chunked prefill, covered in chunked prefill). A small token budget such as 2,048 keeps every step short, so streaming feels smooth, but long prompts take more steps to start answering. A budget above 8,192 improves throughput and time to first token on large GPUs, but a step that includes a big prefill chunk makes every decoding request in that step wait longer for its next token.
When the pool runs out mid-generation, the scheduler preempts the most recently admitted requests: it frees their blocks and later recomputes their cache from the prompt. Recompute is the default in the current engine because it is cheaper than copying blocks to CPU memory. Occasional preemption is fine. Sustained preemption means you admit more work than the pool can hold, and you pay for the same prefill twice.
Prefix caching reuses blocks for identical prompt prefixes such as a long system prompt, and in effect enlarges the pool for chat workloads; the mechanics are in prefix caching in depth.
Launching it
# One GPU, Llama-3.1-8B-Instruct, explicit limits instead of defaults
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--max-model-len 16384 \
--gpu-memory-utilization 0.90 \
--max-num-seqs 256 \
--max-num-batched-tokens 8192 \
--port 8000
# Four GPUs in one node, 70B model sharded with tensor parallelism
vllm serve meta-llama/Llama-3.1-70B-Instruct \
--tensor-parallel-size 4 \
--max-model-len 32768 \
--gpu-memory-utilization 0.90For batch jobs, skip HTTP and use the offline API. Hand the engine every prompt at once and let it schedule; chunking into small client-side batches throws away the continuous batching you came for.
from vllm import LLM, SamplingParams
llm = LLM(
model="meta-llama/Llama-3.1-8B-Instruct",
max_model_len=16384,
gpu_memory_utilization=0.90,
)
params = SamplingParams(temperature=0.2, max_tokens=256)
prompts = [f"Summarise ticket {i} in one line: ..." for i in range(5000)]
# One call; the engine batches internally, you do not chunk by hand
for out in llm.generate(prompts, params):
print(out.request_id, out.outputs[0].text[:80])
Spreading a model over GPUs
vLLM offers three kinds of parallelism, and choosing between them is mostly a question of whether the model fits.
- Tensor parallelism (
--tensor-parallel-size) splits every layer's matrices across GPUs, so each step needs all-reduce communication inside every layer. It cuts per-token latency as well as memory, but it needs a fast link: within one NVLink-connected node it works well; across PCIe it scales poorly. The communication pattern is explained in tensor parallelism in depth. - Pipeline parallelism (
--pipeline-parallel-size) gives each GPU or node a contiguous group of layers. It needs far less bandwidth and is the usual way to span nodes, but a single request crosses every stage in turn, so latency does not improve. - Data parallelism (
--data-parallel-size, or simply more replicas behind a load balancer) copies the whole model. It scales throughput almost linearly and adds no communication, but only works if the model fits on the GPUs one replica owns.
The rule of thumb: use the smallest tensor-parallel size that fits the weights with a healthy KV pool, inside one node; add pipeline stages only when one node cannot hold the model; then add replicas for throughput. Four TP=1 replicas of an 8B model typically out-serve one TP=4 replica, which spends time on all-reduces it never needed.
CUDA graphs, compilation and startup time
At small batch sizes a decode step is a chain of hundreds of short kernels, and the CPU cost of launching them one by one can exceed the GPU time. vLLM compiles the model with torch.compile and captures CUDA graphs for a set of batch sizes, so a whole step replays as a single launch. The price is startup time and some memory. --enforce-eager skips both, which is useful for debugging a model that fails to compile or for fast iteration in development, but costs decode throughput in production.
Startup on a new node is therefore weight download, weight loading, compilation and graph capture, often several minutes for a large model. Keep weights on a local or cached volume and persist vLLM's cache directory between restarts so compilation artefacts are reused. Size Kubernetes probes to match:
# Kubernetes probes for a vLLM pod: loading weights, compiling and capturing
# CUDA graphs can take minutes, so give startup its own generous budget.
startupProbe:
httpGet: {path: /health, port: 8000}
periodSeconds: 10
failureThreshold: 90 # up to 15 minutes to become ready
readinessProbe:
httpGet: {path: /health, port: 8000}
periodSeconds: 5
livenessProbe:
httpGet: {path: /health, port: 8000}
periodSeconds: 15
failureThreshold: 4
Metrics that tell you what the engine is doing
The server exposes Prometheus metrics at /metrics. Five cover most diagnoses:
vllm:num_requests_runningandvllm:num_requests_waiting: a growing waiting count with low KV usage means--max-num-seqsis the limit; with high KV usage, memory is.vllm:kv_cache_usage_perc: sustained values near 1.0 mean the pool is the bottleneck.vllm:num_preemptions: should be near zero in steady state; a rising rate means over-admission.vllm:time_to_first_token_secondsandvllm:inter_token_latency_seconds: the two latencies users feel; the token budget trades one against the other.vllm:prefix_cache_hitsovervllm:prefix_cache_queries: your prefix hit rate, which tells you whether prompt layout changes are worth making.
Alert on queue time (vllm:request_queue_time_seconds) rather than GPU utilisation. A vLLM GPU can report near-100 percent utilisation while users wait in the queue; queue time measures the thing that hurts.
Failure modes and their fixes
- Startup refuses to run because the context does not fit. The pool cannot hold one request of
--max-model-lentokens. Lower the context, raise the memory cap, use an FP8 KV cache, or add tensor parallelism. - Out-of-memory during graph capture or the first large batch. Another process holds GPU memory, or the cap is too close to 1.0. Give the engine the GPU to itself and back the cap off by a few points.
- Preemption storms under load. Too many long requests admitted at once. Lower
--max-num-seqs, and add admission control in front so overload becomes fast rejections instead of slow recomputation. - Latency spikes when long documents arrive. Big prefill chunks are delaying decodes. Lower
--max-num-batched-tokens, or route long-context traffic to its own replicas. - Poor scaling with more GPUs. Tensor parallelism over PCIe, or TP larger than needed. Measure tokens per second per GPU at each TP size; use replicas instead.
What to do next
- Compute bytes per token for your model from its config, and check the result against the capacity line vLLM logs at startup.
- Set
--max-model-len,--gpu-memory-utilization,--max-num-seqsand--max-num-batched-tokensexplicitly, and record why you chose each value. - Load-test with your real prompt and output length distribution, sweeping the token budget, and plot time to first token against inter-token latency.
- Scrape
/metricsand alert on queue time, preemption rate and KV usage near 1.0. - Pick parallelism by the rule: smallest TP that fits, pipeline only across nodes, replicas for throughput.
- Read the paged KV cache article and the serving stacks comparison before deciding whether vLLM is the right engine for your workload.