vLLM is usually introduced through large models split across GPUs with a KV cache that is always too small. Small language models change the picture. A 3B model fits on one modest GPU with memory to spare, a decode step takes milliseconds, and one server handles hundreds of concurrent requests from many tenants or fine-tuned variants. The question stops being whether the model fits and becomes how much useful work each step does.
This article covers vLLM from that angle: the scheduler's token budget, prefix caching, multi-LoRA serving, structured outputs and n-gram speculation. The process layout and the startup memory profile are covered in vLLM on the GPU, and continuous batching and KV-cache mechanics in vLLM continuous batching in production. Flags and fields below were checked against the vLLM documentation in October 2026; vLLM moves fast, so confirm them with vllm serve --help for your version.
What a small model changes
- Memory stops being the limit. A 3B model's 16-bit weights take about 6 GB, so on a 24 GB card most memory becomes KV cache. The binding limits are
--max-num-seqs(how many sequences one step may hold) and the per-step token budget. - Host-side work becomes visible. Scheduling, building inputs, structured-output masks, sampling, detokenising and streaming are a small fraction of a long forward pass but a large share of a short one, so a slow scheduler or Python-heavy sampler shows up directly in tokens per second.
- Workloads get more varied. Small models are cheap to fine-tune, so one server often hosts a base model plus many adapters, with tight output formats such as JSON.
The engine in one paragraph
vLLM's V1 engine puts the scheduler and model executor in an isolated EngineCore loop, separate from the API server that handles HTTP, tokenisation and detokenisation, so CPU-heavy request handling does not stall the loop that feeds the GPU. The same release introduced a persistent batch, which caches input tensors and applies only the differences each step, and piecewise CUDA graphs. All three cut exactly the host overhead that small models expose.
The scheduler's output each step is, in the words of the V1 announcement, a dictionary of the form {request_id: num_tokens}. It does not distinguish prefill from decode: a request that is decoding asks for one token, a request whose prompt is being processed asks for a chunk, and a request that hits the prefix cache skips the tokens that are already cached.
How the scheduler spends a step
Two limits shape every step. --max-num-batched-tokens caps the total tokens processed in one forward pass, and --max-num-seqs caps how many requests can be in it. Running decodes are scheduled first because each needs only one token; the remaining budget goes to waiting prompts, chunked if they do not fit. The sketch below shows the logic. It is a simplification for understanding, not vLLM's source.
# Conceptual shape of one V1 scheduler step (simplified; not vLLM source).
budget = max_num_batched_tokens
plan = {} # request_id -> tokens to process this step
for req in running: # decodes first: one token each
if len(plan) == max_num_seqs or budget == 0:
break
if not kv.allocate(req, 1):
preempt(newest(running)); continue # free blocks, recompute later
plan[req.id] = 1; budget -= 1
for req in waiting: # then prompts, chunked to fit the budget
if len(plan) == max_num_seqs or budget == 0:
break
hit = prefix_cache.longest_full_block_match(req) # cached tokens are skipped
n = min(req.remaining_prompt - hit, budget)
if kv.allocate(req, n):
plan[req.id] = n; budget -= n
run_forward(plan) # one pass over every scheduled tokenA large token budget finishes long prefills in fewer steps, helping time to first token, but slows every step that carries a big chunk, hurting inter-token latency for everyone decoding alongside it. A high --max-num-seqs raises throughput only while the KV cache can hold the sequences; beyond that the scheduler preempts and recomputes, and the work is lost. The startup log reports the KV capacity in tokens; divide it by your typical sequence length to get the concurrency the cache actually supports, then set --max-num-seqs at or below that.
def kv_bytes_per_token(layers, kv_heads, head_dim, dtype_bytes=2):
# one K and one V vector per layer per KV head
return 2 * layers * kv_heads * head_dim * dtype_bytes
# An illustrative 3B-class model with grouped-query attention (not a specific checkpoint)
per_tok = kv_bytes_per_token(layers=28, kv_heads=8, head_dim=128) # 114,688 bytes = 112 KiB
kv_pool = 14 * 2**30 # 14 GiB left for KV
tokens = kv_pool // per_tok # 131,072 tokens
print(per_tok, tokens, tokens // 2048) # 2,048-token requests -> 64 at onceHere the cache holds 131,072 tokens: 64 requests of 2,048 tokens, or 256 of 512. Setting --max-num-seqs to 256 for the longer requests buys preemptions, not concurrency.
Prefix caching: how reuse is detected
Small-model workloads often resend the same long prefix: a system prompt, few-shot examples, a tool schema. Prefix caching reuses the KV blocks computed for it. The V1 announcement states it is enabled by default because its overhead is near zero.
The mechanism is a hash chain over fixed-size blocks. Per vLLM's design document, each block's hash covers its parent block's hash, the block's tokens, and extra values such as the LoRA adapter id, multimodal input hashes and a cache salt, so a block matches only if everything before it matches too. Only full blocks are cached, and free blocks are evicted least-recently-used first. The hash function is configurable with --prefix-caching-hash-algo: sha256 is the default, sha256_cbor is an alternative serialisation, and xxhash and xxhash_cbor are faster but not cryptographic.
import hashlib, pickle
BLOCK = 16
def block_hashes(token_ids, extra=()):
# extra: values that must split the cache, e.g. LoRA adapter id, cache salt
parent, out = b"", []
full = len(token_ids) // BLOCK * BLOCK # only FULL blocks are cacheable
for i in range(0, full, BLOCK):
key = (parent, tuple(token_ids[i:i + BLOCK]), tuple(extra))
parent = hashlib.sha256(pickle.dumps(key)).digest()
out.append(parent)
return out
# A 1,000-token system prompt yields 62 full blocks (992 tokens);
# the last 8 prompt tokens are recomputed for every request.Three practical rules follow from the design:
- Put stable content first, byte for byte. A timestamp, request id or user name in the first line of the system prompt changes the first block's hash and therefore every hash after it, so nothing is reused. Move variable content after the shared prefix.
- Tokens, not characters, must match. A different chat template, a trailing space or a changed tool order produces different tokens. Render prompts with one code path.
- Adapters do not share cache. The LoRA id is part of the hash, which is correct because an adapter changes the keys and values, so a prefix shared by ten adapters is computed and stored ten times.
On multi-tenant servers a shared cache is also a timing side channel: a faster first token reveals that someone recently sent the same prefix. A per-tenant cache salt keeps tenants' blocks apart; check your version's API reference for how it is passed.
Multi-LoRA: many fine-tunes on one base
A LoRA adapter adds small low-rank matrices to some layers of a base model. vLLM can serve many adapters on one copy of the base weights and mix requests for different adapters in the same batch. Start the server with --enable-lora and register adapters with --lora-modules name=path; requests then select an adapter by passing its name as the model. /v1/models lists the base model and every loaded adapter, with each adapter's parent pointing at its base.
--max-loras is how many adapters can be active in one batch (excess requests wait for a slot); --max-lora-rank must cover your largest rank, and higher wastes memory; --max-cpu-loras is how many adapters are cached in host memory. With VLLM_ALLOW_RUNTIME_LORA_UPDATING=True, adapters can be added with POST /v1/load_lora_adapter and removed with POST /v1/unload_lora_adapter; adding "load_inplace": true replaces an adapter under the same name. Runtime loading is an administrative capability, so keep those endpoints behind your own authentication. The adapter side of this, training and packaging, is covered in SLM LoRA serving.
Structured outputs without retries
For classification and extraction the output must parse. vLLM's structured outputs constrain sampling to tokens consistent with a JSON schema, regex, choice list or grammar. In the OpenAI-compatible server, pass {"structured_outputs": {"json": schema}} in the request body, or use response_format with type json_schema. The older guided_json, guided_regex and related fields were removed in v0.12.0; clients still sending them must migrate. The backend is chosen with --structured-outputs-config.backend, whose default auto picks one per request.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")
TICKET_SCHEMA = {
"type": "object",
"properties": {
"category": {"enum": ["billing", "outage", "account", "other"]},
"priority": {"type": "integer", "minimum": 1, "maximum": 4},
"summary": {"type": "string", "maxLength": 200},
},
"required": ["category", "priority", "summary"],
"additionalProperties": False,
}
resp = client.chat.completions.create(
model="support", # the LoRA adapter name, not the base model
messages=[{"role": "system", "content": SYSTEM_PROMPT}, # identical bytes every call
{"role": "user", "content": ticket_text}],
extra_body={"structured_outputs": {"json": TICKET_SCHEMA}},
max_tokens=256,
temperature=0,
)Constraining makes output well-formed, not correct: if a 3B model does not know the category, the schema guarantees a valid wrong answer. Keep schemas small and closed, reuse the same schema across requests, and evaluate accuracy with the constraint on. More on prompting small models for structure is in structured output with SLMs.
N-gram speculation
Decoding is memory-bound at small batch sizes: each step reads all the weights to produce one token per sequence. Speculative decoding proposes several tokens cheaply and verifies them in one forward pass, keeping the ones the model agrees with. The n-gram method needs no draft model. It looks for the most recent tokens elsewhere in the context and proposes what followed them there. vLLM configures it with --speculative-config and the keys method, num_speculative_tokens, prompt_lookup_min and prompt_lookup_max.
It pays off when output copies input: extraction, code edits, answers over retrieved documents. It does little for free-form text and less as batches grow, because a large batch is closer to compute-bound and rejected proposals waste compute. Benchmark at real concurrency;speculative decoding for SLMs covers draft-model alternatives.
Worked example: one GPU, three products
A team serves a 3B instruct model on one 24 GB GPU for three products: ticket triage, SQL generation and billing questions, each with a LoRA adapter. Triage and billing share a 1,200-token system prompt, and triage must return JSON.
The startup log reports about 131,000 tokens of KV cache. Triage requests average 1,500 tokens and SQL 3,000, so a realistic mix holds 60 to 80 sequences; --max-num-seqs 128 works only because many requests share cached prefix blocks. The launch command:
export VLLM_ALLOW_RUNTIME_LORA_UPDATING=True # only if you need runtime load/unload
vllm serve ./models/base-3b-instruct \
--max-model-len 8192 \
--max-num-seqs 128 \
--max-num-batched-tokens 8192 \
--enable-lora \
--lora-modules support=./adapters/support billing=./adapters/billing \
--max-loras 4 --max-lora-rank 16 --max-cpu-loras 16
# separate instance: SQL adapter merged into the weights, so no LoRA, with n-gram speculation
vllm serve ./models/base-3b-sql-merged --max-model-len 8192 \
--speculative-config '{"method": "ngram", "num_speculative_tokens": 4,
"prompt_lookup_min": 2, "prompt_lookup_max": 5}'
# load another adapter later, without a restart
curl -s localhost:8000/v1/load_lora_adapter -H 'Content-Type: application/json' \
-d '{"lora_name": "refunds", "lora_path": "/adapters/refunds"}'The shared system prompt is identical across triage and billing, but because the adapters differ, each adapter caches its own copy of those 75 blocks of 16 tokens: fine at two adapters, wasteful at forty. SQL generation copies table and column names from the prompt, so it suits n-gram speculation, but vLLM's feature compatibility matrix lists LoRA with speculative decoding as unsupported. The team therefore merges the SQL adapter into a copy of the weights and serves it on a second instance with speculation. They track time to first token and inter-token latency per product, preemptions, KV usage and prefix-cache hit rate from /metrics, reading the metric names their version exposes, since names have changed between releases.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| Throughput plateaus while the GPU is far from busy | Host-side work dominates short steps: old release, heavy detokenisation, slow custom logits processors | Upgrade, profile the API server, drop Python per-token hooks |
| Latency spikes and repeated work under load | Preemption: --max-num-seqs exceeds what the KV pool holds | Lower max-num-seqs, cap max-model-len, or quantise the KV cache |
| Prefix hit rate near zero | Variable content at the start of the prompt, or template differences | Move variable content after the shared prefix; render prompts in one place |
| Requests for some adapters queue | More distinct adapters per batch than --max-loras | Raise max-loras, route adapters to instances, merge rarely used adapters |
| Unparseable output after an upgrade | Clients still send guided_json, removed in v0.12.0; do not assume an error, check whether it is ignored | Migrate to structured_outputs; validate output in the client |
| Speculation slower, or rejected at startup | Free-form output, large batches, or combined with LoRA | Measure per workload; check the compatibility matrix |
Operational guidance and trade-offs
- One instance or several per GPU. Several instances isolate products but each reserves memory and duplicates weights; one instance batches better. Split only for isolation.
- Pin versions. Flags, defaults, metric names and request fields change between releases; keep a traffic replay to compare before upgrading.
What to do next
- Read the KV capacity from your startup log and divide by your typical sequence length; set
--max-num-seqsat or below the result. - Audit your prompts for variable content before the shared prefix, then confirm prefix-cache hits in
/metrics. - List your fine-tunes and decide which become LoRA adapters on one base; set
--max-lora-rankto the largest rank you actually use. - Migrate any
guided_*request fields tostructured_outputsand measure accuracy with constraints on. - Benchmark n-gram speculation per workload at peak concurrency and enable it only where it wins.
- Pin the vLLM version and keep a traffic replay to rerun before every upgrade.