LMDeploy is an open-source toolkit from the InternLM team (Shanghai AI Laboratory) for compressing and serving large language models on GPUs. It belongs to the same family as vLLM, SGLang and TensorRT-LLM: it takes a Hugging Face checkpoint, keeps many requests in flight on the GPU at once, and exposes an OpenAI-compatible HTTP API. Its distinguishing parts are a C++/CUDA inference engine called TurboMind, a second engine written in PyTorch for wider model coverage, and a built-in quantization toolkit for 4-bit weights and 8-bit or 4-bit KV caches.
This page explains how a request moves through LMDeploy, how it decides how much GPU memory goes to the KV cache, what its quantization options really trade, and the configuration mistakes that cost teams days. It ends with a worked sizing example for an 8B model on a 24 GB card and a deployment checklist. Flags and defaults below were checked against the LMDeploy documentation in October 2026; the project moves quickly, so confirm them against the version you install.
What LMDeploy is made of
LMDeploy ships as one Python package, lmdeploy, with four pieces you will touch:
- TurboMind, the C++/CUDA engine. It runs a persistent batch loop: requests join and leave the running batch at every decoding step instead of waiting for a whole batch to finish. Other projects call this continuous or in-flight batching.
- The PyTorch engine, written mostly in Python with custom kernels. It supports more model architectures and lands new models sooner, at some cost in per-token overhead.
- The serving layer:
lmdeploy serve api_serverfor an OpenAI-style REST server, and thepipeline()function for offline batch inference inside your own process. - The lite toolkit:
lmdeploy lite auto_awqand related commands that produce 4-bit weight checkpoints offline.
You choose the engine with --backend turbomind or --backend pytorch on the command line, or by passing a TurbomindEngineConfig or PytorchEngineConfig in Python. If TurboMind does not support a model architecture, LMDeploy can fall back to the PyTorch engine. Read the startup log to see which engine actually loaded, because the two have different quantization options and different performance.
Architecture and request flow
Follow one chat request. The client posts messages to /v1/chat/completions. The server applies the model's chat template, turning the message list into a token sequence with the right role markers, and hands it to the engine. The scheduler admits the sequence when enough KV blocks are free to hold its prompt. The first step is prefill: every prompt token passes through the model in parallel and writes its keys and values into blocks. Then decode begins: each step produces one token per running sequence and appends one entry per layer to that sequence's last block, taking a new block when the current one fills.
The KV cache is blocked, the same idea as vLLM's PagedAttention. Memory is divided into fixed-size blocks, 64 tokens each by default (cache_block_seq_len for TurboMind, block_size for the PyTorch engine), and each sequence keeps a table of the blocks it owns. Blocks need not be contiguous, so fragmentation stays low and a finished sequence's blocks return to the pool at once. When the pool runs out mid-decode, the scheduler has to stop some sequences and recompute or resume them later. Users see this as sudden latency spikes, not as errors.
Both engines batch at the step level, so a short request does not wait for a long one to finish. Throughput comes from keeping the batch full. Latency depends on how many sequences share each step and how often a long prefill stalls the decoders behind it.
KV cache memory: a worked sizing example
The most important setting is cache_max_entry_count. Its default, 0.8, means 80% of the GPU memory that is free after the weights load becomes KV cache. It does not mean 80% of total memory, a difference that has changed between releases and still confuses people. The KV cache decides how many tokens can be resident at once, and that decides your concurrency.
Work it out for Llama 3 8B: 32 layers, 8 key-value heads (grouped-query attention) and head dimension 128. One token stores a key and a value per layer per KV head:
bytes_per_token = 2 (K and V) * 32 layers * 8 kv_heads * 128 dims * 2 bytes (fp16)
= 131,072 bytes = 128 KiB
bytes_per_block = 64 tokens * 128 KiB = 8 MiBNow put it on a 24 GB card, such as an L4, with roughly 22 GiB usable:
| Configuration | Weights | Free after load | KV pool (x 0.8) | Resident tokens |
|---|---|---|---|---|
| fp16 weights, fp16 KV | ~15 GiB | ~5 GiB | ~4 GiB | ~32,800 (512 blocks) |
| AWQ 4-bit weights, fp16 KV | ~5.5 GiB | ~15.5 GiB | ~12.4 GiB | ~101,000 (~1,580 blocks) |
| AWQ 4-bit weights, int8 KV | ~5.5 GiB | ~15.5 GiB | ~12.4 GiB | ~200,000 |
These are planning numbers. Activation workspace, CUDA graphs and the runtime all take memory, so read the real block count from the startup log. The shape of the result still holds. If the average active sequence holds 4,000 tokens, the fp16 setup serves about 8 sequences at once, AWQ about 25, and AWQ plus int8 KV about 50. On this card, quantizing the weights matters less for raw speed than for the memory it frees for the cache.
session_len caps prompt plus output for one sequence. Set it to what your product needs, not to the model's advertised maximum. One 128k-token request can take the whole pool and stall everyone else.
Quantization: AWQ weights and KV cache
Weight quantization (W4A16). lmdeploy lite auto_awq applies AWQ (activation-aware weight quantization). It runs a calibration set through the model, finds the weight channels that matter most to activations, scales them to protect them, then stores weights as 4-bit integers in groups of 128 with a scale per group. At run time the kernels dequantize weights to fp16 on the fly, and activations stay in 16-bit, hence W4A16. Decode is mostly limited by memory bandwidth, so reading a quarter of the bytes per weight speeds up each step as well as saving memory. The documentation lists 4-bit inference support for V100 (sm70), Turing (sm75), Ampere (sm80, sm86) and Ada (sm89). Check your version's list before planning an AWQ deployment on any other architecture.
KV cache quantization. quant_policy=8 stores keys and values as int8, and quant_policy=4 as int4. The quantization is online, asymmetric, and applied per head and per token, so no calibration step is needed. The documented results for Llama 2 7B are about 30% more requests per second with int8 KV and about 40% with int4 against an fp16 cache. Accuracy is essentially unchanged with int8, with a small loss with int4. The PyTorch engine also accepts 16 and 17 for fp8 and fp8_e5m2 KV.
In practice, int8 KV is the safe default when you need concurrency. With int4 KV, test on long-context retrieval and on exact-copy tasks such as code edits and quoting IDs, because low-precision keys degrade attention precision first.
Code: quantize, serve and call
Quantize once, offline, on a machine with the GPU type you will serve on:
lmdeploy lite auto_awq meta-llama/Meta-Llama-3-8B-Instruct \
--work-dir ./llama3-8b-awq \
--calib-samples 128 --calib-seqlen 2048 \
--w-bits 4 --w-group-size 128Serve it with an int8 KV cache and a bounded session length:
lmdeploy serve api_server ./llama3-8b-awq \
--backend turbomind --model-format awq \
--quant-policy 8 --cache-max-entry-count 0.8 \
--session-len 8192 --server-port 23333Any OpenAI client can call it. Ask the server for its model name instead of guessing:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:23333/v1", api_key="unused")
model = client.models.list().data[0].id
resp = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": "Explain a blocked KV cache in two sentences."}],
max_tokens=200, temperature=0.2,
)
print(resp.choices[0].message.content, resp.usage)For offline jobs, skip HTTP and use the pipeline in-process:
from lmdeploy import pipeline, TurbomindEngineConfig, GenerationConfig
engine = TurbomindEngineConfig(model_format="awq", quant_policy=8,
session_len=8192, cache_max_entry_count=0.8,
enable_prefix_caching=True, tp=1)
pipe = pipeline("./llama3-8b-awq", backend_config=engine)
gen = GenerationConfig(max_new_tokens=256, temperature=0.0)
for r in pipe(["Summarise: ...", "Classify: ..."], gen_config=gen):
print(r.text)
Prefix caching and tensor parallelism
Prefix caching. enable_prefix_caching is off by default. With it on, blocks whose token content matches an earlier request's prefix are reused instead of recomputed. A long shared system prompt, few-shot examples or a multi-turn chat then skip most of their prefill, and time to first token falls. Matching is per block of identical tokens, so a prefix that differs early gets little reuse. Put the stable text first and the per-request text last.
Tensor parallelism. tp splits each layer across GPUs. Use it when the weights do not fit on one card or when you need lower per-token latency, and keep the GPUs on NVLink, because every layer does an all-reduce. When the model already fits on one GPU, several independent single-GPU replicas behind a load balancer usually give more total throughput than one tensor-parallel instance, and a failure takes out only one replica.
Failure modes
These are the failures that show up most often in production:
- Out of memory at startup on a shared GPU. Another process holds memory, and LMDeploy's share of free memory is then too small for any useful pool, or the next process to start finds nothing left. Give each server its own GPU, or lower
cache_max_entry_counton purpose. - Wrong chat template. A locally renamed or fine-tuned model directory can be matched to the wrong template. The output is fluent but ignores the system prompt or never stops. Compare the tokenized prompt with what the model card specifies before going live.
- Pool exhaustion under long contexts. p99 latency jumps while GPU use looks healthy. The scheduler is evicting and recomputing sequences. Lower
session_len, add int8 KV or cap concurrency at the load balancer. - Silent engine fallback. You planned for TurboMind, the architecture was unsupported, and the PyTorch engine loaded with different speed and options. Assert the backend in your deploy check.
- Calibration mismatch. AWQ calibrated on generic web text can lose more quality on code, maths or another language. Calibrate on samples of your own traffic and evaluate on your own task set, not only on perplexity.
- Version drift. Defaults and flag names change between releases. Pin the package version and store the full launch command with the model artifact.
Operating it
Benchmark with your own traffic shape: the distribution of prompt and output lengths and the arrival rate, not a single fixed prompt. Measure time to first token, inter-token latency and requests per second at the concurrency you expect, and find the point where p99 latency breaks your SLO. That knee, not peak throughput, is your real capacity. Keep a load-balancer cap a little below it.
Roll out a new quantization or version as a canary. Send a few percent of traffic, compare task quality and latency with the old replica, then shift. Track KV block usage, queue depth and whatever eviction or preemption signal your version logs alongside GPU utilisation, because a full pool shows up first in queueing.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| TurboMind vs PyTorch engine | Lower overhead per step, mature AWQ and KV quant path | Narrower model list, new architectures arrive later |
| AWQ 4-bit weights | About 3x smaller weights, faster bandwidth-bound decode | Calibration step, some task-dependent quality loss |
| int8 KV | About 2x resident tokens, documented gain in requests per second | Small extra compute per step |
| int4 KV | About 4x resident tokens | Measurable loss on long-context and exact-copy tasks |
| Higher cache_max_entry_count | More concurrency | Less headroom for activations and other processes |
| Prefix caching | Faster first token on shared prefixes | Memory held by cached blocks, whole-block matching only |
Compared with vLLM, SGLang and TensorRT-LLM, the deciding questions are practical: does the engine support your exact model and GPU, does its quantization path fit your quality bar, and how does it do on your benchmark? Published head-to-head numbers age within months, so run your own.
What to do next
- Write down your model, GPU, target concurrency and length distribution, then compute KV bytes per token with the formula above.
- Start TurboMind with a pinned LMDeploy version, an explicit
--backendand an explicit--session-len. Record the block count it reports. - Try
--quant-policy 8and check quality on your own evaluation set. Only then consider AWQ weights, calibrated on your own data. - Turn on prefix caching if requests share long system prompts, and put stable text first.
- Load-test to the p99 knee and cap concurrency just below it.
- Add deploy checks for engine type, chat template and model name, and canary every change.
- Keep learning: how the main serving stacks compare, paged KV cache internals, continuous batching, SGLang and RadixAttention and TensorRT-LLM.