LMDeploy is an open-source toolkit from the InternLM team (Shanghai AI Laboratory) for compressing and serving large language models on GPUs. It belongs to the same family as vLLM, SGLang and TensorRT-LLM: it takes a Hugging Face checkpoint, keeps many requests in flight on the GPU at once, and exposes an OpenAI-compatible HTTP API. Its distinguishing parts are a C++/CUDA inference engine called TurboMind, a second engine written in PyTorch for wider model coverage, and a built-in quantization toolkit for 4-bit weights and 8-bit or 4-bit KV caches.

This page explains how a request moves through LMDeploy, how it decides how much GPU memory goes to the KV cache, what its quantization options really trade, and the configuration mistakes that cost teams days. It ends with a worked sizing example for an 8B model on a 24 GB card and a deployment checklist. Flags and defaults below were checked against the LMDeploy documentation in October 2026; the project moves quickly, so confirm them against the version you install.

What LMDeploy is made of

LMDeploy ships as one Python package, lmdeploy, with four pieces you will touch:

  • TurboMind, the C++/CUDA engine. It runs a persistent batch loop: requests join and leave the running batch at every decoding step instead of waiting for a whole batch to finish. Other projects call this continuous or in-flight batching.
  • The PyTorch engine, written mostly in Python with custom kernels. It supports more model architectures and lands new models sooner, at some cost in per-token overhead.
  • The serving layer: lmdeploy serve api_server for an OpenAI-style REST server, and the pipeline() function for offline batch inference inside your own process.
  • The lite toolkit: lmdeploy lite auto_awq and related commands that produce 4-bit weight checkpoints offline.

You choose the engine with --backend turbomind or --backend pytorch on the command line, or by passing a TurbomindEngineConfig or PytorchEngineConfig in Python. If TurboMind does not support a model architecture, LMDeploy can fall back to the PyTorch engine. Read the startup log to see which engine actually loaded, because the two have different quantization options and different performance.

Architecture and request flow

LMDeploy: one front end, two engines, one paged KV pool per GPUclientsOpenAI-style HTTPapi_serverdefault port 23333pipeline APIin-process Pythonchat template+ tokenizerTurboMind engineC++/CUDA, persistent batchPyTorch enginePython, broader model listbackend=turbomindbackend=pytorchblocked KV cachefixed-size blocks, share of FREE memory, optional int8/int4 KVlite toolkitAWQ W4A16 offline4-bit weightsWeights, quantization and KV policy are fixed at startup; scheduling and block allocation happen per step.
Request path through LMDeploy. The engine choice, weight format and KV policy are fixed when the process starts.

Follow one chat request. The client posts messages to /v1/chat/completions. The server applies the model's chat template, turning the message list into a token sequence with the right role markers, and hands it to the engine. The scheduler admits the sequence when enough KV blocks are free to hold its prompt. The first step is prefill: every prompt token passes through the model in parallel and writes its keys and values into blocks. Then decode begins: each step produces one token per running sequence and appends one entry per layer to that sequence's last block, taking a new block when the current one fills.

The KV cache is blocked, the same idea as vLLM's PagedAttention. Memory is divided into fixed-size blocks, 64 tokens each by default (cache_block_seq_len for TurboMind, block_size for the PyTorch engine), and each sequence keeps a table of the blocks it owns. Blocks need not be contiguous, so fragmentation stays low and a finished sequence's blocks return to the pool at once. When the pool runs out mid-decode, the scheduler has to stop some sequences and recompute or resume them later. Users see this as sudden latency spikes, not as errors.

Both engines batch at the step level, so a short request does not wait for a long one to finish. Throughput comes from keeping the batch full. Latency depends on how many sequences share each step and how often a long prefill stalls the decoders behind it.

KV cache memory: a worked sizing example

The most important setting is cache_max_entry_count. Its default, 0.8, means 80% of the GPU memory that is free after the weights load becomes KV cache. It does not mean 80% of total memory, a difference that has changed between releases and still confuses people. The KV cache decides how many tokens can be resident at once, and that decides your concurrency.

Work it out for Llama 3 8B: 32 layers, 8 key-value heads (grouped-query attention) and head dimension 128. One token stores a key and a value per layer per KV head:

bytes_per_token = 2 (K and V) * 32 layers * 8 kv_heads * 128 dims * 2 bytes (fp16)
                = 131,072 bytes = 128 KiB
bytes_per_block = 64 tokens * 128 KiB = 8 MiB

Now put it on a 24 GB card, such as an L4, with roughly 22 GiB usable:

ConfigurationWeightsFree after loadKV pool (x 0.8)Resident tokens
fp16 weights, fp16 KV~15 GiB~5 GiB~4 GiB~32,800 (512 blocks)
AWQ 4-bit weights, fp16 KV~5.5 GiB~15.5 GiB~12.4 GiB~101,000 (~1,580 blocks)
AWQ 4-bit weights, int8 KV~5.5 GiB~15.5 GiB~12.4 GiB~200,000

These are planning numbers. Activation workspace, CUDA graphs and the runtime all take memory, so read the real block count from the startup log. The shape of the result still holds. If the average active sequence holds 4,000 tokens, the fp16 setup serves about 8 sequences at once, AWQ about 25, and AWQ plus int8 KV about 50. On this card, quantizing the weights matters less for raw speed than for the memory it frees for the cache.

session_len caps prompt plus output for one sequence. Set it to what your product needs, not to the model's advertised maximum. One 128k-token request can take the whole pool and stall everyone else.

Quantization: AWQ weights and KV cache

Weight quantization (W4A16). lmdeploy lite auto_awq applies AWQ (activation-aware weight quantization). It runs a calibration set through the model, finds the weight channels that matter most to activations, scales them to protect them, then stores weights as 4-bit integers in groups of 128 with a scale per group. At run time the kernels dequantize weights to fp16 on the fly, and activations stay in 16-bit, hence W4A16. Decode is mostly limited by memory bandwidth, so reading a quarter of the bytes per weight speeds up each step as well as saving memory. The documentation lists 4-bit inference support for V100 (sm70), Turing (sm75), Ampere (sm80, sm86) and Ada (sm89). Check your version's list before planning an AWQ deployment on any other architecture.

KV cache quantization. quant_policy=8 stores keys and values as int8, and quant_policy=4 as int4. The quantization is online, asymmetric, and applied per head and per token, so no calibration step is needed. The documented results for Llama 2 7B are about 30% more requests per second with int8 KV and about 40% with int4 against an fp16 cache. Accuracy is essentially unchanged with int8, with a small loss with int4. The PyTorch engine also accepts 16 and 17 for fp8 and fp8_e5m2 KV.

In practice, int8 KV is the safe default when you need concurrency. With int4 KV, test on long-context retrieval and on exact-copy tasks such as code edits and quoting IDs, because low-precision keys degrade attention precision first.

Code: quantize, serve and call

Quantize once, offline, on a machine with the GPU type you will serve on:

lmdeploy lite auto_awq meta-llama/Meta-Llama-3-8B-Instruct \
    --work-dir ./llama3-8b-awq \
    --calib-samples 128 --calib-seqlen 2048 \
    --w-bits 4 --w-group-size 128

Serve it with an int8 KV cache and a bounded session length:

lmdeploy serve api_server ./llama3-8b-awq \
    --backend turbomind --model-format awq \
    --quant-policy 8 --cache-max-entry-count 0.8 \
    --session-len 8192 --server-port 23333

Any OpenAI client can call it. Ask the server for its model name instead of guessing:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:23333/v1", api_key="unused")
model = client.models.list().data[0].id
resp = client.chat.completions.create(
    model=model,
    messages=[{"role": "user", "content": "Explain a blocked KV cache in two sentences."}],
    max_tokens=200, temperature=0.2,
)
print(resp.choices[0].message.content, resp.usage)

For offline jobs, skip HTTP and use the pipeline in-process:

from lmdeploy import pipeline, TurbomindEngineConfig, GenerationConfig

engine = TurbomindEngineConfig(model_format="awq", quant_policy=8,
                               session_len=8192, cache_max_entry_count=0.8,
                               enable_prefix_caching=True, tp=1)
pipe = pipeline("./llama3-8b-awq", backend_config=engine)
gen = GenerationConfig(max_new_tokens=256, temperature=0.0)
for r in pipe(["Summarise: ...", "Classify: ..."], gen_config=gen):
    print(r.text)

Prefix caching and tensor parallelism

Prefix caching. enable_prefix_caching is off by default. With it on, blocks whose token content matches an earlier request's prefix are reused instead of recomputed. A long shared system prompt, few-shot examples or a multi-turn chat then skip most of their prefill, and time to first token falls. Matching is per block of identical tokens, so a prefix that differs early gets little reuse. Put the stable text first and the per-request text last.

Tensor parallelism. tp splits each layer across GPUs. Use it when the weights do not fit on one card or when you need lower per-token latency, and keep the GPUs on NVLink, because every layer does an all-reduce. When the model already fits on one GPU, several independent single-GPU replicas behind a load balancer usually give more total throughput than one tensor-parallel instance, and a failure takes out only one replica.

Failure modes

These are the failures that show up most often in production:

  • Out of memory at startup on a shared GPU. Another process holds memory, and LMDeploy's share of free memory is then too small for any useful pool, or the next process to start finds nothing left. Give each server its own GPU, or lower cache_max_entry_count on purpose.
  • Wrong chat template. A locally renamed or fine-tuned model directory can be matched to the wrong template. The output is fluent but ignores the system prompt or never stops. Compare the tokenized prompt with what the model card specifies before going live.
  • Pool exhaustion under long contexts. p99 latency jumps while GPU use looks healthy. The scheduler is evicting and recomputing sequences. Lower session_len, add int8 KV or cap concurrency at the load balancer.
  • Silent engine fallback. You planned for TurboMind, the architecture was unsupported, and the PyTorch engine loaded with different speed and options. Assert the backend in your deploy check.
  • Calibration mismatch. AWQ calibrated on generic web text can lose more quality on code, maths or another language. Calibrate on samples of your own traffic and evaluate on your own task set, not only on perplexity.
  • Version drift. Defaults and flag names change between releases. Pin the package version and store the full launch command with the model artifact.

Operating it

Benchmark with your own traffic shape: the distribution of prompt and output lengths and the arrival rate, not a single fixed prompt. Measure time to first token, inter-token latency and requests per second at the concurrency you expect, and find the point where p99 latency breaks your SLO. That knee, not peak throughput, is your real capacity. Keep a load-balancer cap a little below it.

Roll out a new quantization or version as a canary. Send a few percent of traffic, compare task quality and latency with the old replica, then shift. Track KV block usage, queue depth and whatever eviction or preemption signal your version logs alongside GPU utilisation, because a full pool shows up first in queueing.

Trade-offs

ChoiceGainCost
TurboMind vs PyTorch engineLower overhead per step, mature AWQ and KV quant pathNarrower model list, new architectures arrive later
AWQ 4-bit weightsAbout 3x smaller weights, faster bandwidth-bound decodeCalibration step, some task-dependent quality loss
int8 KVAbout 2x resident tokens, documented gain in requests per secondSmall extra compute per step
int4 KVAbout 4x resident tokensMeasurable loss on long-context and exact-copy tasks
Higher cache_max_entry_countMore concurrencyLess headroom for activations and other processes
Prefix cachingFaster first token on shared prefixesMemory held by cached blocks, whole-block matching only

Compared with vLLM, SGLang and TensorRT-LLM, the deciding questions are practical: does the engine support your exact model and GPU, does its quantization path fit your quality bar, and how does it do on your benchmark? Published head-to-head numbers age within months, so run your own.

What to do next

  1. Write down your model, GPU, target concurrency and length distribution, then compute KV bytes per token with the formula above.
  2. Start TurboMind with a pinned LMDeploy version, an explicit --backend and an explicit --session-len. Record the block count it reports.
  3. Try --quant-policy 8 and check quality on your own evaluation set. Only then consider AWQ weights, calibrated on your own data.
  4. Turn on prefix caching if requests share long system prompts, and put stable text first.
  5. Load-test to the p99 knee and cap concurrency just below it.
  6. Add deploy checks for engine type, chat template and model name, and canary every change.
  7. Keep learning: how the main serving stacks compare, paged KV cache internals, continuous batching, SGLang and RadixAttention and TensorRT-LLM.
Key takeaway: LMDeploy serves LLMs with a step-level batching engine over a blocked KV cache. The settings that matter most are the KV pool, a share of free memory after the weights load; the session length; and the quantization policy. Size the pool from bytes per token, prefer int8 KV and calibrated AWQ when you need concurrency, pin versions and engine choice, and set capacity from your own p99 latency, not peak throughput.