TensorRT-LLM is NVIDIA's open-source library for running large language models fast on NVIDIA GPUs. It combines hand-tuned attention and GEMM kernels, an in-flight batching scheduler, a paged KV cache, low-precision formats such as FP8 and NVFP4, and multi-GPU parallelism, behind a Python LLM API and an OpenAI-compatible server called trtllm-serve.
Much of what is written about it is out of date. For years the workflow was: convert a checkpoint, compile a TensorRT engine with trtllm-build, then serve the engine. The 1.0 release made a PyTorch-based runtime the stable default, and the 1.2 release removed the TensorRT backend entirely. This article explains the runtime as it is now, from first principles: where GPU memory goes, how the scheduler fills each forward pass, which knobs matter, and how to size and operate a deployment.
What changed, and why it matters
The original design treated a model like a C++ program: you compiled it ahead of time for one GPU type, one precision, one parallel layout and fixed maximum shapes, and got a serialized engine. That delivered excellent kernels but made every change, such as a new batch limit, a new model architecture or a new GPU, a rebuild.
According to the project's release notes, 1.0 promoted the PyTorch backend to the default for both the LLM API and trtllm-serve, and declared the LLM API stable. The 1.2 notes list the TensorRT backend as removed: LLM(backend="tensorrt") raises an error, and trtllm-build, trtllm-refit, trtllm-prune and the per-model convert_checkpoint.py scripts are gone. A Hugging Face checkpoint now loads directly.
The practical consequences: tutorials that start with a build step describe a release line you should be migrating off; if you run prebuilt engines, you are pinned to a pre-1.2 version; and the speed now comes from the custom kernels, CUDA graphs and scheduler inside a PyTorch model definition rather than from a compiled graph. The name stayed, the architecture changed.
The architecture
Requests arrive either through your own Python process calling the LLM API or over HTTP to trtllm-serve, which exposes OpenAI-style completion and chat endpoints. Both feed the same executor. The scheduler decides, every iteration, which requests take part in the next forward pass and how many tokens each contributes. The KV cache manager owns GPU memory for attention keys and values in fixed-size blocks. The model itself is a PyTorch module whose hot paths call optimized kernels, split across GPUs by tensor, pipeline or expert parallelism.
First principles: where the GPU memory goes
An LLM request has two phases. Prefill processes the whole prompt in one pass and is compute-bound. Decode then produces one token per pass per request; each step reads all the weights and the request's entire KV cache, so decode is memory-bandwidth-bound. Batching many requests into one decode step is the main lever for throughput, because the weights are read once for the whole batch.
The limit on batch size is KV cache memory. Each token of each active request stores a key and a value vector in every layer:
def kv_bytes_per_token(layers, kv_heads, head_dim, bytes_per_elem):
# K and V, per layer, per KV head
return 2 * layers * kv_heads * head_dim * bytes_per_elem
# Llama-3-8B: 32 layers, 8 KV heads (grouped-query attention), head_dim 128
bf16 = kv_bytes_per_token(32, 8, 128, 2) # 131072 bytes = 128 KiB
fp8 = kv_bytes_per_token(32, 8, 128, 1) # 65536 bytes = 64 KiB
kv_pool = 55 * 2**30 # bytes left for KV on one 80 GB GPU
print(kv_pool // bf16) # ~450k tokens in BF16
print(kv_pool // fp8) # ~900k tokens in FP8On an 80 GB H100, Llama-3-8B's BF16 weights take about 16 GB. After activations, CUDA graphs and workspace, suppose 55 GB remains for KV. That is roughly 450,000 tokens of cache in BF16, or about 55 concurrent requests at a full 8,192-token context, and twice that with an FP8 KV cache. Almost every tuning decision in this article moves one of these numbers. For a longer treatment of the block allocator, see paged KV cache architecture.
In-flight batching and the scheduler
Static batching waits for a batch to fill and runs it until the longest request finishes, idling slots as short requests end. In-flight batching, also called continuous batching, rebuilds the batch every iteration: finished requests leave, waiting requests join, and prefill work can be mixed with decode work in the same pass.
# Conceptual model of one scheduler iteration (not TensorRT-LLM source)
def schedule(running, waiting, max_batch_size, max_num_tokens, kv):
batch, budget = [], max_num_tokens
for r in running: # decoding requests: 1 token each
if len(batch) == max_batch_size or budget == 0:
break
if not kv.ensure_block_for_next_token(r):
kv.pause(r) # out of blocks: pause, retry later
continue
batch.append((r, 1)); budget -= 1
for r in waiting: # new requests: prefill, maybe chunked
if len(batch) == max_batch_size or budget == 0:
break
reused = kv.match_prefix(r.prompt) # block reuse skips cached prefix
chunk = min(len(r.prompt) - reused, budget)
if not kv.allocate(r, chunk):
break
batch.append((r, chunk)); budget -= chunk
return batch # one forward pass over all of itTwo limits bound each iteration. max_batch_size caps how many requests are in the pass. max_num_tokens caps the total tokens processed, which is what really bounds activation memory and step time. A long prompt that exceeds the remaining token budget can be split with chunked prefill, so one 30,000-token prompt does not stall every decoding user for a whole pass. The sketch above is a model of the behaviour, not the library's code; the general mechanism is covered in continuous batching.
The trade-off is latency against throughput. A larger token budget raises throughput but lengthens each decode step, which users feel as slower inter-token latency. Tune it with a load test at your real prompt and output lengths, not a single-prompt benchmark.
The paged KV cache and block reuse
The KV cache manager splits the reserved memory into blocks of a fixed number of tokens and gives each request a list of blocks, so memory is never reserved for tokens that have not been generated. The main settings live in KvCacheConfig:
free_gpu_memory_fraction: the share of GPU memory still free after loading weights that goes to the KV pool; the documented default is 0.9. Lower it if other allocations run out of memory.max_tokens: an absolute cap; the pool uses the smaller of this and the fraction.enable_block_reuse: on by default. Requests that share a prompt prefix, such as a long system prompt or a document asked several questions, reuse the cached blocks instead of recomputing prefill.host_cache_size: bytes of host memory for offloading evicted blocks, 0 by default. Offloaded blocks can be brought back more cheaply than recomputing a long prefix.
Block reuse only helps when prefixes are byte-identical at the token level. Put stable content first: system prompt, tools, then documents, then the user turn. A timestamp at the top of the system prompt defeats reuse for every request.
Quantization: fewer bytes per weight and per cached token
Quantization cuts both weight memory and bandwidth per decode step. The documentation highlights FP8 on Hopper and later GPUs, NVFP4 on Blackwell and INT4 AWQ as weight-only quantization. The usual pattern is to quantize offline, typically with NVIDIA's Model Optimizer, and serve the resulting checkpoint; NVIDIA publishes pre-quantized checkpoints, such as FP8 variants, that load by name.
Weights and KV cache are separate decisions. FP8 weights halve weight memory and speed up GEMMs on hardware with FP8 tensor cores. An FP8 KV cache halves cache memory, doubling the concurrent tokens in the example above. Weight-only INT4 shrinks memory most but keeps activations in higher precision, so it helps small-batch decode more than large-batch throughput.
Quantization is an accuracy change, not a free speedup. Evaluate the quantized model on your own task set, including long-context and tool-calling cases, before switching traffic.
Parallelism across GPUs
| Strategy | Splits | Communication | Use when |
|---|---|---|---|
| Tensor parallel (TP) | Each layer's matrices across GPUs | All-reduce every layer | Model does not fit on one GPU; stay inside one NVLink domain |
| Pipeline parallel (PP) | Layers into stages | Activations between stages | Spanning nodes where all-reduce per layer is too slow |
| Expert parallel (EP) | MoE experts across GPUs | All-to-all token routing | Mixture-of-experts models with many experts |
TP lowers per-token latency but costs an all-reduce in every layer, so it belongs on GPUs joined by NVLink. If a model fits on one GPU with room for a useful KV pool, several independent single-GPU replicas usually beat one TP group on throughput per GPU. The LLM API takes tensor_parallel_size; trtllm-serve takes --tp_size, --pp_size and --ep_size.
Speculative decoding
Decode leaves compute idle because it is bandwidth-bound. Speculative decoding spends that compute: a cheap drafter proposes several tokens and the target model verifies them in one pass, keeping the prefix it agrees with. TensorRT-LLM supports EAGLE, including EAGLE-3, multi-token prediction (MTP) for models trained with it, and NGram drafting from the prompt itself, configured through classes such as EagleDecodingConfig, MTPDecodingConfig and NGramDecodingConfig. Check the API reference for each class's fields for your version.
Gains depend on acceptance rate. They are largest at small batch sizes and fall as the batch grows, since compute stops being idle. Measure at production concurrency; see speculative decoding architecture for the theory.
Running it: the LLM API and trtllm-serve
For offline batch jobs, call the LLM API directly. Keyword arguments beyond the basics are forwarded to the runtime's argument class, so names can shift between releases; check them against the reference for your version.
from tensorrt_llm import LLM, SamplingParams
from tensorrt_llm.llmapi import KvCacheConfig
kv = KvCacheConfig(
free_gpu_memory_fraction=0.85, # share of free memory, after weights, for KV
enable_block_reuse=True, # reuse KV blocks for shared prompt prefixes
)
llm = LLM(
model="meta-llama/Llama-3.1-8B-Instruct", # HF id or local path, no engine
tensor_parallel_size=1,
kv_cache_config=kv,
max_batch_size=64,
max_num_tokens=8192,
)
params = SamplingParams(temperature=0.2, top_p=0.95, max_tokens=256)
for out in llm.generate(["Explain paged attention in two sentences."], params):
print(out.outputs[0].text)For online serving, trtllm-serve wraps the same runtime. Options that have no flag go in a YAML file passed with --config (also spelled --extra_llm_api_options):
# kv.yaml -- extra LLM API options; explicit CLI flags take precedence
kv_cache_config:
free_gpu_memory_fraction: 0.85
enable_block_reuse: true
$ trtllm-serve meta-llama/Llama-3.1-8B-Instruct \
--host 0.0.0.0 --port 8000 \
--tp_size 1 --max_batch_size 64 --max_num_tokens 8192 \
--config kv.yaml
# any OpenAI-compatible client
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")
r = client.chat.completions.create(
model="meta-llama/Llama-3.1-8B-Instruct",
messages=[{"role": "user", "content": "Hello"}], max_tokens=64)
print(r.choices[0].message.content)
Worked example: a support assistant on one H100
A team serves Llama-3.1-8B-Instruct for a support assistant. Traffic: a 1,500-token system prompt and retrieved context, a 300-token question, 250 output tokens, and a peak of 40 concurrent conversations. The target is under 50 ms per output token at peak.
KV per request is about 2,050 tokens, 256 MB in BF16, so 40 requests need about 10 GB; the 55 GB pool has plenty of headroom. The system prompt is shared, so block reuse removes most of its prefill cost. They set max_batch_size to 64, leaving room over peak, and max_num_tokens to 8,192.
The first load test shows inter-token latency spiking whenever several users paste long logs. Chunked prefill spreads those prompts across iterations; lowering max_num_tokens to 4,096 steadies decode latency at a small throughput cost. Moving to an FP8 checkpoint after an accuracy check frees memory they do not need yet, so they spend it on a second model replica rather than a bigger batch.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| ValueError on startup after an upgrade | Code still passes backend="tensorrt" or an engine path | Load the HF checkpoint directly; drop the backend argument |
| Out of memory during warm-up | KV fraction plus activations exceed the GPU | Lower free_gpu_memory_fraction or max_num_tokens |
| Throughput plateaus well below expectations | KV pool full, requests queued | FP8 KV cache, shorter max context, or more replicas |
| Inter-token latency spikes | Large prefills sharing passes with decode | Chunked prefill, smaller max_num_tokens |
| Prefix reuse has no effect | Prompts differ in their first tokens | Put stable content first; remove per-request headers |
| Quality drop after quantizing | Format too aggressive for the task | Evaluate per format; keep KV at higher precision |
| TP slower than one GPU | All-reduce over PCIe or across nodes | TP only within NVLink; use replicas or PP |
Choosing it, and operating it
TensorRT-LLM is the natural choice when you are on NVIDIA data-centre GPUs and want the newest kernels and formats, such as NVFP4 on Blackwell, as soon as they ship. The costs are NVIDIA lock-in, fast-moving APIs, and container images that must match your driver and CUDA stack. The comparison with vLLM, SGLang and TGI is in LLM serving stacks, and the wider optimization toolbox in LLM inference optimization.
In operation, pin the container image and the model revision together, load-test with production-shaped traffic before every upgrade, alert on queue depth and time to first token as well as GPU utilization, and keep a rollback image. Treat every upgrade as a performance change that needs measurement.
What to do next
- Check your version. If anything calls trtllm-build or loads an engine, plan the move to loading Hugging Face checkpoints.
- Compute KV bytes per token for your model and the concurrent tokens your GPU can hold, in BF16 and FP8.
- Collect real prompt and output length distributions and turn them into a load test.
- Start with defaults, then tune max_num_tokens and max_batch_size against time to first token and inter-token latency.
- Order prompts for prefix reuse, and confirm the gain in time to first token.
- Evaluate one quantized checkpoint on your own tasks before adopting it.
- Try speculative decoding only after measuring acceptance at your production concurrency.