Triton Inference Server (renamed NVIDIA Dynamo Triton in March 2025; the open-source project and container names still say tritonserver) was built for classic inference: a request arrives, a model runs once, one response leaves. Large language models break all three assumptions. A generation request runs for hundreds of forward passes, its cost depends on prompt and output length, and users want tokens streamed as they are produced.
Triton handles this by stepping aside. For LLMs it acts as the server shell (protocols, model repository, health, metrics, multi-model hosting) and hands the generation loop to a backend that owns batching and the KV cache: the vLLM backend or the TensorRT-LLM backend. This article explains that split, walks a working vLLM deployment, shows the TensorRT-LLM knobs that decide throughput and memory, covers multi-GPU modes and the OpenAI-compatible frontend, and ends with failure modes and a checklist. General Triton mechanics such as dynamic batching and ensembles are covered in the Triton backends and scheduling deep dive; this page stays on what changes for LLMs.
Why LLMs change Triton's job
Start from first principles. A transformer generates one token per forward pass. Each pass needs the keys and values of every earlier token, which the server keeps in a KV cache in GPU memory. Throughput comes from running many sequences through each pass together, and because sequences finish at different times, the best servers re-form the batch at every step: finished sequences leave, queued ones join. This is continuous or in-flight batching, explained in continuous batching on GPUs.
Triton's own dynamic batcher cannot do this. It groups whole requests, runs the model once, and returns. So for LLMs the backend receives individual requests and runs its own scheduler inside the model instance. Triton's scheduler becomes a pass-through, and its dynamic batching settings mostly stop mattering.
The second change is the transaction policy. A model declared decoupled may send any number of responses per request, at any time, including after the request call returns. That is how token streaming works: each generated token, or small group of tokens, is a separate response on the same request. Streaming over gRPC uses the bidirectional stream API; over HTTP, Triton's generate extension offers /v2/models/<name>/generate for a single JSON reply and /v2/models/<name>/generate_stream for server-sent events. A client that calls a decoupled model through the plain non-streaming infer API gets an error, which is the first thing most people hit.
Choosing a backend
Three ways to run an LLM in Triton are in common use. Pick by what you already run elsewhere and how much build effort you will accept.
| Path | What runs the model | Build step | Configured in |
|---|---|---|---|
| vLLM backend | vLLM AsyncLLMEngine with PagedAttention | None; loads Hugging Face weights | model.json (engine arguments) |
| TensorRT-LLM, LLM API path | TensorRT-LLM PyTorch runtime | None; loads Hugging Face weights | model.yaml (LLM() arguments) |
| TensorRT-LLM, engine path | Compiled TensorRT engine | Build an engine per model, GPU type and parallel layout | config.pbtxt parameters |
The vLLM backend is the fastest route to a working server and supports what vLLM supports. The TensorRT-LLM backend's source now lives in the TensorRT-LLM repository under triton_backend. Its newer LLM API path removes the engine build; the engine path, served by the inflight_batcher_llm model set, still gives the most control. See TensorRT-LLM explained and vLLM on GPUs for the engines themselves.
Worked example: a streaming vLLM model
Here is a minimal vLLM deployment for an 8B instruct model. The model repository has one model directory, a config, and a version folder holding the engine arguments.
model_repository/
llama8b/
config.pbtxt
1/
model.jsonThe config names the backend and lets vLLM own the devices. KIND_MODEL is required for multi-GPU models, and the custom-metrics parameter turns on vLLM's latency histograms:
backend: "vllm"
instance_group [
{
count: 1
kind: KIND_MODEL
}
]
parameters: {
key: "REPORT_CUSTOM_METRICS"
value: { string_value: "true" }
}The model.json file is passed to the vLLM engine as its arguments:
{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"gpu_memory_utilization": 0.90,
"max_model_len": 8192,
"tensor_parallel_size": 1
}Start the server from the vLLM flavour of the container (tags follow <yy.mm>-vllm-python-py3) and send a request:
tritonserver --model-repository=/models
curl -s -X POST localhost:8000/v2/models/llama8b/generate \
-d '{"text_input": "Explain a KV cache in one sentence.",
"parameters": {"stream": false, "temperature": 0, "max_tokens": 64}}'For streaming from Python, use the gRPC client's stream API and register a callback; every token arrives as its own response:
import tritonclient.grpc as grpcclient
import numpy as np, json
def on_response(result, error):
if error:
print("error:", error)
else:
print(result.as_numpy("text_output")[0].decode(), end="", flush=True)
client = grpcclient.InferenceServerClient("localhost:8001")
client.start_stream(callback=on_response)
prompt = grpcclient.InferInput("text_input", [1], "BYTES")
prompt.set_data_from_numpy(np.array([b"Write a haiku about GPUs."], dtype=object))
stream = grpcclient.InferInput("stream", [1], "BOOL")
stream.set_data_from_numpy(np.array([True]))
sampling = grpcclient.InferInput("sampling_parameters", [1], "BYTES")
sampling.set_data_from_numpy(np.array(
[json.dumps({"temperature": 0.7, "max_tokens": 128}).encode()], dtype=object))
client.async_stream_infer("llama8b", [prompt, stream, sampling])
client.stop_stream() # waits for outstanding responsesTwo details save debugging time. First, the request shape (text_input, a stream flag and sampling parameters) belongs to the backend, not to Triton, and the TensorRT-LLM models use different names; read the model's config before writing clients. Second, gpu_memory_utilization is vLLM's share of the whole GPU. If you co-host an embedding model on the same card, lower it or the second model fails to load.
TensorRT-LLM: in-flight batching and the KV cache
With the LLM API path you set the model and its options in model.yaml, whose keys map directly to TensorRT-LLM's LLM() constructor, for example tensor_parallel_size: 4. With the engine path the same decisions live as parameters in the tensorrt_llm model's config.pbtxt, normally filled from the template files the repository ships:
parameters: { key: "batching_strategy" value: { string_value: "inflight_fused_batching" } }
parameters: { key: "decoupled_mode" value: { string_value: "true" } }
parameters: { key: "kv_cache_free_gpu_mem_fraction" value: { string_value: "0.9" } }
parameters: { key: "enable_kv_cache_reuse" value: { string_value: "true" } }
parameters: { key: "batch_scheduler_policy" value: { string_value: "guaranteed_no_evict" } }
parameters: { key: "max_queue_delay_microseconds" value: { string_value: "0" } }batching_strategy:inflight_fused_batchingenables iteration-level batching;V1disables it and should only appear in comparisons.decoupled_modemust be true for any request that sets the stream flag.kv_cache_free_gpu_mem_fraction(default 0.9) is the share of GPU memory left after the weights load that the KV cache may claim.max_tokens_in_paged_kv_cachecaps it in tokens instead.enable_kv_cache_reusereuses cached blocks for shared prefixes such as a long system prompt.batch_scheduler_policy:max_utilizationpacks greedily and may pause sequences when memory runs out;guaranteed_no_evictadmits a request only if its whole generation is known to fit.exclude_input_in_outputdefaults to false, so responses echo the prompt unless you change it.
Worked sizing. Llama 3.1 8B has 32 layers, 8 KV heads and a head dimension of 128. In FP16 each token needs 2 (K and V) x 32 x 8 x 128 x 2 bytes = 131,072 bytes, or 128 KiB. On an 80 GB GPU the weights take about 16 GB; assume about 62 GB is free after load and runtime overhead. At a fraction of 0.9 the cache gets about 56 GB, or roughly 56e9 / 131,072, about 427,000 tokens. If the average request holds 2,000 tokens of prompt plus output, about 210 sequences fit at once. Under guaranteed_no_evict the scheduler reserves each request's maximum output length, so a client that always sends max_tokens: 4096 cuts concurrency sharply even if answers are short. Cap maximum output tokens at the gateway.
Multiple GPUs: tensor parallelism, leader and orchestrator modes
Models too large for one GPU are split with tensor parallelism. With vLLM, set tensor_parallel_size in model.json and keep KIND_MODEL so Triton does not try to place one instance per GPU; GPU_DEVICE_IDS in the config pins a model to specific GPUs and must list exactly as many devices as the parallel world size.
The TensorRT-LLM backend coordinates GPUs with MPI and has two modes. In leader mode one Triton process runs per GPU; rank 0 serves requests and the others only run their share of the model. It suits Slurm because it does not use MPI_Comm_spawn. In orchestrator mode one Triton process spawns a worker per GPU each model needs. This makes hosting several TensorRT-LLM models in one server simple, but it requires an MPI world size of one, may not work under Slurm, and is single-node only. Multi-node deployments use leader mode.
Separately, Triton's own instance count still matters for small models: two instances of a 1B model on one 80 GB GPU each run their own scheduler and KV cache, which can beat one instance with a bigger batch when prompts are short and requests are bursty. Measure both.
The OpenAI-compatible frontend
Most application code speaks the OpenAI API rather than the KServe v2 protocol. Triton ships an OpenAI-compatible frontend, a Python process that embeds the server and exposes /v1/chat/completions, /v1/completions, /v1/models and a metrics route on port 9000. It needs a tokenizer to apply the model's chat template:
python3 openai_frontend/main.py \
--model-repository /models \
--tokenizer meta-llama/Llama-3.1-8B-Instruct
curl -s localhost:9000/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model": "llama8b", "stream": true,
"messages": [{"role": "user", "content": "Name three uses of a KV cache."}]}'Constraints worth knowing: TensorRT-LLM models work only in orchestrator mode (set TRTLLM_ORCHESTRATOR=1), leader mode is unsupported, and tool calling needs a parser flag for the model family. Treat the frontend as a fast-moving component and test the exact fields your clients send, especially for tools and structured output.
Metrics that matter for LLMs
Triton exposes Prometheus metrics on port 8002. Its general series, such as nv_inference_request_success, nv_inference_queue_duration_us and GPU memory gauges, describe whole requests, which hides what matters for LLMs. With the vLLM backend's custom metrics enabled you also get series prefixed vllm:, including vllm:time_to_first_token_seconds, vllm:time_per_output_token_seconds and token counters. Build dashboards on these:
- Time to first token (TTFT) p50 and p99, which rise with queueing and long prompts.
- Time per output token, which rises with batch size and KV-cache pressure.
- Prompt and generation tokens per second, the real throughput unit.
- Queue depth and rejected requests, the signal to add replicas.
Set SLOs per token, not per request: a 20-token answer and a 2,000-token answer should not share a latency target. Comparing LLM serving stacks covers how engines differ on these axes.
Failure modes
- Non-streaming client against a decoupled model. The infer call fails. Use the generate endpoint, generate_stream or the gRPC stream API.
- Out of memory at load. Two models each assuming they own 90 percent of the GPU. Budget memory explicitly per model.
- Throughput collapses under long outputs. Clients send huge maximum token counts and a no-evict scheduler reserves for all of them. Clamp at the gateway.
- Engine and GPU mismatch. An engine-path model built for one GPU type or parallel layout will not load on another. Tie engine artifacts to GPU SKU and container tag in their names.
- Prompt echoed in output.
exclude_input_in_outputleft at false; clients see the prompt twice. - Slow cold start. Weight download and CUDA graph capture can take minutes. Bake weights into a volume, set readiness on
/v2/health/ready, and do not route traffic before it passes. - Orphaned generations. A client disconnects but the sequence keeps decoding. Check that your client cancels, and watch for generation tokens with no matching delivered responses.
Trade-offs
Triton earns its place when one fleet serves several model types (an LLM, an embedding model, a reranker, a vision model) behind one protocol, one metrics format and one deployment workflow, or when you want TensorRT-LLM's performance with a production server around it. The costs are a second layer of configuration over the engine, backend-specific request shapes, and a lag between engine releases and the containers that package them. If you serve one LLM and nothing else, running vLLM's or TensorRT-LLM's own OpenAI server directly is simpler and usually picks up engine features sooner. See LLM serving architecture for where the server sits in a full stack.
What to do next
- Pick the path: vLLM backend for speed of setup, TensorRT-LLM LLM API for NVIDIA-tuned performance without builds, engine path only when you need its extra control.
- Write the model repository, set decoupled streaming, and confirm one generate and one streaming request work.
- Compute the KV-cache budget per token for your model and decide the memory fraction and maximum output tokens from it.
- Choose the scheduler policy deliberately and load test with your real prompt and output length distribution.
- Enable custom metrics and build TTFT, time per output token and tokens per second dashboards before launch.
- If clients use the OpenAI API, run the OpenAI frontend and test chat templates, streaming, stop sequences and tools end to end.
- Pin container tags, weights and engines together and roll out with a canary model version.