An LLM endpoint can be up, returning 200s and burning GPUs at full power while users stare at a spinner for eight seconds. The request-rate and error-rate dashboards inherited from web services do not see it, because the things that hurt in LLM serving are time to the first token, the gap between tokens, a queue that grows behind a full KV cache, and preemptions that quietly restart work. This article is about measuring those things correctly: which metrics exist, where they come from, how to turn histograms into percentiles without lying to yourself, and how to keep label cardinality from taking down Prometheus.
We use names that we checked against current documentation: the vLLM Prometheus metrics, the OpenTelemetry GenAI semantic conventions (still marked Development, so names can change) and NVIDIA DCGM fields. For the arithmetic behind the latency numbers, read GPU inference latency and the TTFT ledger; this page is about observing them in production.
Four layers of telemetry
Useful LLM telemetry comes from four places, and each answers a different question. The client or gateway sees what the user experiences, including network time, retries and time spent in your own middleware. The inference engine sees why: queue time, prefill and decode time, batch size and KV cache pressure. The GPU layer says whether the hardware is the bottleneck or idle. The product layer turns token counts into cost and turns latency into goodput, the share of requests that met their target.
A common mistake is to pick one layer. Engine metrics alone miss a slow auth hop in the gateway. Gateway metrics alone tell you that p99 TTFT doubled, not that the cause was queueing behind long prompts. Measure at all four, keep the label set consistent (the same model name string everywhere) and you can walk from the symptom down to the cause in one dashboard.
The pipeline
Five latencies, and their metric names
Five latency quantities matter, and they are not interchangeable. Queue time is from arrival at the engine until the scheduler admits the request. Prefill time is the forward pass over the prompt. TTFT is arrival to first output token, so it contains queue plus prefill. Inter-token latency (ITL) is the gap between consecutive streamed tokens, one sample per gap. TPOT (time per output token) is one number per request: decode time divided by the number of tokens after the first. A request with 300 tokens contributes 299 ITL samples and one TPOT sample, so the two histograms weight long answers very differently.
vLLM exposes these as histograms: vllm:request_queue_time_seconds, vllm:request_prefill_time_seconds, vllm:time_to_first_token_seconds, vllm:inter_token_latency_seconds, vllm:request_time_per_output_token_seconds, vllm:request_decode_time_seconds and vllm:e2e_request_latency_seconds. Older releases used vllm:time_per_output_token_seconds for the per-token gap and vllm:gpu_cache_usage_perc for cache use; if your dashboards show no data after an upgrade, renamed metrics are the first thing to check against the engine's own /metrics output.
On the client side, the OpenTelemetry GenAI conventions (we checked v1.37) define gen_ai.client.operation.duration in seconds and gen_ai.client.token.usage in tokens, with gen_ai.operation.name and gen_ai.provider.name required (the provider attribute replaced the older gen_ai.system). The server-side conventions add gen_ai.server.time_to_first_token and gen_ai.server.time_per_output_token. The client set we checked has no TTFT metric, so record client-perceived TTFT under your own namespace rather than borrowing the server name.
Instrumenting the gateway
Here is a gateway wrapper that records duration, client-side TTFT and token usage for a streaming call. It uses an SDK View to set explicit bucket boundaries, because default histogram buckets are built for millisecond web calls and put most LLM latencies in two or three buckets.
import time
from opentelemetry import metrics
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.view import View, ExplicitBucketHistogramAggregation
LAT = [0.05, 0.1, 0.2, 0.4, 0.8, 1.2, 1.6, 2.4, 3.2, 4.8, 6.4, 9.6, 12.8, 25.6, 51.2]
provider = MeterProvider(views=[
View(instrument_name="gen_ai.client.operation.duration",
aggregation=ExplicitBucketHistogramAggregation(LAT)),
View(instrument_name="app.llm.client.ttft",
aggregation=ExplicitBucketHistogramAggregation(LAT)),
]) # add your OTLP metric reader here
metrics.set_meter_provider(provider)
meter = metrics.get_meter("llm.gateway")
duration = meter.create_histogram("gen_ai.client.operation.duration", unit="s")
ttft = meter.create_histogram("app.llm.client.ttft", unit="s")
tokens = meter.create_histogram("gen_ai.client.token.usage", unit="{token}")
def stream_chat(client, model, messages, route):
attrs = {"gen_ai.operation.name": "chat", "gen_ai.provider.name": "vllm",
"gen_ai.request.model": model, "app.route": route}
start = time.perf_counter()
first = chunk = None
try:
for chunk in client.stream(model=model, messages=messages):
if first is None and chunk.text:
first = time.perf_counter()
ttft.record(first - start, attrs)
yield chunk
usage = getattr(chunk, "usage", None) # last chunk carries usage
if usage:
tokens.record(usage.prompt_tokens, {**attrs, "gen_ai.token.type": "input"})
tokens.record(usage.completion_tokens, {**attrs, "gen_ai.token.type": "output"})
except Exception as e:
attrs = {**attrs, "error.type": type(e).__name__}
raise
finally:
now = time.perf_counter()
if first is None: # no token arrived: still a TTFT sample
ttft.record(now - start, attrs)
duration.record(now - start, attrs)Two details matter. The error.type attribute goes on the duration sample, and a call that fails before its first token still records a TTFT sample, so failed calls stay in the latency picture instead of vanishing. And app.route is a bounded label (a handful of product features), not a user or tenant ID; the cardinality section explains why.
Percentiles from histograms
Prometheus stores a histogram as cumulative bucket counters. A percentile is estimated at query time by finding the bucket where the target rank falls and interpolating linearly inside it. The correct pattern aggregates buckets first and takes the quantile last:
# p99 TTFT per model over 5 minutes, across all replicas
histogram_quantile(0.99,
sum by (le, model_name) (rate(vllm:time_to_first_token_seconds_bucket[5m])))
# share of requests with TTFT under 1 s (goodput for a 1 s target);
# only exact if 1.0 is a bucket boundary
sum by (model_name) (rate(vllm:time_to_first_token_seconds_bucket{le="1.0"}[5m]))
/ sum by (model_name) (rate(vllm:time_to_first_token_seconds_count[5m]))
# mean queue time, which is cheap and exact
sum(rate(vllm:request_queue_time_seconds_sum[5m]))
/ sum(rate(vllm:request_queue_time_seconds_count[5m]))Never average percentiles across replicas. If replica A has p99 of 0.4 s on 900 requests and replica B has p99 of 6 s on 100 requests, the average says 3.2 s and the fleet's real p99 is somewhere in B's tail. Summing buckets gives the right answer because buckets are counts and counts add.
Bucket boundaries set your error. If the buckets around your SLO are 1 s and 2.5 s and the true p99 is 1.1 s, interpolation can report anything in that range depending on how samples spread. Put a boundary exactly at each SLO threshold, keep boundaries roughly geometric elsewhere, and check the label names on your engine version (model_name is what vLLM uses at the time of writing). Prometheus native histograms reduce this problem with automatically chosen exponential buckets; if your Prometheus and client library support them, they are worth testing.
Saturation: queue, KV cache, preemption
Latency tells you users are hurting; saturation metrics tell you why and warn you first. The core engine gauges are vllm:num_requests_running, vllm:num_requests_waiting and vllm:kv_cache_usage_perc (a 0 to 1 fraction of KV blocks in use). The counter vllm:num_preemptions counts requests evicted from the batch because the cache ran out; as with every Prometheus counter, the exposed series carries a _total suffix, so query rate(vllm:num_preemptions_total[5m]).
The causal chain is usually the same. Long prompts or long generations fill the KV cache, kv_cache_usage_perc sits near 1, the scheduler stops admitting new requests, num_requests_waiting climbs, queue time rises, and TTFT follows a minute later. If the scheduler must evict running sequences, preemptions rise and the evicted work is recomputed, so throughput drops just when you need it. Alert on waiting requests and preemption rate, not on cache use alone: a full cache with an empty queue is a well-packed engine, which is what you paid for. For how blocks are allocated, see paged KV cache.
Prefix caching has its own pair of counters, vllm:prefix_cache_queries and vllm:prefix_cache_hits, both counted in tokens. Their rate ratio is the hit rate. A drop after a prompt template change is a cost regression that never shows up as an error.
What the GPU layer adds
NVIDIA's DCGM exporter publishes GPU fields to Prometheus. The one everyone graphs, DCGM_FI_DEV_GPU_UTIL, is the fraction of time at least one kernel was running. A GPU running one tiny kernel all the time reads 100 percent. For LLM serving, the profiling fields are more honest: DCGM_FI_PROF_SM_ACTIVE (how many SMs had work), DCGM_FI_PROF_PIPE_TENSOR_ACTIVE (tensor core pipe activity) and DCGM_FI_PROF_DRAM_ACTIVE (memory interface activity). Decode is usually memory-bound, so a decode-heavy pool shows high DRAM activity and modest tensor activity; that is expected, not a problem. Add DCGM_FI_DEV_FB_USED for framebuffer memory, plus power and XID error fields for hardware health. Profiling fields may need to be enabled in the exporter's counter list. Details are in DCGM for GPU fleets.
Cardinality budget
Cardinality is what kills LLM metrics stacks. Every unique combination of label values is a separate time series, and each histogram multiplies that by the bucket count plus two (sum and count). Work an example: 16 buckets plus 2 is 18 series per label set. With 6 models, 3 routes, 2 token types and 40 gateway pods, one histogram is 18 x 6 x 3 x 2 x 40 = 25,920 series. Add a tenant label with 2,000 tenants and it becomes about 52 million, which no ordinary Prometheus will hold.
| Label | Keep on metrics? | Where it goes instead |
|---|---|---|
| model, route, region, pool | yes, bounded and few | - |
| error.type, finish reason | yes, a small enum | - |
| tenant or customer ID | only top N, rest as other | traces, request logs, billing table |
| user ID, request ID, prompt hash | never | traces and logs |
| pod name | drop in recording rules | raw series with short retention |
Per-tenant cost belongs in a billing pipeline that reads token counts from request logs, not in metric labels. Use recording rules to aggregate away pod labels, keep the raw series for days and the aggregated series for months.
Worked incident: the long-prompt launch
A worked incident. At 14:05 an alert fires: p99 TTFT for the chat model is 4.2 s against a 1.5 s target. The gateway's gen_ai.client.operation.duration p99 is up too, so the problem is real and user-facing. On the engine, vllm:request_queue_time_seconds p99 is 3.6 s while prefill p99 is flat at 0.5 s, so the time is spent waiting, not computing. num_requests_waiting went from about 2 to 40 per replica, kv_cache_usage_perc is pinned at 0.98, and the preemption rate jumped from zero to 3 per second.
The token histograms explain it: vllm:request_prompt_tokens p90 went from 1,800 to 14,000 at 13:50, when a document-summarisation feature launched on the same pool. Long prompts hold many more KV blocks, fewer requests fit, the queue grows. DCGM tensor activity actually dropped, because fewer sequences fit in the cache and the effective batch shrank. The fix is routing the long-prompt route to its own pool (or enabling chunked prefill and a tighter max_num_seqs), not adding GPUs blindly. Without engine and token metrics, the team would have seen only slow responses and busy GPUs.
Failure modes
- Averaging percentiles across pods or time windows, which hides the tail you are paging on.
- Default buckets that put every request between 1 s and 10 s into one bucket, so p50 and p99 both interpolate to nonsense.
- Measuring TTFT only on success. Requests that time out before the first token disappear and the percentile improves as the outage worsens.
- Tenant or user IDs as labels, exploding series count until Prometheus runs out of memory.
- Trusting GPU_UTIL as efficiency, which reads 100 percent while the engine is preempting.
- Silent renames after upgrades: a dashboard on a removed metric shows a flat line, which looks healthy.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| More buckets | accurate percentiles | linear growth in series |
| Client-side metrics | true user experience | needs instrumented gateways and SDKs |
| Engine-only metrics | free with vLLM | misses network, auth and retries |
| 15 s scrape interval | catches short spikes | more samples to store |
| Native histograms | low error with few series | newer feature; check tooling support |
What to do next
- Curl the engine's
/metricsendpoint and list the names your version really exposes. - Add explicit buckets with a boundary at every latency SLO threshold, in both gateway and engine.
- Write recording rules for p50, p90, p99 TTFT and ITL per model using
sum by (le, ...). - Alert on
num_requests_waitingand preemption rate; wire them into SLO burn-rate alerts. - Add DCGM profiling fields and put tensor and DRAM activity next to batch size on one panel.
- Audit labels: remove anything unbounded and move per-tenant accounting to logs.
- Record duration for failed and timed-out requests and test that the alert fires during a fake outage.