Most LLM serving dashboards are built by picking every metric the inference engine exports and arranging the results in a grid. The outcome looks impressive and answers nothing. When someone asks why cost per request rose 30 percent last week, or whether slow responses come from queueing or from the GPUs, forty unrelated panels force them to rebuild the reasoning from scratch.

This page shows how to build a KPI dashboard as a metric tree. Each business number at the top decomposes into service indicators, each indicator into engine state, and engine state into hardware activity, so any movement at the top can be followed down to a cause in a few clicks. It covers which series feed each layer for a vLLM deployment with the NVIDIA DCGM exporter, the PromQL and recording rules behind each panel, a worked cost and goodput calculation, and the traps that make a dashboard confidently wrong. Alert rules are covered separately in LLM SLO burn-rate alerts; this page is about the screen people look at.

Advertisement

What a KPI dashboard is for

A KPI dashboard exists to support recurring decisions. For LLM serving there are about four: do we need more or fewer GPUs, is the service meeting its promise to users, is a recent change (new model, new engine version, new batching settings) better or worse, and where is money going. Every panel should map to one of those decisions. A panel that nobody would act on, whatever value it showed, belongs on a debugging page, not on the KPI page.

It also helps to separate audiences. Leadership reads the top layer weekly and cares about cost and quality trends. Service owners read the top two layers daily and care about SLO compliance. On-call engineers start at layer two during an incident and descend. A single dashboard can serve all three if it is ordered top to bottom by layer, with drill-down links between rows rather than one giant flat page.

The four-layer metric tree

One dashboard, four layers: every KPI must drill down to a causeLayer 1: business KPIscost per 1M output tokens, goodput, tokens served per day, requests per tenantLayer 2: service SLIsTTFT p95, inter-token latency p95, error ratio, share of requests inside SLOLayer 3: engine staterequests running and waiting, KV cache usage, preemptions, prefix cache hit ratioLayer 4: hardwareSM active, tensor pipe active, framebuffer used, power, XID errorsAPI gateway logstenant, status, tokensEngine /metricsvllm: histograms, gaugesDCGM exporterDCGM_FI_* fieldsSources at the bottom feed recording rules; panels read the rules, never raw high-cardinality series
The KPI tree. Each layer explains the one above it; the sources at the bottom feed recording rules that every panel reads.

Layer one holds business KPIs: cost per million output tokens, goodput (the share of requests or tokens served inside the latency SLO), tokens served per day and the split by tenant or product. These are the numbers that go into a quarterly review.

Layer two holds service level indicators measured at the boundary the user experiences: time to first token (TTFT), inter-token latency (ITL, the gap between streamed tokens), end-to-end latency for non-streaming calls, and the error ratio. Percentiles, never averages, because users feel the tail. The derivation of these quantities from prefill and decode work is in GPU inference latency from first principles.

Layer three holds engine state, the variables that explain layer two. A TTFT regression is almost always one of: requests waiting in the scheduler queue, prompts getting longer, KV cache running out so requests are preempted and recomputed, or a lower prefix-cache hit ratio. Each of these is a direct engine metric. Why KV cache capacity behaves the way it does is covered in paged KV cache.

Layer four holds hardware: how busy the streaming multiprocessors and tensor cores really are, memory used, power and error events. This layer answers whether the GPUs are saturated (scale out) or idle while latency is bad (a scheduling or configuration problem, not a capacity one).

Advertisement

Where each number comes from

For a vLLM deployment, the engine exposes Prometheus metrics on its HTTP server's /metrics path, labelled with model_name. The names below were checked against the current vLLM metrics documentation. Counters appear with a _total suffix when scraped, which is why the queries use vllm:num_preemptions_total even though the documentation lists vllm:num_preemptions. Engines change names between releases, so read your own /metrics output once before building panels.

LayerSeriesTypeNotes
2vllm:time_to_first_token_secondshistogramTTFT measured inside the engine; excludes gateway and network time
2vllm:inter_token_latency_secondshistogramgap between consecutive output tokens
2vllm:e2e_request_latency_secondshistogramwhole request inside the engine
3vllm:num_requests_running, vllm:num_requests_waitinggaugebatch size and queue depth
3vllm:kv_cache_usage_percgaugea fraction from 0 to 1 despite the name
3vllm:num_preemptionscounterrequests evicted for lack of KV blocks
3vllm:prefix_cache_hits, vllm:prefix_cache_queriescountertokens, not requests
1vllm:prompt_tokens_total, vllm:generation_tokens_totalcounterthroughput and cost denominators
4DCGM_FI_PROF_SM_ACTIVE, DCGM_FI_PROF_PIPE_TENSOR_ACTIVEgaugeprofiling fields from the DCGM exporter
4DCGM_FI_DEV_FB_USED, DCGM_FI_DEV_POWER_USAGEgaugememory in MiB, power in watts

Two things the engine cannot tell you come from the API gateway: which tenant or product sent the request, and requests that never reached the engine because of rate limiting, authentication failures or upstream timeouts. Per-tenant numbers belong in gateway logs or a low-cardinality gateway metric, never as a label on engine histograms. The hardware fields and how the DCGM host engine samples them are explained in the DCGM deep dive.

Recording rules: compute once, read everywhere

Panels should not run raw histogram queries. A histogram_quantile over thirty pods with forty buckets each is expensive, every panel that repeats it is slightly different, and the alert rule ends up computing something subtly unlike the panel. Recording rules fix all three: Prometheus evaluates each expression once per interval, stores the result as a new low-cardinality series, and panels and alerts read that series.

# prometheus rules file: llm-kpi.rules.yml
groups:
- name: llm-kpi
  interval: 30s
  rules:
  # Layer 2: latency percentiles. Aggregate buckets FIRST, then take the quantile.
  - record: model:ttft_seconds:p95_5m
    expr: histogram_quantile(0.95, sum by (le, model_name) (rate(vllm:time_to_first_token_seconds_bucket[5m])))
  - record: model:itl_seconds:p95_5m
    expr: histogram_quantile(0.95, sum by (le, model_name) (rate(vllm:inter_token_latency_seconds_bucket[5m])))

  # Layer 2: fraction of requests with TTFT at or under 1 s. Only exact if "1.0" is a real le boundary.
  - record: model:ttft_within_1s:ratio_5m
    expr: |
      sum by (model_name) (rate(vllm:time_to_first_token_seconds_bucket{le="1.0"}[5m]))
      /
      sum by (model_name) (rate(vllm:time_to_first_token_seconds_count[5m]))

  # Layer 1: throughput in tokens per second, prompt and generation kept apart.
  - record: model:generation_tokens:rate5m
    expr: sum by (model_name) (rate(vllm:generation_tokens_total[5m]))
  - record: model:prompt_tokens:rate5m
    expr: sum by (model_name) (rate(vllm:prompt_tokens_total[5m]))

  # Layer 3: engine pressure.
  - record: model:requests_waiting:max
    expr: max by (model_name) (vllm:num_requests_waiting)
  - record: model:kv_cache_usage:max
    expr: max by (model_name) (vllm:kv_cache_usage_perc)        # 0..1, not 0..100
  - record: model:preemptions:rate5m
    expr: sum by (model_name) (rate(vllm:num_preemptions_total[5m]))
  - record: model:prefix_cache_hit:ratio_5m
    expr: |
      sum by (model_name) (rate(vllm:prefix_cache_hits_total[5m]))
      /
      sum by (model_name) (rate(vllm:prefix_cache_queries_total[5m]))

  # Layer 4: real GPU activity, not "a kernel was running".
  - record: node:sm_active:avg
    expr: avg by (Hostname) (DCGM_FI_PROF_SM_ACTIVE)

The naming convention level:metric:operation records what was aggregated away and how. The critical detail is the order of operations in the percentile rules: rates of buckets are summed across pods by le first, and the quantile is taken last. That gives the true fleet-wide percentile. The opposite order, computing a p95 per pod and then averaging, produces a number with no statistical meaning that is usually lower than the real tail.

Goodput: the KPI that joins cost and quality

Throughput alone rewards a deployment for cramming in so many requests that each one is slow. Goodput counts only work delivered inside the SLO. The cleanest definition is per request: a request is good if its TTFT and its ITL were both within target. Computing that joint condition needs per-request data, which the engine histograms do not keep, so exact goodput comes from request logs at the gateway or engine.

A good approximation comes straight from the histogram buckets. The ratio of the le="1.0" bucket rate to the count rate is the share of requests whose TTFT was at or under one second. This is exact, not estimated, provided 1.0 is a real bucket boundary. Check the le values in your /metrics output; if your target falls between boundaries, use the nearest boundary below it. Do the same for ITL and report the lower share as a conservative goodput figure.

Worked example: one day of one model

Consider a 70B chat model served on 8 GPUs at a blended rate of 2.50 dollars per GPU hour. That rate is illustrative; use your own, including reserved capacity and idle headroom, because idle GPUs cost money too. During one day the generation counter increased by 155 million tokens and the TTFT histogram recorded 410,000 requests, of which 385,400 landed in the one-second bucket.

The daily GPU cost is 8 times 2.50 times 24, which is 480 dollars. Divided by 155 million output tokens that gives 3.10 dollars per million output tokens. TTFT compliance is 385,400 over 410,000, or 94.0 percent. If you treat the 6 percent of slow requests as wasted, the cost per million tokens delivered inside the SLO is 3.10 divided by 0.94, about 3.30 dollars. The script below reproduces the calculation against a live Prometheus.

"""Compute the two layer-1 KPIs from Prometheus for one model over one day."""
import requests

PROM = "http://prometheus:9090/api/v1/query"
MODEL = "llama-70b-chat"
GPU_HOUR_USD = 2.50          # your blended rate: reserved, on-demand and idle capacity included
GPUS = 8                     # GPUs allocated to this model, busy or not

def scalar(q):
    r = requests.get(PROM, params={"query": q}, timeout=10).json()
    res = r["data"]["result"]
    return float(res[0]["value"][1]) if res else 0.0

# increase() over 1d handles counter resets from pod restarts.
gen_tokens = scalar(f'sum(increase(vllm:generation_tokens_total{{model_name="{MODEL}"}}[1d]))')
requests_all = scalar(f'sum(increase(vllm:time_to_first_token_seconds_count{{model_name="{MODEL}"}}[1d]))')
requests_ok = scalar(f'sum(increase(vllm:time_to_first_token_seconds_bucket{{model_name="{MODEL}",le="1.0"}}[1d]))')

cost = GPUS * GPU_HOUR_USD * 24
cost_per_m = cost / (gen_tokens / 1e6) if gen_tokens else float("nan")
ttft_ok = requests_ok / requests_all if requests_all else float("nan")

print(f"generated tokens:            {gen_tokens:,.0f}")
print(f"cost per 1M output tokens:   ${cost_per_m:.2f}")
print(f"requests with TTFT <= 1 s:   {ttft_ok:.1%}")
print(f"cost per 1M tokens in SLO:   ${cost_per_m / ttft_ok:.2f}  (rough: assumes misses are average length)")

Now suppose the following week the cost rises to about 3.65 dollars. The tree tells you where to look. Layer one shows generation tokens fell 15 percent while GPU count stayed fixed, so the denominator shrank: 480 dollars over 131.75 million tokens. Layer two shows TTFT p95 rose. Layer three shows the prompt token rate unchanged but the prefix-cache hit ratio dropped from 0.62 to 0.31. The cause is a client change that stopped reusing a shared system prompt: every request now prefills its full context, the GPUs spend more time on prefill, and fewer output tokens fit in the same hours. The fix is in the client, not the cluster. Deeper cost modelling, including how to price idle capacity, is in LLM cost analysis.

Traps that make panels lie

  • Averaging percentiles. A mean of per-pod p95 values is not a p95. Always aggregate buckets, then quantile.
  • Rate windows shorter than four scrapes. With a 30-second scrape interval, rate(x[1m]) often has two samples or fewer and returns gaps or spikes. Use at least four times the scrape interval; 5m is a sound default.
  • GPU utilisation that means nothing. DCGM_FI_DEV_GPU_UTIL reports the share of time any kernel was running. A GPU running one tiny kernel continuously reads 100 percent. Use DCGM_FI_PROF_SM_ACTIVE and tensor pipe activity for real saturation.
  • Units. The KV cache gauge is 0 to 1, latency histograms are seconds, framebuffer is MiB. Set the panel unit explicitly so 0.85 does not display as 0.85 percent.
  • Mixing prompt and generation tokens. A prompt token costs a fraction of a generated token in GPU time. A single tokens-per-second number that sums both hides the shift in the worked example. Keep them separate.
  • Cardinality. A per-user or per-request-id label on a histogram multiplies its series count by the number of users and can take Prometheus down. Tenants belong in logs or a bounded label set at the gateway.
  • Counter resets. Pods restart and counters return to zero. rate and increase handle this; subtracting raw counter values does not.

Layout and operation

Order the dashboard by layer: a top row of four or five stat panels for layer one with week-over-week deltas, then a row of layer-two time series with the SLO drawn as a threshold line, then collapsed rows for engine and hardware. Add a model_name template variable so one dashboard covers every model, and an annotation stream for deployments, engine upgrades and configuration changes, because most step changes in a KPI line up with one of them.

Give the dashboard an owner and a weekly fifteen-minute review of layers one and two. Version the dashboard JSON and recording rules alongside the serving configuration so a batching change and the panel that judges it are reviewed together.

FailureSymptomFix
Engine metric renamed in an upgradepanels go flat at zero after a deployalert on absent() for each recorded series
Prometheus overloaded by histogram queriesdashboard loads slowly, rule evaluation lagsrecording rules; drop unused buckets with relabelling
Cost KPI ignores idle GPUscost per token looks better than the billdenominate by allocated GPU hours, not busy hours
Goodput threshold between bucketscompliance jumps when buckets changepin the threshold to a boundary and label the panel

What to do next

  1. Curl one engine pod's /metrics and list the exact names, types and le boundaries for TTFT and ITL.
  2. Write down the four decisions the dashboard must support and delete any planned panel that maps to none of them.
  3. Add the recording rules above, adjusted to your names, and point every panel and alert at the recorded series.
  4. Build layer one: cost per million output tokens from allocated GPU hours, TTFT and ITL compliance from buckets, tokens per day split prompt versus generation.
  5. Build layers two to four as collapsed rows with a model variable and deployment annotations.
  6. Replace any DCGM_FI_DEV_GPU_UTIL panel with SM active and tensor pipe active.
  7. Add absent() alerts so a renamed metric cannot silently flatten a KPI.
  8. Schedule a weekly fifteen-minute review and record one sentence per significant movement.
Key takeaway: Build the dashboard as a tree, not a grid: business KPIs on top, service SLIs below, then engine state and GPU hardware, each layer explaining the one above. Feed every panel and alert from shared recording rules, aggregate histogram buckets before taking quantiles, compute goodput from bucket boundaries, price cost per token from allocated rather than busy GPU hours, and review the top two layers weekly.