A web API usually has one latency SLO. An LLM endpoint streams, so a single request has at least three latencies that users feel differently. Time to first token (TTFT) is how long the user waits before anything appears. Time per output token (TPOT) is how fast text flows once it starts. End-to-end latency (E2E) is what matters when a program, not a person, waits for the whole answer. Pick the wrong one, or set it without looking at the GPU, and you get either a fleet that misses its promises or one that is twice the size it needs to be.

This article covers the definition side: which indicators to choose, how to set targets for each traffic class, why 'p90 TTFT and p90 TPOT both meet target' is not the same as '90% of requests are good', and how to turn targets into a request rate with a load sweep. Alerting on the result is covered in LLM SLO burn-rate alerts, scheduling to meet it in SLO-aware scheduling, and the physics behind each latency in GPU inference latency.

Choosing the indicators

An SLI is a measured quantity, an SLO is a target for it over a window, and an SLA is a contract that has penalties attached. For LLM serving, measure these per request:

SLIDefinitionDriven byUsers who feel it
TTFTrequest arrival to first streamed tokenqueueing + prefill (prompt length)chat, code completion
TPOT(E2E - TTFT) / (output tokens - 1)decode batch size, memory bandwidthanyone reading a stream
ITLeach gap between consecutive tokenssame, plus prefill interferencestutter-sensitive UIs, voice
E2Earrival to last tokenTTFT + output length x TPOTagents, batch, tool calls
Error rate5xx, timeouts, truncated streamsoverload, OOM, preemptioneveryone

TPOT and ITL are easy to confuse. vLLM's benchmark computes TPOT per request as (latency - TTFT) / (output_len - 1), which is the average gap for that request, and records ITL as every individual gap. A request with a smooth 40 ms average can still contain a 600 ms stall while a long prompt is being prefilled next to it. TPOT hides that stall, and ITL shows it. If your interface is voice or live captions, the SLO belongs on ITL percentiles.

Measure where the user is. Server-side metrics leave out the load balancer, network and client parsing, and they can leave out the queue if requests wait in front of the engine. Use engine metrics to diagnose and gateway- or client-side timing to define the SLO.

Setting targets per traffic class

Targets come from the product, but they have to be checked against the hardware. Some starting points that can be defended:

  • Interactive chat. People read at roughly 4 to 7 tokens per second, so a TPOT around 50 ms (20 tokens per second) already stays ahead of a reader. Tightening TPOT below what readers can notice costs batch size, and therefore money, for no visible gain. TTFT is the number people notice: a target in the hundreds of milliseconds for short prompts is typical.
  • Code completion. The suggestion is useless if it arrives after the user has typed past it, so the E2E target for short outputs is tight and TPOT barely matters.
  • Agents and batch. No one watches the stream, so E2E or throughput is the SLI and TTFT targets can be loose. Separate this traffic so it cannot consume the interactive budget.
  • Long prompts. Prefill time grows with prompt length, so a single TTFT target either fails long-context requests or is too loose for short ones. Bucket by input length, for example <2k, 2k-16k and >16k tokens, with a target per bucket. Alternatively set the TTFT SLO on queueing delay plus a per-token prefill allowance.
From product promise to a capacity numberTraffic classeschat, code, batchSLIs + targetsTTFT, TPOT, E2E, errorsLoad sweepvllm bench serve --goodputCapacitymax rate at attainmentProduction measurementper-request joint check, histograms with SLO bucket edgesprovisionreviseA target is only real once it has a load sweep behind it and a production measurement in front of it.
The SLO lifecycle: classes and targets come from the product, the load sweep turns them into capacity, and production measurement feeds back into the targets.

Joint attainment and goodput

Teams often write the SLO as 'p90 TTFT under 500 ms and p90 TPOT under 50 ms'. That pair of marginal percentiles does not mean 90% of requests had a good experience. The slow-TTFT requests and the slow-TPOT requests can be different ones, so the fraction of requests that met both targets can be as low as 80%. Joint attainment is what users experience. The word goodput has two related meanings. vLLM's benchmark reports it as completed requests per second that met every listed target. The DistServe paper, which vLLM cites, defines per-GPU goodput as the maximum request rate that can be served while meeting an SLO attainment goal (90% in its examples), divided by the GPUs provisioned. The first is a measurement at one load level. The second is a capacity figure, and it is what the load sweep below estimates. Here is how to compute attainment and the measured rate from request logs:

import numpy as np

def slo_report(reqs, ttft_ms=500.0, tpot_ms=50.0, q=90, duration_s=None):
    """reqs: list of (ttft_ms, e2e_ms, output_tokens)."""
    ttft = np.array([r[0] for r in reqs], float)
    tpot = np.array([(e - t) / (n - 1) if n > 1 else 0.0 for t, e, n in reqs])
    for name, x, slo in (("TTFT", ttft, ttft_ms), ("TPOT", tpot, tpot_ms)):
        lin = np.percentile(x, q)                       # linear interpolation (numpy default)
        near = np.percentile(x, q, method="nearest")
        print(f"p{q} {name}: linear {lin:.1f}  nearest {near:.1f}  target {slo}")
    good = (ttft <= ttft_ms) & (tpot <= tpot_ms)        # per-request, all targets at once
    print(f"joint attainment {good.mean():.0%}")
    if duration_s:
        print(f"goodput {good.sum() / duration_s:.2f} req/s")
    return good

Worked example: two green p90s, 80% good

Ten requests, with targets of TTFT 500 ms and TPOT 50 ms. Each row gives TTFT in ms, E2E in ms and output tokens. TPOT is derived as (E2E - TTFT) / (tokens - 1):

#TTFTE2ETokensTPOTGood?
11806,40020131.1yes
22409,10025135.4yes
33104,31010140.0yes
442014,82030148.0yes
56505,65012141.7no (TTFT)
622013,42020166.0no (TPOT)
73907,79015149.3yes
82608,36018145.0yes
94703,2906147.0yes
1030012,06024149.0yes

Running slo_report on these rows gives three lessons. First, 90% of requests meet the TTFT target and 90% meet the TPOT target, but only 80% meet both, because request 5 and request 6 fail different targets. A dashboard showing two green percentiles would hide a 20% bad-experience rate. Second, the percentile method changes the verdict. Nearest-rank p90 TPOT is 49.3 ms, which passes, while numpy's default linear interpolation, which vLLM's benchmark uses, gives 50.97 ms, which fails. With few samples the method decides the answer, so write it into the SLO definition. Third, a 10-request sample says almost nothing on its own. SLO decisions need thousands of requests per class per window.

From targets to capacity: the load sweep

A target only becomes an operational fact once you know the request rate at which it stops holding. Run a sweep against a production-shaped deployment, with the same model, parallelism, max batch settings and prompt and output length distribution as production:

for rate in 1 2 4 6 8 10 12; do
  vllm bench serve --model "$MODEL" --dataset-name random \
    --random-input-len 2000 --random-output-len 300 \
    --request-rate "$rate" --burstiness 1.0 --num-prompts 1000 \
    --percentile-metrics ttft,tpot,itl,e2el --metric-percentiles 50,90,99 \
    --goodput ttft:500 tpot:50 \
    --save-result --result-filename "sweep_${rate}.json"
done

The --goodput flag takes KEY:VALUE pairs in milliseconds, with keys ttft, tpot and e2el. A request counts as good only if it meets every listed target, and requests with one output token are treated as having zero TPOT. --burstiness 1.0 gives Poisson arrivals. Lower values make arrivals burstier, which is closer to real traffic and well worth a second sweep. For each rate, compute attainment as good requests divided by attempted requests, so that failed requests count as bad, and plot it. Dividing by completed requests silently drops failures, which are most common exactly where the curve matters. Capacity is the highest rate at which attainment is still at or above the target, for example 99%. Provision so that peak traffic sits below that rate with headroom for a failed replica. The capacity estimation article explains where the curve bends and why.

Attainment collapses quickly past the knee. Queueing delay grows without bound as utilisation approaches 1, so TTFT fails long before throughput levels off. That is why throughput-only benchmarks overstate capacity, sometimes by a factor of two.

Measuring in production

In production, record the same per-request fields at the gateway and evaluate the joint check per request. That is one boolean per request, which makes the error budget a simple count. Engine histograms are useful for diagnosis. Recent vLLM versions export vllm:time_to_first_token_seconds, vllm:inter_token_latency_seconds, vllm:request_time_per_output_token_seconds, vllm:e2e_request_latency_seconds and vllm:request_queue_time_seconds. Metric names have changed between versions, so check what your build exports. A histogram can only answer 'what fraction is under 500 ms' exactly if 0.5 s is one of its bucket edges. Otherwise the estimate is interpolated between edges, so align buckets to your targets.

Watch queue time next to TTFT. When TTFT rises and queue time rises with it, you are short of capacity. When TTFT rises and queue time stays flat, prefill got slower, often because prompts got longer or the prefix-cache hit rate dropped.

Write the SLO down as a short document that someone else could implement without asking you questions. It should name the traffic class and how requests are assigned to it. It should give each indicator with its measurement point (gateway, not engine), its threshold, the percentile method or the joint-attainment target (for example 99% of requests good), and the window (for example 28 days, rolling). It should say how errors, timeouts and client cancellations are counted. A cancelled request is usually excluded, but a cancellation after a long wait for the first token is often a TTFT failure the user gave up on, so record the wait time either way. Finally, it should name the error budget: with a 99% target over 28 days, 1% of that class's requests may be bad before releases slow down. That budget is what the error budget policy spends.

Failure modes

  • Marginal percentiles reported as attainment. This overstates user experience, as the worked example shows.
  • Averaging percentiles across replicas or windows. The mean of p99s is not a p99. Merge histograms or raw counts instead.
  • Mixed traffic in one SLO. Batch jobs with long prompts drag down chat TTFT, and the SLO hides which class is hurting. Slice by class and by input-length bucket.
  • Benchmarking with fixed lengths only. Real output lengths have long tails. Replay a sample of production lengths as well as the synthetic sweep.
  • Counting failures as fast. Timeouts and truncated streams have no TPOT. Count them as bad, not as missing.
  • Closed-loop load generators. A fixed-concurrency client slows down when the server does, which hides queueing. Use open-loop arrivals for SLO sweeps.

Trade-offs

Every SLO is a trade against cost. A tighter TPOT means smaller decode batches and more GPUs. A tighter TTFT means more prefill capacity, chunked prefill, or a dedicated prefill pool. Chunked prefill smooths ITL but slightly delays TTFT for the long prompt being chunked. Disaggregating prefill and decode lets you tune each SLO separately, at the cost of moving the KV cache between pools. Per-class SLOs cost routing complexity, but they stop the strictest class from setting the size of the whole fleet. Choose the loosest targets users cannot tell apart from tighter ones, and spend the savings on headroom.

What to do next

  1. List your traffic classes and pick the SLI each one actually feels: TTFT, TPOT, ITL or E2E.
  2. Set targets per class and per input-length bucket, and write the percentile method into the definition.
  3. Define 'good request' as passing every target jointly, and report attainment and goodput, not only marginal percentiles.
  4. Run an open-loop vllm bench serve sweep with --goodput and find the knee. Provision peak below it with headroom.
  5. Instrument the gateway with per-request booleans and align histogram buckets to the targets.
  6. Feed the per-request bad count into burn-rate alerts and repeat the sweep after every model, engine or hardware change.
Key takeaway: An LLM serving SLO needs one indicator per thing users feel: TTFT for waiting, TPOT or ITL for streaming, and E2E for programs. Each traffic class needs its own targets. Judge every request against all its targets at once and report attainment and goodput, because two passing percentiles can hide a 20% bad rate. Then find capacity with an open-loop load sweep and provision below the knee.