A team that serves a large language model usually starts with alerts like 'p99 latency above 5 seconds for 5 minutes' or 'GPU utilisation above 95 percent'. Both page people at 3 a.m. for things users barely notice, and both stay silent through a slow leak that ruins a week. Burn-rate alerting replaces them with one question: at the rate we are failing right now, how fast are we spending the unreliability we agreed to tolerate?

The generic mechanics are covered in Burn-rate SLO alerting architecture. This article applies them to LLM inference, where the usual recipe needs adjusting. A response is a stream that can fail after a 200 status, latency has two parts users feel differently, and the main causes of budget burn, such as queueing behind long prompts, KV-cache pressure and slow autoscaling, are specific to GPU serving.

Advertisement

Why LLM serving breaks the usual recipe

Classic request SLOs assume a request is either answered or not, and that one latency number describes it. An LLM request breaks both assumptions. A streamed completion can return HTTP 200, send 300 tokens and then die when a worker is pre-empted or the connection resets. The status code says success, while the user saw a truncated answer. Latency is really two experiences: time to first token (TTFT), which decides whether the product feels responsive, and the gap between later tokens, which decides whether the text streams smoothly or stutters.

Load is also unusual. Cost per request varies by orders of magnitude with prompt and output length, so 'requests per second' says little about how hard the GPUs are working. One tenant sending 100,000-token prompts can push everyone else's TTFT up without any error appearing. That is why the SLIs below separate availability from TTFT from inter-token latency, and why each gets its own budget.

From GPU server and gateway metrics to a pageInference engineTTFT, ITL histogramsGateway / routeroutcome per requestClient SDK (optional)stream completed?Prometheus scrapecounters + bucketsRecording rulesbad/total per windowFast burn: 14.4x1h AND 5m, pageMedium burn: 6x6h AND 30m, pageSlow burn: 1x3d AND 6h, ticketAlertmanagerroute by severityOn-call runbookqueue, KV cache, deploysbad events must be definedbefore any rule is written
Signals come from the inference engine (latency histograms), the gateway (per-request outcome, including failures after the first byte) and optionally the client. Recording rules turn them into bad-event ratios per window; three alert tiers compare those ratios with burn-rate thresholds.

Choosing SLIs that match what users feel

An SLI is a ratio of good events to valid events. For LLM serving, three cover most of what users notice. Write each as an explicit rule before you touch PromQL.

SLIGood eventWhere to measureTypical source
Availabilityrequest completes its stream with a normal finishgateway or routeryour own counter with an outcome label
TTFTfirst token delivered within threshold Tengine (or gateway for queue time)vllm:time_to_first_token_seconds histogram
Inter-token latencytoken gap within threshold Genginevllm:inter_token_latency_seconds histogram
Per-request decode speedrequest's average time per output token within Genginevllm:request_time_per_output_token_seconds histogram

The last two rows look similar but weight events differently, and that changes what an SLO means. According to the vLLM metrics design notes, vllm:inter_token_latency_seconds records the wall-clock gap between successive outputs, so every gap is one observation and a 2,000-token answer contributes a thousand times more events than a 2-token one. vllm:request_time_per_output_token_seconds records one value per request, (end-to-end latency minus TTFT) divided by (output tokens minus 1). A token-weighted SLO says '99 percent of token gaps are smooth'. A request-weighted one says '99 percent of requests stream smoothly on average'. A few long, stuttering generations can breach the first while the second looks fine. Choose deliberately and write the choice into the SLO document. The older vllm:time_per_output_token_seconds is deprecated in favour of these two, so check which names your version exports.

Engine-side TTFT does not include time spent in your gateway, the load balancer or the client network. If you queue requests in a router before they reach the engine, measure TTFT at the router too, or a queueing incident will hide from the engine histogram.

Advertisement

Defining bad events for a streaming response

This is where most LLM SLOs go wrong, so decide it before writing any rule. The engine counter vllm:request_success_total is labelled with finished_reason (values include stop, length and abort). It counts finished requests by how they ended. It is not an error counter, and an abort can mean either a client that disconnected or a server-side cancellation.

A workable policy, implemented as an outcome label on a gateway counter:

  • Good: the stream ended with a normal finish (stop or length) and the client received the final chunk.
  • Bad: 5xx before the first byte, a stream that ended without a final chunk (mid-stream reset, worker crash, pre-emption that was not recovered), gateway timeouts, and 429 or overload rejections you caused by shedding load.
  • Excluded from valid events: 4xx caused by the caller (bad request, prompt too long, auth failures) and cancellations that the client initiated before any server-side timeout. Count them separately so you can watch for a spike, which can mean that users give up because you are slow.

Rate-limit rejections deserve a conscious decision. If a tenant exceeds a contractual quota, a 429 is correct behaviour and should be excluded. If you shed load because the fleet is saturated, that is unreliability and should burn budget. Use different outcome labels for the two cases.

Burn rate from first principles

An SLO of 99.9 percent over 30 days leaves an error budget of 0.1 percent of valid events. The burn rate is the observed bad ratio divided by that budget fraction. A burn rate of 1 spends exactly the whole budget in 30 days. A burn rate of 14.4 spends it in 30 days divided by 14.4, which is 50 hours. Burning at 14.4 for one hour consumes 14.4 divided by 720 hours, which is 2 percent of the monthly budget.

The Google SRE workbook's multiwindow, multi-burn-rate scheme uses three tiers. Each alert requires both a long window, which proves the burn is significant, and a short window, which proves it is still happening, so the alert resets within minutes of recovery.

TierBurn rateLong / short windowBudget spent when it firesBad ratio at 99.9%Bad ratio at 99%Action
Fast14.41h / 5m2%1.44%14.4%page
Medium66h / 30m5%0.6%6%page
Slow13d / 6h10%0.1%1%ticket

Latency SLOs are usually looser than availability ones. A TTFT SLO of 99 percent under 2 seconds is common, and the table shows the consequence: the fast tier fires only when more than 14.4 percent of requests are slow over an hour. That is intended. A smaller degradation is caught by the medium or slow tier instead.

Recording and alerting rules

Record the bad ratio per window once and reuse it in every alert. The TTFT rule below assumes the threshold is a bucket boundary of the histogram: Prometheus can only count observations at or below an existing le value. Check the exact label string on your /metrics output (for example le="2.0" versus le="2"); a mismatch silently returns no data and the alert never fires. If no boundary matches your threshold, either move the threshold or record your own histogram at the gateway. vLLM labels series with the served model; the rules use model_name, so confirm the key in your version.

groups:
- name: llm-slo-recording
  rules:
  # TTFT SLI: fraction of requests whose first token took longer than 2 s.
  - record: llm:ttft_slow:ratio_rate5m
    expr: |
      1 - (
        sum by (model_name) (rate(vllm:time_to_first_token_seconds_bucket{le="2.0"}[5m]))
        /
        sum by (model_name) (rate(vllm:time_to_first_token_seconds_count[5m]))
      )
  # Repeat for 30m, 1h, 6h with the same expression and a different range.
  - record: llm:ttft_slow:ratio_rate1h
    expr: |
      1 - (
        sum by (model_name) (rate(vllm:time_to_first_token_seconds_bucket{le="2.0"}[1h]))
        /
        sum by (model_name) (rate(vllm:time_to_first_token_seconds_count[1h]))
      )
  # Availability SLI from a gateway counter you own (example name).
  - record: llm:avail_bad:ratio_rate5m
    expr: |
      sum by (model) (rate(llm_gateway_requests_total{outcome=~"error|stream_broken|timeout|shed"}[5m]))
      /
      sum by (model) (rate(llm_gateway_requests_total{outcome!~"client_error|client_cancel"}[5m]))

- name: llm-slo-alerts
  rules:
  - alert: LLMTTFTBudgetFastBurn
    expr: |
      llm:ttft_slow:ratio_rate1h > (14.4 * 0.01)
      and
      llm:ttft_slow:ratio_rate5m > (14.4 * 0.01)
    for: 2m
    labels:
      severity: page
    annotations:
      summary: "{{ $labels.model_name }} is spending TTFT budget at 14x or more"
      runbook: "check queue depth, prompt-length mix, KV cache usage, recent deploys"

The 0.01 is the budget of a 99 percent SLO; an availability SLO at 99.9 percent uses 0.001. The llm_gateway_requests_total counter and its outcome values are an example of a counter you would add to your own gateway, not a metric any engine exports.

Worked example: how long detection really takes

A chat model serves a steady 20 requests per second. Its SLO is 99 percent of requests with TTFT under 2 seconds over 30 days. Thirty days at 20 requests per second is 51,840,000 requests, so the budget is 518,400 slow requests. Normally 0.5 percent of requests are slow, a burn rate of 0.5.

At 14:00 a new tenant starts sending very long prompts. Prefill for those prompts occupies the GPUs, other requests wait, and the slow fraction jumps to 18 percent, a burn rate of 18. When does the fast tier fire? The 1-hour ratio after t minutes is (18t + 0.5(60 - t)) / 60 percent. It crosses 14.4 percent when 17.5t exceeds 834, at about 48 minutes. The 5-minute ratio is already 18 percent, so the page fires around 14:48 (plus the 2-minute hold). By then about 57,600 requests have arrived during the incident, roughly 10,400 of them slow: 2 percent of the monthly budget, exactly what the tier promises.

The general rule is that detection time equals the window times the threshold divided by the actual burn rate. A total outage with every request slow (burn 100) pages after about 9 minutes; a burn of 18 takes 48. If 48 minutes of degraded TTFT is unacceptable for your product, tighten the SLO or add a faster tier, and accept more pages. Do not shorten the long window alone: you trade detection time for alerts on harmless blips.

What actually burns the budget on a GPU fleet

A burn-rate page tells you how bad things are, not why. For LLM serving, the cause usually falls into one of a handful of patterns, each visible in engine gauges.

PatternSLI that burnsWhat you see
Queueing behind long prefillsTTFTvllm:num_requests_waiting climbs; prompt-length mix shifts
KV-cache pressure and pre-emptioninter-token latency, then availabilitycache usage near full; requests pre-empted and recomputed
Capacity lag after a traffic stepTTFTwaiting queue grows while new replicas load weights for minutes
Bad deploy (kernel, quantisation, config)anyburn starts at the rollout timestamp on one model or version
Node or GPU failureavailabilitybroken streams on one host; replicas drop out of the pool
Client-side cancellation stormexcluded events risecancellations spike because users give up waiting

The cancellation row matters because it is where an SLO can lie. If users abandon slow requests and cancellations are excluded, the slowest requests disappear from the TTFT histogram before they finish, and the SLI looks healthier exactly when it is worst. Alert on the cancellation rate separately, or count server-side time-to-cancel as a slow TTFT event. Scheduling decisions that decide who waits in the queue are covered in SLO-aware scheduling for LLM inference, and the components of latency in GPU inference latency.

Low-traffic models and slicing

Burn-rate math needs enough events. A fine-tuned model serving one request a minute produces 5 events in a 5-minute window, so one slow request is a 20 percent bad ratio and a burn of 20. Options: aggregate low-volume models into one SLO, add synthetic probes so every window has a minimum number of events, require a minimum event count in the alert expression, or alert only on the slow tier for those models.

Slice SLOs by what a user experiences, usually model and region, and resist slicing by tenant in the alerting path. With hundreds of tenants, per-tenant alerts multiply false pages. Keep per-tenant SLIs on dashboards and in weekly reports, and page on the aggregate.

Operating the alerts

  1. Write the SLO document first: SLIs, thresholds, windows, valid-event rules, and what happens when the budget is exhausted (for example, freeze risky deploys for that model).
  2. Attach a runbook to every alert that starts with the cause table above: queue depth, prompt-length mix, KV-cache usage, rollout history, host health.
  3. Annotate deploys and autoscaling events on the SLO dashboard so the start of a burn can be matched to a change in seconds.
  4. Review every page weekly: was it real, was it actionable, did a lower tier fire first? Adjust thresholds from evidence, not from one painful night.

Trade-offs

Burn-rate alerting trades immediacy for precision. You will rarely be paged for a two-minute blip, and you will sometimes learn about a moderate degradation tens of minutes after it begins. Token-weighted latency SLIs track what heavy users feel but are dominated by long generations; request-weighted ones are fairer per user but hide stutter inside long answers. And every exclusion rule, such as client errors, cancellations and quota rejections, is a place where a real problem can hide, so each excluded category needs its own volume alert. For budget policy beyond alerting, see error budgets in practice.

What to do next

  1. Write down the good-event definition for availability, including streams that fail after the first byte, and add an outcome label to your gateway counter.
  2. List the bucket boundaries of your TTFT and inter-token histograms and pick SLO thresholds that sit exactly on them.
  3. Decide between token-weighted and request-weighted decode latency, and record the choice in the SLO document.
  4. Deploy the recording rules, then the fast and medium tiers as pages and the slow tier as a ticket.
  5. Replay a past incident and compute its detection time with the formula: window times threshold divided by burn rate.
  6. Add a separate alert on cancellation and rejection rates so excluded events cannot hide an outage.
Key takeaway: An LLM SLO burn-rate alert pages when the service spends its error budget too fast, not when a single number crosses a line. For inference that means defining availability to include streams that break after a 200, measuring TTFT and inter-token latency from histograms whose bucket boundaries match your thresholds, and choosing token- or request-weighting on purpose. Pair a long window with a short one at 14.4x, 6x and 1x, expect detection time to equal window times threshold divided by the real burn rate, and keep a runbook that maps each burn to its usual GPU-side cause: queueing behind long prefills, KV-cache pressure, capacity lag, bad deploys or failing hosts.