An error budget turns a reliability target into something you can spend. If your LLM serving SLO says 99.9 percent of requests succeed within the latency target, then 0.1 percent may fail, and that 0.1 percent is the budget. Spend it on purpose, on rollouts and upgrades that move the product forward, and stop spending when it runs out. Without a budget, every reliability discussion becomes an argument about feelings; with one, it becomes arithmetic.

GPU-backed LLM serving makes budgets both more important and harder. Hardware faults are frequent at fleet scale, a single failed GPU can take down a whole tensor-parallel replica, model and quantisation changes alter quality as well as speed, and a long streaming response can fail halfway through. This article is about the budget itself: the arithmetic, the written policy, attributing spend to causes, and deliberately spending budget on risky GPU changes. Alerting on how fast the budget burns is covered separately in LLM SLO burn rate alerts, which also explains how to define SLIs for time to first token and inter-token latency. This article assumes you have SLIs and asks what to do with them.

The arithmetic of a budget

Start with the arithmetic, because most budget disagreements are really about units. The budget is (1 minus SLO) multiplied by the number of valid events in the window. With a 30-day window, a 99.9 percent target and 1 million valid requests a day, the window holds 30 million events and the budget is 30,000 bad requests. If you think in time instead, 0.1 percent of 30 days is 43.2 minutes of total outage, but request-based budgets are better for LLM serving because traffic is uneven and an outage at 3 a.m. costs far fewer users than one at noon.

LLM serving raises a second unit question: should an event be a request or a token? A request-weighted budget counts a failed one-line answer the same as a failed 4,000-token report. A token-weighted budget counts output tokens lost, which tracks GPU work and cost but lets many short failures hide behind a few long successes. The usual answer is to make the SLO request-weighted, since that is what users experience, and to track token-weighted loss as a secondary measure for capacity and cost.

Be explicit about what counts as a valid event. Requests rejected for exceeding the context length, for quota or for malformed input are not the service's failure and should be excluded. Client cancellations are tricky: a user who gives up waiting is a symptom of slowness, so exclude them only if you count slow time to first token as bad.

Three budgets, not one

One budget per model is rarely enough, because different failures call for different responses. Keep three, each with its own SLO and its own policy.

BudgetBad eventTypical targetMain consumers
Availabilityerror, dropped stream, truncated response99.9%hardware faults, bad deploys, out-of-memory
Latencytime to first token or inter-token latency over threshold99% or 95%queueing, KV-cache pressure, capacity lag
Qualityfailed evaluation probe or validated bad outputset from offline baselinesmodel, quantisation and prompt changes

The quality budget is the one teams skip, and it is the one GPU optimisations spend most quietly. Moving a model from 16-bit weights to an 8-bit or 4-bit format, changing a speculative decoding draft model, or swapping an attention kernel can all keep latency and availability perfect while degrading answers. Measure quality with a fixed set of evaluation prompts sent continuously as synthetic traffic and scored automatically, plus sampled user feedback, and give it a budget like the others: for example, at most 2 percent of probe runs per week may fall below the baseline score.

The error budget policy

The budget only matters if a written policy says what happens as it is spent. Write it before you need it, get it signed by the people who own both the product roadmap and the on-call rotation, and keep it short. A policy that works for most LLM serving teams has three states.

An error budget policy is a state machine driven by budget remainingHealthymore than 50% remainingCaution25% to 50% remainingExhausted0% or less remainingspendspendrecoversrecoversnormal rolloutsdriver upgrades allowedexperiments allowedcanary stages lengthenedone risky change at a timetop spender gets an ownerfeature and model freezeonly fixes and rollbackspostmortem before releaseAttribution ledgerevery bad event is tagged with a cause: deploy, model rollout, driver, hardware, capacity, dependencydeploy 30%hardware 25%capacity 20%model 15%other 10%Example month: share of budget spent by cause. Policy acts on the largest share.
Policy states and their rules (top), and an attribution ledger that tags every bad event with a cause (bottom). The policy's actions target whatever cause holds the largest share.

In the healthy state, more than half the budget remains and teams ship normally. In caution, canary stages run longer, only one risky change is in flight per model at a time, and the largest source of spend gets a named owner with a deadline. In exhausted, new features and new model versions freeze for that model; only fixes, rollbacks and reliability work ship, and each release needs a reviewed postmortem of what consumed the budget. Leaving the exhausted state should require the budget to recover above zero over the rolling window, not just a calendar date.

Two rules keep the policy honest. First, the freeze applies to the service that spent the budget, not to everyone; a hardware-heavy month on one model should not freeze an unrelated model. Second, the policy must say who can grant an exception and require that the exception be written down. Exceptions will happen, and recorded exceptions teach you whether the targets are right.

Attributing spend to causes

A budget tells you how much was spent. Attribution tells you on what, and that is what makes the policy act on the right thing. Tag every bad event, or every incident, with a cause from a fixed list. For GPU LLM serving the list is usually: application deploy, model or quantisation rollout, driver or CUDA upgrade, hardware fault, capacity shortfall, and upstream dependency. The script below computes the remaining budget and the share by cause from a table of bad-event counts.

from collections import Counter

def budget_report(valid_events, slo, bad_events_by_cause):
    """bad_events_by_cause: dict cause -> count of bad events in the window."""
    budget = (1.0 - slo) * valid_events
    spent = sum(bad_events_by_cause.values())
    remaining = 1.0 - spent / budget
    state = ("exhausted" if remaining <= 0
             else "caution" if remaining < 0.5
             else "healthy")
    shares = {c: n / spent for c, n in Counter(bad_events_by_cause).most_common()} if spent else {}
    return {"budget": budget, "spent": spent, "remaining": remaining,
            "state": state, "shares": shares}

report = budget_report(
    valid_events=30_000_000, slo=0.999,
    bad_events_by_cause={"deploy": 5_400, "hardware": 4_500, "capacity": 3_600,
                         "model_rollout": 2_700, "other": 1_800})
# budget 30000, spent 18000, remaining 0.40 -> "caution"; deploy is the largest share

Getting the tags right is the hard part. Deploy and rollout causes can be attributed automatically by joining bad events to the version label on the replica that served them. Hardware causes come from node health signals: GPU driver error events, ECC error counters and NVIDIA DCGM health checks, joined on host and time. Capacity causes show up as queue-time-dominated latency failures during traffic steps. Anything you cannot attribute goes in "other", and if "other" exceeds about a fifth of spend, improve the tagging before trusting the report.

What GPU fleets do to the budget

Several properties of GPU fleets shape how budget is spent, and the policy should account for them.

  • Blast radius scales with parallelism. A model served with tensor parallelism across eight GPUs is one replica; any one of the eight failing takes the whole replica out, and every in-flight stream on it breaks. Larger parallel groups therefore spend budget faster per hardware fault. Keep enough replicas that losing one does not overload the rest, and make the router retry requests that failed before the first token on another replica.
  • Mid-stream failures are expensive. A response that fails after 2,000 tokens cannot be silently retried without the user seeing a restart. Count it as one bad event, but track it separately, because reducing these usually means draining replicas before maintenance rather than killing them.
  • Maintenance spends budget too. Driver upgrades, firmware updates and node reboots should be done by draining traffic first. If draining is done well, maintenance costs almost nothing; if it is not, it shows up in the attribution ledger. Do not exclude planned maintenance from the SLO; users do not care that it was planned.
  • Warm-up after replacement. A replacement replica must load weights, which can take minutes for large models, and its first requests may be slower while caches warm. Budget the capacity lag explicitly when deciding how many spare replicas to keep.

Spending the budget on purpose

The point of a budget is to spend it. Before a risky change, estimate its expected cost and decide whether you can afford it. The expected cost of a canary is roughly traffic share x requests in the canary period x the bad event rate if the change is broken. Suppose you roll a new quantised model to 5 percent of traffic for 2 hours on a service handling 1 million requests a day. The canary sees about 1,000,000 / 24 x 2 x 0.05 = 4,167 requests. If the new build fails 10 percent of them, that is 417 bad requests, about 1.4 percent of a 30,000 budget. Affordable in any state except exhausted.

The same arithmetic tells you when a canary is too big. At 25 percent of traffic for 6 hours, a broken build would cost 6,250 bad requests, a fifth of the monthly budget, so in the caution state the policy should require smaller stages. Automated canary analysis, which compares the canary's SLIs with the baseline's and rolls back on a significant difference, is covered in LLM canary deployments.

Spend budget on experiments too. If a quarter ends with most of the budget unspent, the target may be too loose, or the team may be too cautious about upgrades that would improve throughput. Both are worth discussing at the monthly review.

Worked example: one model, one month

Follow one model through a month. The service handles 1 million valid requests a day, so the 30-day availability budget is 30,000 bad requests. In week one, an application deploy with a tokenizer configuration bug causes 5,400 failures before rollback. In week two, two GPU nodes fail, each taking down an eight-GPU replica; in-flight streams break and the router's retry covers only requests that had not yet streamed, leaving 4,500 bad events. In week three, a traffic spike outruns autoscaling and 3,600 requests time out. In week four, a new model version's canary produces 2,700 failures from an out-of-memory condition on long prompts, and 1,800 other failures remain unexplained.

Total spend is 18,000, leaving 40 percent, so the policy is in caution. The attribution share puts deploys at 30 percent and hardware at 25 percent. The caution rules apply: canaries for the next model version run in smaller stages, and the deploy owner adds a tokenizer configuration check to the pre-deploy tests. Hardware gets a separate action: drain nodes on the first sign of correctable error growth, rather than waiting for failure, and add a spare replica. Neither action needed a debate about whether reliability was "good enough"; the numbers and the policy decided. During incidents, the attribution tags also give responders a shared vocabulary, which pairs well with the practices in running an LLM incident channel.

Failure modes

Error budgets go wrong in predictable ways.

  • No consequences. A budget that is exhausted with no freeze is just a dashboard. Make the policy binding.
  • Freezing everything. Freezing all teams when one model overspends punishes the wrong people and erodes support for the policy. Scope budgets per model or per service.
  • Excluding inconvenient failures. Excluding maintenance, client cancellations or "known issues" makes the numbers look good while users suffer. Every exclusion needs a written reason.
  • Ignoring quality. Without a quality budget, cost and speed optimisations will gradually degrade answers, because every other metric rewards them.
  • Low-traffic models. A model serving a few thousand requests a month has a budget of a few bad requests, and one incident exhausts it. Group low-volume models into a shared budget, or use a longer window.

Trade-offs

Tighter targets mean smaller budgets, slower rollouts and more spare GPUs. Going from 99.9 to 99.99 percent cuts the budget tenfold, which on a GPU fleet usually means extra replicas held idle as failover capacity, an expensive trade when each replica is several accelerators. Looser targets let you ship faster on less hardware but must be acceptable to users. Request-weighted budgets match user experience; token-weighted ones match cost. Per-model budgets give precise accountability but are noisy at low volume. Pick explicitly, write the reasoning into the policy, and revisit it quarterly. If latency budgets keep running out because of queueing, the fix may be in the scheduler rather than the hardware, as discussed in SLO-aware scheduling.

What to do next

Use this checklist to put error budgets to work on your LLM serving fleet.

  1. For each model, set availability, latency and quality SLOs and compute each budget in bad events per 30 days.
  2. Write the valid-event rules, including how cancellations, quota rejections and over-length prompts are treated.
  3. Write a three-state policy with concrete actions and an exception process, and get it signed by product and on-call owners.
  4. Tag every bad event with a cause by joining on replica version and node health signals, and keep "other" below a fifth.
  5. Estimate the budget cost of every risky change, such as a model, quantisation or driver rollout, before starting it.
  6. Drain replicas before maintenance and make the router retry failures that happen before the first token.
  7. Review spend by cause monthly, and adjust targets if the budget is always exhausted or never touched.
Key takeaway: An error budget is (1 minus SLO) times valid events. Keep separate availability, latency and quality budgets per model, bind them to a written three-state policy, and tag every bad event with a cause so the policy targets the real spender. On GPU fleets, account for replica-wide blast radius, mid-stream failures and maintenance, and price each risky rollout before running it.