Classic service level objectives count the requests that failed. An AI application can return HTTP 200 in 800 milliseconds with an answer that is wrong, malformed, a refusal of a reasonable question or ten times more expensive than it should be, and a conventional availability SLO will score it as a success. The dashboards stay green while users leave.
An AI feature needs objectives along four dimensions: whether the user got a usable answer, how long it took, whether the answer was good, and what it cost. Each needs a precise definition of a good event, a way to measure it at production volume and a budget that tells the team when to stop shipping and start fixing. This article builds that set from first principles, then shows how the objectives interact when the system degrades. The general mechanics of error budgets and burn-rate alerting are covered in Error Budgets and Burn-rate SLO alerting; this page concentrates on what is different for AI.
Choose the unit you promise on
Before defining any indicator, decide what counts as one event. For a single-turn feature such as summarisation or classification, the unit is a request. For a chat assistant, it is usually a response turn. For an agent that runs tools over minutes, it is the task: a user asked for something and the agent either delivered it or did not. Counting individual model calls inside an agent run produces indicators that look healthy while tasks fail, because a run with nine successful calls and one fatal one scores 90 percent.
Pick the unit the user experiences, and log one record per unit that carries everything the four indicators need.
{
"event_id": "evt_01J9...",
"unit": "task",
"feature": "support_assistant",
"tenant_tier": "paid",
"started_at": "2026-09-29T10:02:11.402Z",
"ttft_ms": 740,
"total_ms": 6210,
"outcome": "answered",
"schema_valid": true,
"route": "provider_a/eu",
"model_version": "m-2026-08",
"prompt_version": "support-v41",
"degradation_level": 0,
"input_tokens": 3120,
"output_tokens": 412,
"cost_usd": 0.0187,
"sampled_for_quality": true
}The outcome field takes one of answered, error, timeout, invalid_output or refused. Model version, prompt version, route and degradation level are not optional. When a budget starts burning, these fields are how you find out whether a prompt change, a provider incident or a routing decision caused it.
Availability means a usable answer
Define availability as the fraction of units that produced a usable answer. The difference from classic availability is in the bad-event definition:
| Outcome | Counts as | Why |
|---|---|---|
| Transport error or 5xx | Bad | No answer |
| Timeout past the user deadline | Bad | User gave up or saw an error |
| Output fails the schema or parser | Bad | The feature cannot use it |
| Empty or truncated answer | Bad | Nothing usable delivered |
| Refusal of an in-policy request | Bad | User got no help |
| Correct refusal of an out-of-policy request | Good | The system behaved as designed |
| Answered, but wrong | Good here | Counted by the quality SLI instead |
Wrongful refusals are the hardest row, because deciding whether a refusal was justified needs judgment. Run refusals through a classifier or a sample of human review, and count only those judged in-policy as bad. Keep quality out of availability: an answer that is wrong but delivered belongs to the quality indicator, and mixing the two makes both impossible to reason about.
Latency for streamed and agentic work
Mean latency hides everything that matters. Express latency as a fraction of units under a threshold, which turns it into an ordinary good-events-over-total indicator with its own budget. Streaming responses need at least two thresholds: time to first token, which determines whether the product feels responsive, and total time, which determines whether the user waited for a complete answer. A reasonable starting shape is "99 percent of turns show a first token within 2 seconds and 95 percent complete within 15 seconds", tuned to your own traffic.
Total time depends heavily on output length, so a latency regression can be a verbosity regression in disguise. Record output tokens alongside time, and watch time per output token as a diagnostic. For agent tasks, measure end-to-end task time against a threshold appropriate to the task class, and exclude time spent waiting for a human approval, which is not the system's latency.
Quality from sampled, judged events
Quality cannot be read off a status code. It has to be judged, and judging every request is too expensive, so the indicator is computed on a sample. Draw a stratified sample, for example 4 percent of units per feature and tenant tier plus every unit a user flagged, and score each against a written rubric with an LLM judge calibrated against human labels. The calibration step matters: a judge that has not been checked against people drifts with model versions and prompt changes, as discussed in LLM-as-Judge Calibration.
Sampling introduces uncertainty that other indicators do not have, and the SLO must respect it. With a target of 95 percent acceptable and 500 judged samples in a day, the standard error of the measured proportion is the square root of 0.95 times 0.05 divided by 500, about 0.97 percentage points, so a 95 percent confidence interval is roughly plus or minus 1.9 points. A single day that reads 93.5 percent is not clear evidence of a breach. Over a 30-day window with 15,000 samples the interval narrows to about plus or minus 0.35 points, which is precise enough to manage a budget. Alert on quality over longer windows than availability, or increase the sample rate for small features until the interval is useful.
Online signals such as thumbs-down rate, regenerate clicks and escalations to a human are cheap and immediate. Use them as leading indicators that trigger investigation, not as the SLI itself, because they are biased by who bothers to click.
Cost as a guardrail objective
Cost is not usually framed as reliability, but for AI features it fails in the same way: gradually, then suddenly, often because of a change nobody reviewed as a cost change. A prompt that grows by 2,000 tokens of context, an agent loop that takes eight steps instead of five, or a fallback to a more expensive route can double spend without tripping any other alert.
Define the indicator as cost per successful unit, total spend divided by units that were available and judged acceptable, rather than cost per request. The denominator matters: a change that makes requests cheaper but fails more of them should look worse, not better. Set a target, such as a 7-day rolling cost per successful task under $0.04, and track monthly spend against a budget with the same burn-rate logic used for errors. Treat it as a guardrail: breaching it blocks launches that increase cost further, but it does not page someone at night.
Computing indicators and burn rates
All four indicators come from the same event stream, which keeps them consistent and makes it easy to slice by route, version or tenant.
def slis(events, quality_labels, ttft_ms=2000, total_ms=15000):
n = len(events)
usable = [e for e in events if e["outcome"] == "answered" and e["schema_valid"]]
wrongful = sum(1 for e in events
if e["outcome"] == "refused" and quality_labels.get(e["event_id"]) == "wrongful_refusal")
correct_refusals = sum(1 for e in events if e["outcome"] == "refused") - wrongful
availability = (len(usable) + correct_refusals) / n
ttft_ok = sum(1 for e in usable if e["ttft_ms"] <= ttft_ms) / max(len(usable), 1)
total_ok = sum(1 for e in usable if e["total_ms"] <= total_ms) / max(len(usable), 1)
judged = [quality_labels[e["event_id"]] for e in usable if e["event_id"] in quality_labels]
quality = sum(1 for j in judged if j == "acceptable") / max(len(judged), 1)
successes = len(usable) * quality
cost_per_success = sum(e["cost_usd"] for e in events) / max(successes, 1)
return dict(availability=availability, ttft_ok=ttft_ok, total_ok=total_ok,
quality=quality, quality_n=len(judged),
cost_per_success=cost_per_success)
def burn_rate(bad_fraction, slo_target):
return bad_fraction / (1.0 - slo_target)
# Multi-window rule for a 30-day SLO: page if both windows exceed the rate.
PAGE_RULES = [("1h", "5m", 14.4), ("6h", "30m", 6.0)]
TICKET_RULES = [("3d", "6h", 1.0)]A burn rate of 1 consumes the budget exactly over the window; 14.4 sustained for an hour consumes 2 percent of a 30-day budget and would exhaust it in about two days. The paired short window stops the alert firing long after the problem has ended. Use these rules for availability and latency. For quality, where each hour holds only a few dozen judged samples, rely on the 3-day and 30-day views plus the leading online signals.
Worked example: an assistant with four objectives
A support assistant serves about 400,000 tasks per month. The team sets these objectives over a rolling 30 days:
| Objective | Target | Monthly budget |
|---|---|---|
| Usable-answer availability | 99.5% | 2,000 bad tasks |
| First token within 2 s | 99% | 4,000 slow tasks |
| Judged acceptable | 95% | 5% of sampled tasks, about 20,000 tasks by extrapolation |
| Cost per successful task | under $0.04 (7-day) | Guardrail, no budget arithmetic |
In one hour, 556 tasks arrive and 40 of them fail with invalid output after a prompt change. The bad fraction is 7.2 percent, and against a 0.5 percent allowance the burn rate is 14.4, so the one-hour page fires, and the five-minute window confirms that it is still happening. The event records show every failure carries prompt_version support-v42, which makes the rollback decision obvious.
Trading one budget for another
The four objectives are not independent, and graceful degradation is a deliberate trade between them. Suppose the primary provider is unavailable for six hours. Without fallback, roughly 3,300 tasks fail, 167 percent of the monthly availability budget. With fallback to a smaller model, those tasks are answered, but the judged-acceptable rate on that route is 90 percent instead of 96. That produces about 330 unacceptable answers, against a quality allowance of about 20,000 for the month, or under 2 percent of it.
Written out like this, the choice is easy, and the SLOs make it explicit: fallback spends a little quality budget to save far more availability budget. The same framing shows when fallback is the wrong choice. If the fallback route's quality on a task class is so poor that answers mislead users, as with a refund-policy question answered incorrectly, then failing honestly is better, and that task class should be removed from the ladder. Choosing models along the trade-off curve is covered in Quality, Cost and Latency Frontiers for Model Selection; the SLOs tell you which point on that curve each degradation mode is allowed to occupy.
Using the budgets
An SLO that nobody acts on is decoration. Tie budgets to decisions. When the availability or quality budget is exhausted, freeze model, prompt and routing changes that are not fixes until the budget recovers. Require every model or prompt change to pass an offline evaluation gate before release, and watch the quality SLI on a canary slice after it. Review cost per success weekly alongside the other three, because the cheapest way to improve availability, retries and hedging, is also the easiest way to overspend.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
| Counting model calls, not tasks | Green SLO, failed agent runs | Promise on the user-visible unit |
| Invalid output counted as success | Availability fine, feature broken | Schema validity in the good-event definition |
| Quality alerting on tiny samples | Pages on noise | Longer windows; confidence intervals; larger samples |
| Uncalibrated judge | Quality drifts with judge version | Periodic human-labelled calibration set |
| Cost per request instead of per success | Cheaper and worse looks like progress | Divide by successful units |
| No version fields on events | Burn cannot be attributed | Log route, model, prompt and degradation level |
What to do next
- Choose the unit you promise on, request, turn or task, and emit one event record per unit with timings, outcome, versions, route and cost.
- Rewrite your availability definition so that invalid output, empty answers and wrongful refusals count as bad.
- Set latency objectives as fractions under thresholds for time to first token and total time.
- Start a stratified quality sample with a written rubric, a calibrated judge and a sample size large enough for a useful interval.
- Add cost per successful unit as a guardrail with a weekly review.
- Configure multi-window burn-rate alerts for availability and latency, and longer-window views for quality.
- Write down which budget each degradation mode spends, and remove task classes from fallback where quality would mislead users.
- Adopt a budget policy that freezes non-fix changes when a budget is exhausted.