Classic service level objectives count the requests that failed. An AI application can return HTTP 200 in 800 milliseconds with an answer that is wrong, malformed, a refusal of a reasonable question or ten times more expensive than it should be, and a conventional availability SLO will score it as a success. The dashboards stay green while users leave.

An AI feature needs objectives along four dimensions: whether the user got a usable answer, how long it took, whether the answer was good, and what it cost. Each needs a precise definition of a good event, a way to measure it at production volume and a budget that tells the team when to stop shipping and start fixing. This article builds that set from first principles, then shows how the objectives interact when the system degrades. The general mechanics of error budgets and burn-rate alerting are covered in Error Budgets and Burn-rate SLO alerting; this page concentrates on what is different for AI.

Advertisement

Choose the unit you promise on

Before defining any indicator, decide what counts as one event. For a single-turn feature such as summarisation or classification, the unit is a request. For a chat assistant, it is usually a response turn. For an agent that runs tools over minutes, it is the task: a user asked for something and the agent either delivered it or did not. Counting individual model calls inside an agent run produces indicators that look healthy while tasks fail, because a run with nine successful calls and one fatal one scores 90 percent.

Pick the unit the user experiences, and log one record per unit that carries everything the four indicators need.

{
  "event_id": "evt_01J9...",
  "unit": "task",
  "feature": "support_assistant",
  "tenant_tier": "paid",
  "started_at": "2026-09-29T10:02:11.402Z",
  "ttft_ms": 740,
  "total_ms": 6210,
  "outcome": "answered",
  "schema_valid": true,
  "route": "provider_a/eu",
  "model_version": "m-2026-08",
  "prompt_version": "support-v41",
  "degradation_level": 0,
  "input_tokens": 3120,
  "output_tokens": 412,
  "cost_usd": 0.0187,
  "sampled_for_quality": true
}

The outcome field takes one of answered, error, timeout, invalid_output or refused. Model version, prompt version, route and degradation level are not optional. When a budget starts burning, these fields are how you find out whether a prompt change, a provider incident or a routing decision caused it.

Availability means a usable answer

Define availability as the fraction of units that produced a usable answer. The difference from classic availability is in the bad-event definition:

OutcomeCounts asWhy
Transport error or 5xxBadNo answer
Timeout past the user deadlineBadUser gave up or saw an error
Output fails the schema or parserBadThe feature cannot use it
Empty or truncated answerBadNothing usable delivered
Refusal of an in-policy requestBadUser got no help
Correct refusal of an out-of-policy requestGoodThe system behaved as designed
Answered, but wrongGood hereCounted by the quality SLI instead

Wrongful refusals are the hardest row, because deciding whether a refusal was justified needs judgment. Run refusals through a classifier or a sample of human review, and count only those judged in-policy as bad. Keep quality out of availability: an answer that is wrong but delivered belongs to the quality indicator, and mixing the two makes both impossible to reason about.

Advertisement

Latency for streamed and agentic work

Mean latency hides everything that matters. Express latency as a fraction of units under a threshold, which turns it into an ordinary good-events-over-total indicator with its own budget. Streaming responses need at least two thresholds: time to first token, which determines whether the product feels responsive, and total time, which determines whether the user waited for a complete answer. A reasonable starting shape is "99 percent of turns show a first token within 2 seconds and 95 percent complete within 15 seconds", tuned to your own traffic.

Total time depends heavily on output length, so a latency regression can be a verbosity regression in disguise. Record output tokens alongside time, and watch time per output token as a diagnostic. For agent tasks, measure end-to-end task time against a threshold appropriate to the task class, and exclude time spent waiting for a human approval, which is not the system's latency.

Quality from sampled, judged events

Quality cannot be read off a status code. It has to be judged, and judging every request is too expensive, so the indicator is computed on a sample. Draw a stratified sample, for example 4 percent of units per feature and tenant tier plus every unit a user flagged, and score each against a written rubric with an LLM judge calibrated against human labels. The calibration step matters: a judge that has not been checked against people drifts with model versions and prompt changes, as discussed in LLM-as-Judge Calibration.

Sampling introduces uncertainty that other indicators do not have, and the SLO must respect it. With a target of 95 percent acceptable and 500 judged samples in a day, the standard error of the measured proportion is the square root of 0.95 times 0.05 divided by 500, about 0.97 percentage points, so a 95 percent confidence interval is roughly plus or minus 1.9 points. A single day that reads 93.5 percent is not clear evidence of a breach. Over a 30-day window with 15,000 samples the interval narrows to about plus or minus 0.35 points, which is precise enough to manage a budget. Alert on quality over longer windows than availability, or increase the sample rate for small features until the interval is useful.

Online signals such as thumbs-down rate, regenerate clicks and escalations to a human are cheap and immediate. Use them as leading indicators that trigger investigation, not as the SLI itself, because they are biased by who bothers to click.

Cost as a guardrail objective

Cost is not usually framed as reliability, but for AI features it fails in the same way: gradually, then suddenly, often because of a change nobody reviewed as a cost change. A prompt that grows by 2,000 tokens of context, an agent loop that takes eight steps instead of five, or a fallback to a more expensive route can double spend without tripping any other alert.

Define the indicator as cost per successful unit, total spend divided by units that were available and judged acceptable, rather than cost per request. The denominator matters: a change that makes requests cheaper but fails more of them should look worse, not better. Set a target, such as a 7-day rolling cost per successful task under $0.04, and track monthly spend against a budget with the same burn-rate logic used for errors. Treat it as a guardrail: breaching it blocks launches that increase cost further, but it does not page someone at night.

Computing indicators and burn rates

All four indicators come from the same event stream, which keeps them consistent and makes it easy to slice by route, version or tenant.

def slis(events, quality_labels, ttft_ms=2000, total_ms=15000):
    n = len(events)
    usable = [e for e in events if e["outcome"] == "answered" and e["schema_valid"]]
    wrongful = sum(1 for e in events
                   if e["outcome"] == "refused" and quality_labels.get(e["event_id"]) == "wrongful_refusal")
    correct_refusals = sum(1 for e in events if e["outcome"] == "refused") - wrongful
    availability = (len(usable) + correct_refusals) / n
    ttft_ok = sum(1 for e in usable if e["ttft_ms"] <= ttft_ms) / max(len(usable), 1)
    total_ok = sum(1 for e in usable if e["total_ms"] <= total_ms) / max(len(usable), 1)
    judged = [quality_labels[e["event_id"]] for e in usable if e["event_id"] in quality_labels]
    quality = sum(1 for j in judged if j == "acceptable") / max(len(judged), 1)
    successes = len(usable) * quality
    cost_per_success = sum(e["cost_usd"] for e in events) / max(successes, 1)
    return dict(availability=availability, ttft_ok=ttft_ok, total_ok=total_ok,
                quality=quality, quality_n=len(judged),
                cost_per_success=cost_per_success)

def burn_rate(bad_fraction, slo_target):
    return bad_fraction / (1.0 - slo_target)

# Multi-window rule for a 30-day SLO: page if both windows exceed the rate.
PAGE_RULES = [("1h", "5m", 14.4), ("6h", "30m", 6.0)]
TICKET_RULES = [("3d", "6h", 1.0)]

A burn rate of 1 consumes the budget exactly over the window; 14.4 sustained for an hour consumes 2 percent of a 30-day budget and would exhaust it in about two days. The paired short window stops the alert firing long after the problem has ended. Use these rules for availability and latency. For quality, where each hour holds only a few dozen judged samples, rely on the 3-day and 30-day views plus the leading online signals.

One event record per request feeds four SLI families and three consumersAI applicationrequest / taskEvent recordtimings, route, costemitSamplerstratified sampleJudge + humansrubric, calibrationAvailability SLIusable answersLatency SLITTFT, end-to-endQuality SLIjudged acceptableCost guardrailcost per successlabelsBurn-rate alertspage or ticketBudget dashboardper SLO, per routeRelease gatefreeze when exhaustedEvery event keeps route, model version, prompt version and degradation levelso a budget burn can be traced to the change that caused it
One event record per unit feeds all four indicators. Quality labels come from a stratified sample scored by a calibrated judge. Budgets drive alerts, dashboards and the release gate.

Worked example: an assistant with four objectives

A support assistant serves about 400,000 tasks per month. The team sets these objectives over a rolling 30 days:

ObjectiveTargetMonthly budget
Usable-answer availability99.5%2,000 bad tasks
First token within 2 s99%4,000 slow tasks
Judged acceptable95%5% of sampled tasks, about 20,000 tasks by extrapolation
Cost per successful taskunder $0.04 (7-day)Guardrail, no budget arithmetic

In one hour, 556 tasks arrive and 40 of them fail with invalid output after a prompt change. The bad fraction is 7.2 percent, and against a 0.5 percent allowance the burn rate is 14.4, so the one-hour page fires, and the five-minute window confirms that it is still happening. The event records show every failure carries prompt_version support-v42, which makes the rollback decision obvious.

Trading one budget for another

The four objectives are not independent, and graceful degradation is a deliberate trade between them. Suppose the primary provider is unavailable for six hours. Without fallback, roughly 3,300 tasks fail, 167 percent of the monthly availability budget. With fallback to a smaller model, those tasks are answered, but the judged-acceptable rate on that route is 90 percent instead of 96. That produces about 330 unacceptable answers, against a quality allowance of about 20,000 for the month, or under 2 percent of it.

Written out like this, the choice is easy, and the SLOs make it explicit: fallback spends a little quality budget to save far more availability budget. The same framing shows when fallback is the wrong choice. If the fallback route's quality on a task class is so poor that answers mislead users, as with a refund-policy question answered incorrectly, then failing honestly is better, and that task class should be removed from the ladder. Choosing models along the trade-off curve is covered in Quality, Cost and Latency Frontiers for Model Selection; the SLOs tell you which point on that curve each degradation mode is allowed to occupy.

Using the budgets

An SLO that nobody acts on is decoration. Tie budgets to decisions. When the availability or quality budget is exhausted, freeze model, prompt and routing changes that are not fixes until the budget recovers. Require every model or prompt change to pass an offline evaluation gate before release, and watch the quality SLI on a canary slice after it. Review cost per success weekly alongside the other three, because the cheapest way to improve availability, retries and hedging, is also the easiest way to overspend.

Failure modes

FailureSymptomFix
Counting model calls, not tasksGreen SLO, failed agent runsPromise on the user-visible unit
Invalid output counted as successAvailability fine, feature brokenSchema validity in the good-event definition
Quality alerting on tiny samplesPages on noiseLonger windows; confidence intervals; larger samples
Uncalibrated judgeQuality drifts with judge versionPeriodic human-labelled calibration set
Cost per request instead of per successCheaper and worse looks like progressDivide by successful units
No version fields on eventsBurn cannot be attributedLog route, model, prompt and degradation level

What to do next

  1. Choose the unit you promise on, request, turn or task, and emit one event record per unit with timings, outcome, versions, route and cost.
  2. Rewrite your availability definition so that invalid output, empty answers and wrongful refusals count as bad.
  3. Set latency objectives as fractions under thresholds for time to first token and total time.
  4. Start a stratified quality sample with a written rubric, a calibrated judge and a sample size large enough for a useful interval.
  5. Add cost per successful unit as a guardrail with a weekly review.
  6. Configure multi-window burn-rate alerts for availability and latency, and longer-window views for quality.
  7. Write down which budget each degradation mode spends, and remove task classes from fallback where quality would mislead users.
  8. Adopt a budget policy that freezes non-fix changes when a budget is exhausted.
Key takeaway: An AI feature is reliable only if users get usable, good answers in time at a sustainable cost. Promise on the unit users experience, count invalid output and wrongful refusals as failures, express latency as fractions under thresholds, measure quality on a calibrated and statistically honest sample, guard cost per successful unit, alert on burn rate, and use the budgets to decide when to fall back, when to fail honestly and when to stop shipping.