A service-level objective is a promise you can measure: a target fraction of good events over a window, for example 99% of agent invocations completing successfully within 60 seconds over 30 days. The missing 1% is the error budget, and spending it is how a team decides between shipping changes and fixing reliability. For a REST endpoint the definition is routine. For an LLM agent it is not, because one request fans out into several model calls and tool calls, the same input can take a different path each time, latency ranges from two seconds to two minutes, and a response can arrive on time, without errors, and still be wrong.
This article defines SLIs that hold up for agents, shows how to compute them in an ADK Java application with a plugin and OpenTelemetry, works through the error budget arithmetic and burn-rate alerting, and covers the quality SLI that ordinary SLO practice lacks. The general theory of SLIs and budgets is in SLIs, SLOs, and Error Budgets; here the focus is what changes when the service is an agent.
Why agent SLOs are different
Three properties of agents break naive SLOs.
Variable work per request. A support agent may answer from memory in one model call or run six tool calls and three model calls. One latency threshold for all requests is either too loose for the simple ones or permanently violated by the complex ones. Segment by task class, or set per-step thresholds as well as an end-to-end one.
Errors that are not exceptions. An agent that hits its step limit, loops on a failing tool, or returns a polite refusal to a legitimate request has failed the user without throwing. These outcomes must be classified as bad events explicitly, or your availability SLI will read 99.9% while users complain.
Correctness is probabilistic. You cannot check every answer cheaply at request time. Quality has to be measured by sampling and grading, which gives a noisy estimate on a delay. It needs its own SLO with a longer window, not a slot in the availability ratio.
Choosing agent SLIs
A workable SLI set for a production agent has four to five members. Each one is a ratio of good events to valid events, so it can carry a target and a budget.
| SLI | Good event | Valid event | Typical target |
|---|---|---|---|
| Completion | Invocation ends with a final response, no error, no step-limit hit | All invocations except client cancellations | 99% to 99.5% |
| Latency | Completion within the threshold for its task class | Completed invocations | 95% within p95 of a healthy week |
| First response | First streamed event within N seconds | Streaming invocations | 99% within 3 s |
| Quality | Graded pass on the task rubric | Sampled, graded invocations | 90% to 97%, task dependent |
| Cost | Invocation under its token or cost ceiling | All invocations | 99% under ceiling |
Start with completion and latency; they come straight from instrumentation. Add quality once you have an evaluator you trust. The cost SLI catches the runaway agent that loops for 40 steps and is easy to overlook, because it neither errors nor runs slowly enough to breach latency on its own. The SLO categories for AI applications in general are compared in SLOs for AI Applications.
Exclusions must be written down. Client cancellations, requests rejected by input validation and load tests are usually not valid events. Provider outages are valid events: users felt them, and the budget is what tells you whether to add a fallback model.
Architecture and data flow
The data flow has two halves. Every invocation passes through a plugin that records outcome, duration and step count as metrics, which feed completion, latency and cost SLIs in near real time. A sample of traces is sent to an evaluator that grades the response, which feeds the quality SLI on a delay of minutes to hours.
Measuring with an ADK Java plugin
ADK Java exposes a plugin interface, com.google.adk.plugins.Plugin, with run-level callbacks that receive the InvocationContext: beforeRunCallback, onEventCallback, afterRunCallback and onRunErrorCallback, plus agent, model and tool callbacks. BasePlugin supplies a name and no-op defaults. One plugin instance serves every concurrent invocation, so per-invocation state must be keyed by invocationId() and removed when the run ends, on both the success and the error path, or it leaks.
import com.google.adk.agents.InvocationContext;
import com.google.adk.events.Event;
import com.google.adk.plugins.BasePlugin;
import com.google.genai.types.Content;
import io.opentelemetry.api.common.AttributeKey;
import io.opentelemetry.api.common.Attributes;
import io.opentelemetry.api.metrics.DoubleHistogram;
import io.opentelemetry.api.metrics.LongCounter;
import io.opentelemetry.api.metrics.Meter;
import io.reactivex.rxjava3.core.Completable;
import io.reactivex.rxjava3.core.Maybe;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.atomic.AtomicInteger;
public final class SloPlugin extends BasePlugin {
private record RunState(long startNanos, AtomicInteger events) {}
private static final AttributeKey<String> AGENT = AttributeKey.stringKey("agent");
private static final AttributeKey<String> OUTCOME = AttributeKey.stringKey("outcome");
private final Map<String, RunState> runs = new ConcurrentHashMap<>();
private final Map<String, Double> latencyThresholdSec; // per agent, from config
private final int maxEvents;
private final LongCounter invocations;
private final DoubleHistogram duration;
public SloPlugin(Meter meter, Map<String, Double> latencyThresholdSec, int maxEvents) {
super("slo");
this.latencyThresholdSec = latencyThresholdSec;
this.maxEvents = maxEvents;
this.invocations = meter.counterBuilder("agent.slo.invocations").build();
this.duration = meter.histogramBuilder("agent.slo.duration").setUnit("s").build();
}
@Override
public Maybe<Content> beforeRunCallback(InvocationContext ctx) {
runs.put(ctx.invocationId(), new RunState(System.nanoTime(), new AtomicInteger()));
return Maybe.empty();
}
@Override
public Maybe<Event> onEventCallback(InvocationContext ctx, Event event) {
RunState s = runs.get(ctx.invocationId());
if (s != null) s.events().incrementAndGet();
return Maybe.empty(); // never alter the event stream
}
@Override
public Completable afterRunCallback(InvocationContext ctx) {
return Completable.fromAction(() -> finish(ctx, null));
}
@Override
public Completable onRunErrorCallback(InvocationContext ctx, Throwable error) {
return Completable.fromAction(() -> finish(ctx, error));
}
private void finish(InvocationContext ctx, Throwable error) {
RunState s = runs.remove(ctx.invocationId());
if (s == null) return; // already recorded on the other path
String agent = ctx.agent().name();
double sec = (System.nanoTime() - s.startNanos()) / 1e9;
String outcome;
if (error != null) outcome = "error";
else if (s.events().get() > maxEvents) outcome = "step_limit";
else if (sec > latencyThresholdSec.getOrDefault(agent, 60.0)) outcome = "slow";
else outcome = "good";
Attributes attrs = Attributes.of(AGENT, agent, OUTCOME, outcome);
invocations.add(1, attrs);
duration.record(sec, Attributes.of(AGENT, agent));
}
}Notes on the design. The outcome label has a handful of fixed values, so metric cardinality stays bounded; never put the invocation ID or user ID on a metric. Events are a proxy for steps, since each model response and tool result is emitted as an event; calibrate maxEvents from the event counts of a healthy week rather than guessing. If the agent streams, partial response chunks arrive as events too, so calibrate the limit separately for streaming and non-streaming runs. Refusals and wrong answers are not visible here; they belong to the quality pipeline. How the plugin is registered depends on your ADK version and on whether you construct a runner or an app object, so follow the current ADK Java documentation for that step. Span-level detail for each invocation is covered in ADK Java observability architecture.
Error budgets and burn-rate alerts
Worked example. A support agent handles 600,000 invocations per 30-day window, about 833 an hour. The completion-within-latency SLO is 99%. The error budget is 1% of 600,000, which is 6,000 bad invocations per window.
Burn rate is the observed bad fraction divided by the budgeted bad fraction. A burn rate of 1 spends exactly the budget over 30 days; a burn rate of 14.4 spends 2% of the 30-day budget in one hour, because 14.4 hours of budget-rate errors divided by 720 hours in the window is 0.02.
Now a tool backend starts timing out. In the last hour the agent handled 900 invocations and 150 were bad, a bad fraction of 16.7% and a burn rate of 16.7. That hour consumed 150 of the 6,000 budget, 2.5%. At this rate the whole budget lasts 720 / 16.7, about 43 hours. That is a page.
# fast burn: 2% of a 30-day budget in 1 hour, confirmed over 5 minutes
(
sum(rate(agent_slo_invocations_total{outcome!="good"}[1h]))
/ sum(rate(agent_slo_invocations_total[1h]))
) > (14.4 * 0.01)
and
(
sum(rate(agent_slo_invocations_total{outcome!="good"}[5m]))
/ sum(rate(agent_slo_invocations_total[5m]))
) > (14.4 * 0.01)
# slow burn: 5% of the budget in 6 hours (burn 6), confirmed over 30 minutes -> ticketThe short window stops the alert from firing long after recovery; the long window stops a two-minute blip from paging. A 6x burn over 6 hours spends 5% of the budget and should open a ticket; a 1x burn over three days spends 10% and belongs in the weekly review. The metric name above is the Prometheus form of the plugin's counter; adjust it to your exporter. The same multiwindow pattern applied to GPU serving SLIs is in LLM SLO Burn Rate Alerts.
The quality SLI
The quality SLI samples finished invocations, grades them against a task rubric and reports passes over graded. The grader can be deterministic checks (did the agent call the refund tool with the right order ID), an LLM judge, or both, but it must be calibrated against human labels before its numbers drive decisions. Keep the rubric and grader version in the metric labels so a grader change is not mistaken for a quality change.
Sample size sets what you can detect. Grading 2% of 600,000 invocations gives 12,000 graded tasks per window, about 400 a day. A 95% confidence interval on a daily pass rate near 93% is roughly plus or minus 2.5 points, so day-to-day wiggles of a point are noise. Alert on quality over windows of a day or more, and compare against a fixed offline evaluation set after every prompt or model change, as in dataset-driven evaluation in ADK Java. The online SLI and the offline set answer different questions: the set tells you whether a change is safe to ship, and the SLI tells you whether production traffic, which drifts away from any fixed set, is still being served well. Oversample rare but important task classes so they are not invisible in the aggregate.
Error budget policy
An SLO without a policy is a dashboard. Agree in advance what happens as the budget is spent:
- Budget healthy. Ship prompt, model and tool changes normally, behind the usual offline evaluation gate.
- Over 50% spent with more than half the window left. Riskier changes such as model swaps need an explicit owner sign-off and a canary.
- Budget exhausted. Freeze non-reliability changes to that agent until the burn rate is below 1, and prioritise the top budget consumers: a flaky tool, a slow provider, a loop-prone prompt.
- Quality SLO breached. Roll back the last prompt or model change first, investigate second; quality regressions rarely fix themselves.
Failure modes
- One SLO for every agent. A fast FAQ agent hides a slow research agent's misses. Set SLOs per agent or task class.
- Counting retries as good. If the client retries a failed invocation and succeeds, the first failure still happened; count at the invocation, not the user session, or define the SLI at session level deliberately.
- Targets from aspiration. Setting 99.9% when the provider alone delivers less makes the budget permanently negative. Start from a measured baseline.
- High-cardinality labels. User or session IDs on metrics explode storage; keep those in traces.
- Judge drift. Upgrading the judge model moves the quality SLI with no change in the agent. Version the grader and re-baseline.
What to do next
- List your agents and task classes; for each, write the user-visible promise in one sentence.
- Turn each promise into SLIs with explicit good, valid and excluded events.
- Add a run-level plugin that records outcome, duration and step counts with bounded labels.
- Measure two healthy weeks, then set targets slightly below that baseline.
- Configure fast and slow multiwindow burn-rate alerts on the completion and latency SLIs.
- Stand up a sampled, versioned evaluator for the quality SLI and calibrate it against human labels.
- Write and agree the error budget policy before the first budget is spent.