Your agent platform promises each tenant a level of service in their contract, and pays them back in credits when it misses. That promise is a Service Level Agreement. It is not the same thing as the service level objectives your team uses to run the platform, and treating the two as one leads to one of two mistakes: an SLA so strict that every bad afternoon costs money, or one so loose that it means nothing.

This article covers the SLA layer for an ADK Java platform serving many tenants: what an agent SLA can honestly promise, how to measure it per tenant, how to classify failures so exclusions are defensible, how to compute availability and credits, how service tiers interact with isolation, and what the monthly report should contain. Internal SLIs, error budgets and burn-rate alerting are covered in SLO definition for agents; this page builds on them.

SLI, SLO and SLA are three different things

SLISLOSLA
What it isA measurementAn internal target for the measurementA contractual promise with a remedy
AudienceEngineersEngineers and productCustomers, sales and legal
Typical targetn/aStricter, for example 99.95%Looser, for example 99.9%
Breach meansn/aFreeze risky changes, fix reliabilityPay credits, explain in writing
WindowAnyRolling 28 or 30 daysCalendar month, per tenant

The gap between SLO and SLA is deliberate: the SLO alarms your team long before the contract is at risk. The figures in the table are examples, not recommendations. Two details distinguish an SLA in practice. It is computed per tenant, because a regional outage that hit one large tenant hard is a breach for them even if the fleet average looks fine. And it is computed over a calendar month with fixed rules that both sides can recompute, because it ends up on an invoice.

What an agent SLA can honestly promise

An agent run is a long, multi-step process with a model provider and tools in the loop. Promise only what you control and can measure:

  • Availability of the run API: a run that the platform accepted completes without a platform-caused error. This is the core of most SLAs.
  • Responsiveness: time to the first event, a percentile over the month. Total run duration depends on the task and the tenant's tools, so it makes a poor promise.
  • Support response times by severity, which are process promises rather than telemetry.

Do not promise answer quality in an SLA; it cannot be measured objectively per run, and a dispute about it has no resolution. The harder question is the model provider. Your availability can never exceed what your providers deliver unless you can fail over between them. Read each provider's own SLA, note what it covers and what it excludes, and either set your promise below the dependency chain or put provider outages in your exclusions. Both are honest; silently promising more than your dependencies is not.

Responsiveness needs the same precision. Define the start (the gateway accepted the request) and the end (the first event reached the client), compute the percentile per tenant over the whole month rather than averaging daily percentiles, which is not mathematically valid, and state the minimum number of runs below which the latency promise does not apply, because a percentile over a handful of runs is noise.

Per-tenant measurement architecture

Per-tenant SLA pipeline: measure at the edge and in the runner, compute monthly, report and creditGatewayadmission outcomeRunner + SlaPluginrun outcome per tenantSLA event logappend-only, per runMinute bucketstenant x minuteMonthly calculatorbad minutes, exclusionsExclusion registermaintenance, tenant-causedCreditsapplied to next invoiceTenant reportavailability, latency, incidentsInternal SLOs and burn-rate alertsstricter than the SLA, page the on-callThe SLA is computed from the same events as the SLO,with contract rules applied on top.
Measurement is shared with the internal SLO pipeline; the SLA layer adds tenant scoping, contract exclusions and credits.

Two measurement points are needed. The gateway sees requests that never reach a runner: admission rejections, authentication failures and timeouts before dispatch. The runner sees how each accepted run ended. Both append to one event log keyed by tenant id, run id and the minute the run started. A batch job folds the log into tenant-by-minute buckets, and a monthly calculator applies the contract rules. Keep the raw log for the dispute period in your contract, because a tenant who questions a number deserves to see the runs behind it.

Classifying run outcomes in a plugin

In ADK Java, the run-level plugin callbacks are the natural hook: beforeRunCallback, afterRunCallback and onRunErrorCallback each receive the InvocationContext. If each tenant has its own runner, as in Agent Tenant Isolation, the plugin holds the tenant id as a field instead of trusting the request. TenantScope, SlaSink and the exception types are your own classes.

import com.google.adk.agents.InvocationContext;
import com.google.adk.events.Event;
import com.google.adk.plugins.BasePlugin;
import com.google.genai.types.Content;
import io.reactivex.rxjava3.core.Completable;
import io.reactivex.rxjava3.core.Maybe;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;

public final class SlaPlugin extends BasePlugin {
  public enum Outcome { GOOD, PLATFORM_ERROR, EXCLUDED_TENANT_LIMIT, EXCLUDED_TENANT_DEPENDENCY }
  public record SlaRecord(String tenantId, String runId, long startMinute, Outcome outcome, long millis) {}

  private final TenantScope t;
  private final SlaSink sink;                       // append-only, durable
  private final Map<String, Long> started = new ConcurrentHashMap<>();
  private final Set<String> errored = ConcurrentHashMap.newKeySet();

  public SlaPlugin(TenantScope t, SlaSink sink) {
    super("sla_" + t.tenantId());
    this.t = t;
    this.sink = sink;
  }

  @Override
  public Maybe<Content> beforeRunCallback(InvocationContext ctx) {
    started.put(ctx.invocationId(), System.currentTimeMillis());
    return Maybe.empty();
  }

  @Override
  public Maybe<Event> onEventCallback(InvocationContext ctx, Event event) {
    if (event.errorCode().isPresent() || event.errorMessage().isPresent())
      errored.add(ctx.invocationId());              // the run may still complete normally
    return Maybe.empty();                           // never alter the event stream
  }

  @Override
  public Completable afterRunCallback(InvocationContext ctx) {
    return Completable.fromAction(() -> record(ctx,
        errored.contains(ctx.invocationId()) ? Outcome.PLATFORM_ERROR : Outcome.GOOD));
  }

  @Override
  public Completable onRunErrorCallback(InvocationContext ctx, Throwable error) {
    return Completable.fromAction(() -> record(ctx, classify(error)));
  }

  static Outcome classify(Throwable e) {
    if (e instanceof TenantQuotaExceeded) return Outcome.EXCLUDED_TENANT_LIMIT;
    if (e instanceof TenantToolFailure) return Outcome.EXCLUDED_TENANT_DEPENDENCY;
    return Outcome.PLATFORM_ERROR;                  // unknown errors count against us
  }

  private void record(InvocationContext ctx, Outcome o) {
    errored.remove(ctx.invocationId());
    Long t0 = started.remove(ctx.invocationId());   // remove on both paths, or the maps leak
    if (t0 == null) return;
    long now = System.currentTimeMillis();
    sink.append(new SlaRecord(t.tenantId(), ctx.invocationId(), t0 / 60_000, o, now - t0));
  }
}

In the current adk-java runner, afterRunCallback is chained after the event stream completes, so it fires only on success, while a stream error goes to onRunErrorCallback. A run can still complete normally with an error recorded on one of its events, which is why the plugin also inspects errorCode() and errorMessage(). Two more traps: a plugin that returns a fallback response from onModelErrorCallback hides a provider error from this one, so record the error there too; and a run short-circuited by another plugin's beforeRunCallback still completes, so count admission rejections at the gateway instead.

Three rules make the data defensible. Unknown errors default to the platform's account, so an exclusion always needs a specific, typed cause. Tenant-caused failures, such as the tenant's own quota (see quota management) or the tenant's own tool backend returning errors, are recorded rather than dropped, so the report can show them. And runs that never finish do not fire either callback, so a sweeper must close any entry older than your run timeout as a platform error; otherwise hangs, the worst failures, vanish from the numbers.

Availability and credit maths

Request-based availability, good runs over all runs, is simple but lets a burst of traffic dominate the month. Many SLAs use minute-based availability instead: a minute is bad if the platform error rate among runs started in it exceeds a threshold, and monthly availability is the share of eligible minutes that were not bad. Minutes with no traffic count as good, which is generous to the provider; some contracts instead require a minimum number of runs in a minute before it is counted at all. Write down whichever you choose. The calculator below implements the minute-based rule; the credit tiers are illustrative, not a recommendation.

from dataclasses import dataclass

@dataclass
class Minute:
    total: int       # eligible runs started in this minute (exclusions removed)
    bad: int         # of those, runs with a platform-caused outcome
    excluded: bool   # inside an announced maintenance window, for example

# (availability floor in percent, credit in percent of the monthly fee). Illustrative.
CREDIT_TIERS = [(99.9, 0), (99.0, 10), (95.0, 25), (0.0, 50)]

def minute_is_bad(m, error_threshold=0.05):
    return m.total > 0 and m.bad / m.total > error_threshold

def monthly_availability(minutes):
    eligible = [m for m in minutes if not m.excluded]
    bad = sum(1 for m in eligible if minute_is_bad(m))
    return 100.0 * (len(eligible) - bad) / len(eligible), bad, len(eligible)

def credit_percent(availability):
    for floor, pct in CREDIT_TIERS:
        if availability >= floor:
            return pct

Worked example. A 30-day month has 43,200 minutes. A tenant on a 99.9% promise can therefore have about 43 bad minutes. This month the calculator finds 61 bad minutes, but 12 fell inside a maintenance window announced in line with the contract, so those minutes are excluded. That leaves 43,188 eligible minutes and 49 bad ones. Availability is (43,188 - 49) / 43,188, which is 99.887%: below 99.9 and above 99.0, so the illustrative tier gives a 10% credit. On an $8,000 monthly fee that is $800, applied to the next invoice through the same billing pipeline described in Agent Tenant Billing.

Exclusions that survive an audit

Exclusions are where SLA trust is won or lost. Typical categories are announced maintenance, failures caused by the tenant (exceeding their own limits, malformed input, their own tools failing), suspension for non-payment or abuse, and events beyond reasonable control. Each excluded minute or run must point to a record: a maintenance notice with its send time, an exception type with the run ids, a suspension entry in the tenant directory. Keep an exclusion register that the calculator joins against, never a manual adjustment applied after the fact.

Service tiers need scheduling differences

Selling a Gold tier at 99.95% and a Standard tier at 99.5% only makes sense if the platform actually treats them differently when resources are short. Options, from cheapest to strongest: priority in the admission queue, so Standard runs are shed first under overload; separate runner pools per tier, so a Standard surge cannot take Gold capacity; and dedicated capacity with its own provider quota for the top tier. Whatever you choose, the SLA calculation should still be per tenant, and capacity planning should check that the Gold pool can absorb a provider brownout without breaching. A tier name without a scheduling difference is marketing, and the credits will show it.

The monthly report

  • Monthly availability and the eligible, bad and excluded minute counts.
  • Time-to-first-event percentiles for the month.
  • Every incident that produced bad minutes, with start, end and a one-line cause.
  • Excluded minutes and runs, each with its category and reference.
  • Credit due and the invoice it will appear on.
  • How to dispute, and the deadline in the contract.

Generate the report from the calculator output, never by hand, and include a link to the run-level export so the tenant's engineers can recompute any figure. When a dispute arrives, the conversation should be about which rule applies to which runs, not about whose numbers are right.

Failure modes

  • Fleet-wide numbers in a per-tenant contract: the average hides the tenant who was hit. Compute per tenant from the start.
  • Hangs not counted: runs that never call back disappear. Sweep open entries.
  • Over-broad exclusions: classifying every model-provider error as tenant-caused inflates the number and fails the first audit.
  • Clock and window disagreement: define the time zone and the month boundaries in the contract and use the same ones in code.
  • SLA tighter than SLO: the team learns about breaches from invoices. Keep the internal target stricter.
  • Credits computed by hand: slow, inconsistent and hard to defend. Automate them from the event log.

Trade-offs

ChoiceGainCost
Minute-based availabilityBurst-resistant, easy to explainThreshold and empty-minute rules to agree
Request-based availabilityDirect link to user experienceBusy minutes dominate the month
Provider outages excludedPromise stays within your controlWeaker promise for the customer
Per-tier runner poolsReal differentiation under loadIdle capacity in the top tier

What to do next

  1. Write down the SLI, the window, the minute rule and the exclusion categories in one page that legal and engineering both sign.
  2. Add the SLA plugin and a gateway emitter, both writing to one append-only log keyed by tenant and run id.
  3. Add a sweeper that closes runs older than the timeout as platform errors.
  4. Build the exclusion register and make the calculator join against it.
  5. Recompute last month for your three largest tenants and compare with your SLO dashboards.
  6. Check each model provider's published SLA against your promise and decide: lower the promise or exclude provider outages.
  7. Generate the monthly report automatically and send credits through billing.
Key takeaway: An SLA is a per-tenant, calendar-month promise with a remedy, sitting below a stricter internal SLO. Promise availability and responsiveness, not answer quality, and stay within what your model providers deliver. Classify every run outcome with typed causes, count hangs, back every exclusion with a record, and compute availability and credits automatically from the same events your team already monitors.