Writing one service level objective is easy: pick a metric, pick a number, draw a line on a dashboard. Getting forty teams to own SLOs that actually change what they work on is an organisational system with a technical backbone. Most SLO initiatives stall in one of two places. They either stay a dashboard nobody looks at, or they turn into a paging firehose that on-call engineers learn to ignore.
This article treats an SLO program as architecture. It covers the components you need, the data flow from telemetry to decisions, the arithmetic behind targets and alerts, and a phased rollout with explicit entry and exit criteria for each phase. The mechanics of spending a budget, such as freezes and release gates, are covered in the error budget guide. This page is about getting an organisation to the point where those policies can work.
What the program consists of
An SLO program has six components, and a missing one is usually the reason a rollout stalls.
- Specifications: one version-controlled file per service declaring its SLIs, objectives, window and owner.
- An SLI pipeline: telemetry that measures good and total events where users experience them.
- A generator: CI that validates specs and renders recording rules, alerts and dashboards, so that nothing is hand-written.
- Alerting: burn-rate alerts routed to the owning team, each with a runbook.
- Reporting: budget remaining per service, per window, visible to engineering and product leadership.
- Governance: an agreed budget policy and a regular review that changes targets and priorities.
Choosing SLIs from user journeys
Start from what users do, not from what you already measure. For an online shop, the critical journeys might be browse, add to cart and check out. For each journey, ask what "good" means to a user: the request succeeded, and it was fast enough. That yields the two SLI types that cover most services: availability (good events divided by valid events) and latency (requests faster than a threshold divided by valid requests). Measure as close to the user as you can. A load balancer or API gateway log sees failures that the service's own metrics miss, such as a pod that is not running at all.
Worked example. The checkout service handles 10 million requests in a 30-day window. An availability objective of 99.9 percent leaves an error budget of 0.1 percent: 10,000 failed requests. Expressed as time, 0.1 percent of 30 days is 43.2 minutes of total outage. A latency objective of 99 percent of requests under 300 ms allows 100,000 slow requests. Now check the target against history. If checkout achieved 99.95 percent over the last three months, 99.9 leaves room for normal change. If it achieved 99.7, setting 99.9 means the budget is exhausted from the first day and the SLO carries no information. Set it at 99.5 and plan the reliability work needed to raise it.
Resist two temptations. Do not set the objective to current performance plus a nine because it sounds good, since every nine costs roughly ten times the engineering effort of the one before. And do not start with ten SLIs per service. Two or three per critical journey is plenty, and each extra one dilutes attention.
SLOs as code
The spec is the contract, so keep it in Git next to the service or in a central repository, review it like code, and generate everything else from it. Here is a minimal in-house format.
# slos/checkout.yaml (in-house format; tools such as Sloth and OpenSLO use similar ideas)
service: checkout
owner: team-payments
tier: 1
window_days: 30
slos:
- name: availability
objective: 99.9
sli:
good: sum(rate(http_requests_total{job="checkout",code!~"5.."}[{{window}}]))
total: sum(rate(http_requests_total{job="checkout"}[{{window}}]))
- name: latency
objective: 99.0
sli: # requests under 300 ms, from a histogram bucket
good: sum(rate(http_request_duration_seconds_bucket{job="checkout",le="0.3"}[{{window}}]))
total: sum(rate(http_request_duration_seconds_count{job="checkout"}[{{window}}]))A small generator validates each spec and renders Prometheus recording rules and alerts. The recording rules precompute the error ratio over each window so that alert evaluation is cheap.
import yaml
WINDOWS = ["5m", "30m", "1h", "6h", "3d"]
# (long window, short window, burn rate, severity): SRE workbook multiwindow table
ALERTS = [("1h", "5m", 14.4, "page"), ("6h", "30m", 6, "page"), ("3d", "6h", 1, "ticket")]
def render(spec: dict) -> dict:
rules, alerts = [], []
for slo in spec["slos"]:
name = f'{spec["service"]}_{slo["name"]}'
budget = 1 - slo["objective"] / 100 # 99.9 -> 0.001
for w in WINDOWS:
good = slo["sli"]["good"].replace("{{window}}", w)
total = slo["sli"]["total"].replace("{{window}}", w)
rules.append({"record": f"slo:error_ratio:{w}",
"expr": f"1 - ({good}) / ({total})",
"labels": {"slo": name}})
for long_w, short_w, rate, sev in ALERTS:
thr = rate * budget
alerts.append({
"alert": f"{name}_burn_{long_w}",
"expr": (f'slo:error_ratio:{long_w}{{slo="{name}"}} > {thr:.6f} and '
f'slo:error_ratio:{short_w}{{slo="{name}"}} > {thr:.6f}'),
"labels": {"severity": sev, "team": spec["owner"]},
"annotations": {"runbook": f"runbooks/{name}.md"}})
return {"groups": [{"name": f'slo-{spec["service"]}', "rules": rules + alerts}]}
def validate(spec: dict) -> None:
assert spec.get("owner"), "every SLO needs an owning team"
for slo in spec["slos"]:
assert 90 <= slo["objective"] < 100, f'{slo["name"]}: objective out of range'The generator is where you enforce program rules: every SLO has an owner, objectives are within sane bounds, and every alert links a runbook (see the runbook guide). Open-source generators such as Sloth follow the same pattern, and the OpenSLO project defines a vendor-neutral spec format. Whichever you choose, the principle holds: nobody hand-edits a rule file.
Alerting on burn rate
Alerting when the error ratio crosses the objective is too noisy for short blips and too slow for real incidents. Instead, alert on burn rate: how fast the budget is being consumed relative to a steady spend that would use it exactly over the window. A burn rate of 1 uses the whole 30-day budget in 30 days. A burn rate of 14.4 sustained for one hour uses 14.4 x 1/720 = 2 percent of the budget in that hour.
| Budget consumed | Long window | Short window | Burn rate | Action |
|---|---|---|---|---|
| 2% | 1 hour | 5 minutes | 14.4 | Page |
| 5% | 6 hours | 30 minutes | 6 | Page |
| 10% | 3 days | 6 hours | 1 | Ticket |
This multiwindow, multi-burn-rate scheme comes from the Google SRE Workbook. The long window gives significance, so a two-minute blip does not page. The short window, both conditions being required, makes the alert stop firing soon after the problem is fixed. For checkout at 99.9 percent, the fast page fires when the error ratio exceeds 14.4 x 0.001 = 1.44 percent over both the last hour and the last five minutes.
Low-traffic services break this arithmetic. At 100 requests an hour, a single failure is a 1 percent error ratio. Either aggregate low-traffic endpoints into one SLO, add synthetic probes to create a steady baseline, or alert only on the slower windows.
Rolling it out in phases
A program-wide launch fails because every team hits the same problems at once. Roll out in phases, each with a clear gate.
| Phase | Scope | Exit criteria |
|---|---|---|
| 0. Foundations | Platform team | Spec format, generator in CI, recording rules at scale, one dashboard template |
| 1. Pilot | 2 to 3 willing tier-1 services | SLOs reviewed with product owners, 4 weeks of data, targets adjusted at least once |
| 2. Shadow alerts | Pilot services | Burn-rate alerts go to a channel, not a pager; each firing is triaged as real or false |
| 3. Paging cutover | Pilot services | False page rate acceptable to on-call; old threshold alerts removed |
| 4. Policy | Pilot services | Budget policy signed by engineering and product; first budget-driven decision made |
| 5. Scale out | All tier-1, then tier-2 | Each team completes phases 1 to 4 with a platform partner; coverage tracked |
Two gates matter most. Shadow alerting (phase 2) is where you discover that an SLI counts health checks as traffic or that a latency histogram's buckets do not include your threshold. Discovering that on a live pager destroys trust. Removing old threshold alerts at cutover (phase 3) is what makes the program a net reduction in noise. If SLO alerts are only added and never replace anything, on-call load rises and teams resist.
Tie the rollout to change processes that already exist. For example, releases through canary analysis can use the SLO's error ratio as the canary metric, which makes the SLI earn its keep during deployments, not just during incidents.
Governance: turning numbers into decisions
An SLO that never changes a decision is decoration. Hold a monthly review per group of services with engineering and product leads. It looks at three things: budget remaining for each SLO, the incidents that consumed budget, and targets that were never at risk or always breached. Out of that review come concrete outcomes: a target adjusted, a reliability project prioritised, or a feature launch delayed under the budget policy. Keep a written log of those decisions, because the log is the evidence that the program works.
Assign ownership explicitly. Every SLO belongs to one team, recorded in the spec and in the service catalog. SLOs for shared dependencies, such as a database platform or an identity service, are owned by the platform team. Their consumers treat those SLOs as the inputs for their own objectives: a service cannot sensibly promise 99.99 percent on top of a dependency that promises 99.9 without redundancy.
Scaling the platform
Several data-path problems appear at program scale. Thirty-day windows computed from raw series are expensive, so compute ratios through recording rules and keep them for longer than the raw data. A three-day burn window needs three days of recorded history on the alerting path. High-cardinality labels in SLI queries multiply rule cost, so aggregate before recording. The monitoring system is itself a dependency: if it is down, you have no alerts, so give it an independent dead-man's-switch alert that fires when expected rule evaluations stop arriving. Finally, render dashboards from the same specs so that every service's SLO page looks the same, which makes the monthly review fast.
Failure modes and trade-offs
- SLO theatre. Specs exist, dashboards are green, and nobody has ever changed a plan because of them. Measure the program by decisions logged, not by SLO count.
- Targets set at aspiration. A target the service has never met is permanently breached and quickly ignored. Start from history.
- Measuring in the wrong place. Server-side metrics miss load balancer rejections, DNS failures and dead pods. Prefer edge or client measurements for availability.
- Alert sprawl. Adding burn-rate alerts without deleting old threshold alerts doubles on-call load.
- Too many SLOs. Twelve objectives per service means nobody knows which one matters. Two or three per critical journey is enough.
- Trade-off: central versus team-owned. A central SRE team can move fast but produces SLOs that teams do not believe. Team-owned SLOs are credible but inconsistent. The platform-plus-partner model in the phase table is the usual compromise.
- Trade-off: strictness versus adoption. Enforcing budget freezes from day one gets you resistance. Introduce the policy only after teams trust the numbers.
What to do next
- List the three to five most important user journeys and the services behind each.
- Choose two or three pilot services with willing owners, and draft availability and latency SLIs from 90 days of history.
- Build the spec format and a generator in CI before writing any rules by hand.
- Run burn-rate alerts in shadow mode for at least two weeks and triage every firing.
- Cut over to paging, delete the threshold alerts the SLO replaces, and write a runbook for each alert.
- Get the budget policy signed, start the monthly review, and log every decision it makes.