A service level objective is a target for the fraction of events that go well, measured over a window: 99.9% of checkout requests succeed over 28 days. The target is the easy part. SLO engineering is everything around it: deciding which events count and where to count them, turning the objective into an error budget, computing that budget cheaply and correctly, alerting when it burns, and agreeing in advance what the organisation will do when it runs out. Teams that stop at the number get a dashboard nobody acts on. Teams that build the whole pipeline get a shared currency for trading reliability against feature velocity.
This article builds that pipeline from first principles with a worked example, a checkout API. Multi-window burn-rate alerting is a large topic of its own and is covered in burn-rate SLO alerting architecture; here the focus is on everything that feeds it and everything that consumes it.
The pipeline at a glance
Six artefacts make up an SLO system. The SLI is a ratio: good events divided by valid events, where valid excludes things the service is not responsible for, such as client-cancelled requests or 4xx caused by bad input. The SLO pairs that SLI with an objective and a window. The error budget is the permitted number of bad events, one minus the objective times valid events. Recording rules compute SLI rates continuously so that queries stay cheap. Alerts page when the budget burns fast enough to matter. The error-budget policy is a written agreement that says what happens at each level of consumption.
Treat the spec as source code. Every other artefact, recording rules, alerts, dashboards, is generated from it, reviewed in pull requests and versioned. The moment someone hand-edits an alert threshold, the alert and the objective start to disagree, and nobody notices until an incident.
Choosing SLIs that reflect users
Start from user journeys, not from metrics you happen to have. For checkout the journeys are 'load the cart', 'submit payment', 'see confirmation'. For each, ask what a user would call a failure: an error, or an answer so slow they gave up. That yields the two workhorse SLI types, availability (non-error responses over valid responses) and latency (responses faster than a threshold over valid responses). Pipelines use freshness and correctness; storage uses durability.
Where you measure changes what you see. Server-side metrics miss requests that never reached the server: a crashed pod, a bad deploy of the ingress, a DNS failure. Load-balancer logs catch those but not failures before the load balancer. Synthetic probes, discussed in synthetic monitoring, see the whole path but sample a tiny, unrepresentative slice of traffic. Real-user measurement sees everything but is noisy. A sound default is to measure availability at the load balancer, latency at the server with a histogram, and to add a synthetic probe for low-traffic paths.
Keep the set small: two or three SLIs per service, each tied to a journey. Twenty SLIs mean twenty budgets, and when everything has a budget nothing does.
Error-budget arithmetic, worked
Take an objective of 99.9% over a rolling 28-day window. The budget is 0.1% of valid events. If checkout serves 50 million valid requests in 28 days, the budget is 50,000 failed requests. Expressed as time for a service that is either fully up or fully down, 0.1% of 28 days is 40.32 minutes. The event form is the one to use, because real incidents are partial: a bad deploy that fails 20% of requests for 30 minutes at 1,240 requests per second spends 0.2 x 1,800 x 1,240 = 446,400 events. That single incident would exceed the budget almost nine times over.
Each additional nine divides the budget by ten. At 99.99% the same service may fail 5,000 requests a month, or about four minutes of full outage, which is shorter than most detection-plus-rollback times. Before choosing an objective, check it against your measured mean time to detect and recover: an objective whose budget is smaller than one ordinary incident is a promise you have already broken.
Burn rate is the speed of spending: the observed error ratio divided by the budget ratio. An error ratio of 1.44% against a 0.1% budget is a burn rate of 14.4, which would spend 28 days of budget in under two days. That single number is what burn-rate alerts evaluate over paired long and short windows.
| Objective | Budget ratio | Full-outage time per 28 days | Bad events per 50M |
|---|---|---|---|
| 99% | 1% | 6 h 43 min | 500,000 |
| 99.5% | 0.5% | 3 h 22 min | 250,000 |
| 99.9% | 0.1% | 40.3 min | 50,000 |
| 99.95% | 0.05% | 20.2 min | 25,000 |
| 99.99% | 0.01% | 4.0 min | 5,000 |
Latency SLIs from histograms
A latency SLI counts requests faster than a threshold, which is not the same as a percentile. 'The 99th percentile is under 300 ms' cannot be summed across instances or windows; 'at least 99% of requests finish within 300 ms' can, because it is just good events over valid events. With a Prometheus classic histogram the good count is the cumulative bucket at the threshold, so the threshold must be an actual bucket boundary. If your buckets are 0.25 and 0.5 and your threshold is 0.3, you cannot compute it exactly, and interpolating is guessing. Choose bucket boundaries from your SLO thresholds, as discussed in histogram architecture. Native histograms relax this: histogram_fraction estimates the fraction below any value, with an approximation error you should understand before paging on it.
groups:
- name: slo-checkout-latency
interval: 30s
rules:
# Latency SLI: a request is good if it finished within 300 ms (bucket le="0.3" must exist)
- record: slo:latency_good:rate5m
expr: sum(rate(http_server_duration_seconds_bucket{service="checkout-api",le="0.3"}[5m]))
- record: slo:latency_valid:rate5m
expr: sum(rate(http_server_duration_seconds_count{service="checkout-api"}[5m]))
# Budget remaining over a rolling 28 days, traffic-weighted (ratio of sums, not mean of ratios)
- record: slo:latency_budget_remaining:ratio28d
expr: |
1 - (
1 - sum_over_time(slo:latency_good:rate5m[28d])
/ sum_over_time(slo:latency_valid:rate5m[28d])
) / (1 - 0.99)The budget query uses a ratio of sums over 28 days, not an average of five-minute ratios. Averaging ratios gives a quiet 3 a.m. window the same weight as the peak hour, so a burst of errors at night looks far worse than it was and a peak-hour incident looks milder. Recording rules at five minutes keep a 28-day query cheap enough to put on a dashboard; see recording rules for naming and evaluation cost.
SLOs as code
Tools such as Sloth generate the full set of recording rules and multi-window burn-rate alerts from a short spec, so you declare intent and review it like any other code. The spec for the checkout availability SLO looks like this; the {{.window}} placeholder is filled in with each window the generator needs.
version: "prometheus/v1"
service: "checkout-api"
labels:
owner: "payments-team"
tier: "1"
slos:
- name: "requests-availability"
objective: 99.9
description: "Non-5xx responses for checkout requests, measured at the load balancer."
sli:
events:
error_query: sum(rate(lb_requests_total{service="checkout-api",code=~"5.."}[{{.window}}]))
total_query: sum(rate(lb_requests_total{service="checkout-api"}[{{.window}}]))
alerting:
name: "CheckoutAvailability"
page_alert:
labels:
severity: "page"
ticket_alert:
labels:
severity: "ticket"Running sloth generate -i checkout.yml emits Prometheus rule files that you commit and deploy through the same pipeline as application config. Put the spec next to the service code, require the owning team's review for changes, and add a CI check that fails if the generated rules differ from what is committed. OpenSLO offers a vendor-neutral spec format for the same idea; choose one format per organisation so dashboards and reports can be generated uniformly.
Dependencies, windows and low traffic
Serial dependencies multiply. A request that must pass through three services each at 99.9% succeeds at best 0.999 cubed, about 99.7%, if their failures are independent. So a user-facing SLO of 99.9% on top of three 99.9% hard dependencies is not achievable without retries, caching or graceful degradation. Write down each critical dependency's SLO and do this multiplication before committing to yours.
Rolling windows (the last 28 days) give a continuous signal and suit alerting and release decisions. Calendar windows (this month) align with reporting and contracts but reset abruptly, so a team that burned the budget on day 29 is clean again on day 1. Use rolling windows for engineering and keep calendar reports for customers. A 28-day window contains exactly four of each weekday, which avoids weekend bias.
Low-traffic services break the arithmetic. With 3,000 requests a month, a 99.9% budget is three failures; one unlucky retry storm spends it, and alerts flap on single events. Options: widen the window, merge related low-traffic endpoints into one SLI, add synthetic traffic so the SLI has a steady denominator, or pick a lower objective that matches what you can measure.
The error-budget policy
The policy is the part that makes SLOs change behaviour, and it must be agreed before it is needed, by the service team, its product owner and the people who run incidents. A workable policy for checkout:
- Budget above 50%: ship normally; spend budget on experiments and risky migrations deliberately.
- Budget between 0 and 50%: every release needs a rollback plan tested in staging; the top postmortem action item is scheduled this sprint.
- Budget exhausted: feature releases pause except for reliability fixes and security patches until the rolling budget recovers above 10%; the team lead reviews every exception.
- A single incident spending more than 20% of budget: requires a postmortem with owners and dates, linked from the SLO dashboard.
- Budget consistently unspent: the objective is probably too strict or the team is under-shipping; review the objective at the quarterly SLO review.
Pair the policy with runbooks so that the person paged by a burn alert knows the first three things to check. On-call runbooks are a separate artefact, but the SLO dashboard should link to them.
Failure modes
- Measuring the wrong place. Server metrics read 100% while the ingress is down. Measure availability at or before the edge.
- Vanity objectives. 99.99% chosen because it sounds good, with a budget smaller than one normal rollback. It is broken every month and ignored.
- Unbounded valid set. Counting 4xx from bots as bad events lets an attacker spend your budget; counting all 4xx as valid-but-good hides your own validation bugs. Decide explicitly.
- Hand-edited rules. Alerts drift from the spec; generate everything.
- Label churn. Renaming a metric label silently zeroes the SLI; alert when the valid-event rate drops to zero.
- Policy without teeth. An exhausted budget with no consequence teaches everyone that the number is decorative.
What to do next
- List your service's two or three most important user journeys and write one availability or latency SLI for each, stating the measurement point.
- Compute the error budget in events and in minutes, and compare it with your measured detect-plus-recover time.
- Set histogram bucket boundaries at your latency thresholds.
- Write the SLO spec as code, generate rules and alerts from it, and add a CI check against drift.
- Multiply your critical dependencies' SLOs and adjust the objective or the architecture.
- Draft the error-budget policy, get product and engineering sign-off, and schedule a quarterly review.