Three terms carry most of reliability engineering. A service level indicator (SLI) is a measurement of how well the service treated its users, almost always expressed as a ratio: good events divided by valid events. A service level objective (SLO) is a target for that ratio over a window, such as 99.5% of uploads succeed in each calendar month. The error budget is what the objective allows to go wrong: one minus the objective, multiplied by the volume of valid events. The work is in the details those definitions hide: which events you counted, where, which you discarded as invalid, and how you account for the budget as it drains.
This article is about that measurement and accounting work. Choosing SLIs from user journeys, histogram-based latency SLIs, SLOs as code and the error-budget policy are covered in the SLO engineering article, and turning budget consumption into pages is covered in burn-rate alerting. Here we take one service, a file-upload API with a 99.5% monthly objective, and build its SLI from raw data, decide every edge case, compute the budget, and keep a ledger of where the budget went. By the end you should be able to defend each number in an SLO report to a sceptical engineer.
The three definitions, made precise
An SLI needs four parts to be unambiguous: the event (an HTTP request to a named route, a message processed, a minute of a probe), the validity rule (which events the service is accountable for), the goodness rule (what success means, including any latency threshold), and the measurement point (whose logs or metrics produce the counts). Leave any of the four implicit and two engineers will compute different numbers from the same incident, and the SLO stops being a shared fact.
An SLO adds an objective and a window. The objective is a fraction, not a percentile: 'at least 99.5% of valid uploads are good'. The window is either rolling (the last 28 or 30 days, recomputed continuously) or calendar (each month, resetting on the first). The error budget follows mechanically. For our upload API, which receives about 6 million valid uploads in a typical month, a 99.5% objective allows 0.5% of 6,000,000, or 30,000 bad uploads. If the service were simply up or down, 0.5% of a 30-day month is 216 minutes, about three and a half hours. That one bridge from events to minutes is the only arithmetic this page needs; the rest is about getting the 6,000,000 and the 30,000 right.
Where you measure decides what you can see
The same upload can be observed at five places, and each sees a different truth. Server-side metrics are cheap and precise but blind to anything that never reached the process: a crashed pod, a full accept queue, a misconfigured route. Load-balancer logs see every request that reached your edge, including the 502s and 504s your service never knew about; this is the usual default and the one our example uses. CDN or edge logs add TLS failures and requests rejected before the origin. Client-side telemetry (real user monitoring or a mobile SDK) sees DNS failures, network timeouts and retries the user sat through, but arrives late, is sampled, and is skewed by bad networks you do not control. Synthetic probes run a scripted journey from outside on a schedule; they see everything end to end but represent a robot, not your traffic mix.
The rule of thumb is to measure as close to the user as your data quality allows, then cover the blind spot with a second signal. For the upload API that means the SLI comes from load-balancer logs, and a synthetic probe uploading a small file every minute from two regions catches the case where the load balancer itself, or DNS, is broken and therefore logging nothing. A load balancer that receives no requests reports a perfect SLI; that is the most dangerous failure in this whole subject.
Deciding valid and good, edge case by edge case
Most SLO disputes are really disputes about the validity and goodness rules. Write them down as a table, review them with the people who will be held to the number, and treat changes to the table like code changes.
| Event | Valid? | Good? | Reasoning |
|---|---|---|---|
| 2xx response under 10 s | Yes | Yes | The upload worked within the latency the product promises |
| 2xx response over 10 s | Yes | No | The user experienced a failure even though the status says success |
| 5xx, including load-balancer 502/503/504 | Yes | No | Service or infrastructure failure |
| 429 rate limited | Usually yes | Usually no | Valid and bad if the limit protects you from your own capacity shortfall; invalid if the client exceeded a documented quota |
| 400, 401, 403, 404, 413 | No | n/a | Client error; the service behaved correctly. Watch for a 4xx spike caused by your own bad deploy |
| 499 client closed request | Depends | No if valid | Count as valid and bad if the client gave up after a long wait; drop if it closed in under a second |
| Health checks, synthetic probes | No | n/a | Not users; keep them in their own SLI |
| Known bots and load tests | No | n/a | Filter by user agent or tag; they distort both numerator and denominator |
| Client retry of a failed upload | Yes | Own outcome | Each attempt is an event; a separate journey SLI can count uploads that eventually succeeded |
Two traps deserve emphasis. First, timeouts that never log: if the client gives up and the server finishes later, the server logs a 200 and the user saw a failure. Measuring at the load balancer with an idle timeout shorter than the client timeout, or using client telemetry, closes the gap. Second, a deploy that turns valid requests into 400s will vanish from an SLI that excludes 4xx. Alert on the 4xx rate separately, or exclude only the specific codes and routes you have reasoned about rather than the whole class.
Request-based versus bad-minute SLIs
A request-based SLI counts events: good uploads over valid uploads. It weights each user interaction equally, so an outage at peak burns budget faster than the same outage at 4 a.m., which matches user impact. It needs enough traffic for the ratio to mean something.
A time-window or bad-minute SLI divides time into slices, marks each slice good or bad by a rule (for example, 'a minute is bad if more than 1% of its requests failed or its p95 latency exceeded 10 s'), and counts good minutes over total minutes. It is easy to explain to customers, and it is what many SLAs use, but it hides partial degradation below the threshold and treats a quiet minute the same as a busy one.
| Property | Request-based | Bad-minute |
|---|---|---|
| Weights by user impact | Yes | No, every minute counts equally |
| Works at low traffic | Poorly: one failure in ten requests is 10% | Better, especially with synthetic probes filling empty minutes |
| Partial degradation | Counted proportionally | Invisible below the per-minute threshold |
| Easy to put in a contract | Moderately | Yes, minutes of downtime are intuitive |
| Budget unit | Bad events | Bad minutes |
Our upload API uses request-based for its internal SLO and reports bad minutes only where a customer contract demands it. If you run both, compute them from the same validity rules so they never disagree about what an event was.
Building the ratio from logs
Logs are the most transparent source because every event is visible and every rule is a line of SQL anyone can review. Assume load-balancer access logs land in a table with a timestamp, route, status, duration and user agent:
-- Monthly SLI and budget for the upload API, from load-balancer logs.
WITH events AS (
SELECT
ts,
status,
duration_ms,
CASE
WHEN user_agent LIKE 'synthetic-probe/%' THEN FALSE
WHEN user_agent LIKE 'loadtest/%' THEN FALSE
WHEN status IN (400, 401, 403, 404, 413) THEN FALSE
WHEN status = 499 AND duration_ms < 1000 THEN FALSE
ELSE TRUE
END AS is_valid,
(status BETWEEN 200 AND 299 AND duration_ms <= 10000) AS is_good
FROM lb_access_log
WHERE route = 'POST /v1/uploads'
AND ts >= DATE '2026-09-01' AND ts < DATE '2026-10-01'
)
SELECT
COUNT(*) FILTER (WHERE is_valid) AS valid,
COUNT(*) FILTER (WHERE is_valid AND is_good) AS good,
1.0 * COUNT(*) FILTER (WHERE is_valid AND is_good)
/ NULLIF(COUNT(*) FILTER (WHERE is_valid), 0) AS sli,
0.005 * COUNT(*) FILTER (WHERE is_valid) AS budget_events,
COUNT(*) FILTER (WHERE is_valid AND NOT is_good) AS spent_events
FROM events;Suppose September returns 6,120,000 valid uploads and 6,094,300 good ones. The SLI is 99.58%, the budget is 30,600 bad uploads, and 25,700 were spent, leaving 4,900, or 16% of the budget. The budget comes from actual valid volume, not a forecast, so a traffic surge enlarges it in events but not as a fraction.
Building the same ratio from metrics
Logs are exact but slow and expensive to query continuously. For dashboards and alerts, build the same rules into counters. The rules must match the SQL; differences between the two are a classic source of 'the dashboard says green, the report says red'.
# Prometheus recording rules for the upload SLI (5m rates, aggregated later).
groups:
- name: upload-sli
rules:
- record: upload:valid:rate5m
expr: |
sum(rate(lb_requests_total{route="POST /v1/uploads",
code!~"400|401|403|404|413",
client_class="user"}[5m]))
- record: upload:good:rate5m
expr: |
sum(rate(lb_request_duration_seconds_bucket{route="POST /v1/uploads",
code=~"2..", le="10.0",
client_class="user"}[5m]))
# 30-day SLI: ratio of sums, never an average of ratios.
# sum_over_time(upload:good:rate5m[30d]) / sum_over_time(upload:valid:rate5m[30d])The good counter uses the histogram bucket at the 10-second boundary, which only works if a bucket boundary sits exactly at your threshold (Prometheus 3 stores bucket bounds as floats, hence 10.0); see histograms for why. The client_class label is assigned at the edge from the user-agent rules, so the bot and probe filtering happens once, upstream, instead of being repeated in every query. Precomputing these series, as described in recording rules, keeps month-long queries cheap. The 499 duration rule is not expressible in labels; emit a separate counter if it matters.
Windows and reporting
A rolling window answers 'how are users being treated right now' and is what alerts and release decisions should use: there is no cliff at month end, and a bad week stays visible for exactly the window length. A calendar window answers 'did we meet the objective this month' and is what reports and contracts use, because it maps to billing periods and resets cleanly. Label every chart with its window: a rolling budget can be nearly empty on the second of the month while the calendar budget is full.
A monthly SLO report needs six numbers: valid events, good events, SLI, objective, budget in events and budget remaining, plus the ledger summary below and one sentence of interpretation.
The error budget ledger
A budget that only shows 'remaining: 16%' tells you that something went wrong but not what to fix. Keep a ledger: every significant spend attributed to a cause, with the event count and a link to the incident or change. The attribution can be semi-automatic, by joining bad events in time against deploy events, incident windows and dependency status, then confirmed by a human at the weekly review.
| Date | Cause | Category | Bad uploads | Share of budget |
|---|---|---|---|---|
| Sep 4 | Object store throttling in one region | Dependency | 11,200 | 36.6% |
| Sep 12 | Deploy 4.18 leaked file handles | Change | 8,900 | 29.1% |
| Sep 12-30 | Background error rate, unattributed | Baseline | 4,100 | 13.4% |
| Sep 21 | Certificate rotation restarted pods | Operations | 1,500 | 4.9% |
| Total | 25,700 | 84.0% |
Categories turn the ledger into priorities. Here, dependency and change failures dominate, which argues for regional failover on the object store and for canarying deploys, not for a general reliability push. The baseline line matters too: if unexplained background errors alone consume a third of the budget, the objective is too tight for the current system or there is a slow bug nobody has looked for.
SLOs, SLAs and the margin between them
A service level agreement is a contract: it names a measurement, a threshold and a consequence, usually service credits. The SLO is your internal target and should be stricter than the SLA so that you notice and react long before money changes hands. If the contract promises 99.0% monthly availability, an internal SLO of 99.5% gives a margin of half a percent, which for our volumes is roughly 30,000 uploads of warning. Keep the SLA's definitions in the same validity table as the SLO so the two cannot drift apart.
Failure modes
- The silent edge. The measurement point stops receiving traffic and the SLI reads 100%. Cover with synthetic probes and an alert on valid-event volume dropping.
- Rule drift. Dashboards, alerts and the monthly report each re-implement the validity rules slightly differently. Define them once, upstream, and test them.
- Averaging ratios. Averaging daily SLIs weights a quiet Sunday like a busy Monday. Always divide summed good by summed valid.
- Low traffic. A handful of requests makes the ratio swing wildly. Lengthen the window, use bad minutes, or lean on probes.
- Excluding too much. Each exclusion makes the number look better and the user experience no better. Review the validity table every quarter.
What to do next
- For your most important journey, write the four-part SLI definition: event, validity rule, goodness rule and measurement point.
- Build the validity table, including 429, 499, bots, probes and retries, and get the owning team to agree to it.
- Compute last month's SLI from logs with a reviewed query, then build matching counters and check the two agree to within a rounding error.
- Add a synthetic probe and an alert on valid-event volume so a dead measurement point cannot report perfection.
- Start a budget ledger, attribute every spend over 5% of the budget, and review it weekly with categories.
- If you have an SLA, record its definitions next to the SLO and confirm your objective leaves a margin above it.