Three terms carry most of reliability engineering. A service level indicator (SLI) is a measurement of how well the service treated its users, almost always expressed as a ratio: good events divided by valid events. A service level objective (SLO) is a target for that ratio over a window, such as 99.5% of uploads succeed in each calendar month. The error budget is what the objective allows to go wrong: one minus the objective, multiplied by the volume of valid events. The work is in the details those definitions hide: which events you counted, where, which you discarded as invalid, and how you account for the budget as it drains.

This article is about that measurement and accounting work. Choosing SLIs from user journeys, histogram-based latency SLIs, SLOs as code and the error-budget policy are covered in the SLO engineering article, and turning budget consumption into pages is covered in burn-rate alerting. Here we take one service, a file-upload API with a 99.5% monthly objective, and build its SLI from raw data, decide every edge case, compute the budget, and keep a ledger of where the budget went. By the end you should be able to defend each number in an SLO report to a sceptical engineer.

Advertisement

The three definitions, made precise

An SLI needs four parts to be unambiguous: the event (an HTTP request to a named route, a message processed, a minute of a probe), the validity rule (which events the service is accountable for), the goodness rule (what success means, including any latency threshold), and the measurement point (whose logs or metrics produce the counts). Leave any of the four implicit and two engineers will compute different numbers from the same incident, and the SLO stops being a shared fact.

An SLO adds an objective and a window. The objective is a fraction, not a percentile: 'at least 99.5% of valid uploads are good'. The window is either rolling (the last 28 or 30 days, recomputed continuously) or calendar (each month, resetting on the first). The error budget follows mechanically. For our upload API, which receives about 6 million valid uploads in a typical month, a 99.5% objective allows 0.5% of 6,000,000, or 30,000 bad uploads. If the service were simply up or down, 0.5% of a 30-day month is 216 minutes, about three and a half hours. That one bridge from events to minutes is the only arithmetic this page needs; the rest is about getting the 6,000,000 and the 30,000 right.

Where you measure decides what you can see

Where an SLI can be measured, and what each point cannot seeUser / clientRUM, mobile SDKCDN / edgeedge logsLoad balanceraccess logs, metricsServiceserver metricsSynthetic probescripted journeyoutside-inSLI = good events / valid eventsper journey, per windowcountsError budget ledgerspend attributed by causeSLO reportobjective, window, remainingClient sees DNS, TLS and network failures; the load balancer sees requests that reached it;the service sees only requests it accepted. Measure as close to the user as your data allows.Synthetic probes give a signal when real traffic is too low to be meaningful.
Measurement points from the client inward. Each point is blind to failures that happen before traffic reaches it.

The same upload can be observed at five places, and each sees a different truth. Server-side metrics are cheap and precise but blind to anything that never reached the process: a crashed pod, a full accept queue, a misconfigured route. Load-balancer logs see every request that reached your edge, including the 502s and 504s your service never knew about; this is the usual default and the one our example uses. CDN or edge logs add TLS failures and requests rejected before the origin. Client-side telemetry (real user monitoring or a mobile SDK) sees DNS failures, network timeouts and retries the user sat through, but arrives late, is sampled, and is skewed by bad networks you do not control. Synthetic probes run a scripted journey from outside on a schedule; they see everything end to end but represent a robot, not your traffic mix.

The rule of thumb is to measure as close to the user as your data quality allows, then cover the blind spot with a second signal. For the upload API that means the SLI comes from load-balancer logs, and a synthetic probe uploading a small file every minute from two regions catches the case where the load balancer itself, or DNS, is broken and therefore logging nothing. A load balancer that receives no requests reports a perfect SLI; that is the most dangerous failure in this whole subject.

Advertisement

Deciding valid and good, edge case by edge case

Most SLO disputes are really disputes about the validity and goodness rules. Write them down as a table, review them with the people who will be held to the number, and treat changes to the table like code changes.

EventValid?Good?Reasoning
2xx response under 10 sYesYesThe upload worked within the latency the product promises
2xx response over 10 sYesNoThe user experienced a failure even though the status says success
5xx, including load-balancer 502/503/504YesNoService or infrastructure failure
429 rate limitedUsually yesUsually noValid and bad if the limit protects you from your own capacity shortfall; invalid if the client exceeded a documented quota
400, 401, 403, 404, 413Non/aClient error; the service behaved correctly. Watch for a 4xx spike caused by your own bad deploy
499 client closed requestDependsNo if validCount as valid and bad if the client gave up after a long wait; drop if it closed in under a second
Health checks, synthetic probesNon/aNot users; keep them in their own SLI
Known bots and load testsNon/aFilter by user agent or tag; they distort both numerator and denominator
Client retry of a failed uploadYesOwn outcomeEach attempt is an event; a separate journey SLI can count uploads that eventually succeeded

Two traps deserve emphasis. First, timeouts that never log: if the client gives up and the server finishes later, the server logs a 200 and the user saw a failure. Measuring at the load balancer with an idle timeout shorter than the client timeout, or using client telemetry, closes the gap. Second, a deploy that turns valid requests into 400s will vanish from an SLI that excludes 4xx. Alert on the 4xx rate separately, or exclude only the specific codes and routes you have reasoned about rather than the whole class.

Request-based versus bad-minute SLIs

A request-based SLI counts events: good uploads over valid uploads. It weights each user interaction equally, so an outage at peak burns budget faster than the same outage at 4 a.m., which matches user impact. It needs enough traffic for the ratio to mean something.

A time-window or bad-minute SLI divides time into slices, marks each slice good or bad by a rule (for example, 'a minute is bad if more than 1% of its requests failed or its p95 latency exceeded 10 s'), and counts good minutes over total minutes. It is easy to explain to customers, and it is what many SLAs use, but it hides partial degradation below the threshold and treats a quiet minute the same as a busy one.

PropertyRequest-basedBad-minute
Weights by user impactYesNo, every minute counts equally
Works at low trafficPoorly: one failure in ten requests is 10%Better, especially with synthetic probes filling empty minutes
Partial degradationCounted proportionallyInvisible below the per-minute threshold
Easy to put in a contractModeratelyYes, minutes of downtime are intuitive
Budget unitBad eventsBad minutes

Our upload API uses request-based for its internal SLO and reports bad minutes only where a customer contract demands it. If you run both, compute them from the same validity rules so they never disagree about what an event was.

Building the ratio from logs

Logs are the most transparent source because every event is visible and every rule is a line of SQL anyone can review. Assume load-balancer access logs land in a table with a timestamp, route, status, duration and user agent:

-- Monthly SLI and budget for the upload API, from load-balancer logs.
WITH events AS (
  SELECT
    ts,
    status,
    duration_ms,
    CASE
      WHEN user_agent LIKE 'synthetic-probe/%'          THEN FALSE
      WHEN user_agent LIKE 'loadtest/%'                 THEN FALSE
      WHEN status IN (400, 401, 403, 404, 413)          THEN FALSE
      WHEN status = 499 AND duration_ms < 1000          THEN FALSE
      ELSE TRUE
    END AS is_valid,
    (status BETWEEN 200 AND 299 AND duration_ms <= 10000) AS is_good
  FROM lb_access_log
  WHERE route = 'POST /v1/uploads'
    AND ts >= DATE '2026-09-01' AND ts < DATE '2026-10-01'
)
SELECT
  COUNT(*) FILTER (WHERE is_valid)                         AS valid,
  COUNT(*) FILTER (WHERE is_valid AND is_good)             AS good,
  1.0 * COUNT(*) FILTER (WHERE is_valid AND is_good)
      / NULLIF(COUNT(*) FILTER (WHERE is_valid), 0)        AS sli,
  0.005 * COUNT(*) FILTER (WHERE is_valid)                 AS budget_events,
  COUNT(*) FILTER (WHERE is_valid AND NOT is_good)         AS spent_events
FROM events;

Suppose September returns 6,120,000 valid uploads and 6,094,300 good ones. The SLI is 99.58%, the budget is 30,600 bad uploads, and 25,700 were spent, leaving 4,900, or 16% of the budget. The budget comes from actual valid volume, not a forecast, so a traffic surge enlarges it in events but not as a fraction.

Building the same ratio from metrics

Logs are exact but slow and expensive to query continuously. For dashboards and alerts, build the same rules into counters. The rules must match the SQL; differences between the two are a classic source of 'the dashboard says green, the report says red'.

# Prometheus recording rules for the upload SLI (5m rates, aggregated later).
groups:
- name: upload-sli
  rules:
  - record: upload:valid:rate5m
    expr: |
      sum(rate(lb_requests_total{route="POST /v1/uploads",
                                 code!~"400|401|403|404|413",
                                 client_class="user"}[5m]))
  - record: upload:good:rate5m
    expr: |
      sum(rate(lb_request_duration_seconds_bucket{route="POST /v1/uploads",
                                 code=~"2..", le="10.0",
                                 client_class="user"}[5m]))
# 30-day SLI: ratio of sums, never an average of ratios.
#   sum_over_time(upload:good:rate5m[30d]) / sum_over_time(upload:valid:rate5m[30d])

The good counter uses the histogram bucket at the 10-second boundary, which only works if a bucket boundary sits exactly at your threshold (Prometheus 3 stores bucket bounds as floats, hence 10.0); see histograms for why. The client_class label is assigned at the edge from the user-agent rules, so the bot and probe filtering happens once, upstream, instead of being repeated in every query. Precomputing these series, as described in recording rules, keeps month-long queries cheap. The 499 duration rule is not expressible in labels; emit a separate counter if it matters.

Windows and reporting

A rolling window answers 'how are users being treated right now' and is what alerts and release decisions should use: there is no cliff at month end, and a bad week stays visible for exactly the window length. A calendar window answers 'did we meet the objective this month' and is what reports and contracts use, because it maps to billing periods and resets cleanly. Label every chart with its window: a rolling budget can be nearly empty on the second of the month while the calendar budget is full.

A monthly SLO report needs six numbers: valid events, good events, SLI, objective, budget in events and budget remaining, plus the ledger summary below and one sentence of interpretation.

The error budget ledger

A budget that only shows 'remaining: 16%' tells you that something went wrong but not what to fix. Keep a ledger: every significant spend attributed to a cause, with the event count and a link to the incident or change. The attribution can be semi-automatic, by joining bad events in time against deploy events, incident windows and dependency status, then confirmed by a human at the weekly review.

DateCauseCategoryBad uploadsShare of budget
Sep 4Object store throttling in one regionDependency11,20036.6%
Sep 12Deploy 4.18 leaked file handlesChange8,90029.1%
Sep 12-30Background error rate, unattributedBaseline4,10013.4%
Sep 21Certificate rotation restarted podsOperations1,5004.9%
Total25,70084.0%

Categories turn the ledger into priorities. Here, dependency and change failures dominate, which argues for regional failover on the object store and for canarying deploys, not for a general reliability push. The baseline line matters too: if unexplained background errors alone consume a third of the budget, the objective is too tight for the current system or there is a slow bug nobody has looked for.

SLOs, SLAs and the margin between them

A service level agreement is a contract: it names a measurement, a threshold and a consequence, usually service credits. The SLO is your internal target and should be stricter than the SLA so that you notice and react long before money changes hands. If the contract promises 99.0% monthly availability, an internal SLO of 99.5% gives a margin of half a percent, which for our volumes is roughly 30,000 uploads of warning. Keep the SLA's definitions in the same validity table as the SLO so the two cannot drift apart.

Failure modes

  • The silent edge. The measurement point stops receiving traffic and the SLI reads 100%. Cover with synthetic probes and an alert on valid-event volume dropping.
  • Rule drift. Dashboards, alerts and the monthly report each re-implement the validity rules slightly differently. Define them once, upstream, and test them.
  • Averaging ratios. Averaging daily SLIs weights a quiet Sunday like a busy Monday. Always divide summed good by summed valid.
  • Low traffic. A handful of requests makes the ratio swing wildly. Lengthen the window, use bad minutes, or lean on probes.
  • Excluding too much. Each exclusion makes the number look better and the user experience no better. Review the validity table every quarter.

What to do next

  1. For your most important journey, write the four-part SLI definition: event, validity rule, goodness rule and measurement point.
  2. Build the validity table, including 429, 499, bots, probes and retries, and get the owning team to agree to it.
  3. Compute last month's SLI from logs with a reviewed query, then build matching counters and check the two agree to within a rounding error.
  4. Add a synthetic probe and an alert on valid-event volume so a dead measurement point cannot report perfection.
  5. Start a budget ledger, attribute every spend over 5% of the budget, and review it weekly with categories.
  6. If you have an SLA, record its definitions next to the SLO and confirm your objective leaves a margin above it.
Key takeaway: An SLI is good events over valid events, and every word in that sentence needs a written rule: which events, which are valid, what is good, and where they are counted. Measure as close to the user as data allows and cover the blind spot with probes, compute the ratio identically from logs and metrics, report on calendar windows but decide on rolling ones, and keep a ledger that says where the budget went so the next reliability investment is aimed rather than guessed.