Most SLO guides start with an HTTP service: count requests, count the good ones, divide. That works for an API, and this site covers it in SLO engineering and SLIs, SLOs and error budgets. But a large share of what breaks in production has no request to count. A data pipeline that silently stops loading still answers every dashboard query in 40 ms, with yesterday's numbers. A queue whose workers crashed returns 202 Accepted to every producer. A storage system that loses one object in a million passes every latency check. A nightly job that never started emits no errors at all.

A service level indicator is a measurement of one property users care about, expressed as the proportion of good events (or good time) out of valid events. A service level objective is a target for that proportion over a window, such as 99.5% of minutes fresh over 28 days. This article is about the hard half: finding the right event for each kind of system, measuring it from the right place, and then proving, with data, that the indicator turns bad when users are hurt and stays good when they are not. An SLI that has never been checked against user pain is a guess with a dashboard.

Advertisement

Specification first, implementation second

Separate two things that teams usually blur. The SLI specification is a sentence about user experience: the share of orders that appear in the finance warehouse within 15 minutes of being placed. The SLI implementation is how you approximate it: the age of the newest committed row, sampled every minute by Prometheus. The specification rarely changes; the implementation changes whenever the architecture does, and every implementation is lossy in some way you should be able to name.

The way to find the specification is to ask what the user does with the output and when they would notice it was wrong. For a pipeline, the reader notices stale or missing data, not slow batches. For a queue, the person waiting for an email notices the total time from clicking send, not how long the worker spent. For storage, the owner notices data that cannot be read back, possibly years later. Each answer points to a different event, a different clock and a different measurement point.

Where the SLI is measured depends on the shape of the system, not on the toolingData pipelinesource to warehouseQueue and workersenqueue to doneStoragewrite, keep, read backScheduled jobone run per periodFreshnessnow minus committedwatermark, per minuteTimely completiondone within T of enqueueplus age of oldest itemDurabilityprobe objects read backand scrub mismatchesOn-time successper run, not per minutea missed run is badValidationcompare bad SLI periods with tickets, incidents and RUMAn SLI is only finished when it is bad during the periods users were hurt, and good otherwise.Precision: of the bad periods, how many hurt someone. Recall: of the hurt periods, how many the SLI saw.
Four system shapes and the indicator that matches each. All four feed one validation step, which checks the indicator against independent evidence of user pain.

A menu by system shape

ShapeWhat users noticeSLI specificationUsual implementation
Data pipelineOld, missing or wrong dataFresh minutes; complete partitions; rows passing checksWatermark gauge; row counts against source; validation job
Queue and workersWaiting too long, or neverItems finished within T of enqueueProducer timestamp checked at completion, plus oldest-item age
StorageData cannot be read backProbe objects read back intactWrite-read probes and background checksum scrubbing
Scheduled jobReport late or absentRuns that succeed before their deadlineOne event per expected run, including runs that never started
Client appSlow screens, failed actionsSessions or actions without errors or long waitsReal user monitoring beacons, sampled

The common thread is that the event is defined from the consumer's side. Where the consumer cannot report, you synthesize the consumer: a probe that writes and reads, a check that compares counts, a timestamp stamped at the start of the user's wait.

Advertisement

Data pipelines: freshness, completeness and correctness

Freshness is the age of the newest data a reader can query. Implement it with a watermark: after each commit, the pipeline exports the event time of the newest record it has made visible. Freshness is then the current time minus that watermark. Because a pipeline has no requests, use a time-based SLI: each minute is good if freshness is under the threshold. Measuring event time rather than processing time is the point. A pipeline that is busily reprocessing an hour-old backlog has a recent processing timestamp and stale data.

Completeness catches the pipeline that is fresh but lossy, for example one that drops a partition after a schema change. Compare per-hour row counts at the source with counts at the destination after a settle period, and count an hour as good when they match within a tolerance you justify, such as 0.1% for late arrivals. Correctness is the hardest: run validation queries (non-null keys, totals that reconcile with the ledger, values in range) and count partitions that pass. Correctness checks are slow, so they often run daily and feed a separate, longer-window SLO.

groups:
  - name: sli-non-request
    interval: 1m
    rules:
      # Pipeline freshness, worst region: max here would hide one stalled region.
      - record: sli:pipeline_freshness_seconds
        expr: time() - min by (pipeline) (max by (pipeline, region) (pipeline_committed_watermark_seconds))
      - record: sli:pipeline_fresh:bool
        expr: sli:pipeline_freshness_seconds <= bool 900
      # Queue: age of the oldest item still waiting, the signal stuck work cannot hide from.
      - record: sli:queue_oldest_age_seconds
        expr: time() - min by (queue) (queue_oldest_enqueued_timestamp_seconds)
      # Queue: completed items that finished within their target, as a 5m ratio.
      - record: sli:queue_timely:ratio_rate5m
        expr: |
          sum by (queue) (rate(jobs_completed_total{outcome="ok", within_slo="true"}[5m]))
          /
          sum by (queue) (rate(jobs_completed_total[5m]))
  - name: sli-missing-data
    rules:
      # A pipeline that stops exporting has no freshness series at all, so averages stay green.
      - alert: PipelineWatermarkMissing
        expr: absent(pipeline_committed_watermark_seconds{pipeline="orders_to_warehouse"})
        for: 10m

Note the last rule. A time-based SLI computed from a recording rule has a blind spot: if the exporter dies, the series disappears, the ratio is computed from whatever samples remain and the SLO looks healthy. Treat absence as badness, either with an alert on absent() or by computing the good-minutes ratio against the number of minutes in the window rather than the number of samples.

Queues and asynchronous work

For asynchronous work the user's wait starts when the producer enqueues and ends when the effect is visible. Stamp the enqueue time on the message at the producer, and at completion record whether the item finished successfully within its target. That gives a clean ratio of good completions to all completions. Worker-side processing time is the wrong clock: when workers are scarce, items spend minutes waiting and seconds being processed, and only the waiting shows up for users.

import time
from prometheus_client import Counter, Gauge

TARGET_SECONDS = {"email": 120, "invoice_pdf": 600}
JOBS = Counter("jobs_completed_total", "Finished jobs", ["queue", "outcome", "within_slo"])
OLDEST = Gauge("queue_oldest_enqueued_timestamp_seconds", "Enqueue time of oldest waiting job", ["queue"])

def on_finished(job, ok: bool) -> None:
    # enqueued_at is stamped by the producer, so the latency includes time spent waiting,
    # which is the part users feel and the part worker-side timers never see.
    end_to_end = time.time() - job.enqueued_at
    timely = ok and end_to_end <= TARGET_SECONDS[job.queue]
    JOBS.labels(job.queue, "ok" if ok else "error", "true" if timely else "false").inc()

def export_backlog(queue_name: str, broker) -> None:
    head = broker.peek_oldest(queue_name)          # None when the queue is empty
    OLDEST.labels(queue_name).set(head.enqueued_at if head else time.time())

Completion-based ratios have a structural hole: items that never complete are never counted. If every worker is wedged, the completion rate drops to zero and the ratio becomes undefined, which many dashboards render as no data rather than as an outage. Pair the ratio with the age of the oldest waiting item. That number keeps growing while work is stuck, needs no completions to exist, and converts naturally into a time-based SLI: a minute is bad when the oldest item is older than the target. Queue depth is a poor substitute, because a deep queue that drains fast is fine and a shallow queue that does not drain is not.

Storage: durability and read-after-write

Durability failures are rare, silent and discovered late, so request counting cannot see them. Two mechanisms make them measurable. Probes write known objects continuously across every shard, region and storage class, read them back after intervals from minutes to months, and verify a checksum; each read-back is one SLI event. Scrubbing reads stored data in the background and compares it with recorded checksums; every mismatch or unreadable block is a bad event, whether or not a replica later repaired it. A repaired mismatch is still evidence that the margin is shrinking.

Read-after-write consistency is a separate property users notice immediately: they save, reload and see the old value. A probe that writes then reads from a different client within one second, counting a stale read as bad, turns that complaint into a number. Keep durability and freshness SLIs separate from availability; folding them into one ratio lets millions of fast reads drown a handful of lost objects.

Scheduled jobs and clients

For a job that should run once per period, the event is the expected run, not the run that happened. Generate the list of expected runs from the schedule, then mark each good only if it completed successfully before its deadline. A job that never started counts as bad because the expected run exists without a matching completion. Counting only the runs that reported back is the most common way a cron SLO stays at 100% while the report is missing.

Client measurement covers what the server cannot see: DNS, TLS, mobile networks, rendering and retries the app performed silently. Real user monitoring beacons give a session- or action-based SLI, such as the share of checkout attempts that completed without an error dialog within five seconds. Beacons are sampled and biased toward users whose browsers managed to send them, so use them to validate server-side SLIs rather than as the only source.

Worked example: a billing export

A team runs an event pipeline that loads orders into a warehouse, and a nightly job that produces invoices from it. Finance needs orders visible within 15 minutes and invoices ready by 06:00. The SLIs are: fresh minutes (freshness under 900 seconds) with an SLO of 99.5% over 28 days, which allows 201.6 bad minutes; complete hours with an SLO of 99.9%; and on-time invoice runs with an SLO of 27 out of 28.

In the first month the freshness SLI reads 99.92% while finance files two tickets about missing orders. Validation explains it: a connector change dropped one region's orders for six hours while the other regions kept the watermark moving, because the watermark was a maximum across regions. The fix is a per-region watermark with freshness taken as the oldest region, which would have shown 360 bad minutes and consumed the whole budget. The completeness SLI would also have caught it, because that region's counts stopped matching. This is the typical outcome of a first validation pass: the implementation was averaging away the failure.

Proving the SLI tracks user pain

Validation treats the SLI as a classifier. Collect independent evidence of user pain for a review window, such as support tickets with timestamps, incident impact periods and RUM error spikes, labelled by someone who did not look at the SLI. Convert both into sets of bad minutes and compute two numbers. Precision is the share of SLI-bad minutes in which users were hurt; low precision means the SLO burns budget and pages people for nothing. Recall is the share of hurt minutes the SLI flagged; low recall means users suffer while everything is green, which is worse.

def validate_sli(sli_bad: set[int], pain: set[int]) -> dict:
    """Both sets hold minute indexes over the same review window (for example 90 days).
    sli_bad: minutes the SLI counted as bad. pain: minutes with tickets, incident
    impact or a RUM error spike, labelled by people who never looked at the SLI."""
    hit = len(sli_bad & pain)
    precision = hit / len(sli_bad) if sli_bad else 1.0
    recall = hit / len(pain) if pain else 1.0
    return {
        "precision": round(precision, 3),           # low: the SLI burns budget on nothing
        "recall": round(recall, 3),                 # low: users suffer while the SLI is green
        "missed_minutes": sorted(pain - sli_bad)[:20],
        "false_minutes": sorted(sli_bad - pain)[:20],
    }

Then read the lists, not the ratios. Every missed minute points to a gap in the implementation, such as a maximum that should be a minimum, an absent series or an unmeasured region. Every false minute points to a threshold that is too tight or an event that users do not perceive. Run this review quarterly and after each major incident. Recall below about 0.8 is worth fixing before anyone tunes alert thresholds, because burn-rate alerts can only be as good as the indicator they burn.

Failure modes

  • Aggregation hides the failure. Maximum watermarks, global averages and fleet-wide ratios let one broken region or tenant disappear. Take the worst slice for freshness, and keep per-slice SLIs for large customers.
  • Absence reads as health. Missing series, no completions and jobs that never started all produce no bad events unless you define expected events independently.
  • Processing time instead of user time. Worker timers and batch durations miss queueing and backlog, the parts users feel most.
  • Never validated. An SLI adopted because a template suggested it, and never compared with tickets, will drift away from what users experience as the system changes.

What to do next

  1. List every system you own that has no request path: pipelines, queues, stored data and scheduled jobs.
  2. For each, write a one-sentence SLI specification from the consumer's point of view before choosing any metric.
  3. Export a per-slice event-time watermark from each pipeline and record freshness as the worst slice.
  4. Stamp enqueue time at producers, record timely completion at workers, and add an oldest-item age SLI.
  5. Generate expected runs for scheduled jobs from their schedules and count missing runs as bad.
  6. Add absent() alerts or window-based denominators so a silent exporter cannot keep an SLO green.
  7. Run the precision and recall review against 90 days of tickets and incidents, read every missed minute, and fix the implementation before tuning alerts. Use the SLA guide when an external commitment depends on the result.
Key takeaway: Systems without requests still need SLIs, defined from the consumer side: freshness, completeness and correctness for pipelines; timely completion and oldest-item age for queues; probes and scrubbing for durability; expected runs for scheduled jobs; and client beacons for what servers cannot see. Write the specification before the implementation, take the worst slice rather than an average, treat missing data as bad, and prove the indicator by measuring its precision and recall against independent evidence of user pain.