The four golden signals, latency, traffic, errors and saturation, come from the monitoring chapter of Google's Site Reliability Engineering book, which advised that if you can only measure four metrics of a user-facing system, measure these. A decade on, almost every dashboard claims to show them, and many show them wrongly: an average latency that hides the tail, an error rate that excludes the failures users notice most, and a CPU graph standing in for saturation on a service whose real limit is a connection pool.

This page revisits each signal with the original intent in hand and asks what it means for systems the book did not centre on: queue consumers, streaming model servers and services that autoscale. How the signals relate to the RED and USE methods is covered in RED vs USE, and turning a signal into an SLO is covered in SLIs, SLOs and error budgets. The subject here is measuring the four signals correctly, so they explain an incident instead of decorating it.

Advertisement

What the book actually said

The original text is more precise than its popular summary. Latency is the time it takes to serve a request, and the book insists on distinguishing the latency of successful requests from that of failed ones: a fast error, such as an immediate 500 when a database connection is lost, pulls the average down and makes a broken service look quick, while a slow error is worse than a fast one. Traffic is a measure of demand in a system-specific unit: requests per second for a web service, broken down by type, or transactions or bytes for other systems.

Errors are the rate of failed requests, and the book names three kinds: explicit failures such as HTTP 500s, implicit failures such as a 200 response with the wrong content, and failures by policy, such as any response slower than a committed threshold. Saturation is how full the service is, with emphasis on the most constrained resource. The book notes that many systems degrade before they reach 100 percent utilisation, that latency increases are often a leading indicator of saturation, and that saturation also covers predictions, such as a database filling its disk in four hours.

Each of those details is routinely lost.

Where you measure changes what you see

A request passes through a client, usually a load balancer or gateway, the service itself and its dependencies, and each vantage point sees a different system. Server-side handler metrics miss the time a request spent queued in the load balancer or waiting for a worker thread, and they miss requests that never reached a handler: a 502 from a proxy because every backend was down produces no server-side error at all.

Where you measure decides what latency and errors you seeClienttrue user latencyLoad balancerincludes queueingServicehandler time onlyDependencyDB, cache, modelSaturationpool, queue, KV cachesees DNS, TLS, retriessees 502 / 504 theservice never loggedclient-side view ofthe dependencyGolden signals per service, from the edge closest to the user you can instrumentLatency: histogram, split by outcome | Traffic: in the unit that costsErrors: explicit + implicit + by policy | Saturation: the scarcest resource, as headroomEach signal from two vantage points where you can: server-side and edge
Each vantage point sees a different system. The load balancer sees errors the service never logged; the service sees only handler time; saturation lives in the resources between them.

The practical rule is to measure each signal at the outermost point you control, usually the gateway or load balancer, and also at the service, then compare. A gap between the two is a diagnosis in itself: gateway latency far above handler latency means queueing in front of the handler, and gateway errors with a clean service error rate mean requests are failing before they arrive. OpenTelemetry's stable HTTP semantic conventions help here, because a server records http.server.request.duration as a histogram in seconds with attributes such as http.route and http.response.status_code, and a client records the matching http.client.request.duration, so both sides use the same names and units.

Advertisement

Latency: distributions, split by outcome

Latency is a distribution, so record it as a histogram and read percentiles, never an average. A mean of 120 ms is compatible with every user waiting 120 ms and with 98 percent waiting 50 ms while two percent wait three seconds, and those are different incidents. Percentiles cannot be averaged across instances, which is why the histogram buckets are the thing to aggregate; the mechanics, and how bucket boundaries bound your error, are covered in latency histograms and quantile estimation. Put a bucket boundary at every threshold you alert on, so a policy like "under 300 ms" is counted exactly rather than interpolated.

Split by outcome, as the book prescribed, by keeping the status as a label on the histogram. Graph p50 and p99 of successful requests, and separately the latency of failures. During a dependency outage the success latency may rise while the failure latency drops to near zero, and only the split shows both. Split by route as well, since a mixed p99 of a 5 ms health check and a 2 s report is meaningless.

Traffic: the unit that costs

Requests per second is a convenient default and often the wrong unit. Traffic should be measured in whatever drives cost and saturation: rows scanned for a query service, bytes for a storage or streaming service, messages for a queue, and tokens for a model server, where one request may be ten tokens or a hundred thousand. When work per request varies widely, show two traffic series, request rate and work rate, because a capacity incident can come from a flat request rate carrying heavier requests.

Traffic is also the context every other signal needs. An error ratio of five percent on ten requests a minute is a different situation from five percent on ten thousand, and a sudden drop in traffic is itself an error signal: if checkout requests fall by half at noon on a weekday, something upstream is failing silently even though every graph you own is green.

Errors: explicit, implicit and by policy

Explicit errors are the easy part, but decide deliberately which status codes count. Server errors count; most client errors do not, because a 404 for a mistyped URL is not your failure. Some 4xx codes are your failure in disguise, such as a 429 from your own rate limiter, so track them as their own series.

Implicit errors need a check that knows what correct looks like: a search that returns zero results for a query that always has results, a response missing a required field, a payment that succeeded but was never recorded. These usually come from synthetic probes or from validation in the client. Policy errors turn latency into errors: count every response slower than the threshold you promised as failed. That is what makes an error ratio match user experience, since users treat a ten-second page as broken. The recording rule below folds policy errors into the ratio using the 1 s bucket boundary.

Saturation: headroom and time to exhaustion

Saturation is the signal most often faked with CPU. The book's wording is the most constrained resource, and for many services that is not CPU but a database connection pool, a thread pool, file descriptors, a queue, memory, disk, an API quota or, for a model server, accelerator memory for the KV cache. Find it by load testing until latency bends, then asking what ran out first; that resource, expressed as a fraction of its limit, is your saturation signal.

Two forms are worth graphing. Utilisation of the scarce resource tells you how close you are now; latency typically starts to rise well before 100 percent because of queueing, so alert on a threshold such as 80 percent sustained rather than on exhaustion. Time to exhaustion handles resources that fill steadily, such as disks, quotas or a growing backlog: divide the remaining capacity by the current rate of consumption and page when the answer is shorter than the time a human needs to act.

Autoscaling hides saturation. When an autoscaler adds replicas, per-instance utilisation stays flat while the system consumes headroom elsewhere: replica count approaches its maximum, the node pool fills, or the shared database behind the replicas saturates. For an autoscaled service, treat replicas as a fraction of the maximum as a saturation signal, and measure the shared dependencies, not just the instances.

Systems the book did not centre on

For queue consumers there is no request latency the user sees directly; the user waits for the work to be done. Latency becomes end-to-end time from enqueue to completion, best approximated by the age of the oldest unprocessed message; traffic is enqueue and dequeue rates; errors include messages sent to a dead-letter queue; saturation is backlog depth relative to drain rate, which is also time to exhaustion run backwards: at the current drain rate, how long until the backlog is gone.

Streaming model servers break the single-number latency. A user perceives the time to the first token and then the pace of the rest, so record time to first token and time per output token as separate histograms, plus end-to-end duration. Traffic needs input and output tokens as well as requests. Errors include truncated or empty generations and content-filter refusals your product counts as failures. Saturation is usually accelerator memory available for the KV cache and the depth of the scheduler's waiting queue, not GPU utilisation, which can read high while the server is starved for memory.

The signals as code

Signals that live only in a dashboard drift. Define them as recording rules next to the service, with one rule per signal, so dashboards, alerts and SLOs all read the same series. The example below is for a checkout service instrumented with a Prometheus histogram; the names are illustrative and follow whatever your exporter produces.

groups:
- name: golden-signals-checkout
  interval: 30s
  rules:
  # Traffic: requests per second by route
  - record: route:http_server_requests:rate5m
    expr: sum by (route) (rate(http_server_request_duration_seconds_count{service="checkout"}[5m]))

  # Errors: explicit 5xx plus policy errors (slower than 1s) as a ratio
  - record: route:http_server_errors:ratio5m
    expr: |
      (
        sum by (route) (rate(http_server_request_duration_seconds_count{service="checkout", status=~"5.."}[5m]))
      + sum by (route) (rate(http_server_request_duration_seconds_count{service="checkout", status!~"5.."}[5m]))
      - sum by (route) (rate(http_server_request_duration_seconds_bucket{service="checkout", status!~"5..", le="1"}[5m]))
      )
      / sum by (route) (rate(http_server_request_duration_seconds_count{service="checkout"}[5m]))

  # Latency: p99 of successful requests only
  - record: route:http_server_latency_ok:p99_5m
    expr: |
      histogram_quantile(0.99,
        sum by (route, le) (rate(http_server_request_duration_seconds_bucket{service="checkout", status!~"5.."}[5m])))

  # Saturation: DB pool in use as a fraction of its size, and time to disk full
  - record: service:db_pool:utilisation
    expr: max(db_pool_connections_in_use{service="checkout"} / db_pool_connections_max{service="checkout"})
  - record: node:disk:hours_to_full
    expr: |
      node_filesystem_avail_bytes{mountpoint="/data"}
      / clamp_min(-deriv(node_filesystem_avail_bytes{mountpoint="/data"}[6h]), 1) / 3600

The error ratio counts 5xx responses plus non-5xx responses slower than one second, which relies on a bucket boundary at exactly 1 s. Latency is taken over successful requests only. The disk rule divides free space by its rate of decline over six hours to give hours to full; the clamp avoids dividing by zero when usage is flat or falling. Lay out every service dashboard in the same order: traffic, errors, latency, saturation. Page on the error ratio and latency through burn-rate alerts, described in burn-rate alerting, and on time to exhaustion; leave raw traffic and utilisation as context, not pages.

Worked incident

At 14:05 the checkout p99 for successful requests rises from 280 ms to 1.4 s while traffic is flat. The error ratio climbs to six percent, but almost all of it is policy errors: requests completing slowly, not failing. Gateway and service latency rise together, so the time is spent inside the handler, not in front of it. The CPU graph is unremarkable, at 40 percent.

The saturation panel tells the story: database pool utilisation has been at 100 percent since 14:03. A deploy at 14:01 added a query to the order-confirmation path that holds a connection while calling an external fraud service, so every request now holds a pooled connection for the duration of a slow network call. Requests queue for connections, which is why latency rose before any explicit errors appeared, exactly the leading-indicator behaviour the book describes. Rolling back restores the pool to 35 percent and p99 to 290 ms. A CPU-based saturation panel would have shown nothing.

Traps

TrapEffectFix
Average latencyTail regressions invisibleHistogram; graph p50 and p99
Latency not split by outcomeFast failures make outages look fastStatus label on the histogram; success-only latency
Server-only errorsProxy 502s and timeouts missingMeasure at the gateway too and compare
CPU as saturationPool or queue exhaustion missedLoad test to find the scarce resource
Per-instance view under autoscalingHeadroom loss hiddenReplicas as fraction of max; shared dependencies
Unbounded labelsSeries explosion and slow queriesLabel by route template, never raw path or user id

Trade-offs

Four signals are a floor, not a complete monitoring strategy: they tell you that a service is unhealthy and roughly where to look, and traces and logs explain why. Measuring at both gateway and service doubles the series and the dashboards to keep consistent, in exchange for the gap analysis that localises an incident in minutes. Policy errors make the error ratio honest but couple it to a threshold you must keep in step with the SLO.

What to do next

  1. For each critical service, write down the unit of traffic and the most constrained resource; confirm the latter with a load test.
  2. Replace average latency panels with p50 and p99 from histograms, split by success and failure, with bucket boundaries at your thresholds.
  3. Add policy errors to the error ratio and track your own 429s and auth failures as separate series.
  4. Collect the same signals at the gateway and the service, and add a panel for the difference.
  5. Add a time-to-exhaustion signal for disks, quotas and backlogs, and replicas-to-maximum for autoscaled services.
  6. For queues and model servers, adopt the adapted signals: age of oldest message, time to first token, KV cache headroom.
  7. Move all four signals into recording rules and lay out every service dashboard in the same order.
Key takeaway: The golden signals still work, provided they are measured the way they were defined. Record latency as a histogram split by outcome, measure traffic in the unit that drives cost, count explicit, implicit and policy errors, and express saturation as headroom on the scarcest resource plus time to exhaustion. Measure at the gateway as well as the service, adapt the signals for queues, streaming model servers and autoscaled fleets, and keep them as recording rules so every dashboard and alert reads the same numbers.