Two short acronyms dominate how teams decide what to put on a dashboard. The USE method, from Brendan Gregg, says: for every resource, check utilization, saturation and errors. The RED method, from Tom Wilkie, says: for every service, measure the rate of requests, the errors among them and the duration they take. They are often presented as rivals, and people argue about which one to adopt. That framing is wrong. They look at different objects. USE looks at the things that do the work, such as CPUs, disks, links and pools. RED looks at the work itself as it flows between services.

This article defines each method precisely, shows how to measure both with Prometheus, and works through an incident where a RED alert fired and a USE sweep found the cause. You will leave with a dashboard layout, an alerting rule of thumb and a list of measurement traps that make either method lie. For the metrics pipeline underneath, see the metrics architecture article.

Advertisement

Two methods, two kinds of object

A resource is something with finite capacity that work competes for: a CPU core, a memory pool, a disk, a network interface, a database connection pool or a worker thread pool. A service is something that accepts requests and returns responses: an HTTP endpoint, a gRPC method or a queue consumer. Resources have capacity and can be full. Services have callers who can be disappointed.

USE was designed for performance analysis of systems: start from a list of every resource and check three things for each, so that you do not miss a bottleneck because you only looked at the familiar ones. RED was designed for microservice fleets, where the operator owns dozens of request-driven services and needs one uniform view of whether each is serving its callers well. Google's four golden signals (latency, traffic, errors, saturation) are close to RED with one USE idea, saturation, added to the service view.

RED watches the work flowing through; USE watches the things doing the workClientsrequests incheckout serviceRED: rate, errors, durationpayments serviceRED at its boundaryreq/scallsCPUutil, run queue, -memoryused, swap/PSI, OOMdisk / networkbusy, queue, errorsDB connection poolin use, waiters, timeoutsUSE: utilization, saturation, errors for every resource (hardware and software)Diagnosis loopRED alert says users hurt -> USE sweep says which resource is the bottleneckRED answers 'is the service healthy for its callers?' USE answers 'why not?'
RED is measured at service boundaries, where requests enter and leave. USE is measured on each resource the service consumes, including software resources such as connection pools.

Precise definitions

SignalDefinitionTypical unit
Utilization (USE)Fraction of time the resource was busy, or fraction of its capacity in usepercent of time or of capacity
Saturation (USE)Work the resource could not serve yet: queue length, wait time, throttlingqueue depth, seconds waiting
Errors (USE)Error events at the resource: device errors, drops, allocation failurescount per second
Rate (RED)Requests the service handled per secondrequests per second
Errors (RED)Requests that failed, as a rate or fraction of the ratefailures per second or ratio
Duration (RED)Distribution of time to serve a requestseconds, as percentiles

Two subtleties matter. First, utilization comes in two flavours. A disk is time-based: it is either servicing an IO or not. A memory pool is capacity-based: it is 70 percent full. A resource can be 100 percent time-utilized and still accept more work if it serves requests in parallel, which is why disk busy time misleads on modern SSDs. Second, RED duration must be a distribution. An average hides the tail, and the tail is what users notice and what reveals queueing.

Advertisement

When each one fits

SituationStart withWhy
Is a user-facing service healthy right now?REDIt measures what callers experience
Paging alertsRED (via SLOs)Symptoms page; causes do not need to
Why is latency rising?USEQueueing at some resource is the usual cause
Capacity planningUSEUtilization trends predict when a resource runs out
Batch jobs and databasesUSE plus throughputThere may be no request boundary to measure
Queue consumersRED on messages plus consumer lagLag is saturation for the pipeline

The combined rule is simple: alert on RED, diagnose with USE. RED tells you that callers are hurt and roughly how badly. USE tells you which resource is at its limit. Neither alone is enough. A RED dashboard can show a latency spike without telling you that a connection pool has ten waiters. A USE dashboard can show a CPU at 95 percent while every caller is perfectly happy.

RED in practice with Prometheus

Instrument each service at its boundary with one counter of requests, labelled by status class, and one histogram of duration. Many frameworks and OpenTelemetry instrumentations emit these already, and they can also be derived from traces, as the span metrics article shows. The queries below compute the three signals.

# Rate: requests per second, per service and route
sum by (service, route) (rate(http_server_requests_total[5m]))

# Errors: fraction of requests that failed (5xx only; 4xx is usually the caller's fault)
sum by (service) (rate(http_server_requests_total{code=~"5.."}[5m]))
  /
sum by (service) (rate(http_server_requests_total[5m]))

# Duration: p99 from a histogram, aggregated across instances BEFORE the quantile
histogram_quantile(0.99,
  sum by (service, le) (rate(http_server_request_duration_seconds_bucket[5m])))

# Duration: mean, useful next to the tail for spotting bimodal behaviour
sum by (service) (rate(http_server_request_duration_seconds_sum[5m]))
  /
sum by (service) (rate(http_server_request_duration_seconds_count[5m]))

Note the order in the quantile query: sum the bucket rates across instances first, then compute the quantile. Averaging per-instance p99 values is mathematically meaningless. Decide deliberately what counts as an error. Server faults (5xx) and timeouts are errors. Most 4xx responses are the caller's mistake and should be tracked separately, or a client bug will page the server team. For low-traffic services, a 5-minute error ratio swings wildly with one failure; alert on burn rate against an SLO instead, as described in the SLIs, SLOs and error budgets article.

USE in practice: the resource checklist

USE only works if the resource list is complete, so write it down. The table gives each signal a metric from node_exporter or cAdvisor and the Linux command that answers the same question on a live host.

ResourceUtilizationSaturationErrors
CPUnode_cpu_seconds_total, mpstatrun queue (vmstat r), PSI cpu, CFS throttlingrare; machine check logs
Memoryused vs total, working set vs limitPSI memory, major faults, swappingOOM kills, allocation failures
Disknode_disk_io_time_seconds_total, iostat -xweighted IO time, iostat aqu-szdevice errors in dmesg
Networkbytes vs link speed, sar -n DEVdrops, retransmits, socket backlogerrs, drops counters
Conn / thread poolsin use vs sizewaiters, wait timecheckout timeouts
# CPU utilization per node (fraction of time not idle)
1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m]))

# CPU saturation: share of time some task waited for CPU (Linux PSI)
rate(node_pressure_cpu_waiting_seconds_total[5m])

# Container CPU saturation: fraction of CFS periods in which the cgroup was throttled
rate(container_cpu_cfs_throttled_periods_total[5m])
  / rate(container_cpu_cfs_periods_total[5m])

# Memory saturation: PSI stall time and major page faults
rate(node_pressure_memory_waiting_seconds_total[5m])
rate(node_vmstat_pgmajfault[5m])

# Disk: busy time (utilization) and average queue depth (saturation)
rate(node_disk_io_time_seconds_total[5m])
rate(node_disk_io_time_weighted_seconds_total[5m])

# Network errors and drops
rate(node_network_receive_errs_total[5m]) + rate(node_network_receive_drop_total[5m])

Pressure stall information (PSI), exposed by Linux 4.20 and later, is the most useful saturation signal for CPU, memory and IO, because it measures directly the share of time tasks were stalled waiting for that resource. In containers, CPU saturation shows up as CFS throttling: a pod can be throttled heavily while node CPU sits at 40 percent, because its quota, not the machine, is the resource that ran out. Memory errors appear as OOM kills, so track container restarts with reason OOMKilled alongside working set against the limit.

USE for software resources

Gregg's original list was hardware, but the most common bottlenecks in services are software: connection pools, thread pools, semaphores, rate limiters, and bounded queues. Each has a capacity, so each has utilization, saturation and errors. Few libraries export all three by default, which is why pool exhaustion is so often diagnosed late. Instrument them yourself.

from prometheus_client import Gauge, Counter, Histogram
import time

POOL_SIZE   = Gauge("db_pool_size", "Configured connections")
POOL_IN_USE = Gauge("db_pool_in_use", "Connections checked out")          # utilization
POOL_WAIT   = Gauge("db_pool_waiters", "Callers waiting for a connection")  # saturation
WAIT_TIME   = Histogram("db_pool_wait_seconds", "Time spent waiting for a connection")
POOL_ERRORS = Counter("db_pool_errors_total", "Checkout failures", ["reason"])  # errors

def checkout(pool, timeout=0.5):
    POOL_WAIT.inc()
    start = time.monotonic()
    try:
        conn = pool.acquire(timeout=timeout)
    except TimeoutError:
        POOL_ERRORS.labels(reason="timeout").inc()
        raise
    finally:
        POOL_WAIT.dec()
        WAIT_TIME.observe(time.monotonic() - start)
    POOL_IN_USE.inc()
    return conn

def release(pool, conn):
    pool.release(conn)
    POOL_IN_USE.dec()

A pool at 100 percent utilization with zero waiters is fine: it is exactly sized. A pool with waiters is saturated, and every waiting millisecond is added directly to request duration. The wait-time histogram is the bridge between the two methods, since it is a USE saturation signal whose units match RED duration.

A worked incident

At 14:05 the checkout service's p99 latency alert fires: p99 duration has moved from 180 ms to 650 ms, while the request rate is a normal 900 per second and the error ratio has risen from 0.05 to 0.6 percent, all timeouts. RED has done its job: callers are hurt, and the shape (latency up, rate flat, errors as timeouts) points to waiting rather than crashing.

Now the USE sweep, resource by resource. Node CPU utilization is 55 percent and CPU PSI is near zero, so CPU is not it. Memory working set is 60 percent of the limit with no major faults. Disk and network are idle. Then the software resources: the database connection pool shows 20 of 20 connections in use, around 35 waiters and a p99 wait of 480 ms, just under the pool's 500 ms checkout timeout. By Little's law, 35 waiters at 900 requests per second is a mean wait of about 39 ms, so most requests wait briefly while an unlucky tail waits almost until the timeout, and the requests that reach it are the timeouts RED saw. The pool wait accounts for almost all of the added latency.

The question becomes why connections are held longer. Each request holds a connection for its queries and, by mistake, also during a call to the payments service, about 16 ms in total. At 900 requests per second that keeps 900 × 0.016 = 14.4 connections busy on average, 72 percent of the pool of 20. The payments call has slowed from 6 ms to 11 ms, so each request now holds its connection for about 21 ms, and the average demand rises to 900 × 0.021 = 18.9 connections, 95 percent of the pool. A 5 ms slowdown in another service did this, because waiting time grows explosively as utilization approaches 100 percent and every burst above the average now queues. The fix is twofold: release the database connection before the remote call, and add a timeout on the payments call. The lesson is the pattern: RED located the symptom, USE located the saturated resource, and Little's law (in-flight equals rate times time held) explained the size of the effect.

Dashboard and alert layout

  • Top row per service: request rate, error ratio and p50/p99 duration, one panel each, with the SLO target drawn on the duration panel.
  • Second row: saturation for the resources this service depends on, including pool waiters, CPU throttling, PSI and queue lag. Put saturation above utilization, because saturation is what turns into latency.
  • Third row: utilization and errors per resource, for capacity planning and for completeness.
  • Page on RED-based SLO burn rates. Ticket, do not page, on USE thresholds such as sustained utilization above 80 percent or growing saturation, unless a resource is about to cause an outage, such as a disk that will be full in hours.
  • Link each RED panel to the USE panels of the resources behind it, so the diagnosis path is one click.

Measurement traps

  • Averaging away bursts. A CPU at 60 percent averaged over a minute may be at 100 percent for ten seconds of it, and the requests in those seconds queue. Use short windows or saturation metrics, not long-window utilization.
  • Disk busy time on parallel devices. An NVMe drive can show 100 percent busy while serving far below its capacity, because it handles many IOs at once. Look at queue depth and latency instead.
  • Quantiles of quantiles. Averaging per-pod p99 values, or computing p99 over too few requests, produces a number with no meaning.
  • Rate without a denominator. An error count of 50 per second means nothing without the request rate beside it.
  • Missing resources. USE is only as good as its list. Pools, file descriptors, ephemeral ports and quotas are the usual omissions.
  • Health checks in the numbers. Load balancer probes inflate rate and deflate duration. Exclude them by route.

Trade-offs

RED is cheap to adopt uniformly, maps directly onto SLOs and needs only two metrics per service, but it says nothing about cause and nothing about capacity headroom. USE is thorough and catches bottlenecks before users notice, but it needs a resource inventory, more instrumentation and more expertise to read, and it can generate noisy alerts if used for paging. Most teams should treat RED as the contract with callers and USE as the engineering view of the machine and the software resources inside it. Adding saturation to the service view, as the golden signals do, is the cheapest bridge between them.

What to do next

  1. List your services and confirm each emits a request counter with a status label and a duration histogram; add them where missing.
  2. Write the four RED queries above for each service and build the top dashboard row.
  3. Write a resource inventory per service, including every connection pool, thread pool and bounded queue.
  4. Instrument utilization, saturation and errors for each software resource, starting with database pools.
  5. Add PSI and CFS throttling panels for every host and pod.
  6. Move paging to RED-based SLO burn-rate alerts and turn USE threshold alerts into tickets.
  7. Rehearse one incident with the team using the RED-then-USE path, and fix whatever link or panel was missing. For query patterns at scale, continue with the Prometheus deep dive.
Key takeaway: RED and USE are not rivals; they look at different things. RED measures the requests a service handles, which is what callers feel, so it belongs in alerts and SLOs. USE measures every resource that does the work, including connection pools and queues, so it finds the bottleneck behind a RED symptom and warns before capacity runs out. Alert on RED, diagnose with USE, put saturation where you will see it first, and keep the resource list complete.