Capacity analysis answers one question with evidence: how much traffic can this deployment carry while still meeting its latency objectives, and which resource gives out first when it cannot? Two neighbouring jobs are easy to confuse with it. Capacity estimation predicts throughput from model size, memory and bandwidth before you own a replica; the serving capacity calculator does that from first principles. Capacity planning decides how many GPUs a fleet should buy next quarter; the capacity planning article covers the demand ledger and buy trigger. Analysis sits between them. It takes a real replica, a real request mix and real telemetry, and produces a rated number that planning can multiply.

This article shows how to turn a load-test sweep into goodput at SLO, how to read engine metrics to name the binding resource, how to check the measurement against a model ceiling, and how to convert a per-replica rating into a replica count with honest redundancy. A worked example runs through every step with numbers.

What a capacity analysis must produce

A capacity analysis has four outputs, and a report missing any of them is incomplete. First, a rated rate: requests per second per replica at which the SLO is met for the agreed fraction of requests. Second, the binding resource: the first thing that saturates, plus the second one, because fixing the first only buys the distance to the second. Third, the conditions the number holds under: request length distribution, prefix-cache hit rate, engine version, parallelism, quantisation. Fourth, the replica count for forecast demand, with the redundancy policy written down.

The conditions matter: a rating measured with 500-token prompts says little about a retrieval workload with 6,000-token prompts, because prefill cost and KV footprint both grow with context.

Capacity analysis: measured evidence in, a defensible replica count outProduction traceslength mix, arrivalsLoad-test sweepopen loop, one replicaEngine telemetryKV, queue, preemptionGoodput at SLOrated req/s per replicaBinding resourcewhat fails first, and nextDemand forecastpeak rate x burst factorReplica countceil(demand / rated) + redundancyRe-analysis triggersmodel, engine, mix driftThe sweep produces a number; telemetry explains it; the forecast multiplies it.
The analysis pipeline. Traces define the workload, the sweep measures one replica, telemetry explains where it saturates, and the forecast turns the rating into a fleet size.

Goodput at SLO, not throughput

Raw throughput is the wrong headline. A replica pushed past its knee keeps completing requests, so achieved throughput stays flat or even rises slightly, while every request waits longer and longer. Goodput counts only requests that met the SLO: achieved rate multiplied by the fraction of requests whose time to first token (TTFT) and inter-token latency (ITL) were both inside their targets. Goodput rises with load, peaks, then collapses as queueing pushes most requests over the line.

Define the SLO per request, not per percentile of a time window, so goodput can be computed from a request log. A typical interactive definition is TTFT at or under 2 seconds and p95 ITL within the request at or under 60 ms. The rated rate is then the highest offered load at which at least 99 percent of requests are good. Choose the attainment target to match the published SLO; a 99 percent objective cannot be supported by a rating taken at 95 percent attainment.

The sweep that feeds this must be open loop. Closed-loop generators wait for each response before sending the next request, so they slow down exactly when the server does and never show overload. The load testing article covers Poisson arrivals, coordinated omission, warm-up and prefix-cache control; this analysis assumes those were done properly.

Reading a sweep

Here is a sweep for one four-GPU replica of a 70B-class model, driven with prompt and output lengths resampled from a week of production logs (median prompt about 2,000 tokens, median output about 400). Each step ran for ten minutes after a two-minute warm-up. The numbers are illustrative but shaped like real runs; your own sweep replaces them.

Offered req/sAchievedTTFT p95ITL p95Running (mean)Waiting p95KV usage p95Good %
1.01.000.31 s28 ms1000.09100
2.02.000.38 s33 ms2300.21100
3.03.000.52 s39 ms4500.4299.6
4.04.000.95 s47 ms7010.6698.1
4.54.461.9 s55 ms9040.8491.5
5.04.716.8 s63 ms101190.9541
6.04.7321 s64 ms102750.973

Three regions are visible. Up to 3 req/s the replica is underloaded: latency creeps up gently and nothing queues. Between 4 and 4.5 is the knee: TTFT doubles, a queue appears and attainment falls below target. From 5 onward the replica is saturated: achieved throughput is pinned near 4.7, the waiting queue grows without bound and goodput collapses. A finer sweep between 3 and 4 found 99.1 percent attainment at 3.6 req/s, so the rating is 3.6.

Offered load against achieved throughput and goodput for one replica (sweep below)1.02.03.04.04.55.06.0012345offered load (requests per second)req/srated 3.6 (99% of requests meet SLO)achieved throughputgoodput
Achieved throughput plateaus near 4.7 req/s, but goodput peaks lower and then collapses. The rating sits left of the peak, where attainment is still at target.

Run a sanity check with Little's law before trusting any row: requests in flight should equal arrival rate times mean time in system. At 3 req/s, mean duration is roughly 0.4 s of TTFT plus 400 tokens at about 35 ms, or 14.4 s, so about 43 requests should be running; the engine reported 45. If the two disagree by more than ten or fifteen percent, the generator is not offering what it claims or the length mix is not what you think.

Naming the binding resource

The sweep says where the replica breaks; engine telemetry says why. vLLM exposes the gauges you need on its Prometheus endpoint: vllm:num_requests_running, vllm:num_requests_waiting, vllm:kv_cache_usage_perc (1.0 means every KV block is in use), the vllm:num_preemptions counter, and histograms such as vllm:time_to_first_token_seconds and vllm:inter_token_latency_seconds. Names have changed between releases, so read your own /metrics output before writing queries. GPU-side utilisation and memory bandwidth come from DCGM.

ResourceSignal that it bindsTypical fix
KV cache memoryKV usage near 0.9 or above, preemptions rising, running count flatShorter max context, FP8 KV, more GPUs per replica, prefix sharing
Prefill computeTTFT climbs and a queue forms while KV usage is still moderateTune the chunked-prefill token budget, separate prefill from decode
Decode bandwidthITL rises steadily with running count, no queue yetSmaller batch cap, quantised weights, speculative decoding
Scheduler or CPUGPU busy fraction drops while requests waitFaster tokenisation, more API workers, newer engine
Admission limitsWaiting grows while KV usage and GPU busy are lowRaise max concurrent sequences or batched-token caps
# Saturation panel, one replica. Counters are exposed with a _total suffix by the
# Prometheus client; confirm the exact names on your /metrics page.
max_over_time(vllm:kv_cache_usage_perc[5m])
quantile_over_time(0.95, vllm:num_requests_waiting[5m])
rate(vllm:num_preemptions_total[5m])
histogram_quantile(0.95, sum by (le) (rate(vllm:time_to_first_token_seconds_bucket[5m])))
histogram_quantile(0.95, sum by (le) (rate(vllm:inter_token_latency_seconds_bucket[5m])))

Read the example sweep with this table. At 4.5 req/s TTFT is already near its 2-second limit and four requests wait, while KV usage is 0.84 and preemptions are zero. Memory is not yet full; new prompts are waiting for prefill time that decode steps are also competing for. So the first binding resource is prefill time-sharing. At 5 req/s KV usage reaches 0.95 and the running count stops growing at about 101: the second wall is KV memory, roughly 10 to 15 percent further out. That gap is the most useful finding in the report. Retuning chunked prefill can move the rating up by at most that margin before memory caps it, so a team expecting a 2x gain from scheduler tuning alone will be disappointed.

Reconciling against a model ceiling

Compare the measured rating with a first-principles ceiling before publishing it. Estimate the ceiling the way KV cache sizing and the serving calculator do: KV bytes per token times mean context gives the maximum concurrent sequences, Little's law turns that into a maximum rate, and the time each request spends in prefill plus its share of decode steps gives a time-sharing bound. The measured rating should land below the smallest ceiling, typically at 60 to 90 percent of it.

If the measurement is above the ceiling, the model is wrong: prefix caching is probably hiding prefill work, or the length mix in the test is shorter than production. If it is below half, something is misconfigured: a batch cap set too low, KV cache memory fraction left at a conservative value, tensor-parallel communication over a slow link, or a tokeniser running on one CPU core. Either gap is a finding, not noise.

From a rating to a replica count

Converting a rating into replicas needs three more inputs: the peak arrival rate you must serve, a burst factor for arrivals faster than the hourly average, and a redundancy policy. Express redundancy as the failure you must survive without breaching the SLO, not as a percentage, because a percentage hides which failure it covers.

import math

def rated_rate(sweep, target=0.99):
    # sweep: list of (offered_rps, good_fraction), sorted by offered_rps.
    # Linear interpolation between the last passing and first failing point.
    prev = None
    for rps, good in sweep:
        if good < target:
            if prev is None:
                return 0.0
            r0, g0 = prev
            return r0 + (rps - r0) * (g0 - target) / (g0 - good)
        prev = (rps, good)
    return sweep[-1][0]  # never failed: the sweep did not go far enough

def replicas(peak_rps, burst, rated, zones=1, survive="node"):
    design = peak_rps * burst
    base = math.ceil(design / rated)
    if survive == "node":
        return base + 2                      # N+2: one failure plus one deploy drain
    if survive == "zone":
        per_zone = math.ceil(base / (zones - 1))
        return per_zone * zones              # remaining zones carry full load
    return base

sweep = [(3.0, 0.996), (3.6, 0.991), (4.0, 0.981), (4.5, 0.915)]
print(round(rated_rate(sweep), 2))           # 3.64; publish the measured 3.6

Worked example: sizing next quarter

The service above forecasts a peak-hour average of 52 req/s next quarter. One-minute arrival rates in last month's logs peaked at 1.25 times their hourly average, so the design rate is 52 x 1.25 = 65 req/s. At the measured 3.6 req/s per replica that is ceil(65 / 3.6) = ceil(18.06) = 19 replicas just to serve the peak minute.

Redundancy then forks into two honest options. N+2 for node failure adds one replica for an unplanned GPU or host failure and one for a rolling deploy draining a replica: 21 replicas, 84 GPUs. Zone survival across three zones requires the two surviving zones to carry 19 replicas between them, so 10 per zone and 30 in total, 120 GPUs. The second option costs 36 GPUs more to cover an event that may happen once a year. Writing both down, with their GPU cost, turns an argument about headroom percentages into a business decision about risk.

Finally, record the conditions: engine version, tensor-parallel degree 4, BF16 weights, the length distribution file, prefix-cache hit rate during the test (11 percent), and the SLO definition. Attach the binding-resource finding: prefill first at about 4.3 req/s, KV memory at about 4.8. If the team wants more per replica, the cheapest next experiment is the chunked-prefill budget, with an expected ceiling of about 4.5.

Keeping the rating honest over time

A rating decays. Re-run the analysis when any of these change: model weights or quantisation, engine version, parallelism, maximum context length, or the request length mix by more than about 20 percent at the median or p90. Between sweeps, track production headroom: the ratio of observed peak-minute rate per replica to the rated rate. Alert at 0.8, act at 0.9. Also alert when production telemetry disagrees with the sweep, for example KV usage at 0.7 while the sweep predicted 0.45 at the same rate; that usually means prompts grew, and the rating is already stale.

Failure modes

  • Rating from synthetic prompts. Fixed-length prompts hide the long tail that fills KV memory. Resample lengths from production.
  • Warm prefix cache in the test, cold in production (or the reverse). Report the hit rate and match it.
  • Throughput at saturation published as capacity. The 4.7 req/s plateau is not a rate the SLO can survive.
  • Averages instead of attainment. A mean TTFT of 0.8 s can coexist with 15 percent of requests over 2 s.
  • One binding resource reported. Without the second wall, optimisation plans overpromise.
  • Redundancy as a percentage. 20 percent headroom does not say whether a zone loss is survivable.

Trade-offs

ChoiceGainCost
Higher attainment target (99.9 vs 99)Fewer SLO breachesLower rating, more replicas
Zone redundancy vs N+2Survives a zone outageRoughly 40 percent more GPUs in the example
Larger replicas (TP 8 vs 4)More KV room, longer contextsCoarser scaling steps, more communication
Long sweeps per stepStable tailsHours of GPU time per analysis
Rating per workload classAccurate per tenantMore sweeps to maintain

What to do next

  1. Write the SLO as a per-request rule (TTFT and ITL thresholds) and pick an attainment target.
  2. Export a week of production prompt and output lengths, and build the load generator's length sampler from it.
  3. Run an open-loop sweep on one replica with at least six offered rates spanning underload to saturation.
  4. Compute goodput per step from the request log and find the rated rate by a fine sweep around the knee.
  5. Check each row with Little's law, then name the first and second binding resources from engine metrics.
  6. Compare the rating with a first-principles ceiling and explain any gap larger than about 40 percent.
  7. Compute replicas for both N+2 and zone survival, with GPU counts, and let the service owner choose.
  8. Publish the conditions with the rating, set headroom alerts at 0.8 and 0.9, and list the re-sweep triggers.
Key takeaway: Capacity analysis turns a real replica, a real request mix and real telemetry into a rated rate: the highest load at which the SLO attainment target still holds. Measure it with an open-loop sweep, judge it by goodput rather than throughput, name both the first and second binding resources, check it against Little's law and a model ceiling, and size the fleet with redundancy expressed as the failure you must survive.