Capacity analysis answers one question with evidence: how much traffic can this deployment carry while still meeting its latency objectives, and which resource gives out first when it cannot? Two neighbouring jobs are easy to confuse with it. Capacity estimation predicts throughput from model size, memory and bandwidth before you own a replica; the serving capacity calculator does that from first principles. Capacity planning decides how many GPUs a fleet should buy next quarter; the capacity planning article covers the demand ledger and buy trigger. Analysis sits between them. It takes a real replica, a real request mix and real telemetry, and produces a rated number that planning can multiply.
This article shows how to turn a load-test sweep into goodput at SLO, how to read engine metrics to name the binding resource, how to check the measurement against a model ceiling, and how to convert a per-replica rating into a replica count with honest redundancy. A worked example runs through every step with numbers.
What a capacity analysis must produce
A capacity analysis has four outputs, and a report missing any of them is incomplete. First, a rated rate: requests per second per replica at which the SLO is met for the agreed fraction of requests. Second, the binding resource: the first thing that saturates, plus the second one, because fixing the first only buys the distance to the second. Third, the conditions the number holds under: request length distribution, prefix-cache hit rate, engine version, parallelism, quantisation. Fourth, the replica count for forecast demand, with the redundancy policy written down.
The conditions matter: a rating measured with 500-token prompts says little about a retrieval workload with 6,000-token prompts, because prefill cost and KV footprint both grow with context.
Goodput at SLO, not throughput
Raw throughput is the wrong headline. A replica pushed past its knee keeps completing requests, so achieved throughput stays flat or even rises slightly, while every request waits longer and longer. Goodput counts only requests that met the SLO: achieved rate multiplied by the fraction of requests whose time to first token (TTFT) and inter-token latency (ITL) were both inside their targets. Goodput rises with load, peaks, then collapses as queueing pushes most requests over the line.
Define the SLO per request, not per percentile of a time window, so goodput can be computed from a request log. A typical interactive definition is TTFT at or under 2 seconds and p95 ITL within the request at or under 60 ms. The rated rate is then the highest offered load at which at least 99 percent of requests are good. Choose the attainment target to match the published SLO; a 99 percent objective cannot be supported by a rating taken at 95 percent attainment.
The sweep that feeds this must be open loop. Closed-loop generators wait for each response before sending the next request, so they slow down exactly when the server does and never show overload. The load testing article covers Poisson arrivals, coordinated omission, warm-up and prefix-cache control; this analysis assumes those were done properly.
Reading a sweep
Here is a sweep for one four-GPU replica of a 70B-class model, driven with prompt and output lengths resampled from a week of production logs (median prompt about 2,000 tokens, median output about 400). Each step ran for ten minutes after a two-minute warm-up. The numbers are illustrative but shaped like real runs; your own sweep replaces them.
| Offered req/s | Achieved | TTFT p95 | ITL p95 | Running (mean) | Waiting p95 | KV usage p95 | Good % |
|---|---|---|---|---|---|---|---|
| 1.0 | 1.00 | 0.31 s | 28 ms | 10 | 0 | 0.09 | 100 |
| 2.0 | 2.00 | 0.38 s | 33 ms | 23 | 0 | 0.21 | 100 |
| 3.0 | 3.00 | 0.52 s | 39 ms | 45 | 0 | 0.42 | 99.6 |
| 4.0 | 4.00 | 0.95 s | 47 ms | 70 | 1 | 0.66 | 98.1 |
| 4.5 | 4.46 | 1.9 s | 55 ms | 90 | 4 | 0.84 | 91.5 |
| 5.0 | 4.71 | 6.8 s | 63 ms | 101 | 19 | 0.95 | 41 |
| 6.0 | 4.73 | 21 s | 64 ms | 102 | 75 | 0.97 | 3 |
Three regions are visible. Up to 3 req/s the replica is underloaded: latency creeps up gently and nothing queues. Between 4 and 4.5 is the knee: TTFT doubles, a queue appears and attainment falls below target. From 5 onward the replica is saturated: achieved throughput is pinned near 4.7, the waiting queue grows without bound and goodput collapses. A finer sweep between 3 and 4 found 99.1 percent attainment at 3.6 req/s, so the rating is 3.6.
Run a sanity check with Little's law before trusting any row: requests in flight should equal arrival rate times mean time in system. At 3 req/s, mean duration is roughly 0.4 s of TTFT plus 400 tokens at about 35 ms, or 14.4 s, so about 43 requests should be running; the engine reported 45. If the two disagree by more than ten or fifteen percent, the generator is not offering what it claims or the length mix is not what you think.
Naming the binding resource
The sweep says where the replica breaks; engine telemetry says why. vLLM exposes the gauges you need on its Prometheus endpoint: vllm:num_requests_running, vllm:num_requests_waiting, vllm:kv_cache_usage_perc (1.0 means every KV block is in use), the vllm:num_preemptions counter, and histograms such as vllm:time_to_first_token_seconds and vllm:inter_token_latency_seconds. Names have changed between releases, so read your own /metrics output before writing queries. GPU-side utilisation and memory bandwidth come from DCGM.
| Resource | Signal that it binds | Typical fix |
|---|---|---|
| KV cache memory | KV usage near 0.9 or above, preemptions rising, running count flat | Shorter max context, FP8 KV, more GPUs per replica, prefix sharing |
| Prefill compute | TTFT climbs and a queue forms while KV usage is still moderate | Tune the chunked-prefill token budget, separate prefill from decode |
| Decode bandwidth | ITL rises steadily with running count, no queue yet | Smaller batch cap, quantised weights, speculative decoding |
| Scheduler or CPU | GPU busy fraction drops while requests wait | Faster tokenisation, more API workers, newer engine |
| Admission limits | Waiting grows while KV usage and GPU busy are low | Raise max concurrent sequences or batched-token caps |
# Saturation panel, one replica. Counters are exposed with a _total suffix by the
# Prometheus client; confirm the exact names on your /metrics page.
max_over_time(vllm:kv_cache_usage_perc[5m])
quantile_over_time(0.95, vllm:num_requests_waiting[5m])
rate(vllm:num_preemptions_total[5m])
histogram_quantile(0.95, sum by (le) (rate(vllm:time_to_first_token_seconds_bucket[5m])))
histogram_quantile(0.95, sum by (le) (rate(vllm:inter_token_latency_seconds_bucket[5m])))Read the example sweep with this table. At 4.5 req/s TTFT is already near its 2-second limit and four requests wait, while KV usage is 0.84 and preemptions are zero. Memory is not yet full; new prompts are waiting for prefill time that decode steps are also competing for. So the first binding resource is prefill time-sharing. At 5 req/s KV usage reaches 0.95 and the running count stops growing at about 101: the second wall is KV memory, roughly 10 to 15 percent further out. That gap is the most useful finding in the report. Retuning chunked prefill can move the rating up by at most that margin before memory caps it, so a team expecting a 2x gain from scheduler tuning alone will be disappointed.
Reconciling against a model ceiling
Compare the measured rating with a first-principles ceiling before publishing it. Estimate the ceiling the way KV cache sizing and the serving calculator do: KV bytes per token times mean context gives the maximum concurrent sequences, Little's law turns that into a maximum rate, and the time each request spends in prefill plus its share of decode steps gives a time-sharing bound. The measured rating should land below the smallest ceiling, typically at 60 to 90 percent of it.
If the measurement is above the ceiling, the model is wrong: prefix caching is probably hiding prefill work, or the length mix in the test is shorter than production. If it is below half, something is misconfigured: a batch cap set too low, KV cache memory fraction left at a conservative value, tensor-parallel communication over a slow link, or a tokeniser running on one CPU core. Either gap is a finding, not noise.
From a rating to a replica count
Converting a rating into replicas needs three more inputs: the peak arrival rate you must serve, a burst factor for arrivals faster than the hourly average, and a redundancy policy. Express redundancy as the failure you must survive without breaching the SLO, not as a percentage, because a percentage hides which failure it covers.
import math
def rated_rate(sweep, target=0.99):
# sweep: list of (offered_rps, good_fraction), sorted by offered_rps.
# Linear interpolation between the last passing and first failing point.
prev = None
for rps, good in sweep:
if good < target:
if prev is None:
return 0.0
r0, g0 = prev
return r0 + (rps - r0) * (g0 - target) / (g0 - good)
prev = (rps, good)
return sweep[-1][0] # never failed: the sweep did not go far enough
def replicas(peak_rps, burst, rated, zones=1, survive="node"):
design = peak_rps * burst
base = math.ceil(design / rated)
if survive == "node":
return base + 2 # N+2: one failure plus one deploy drain
if survive == "zone":
per_zone = math.ceil(base / (zones - 1))
return per_zone * zones # remaining zones carry full load
return base
sweep = [(3.0, 0.996), (3.6, 0.991), (4.0, 0.981), (4.5, 0.915)]
print(round(rated_rate(sweep), 2)) # 3.64; publish the measured 3.6
Worked example: sizing next quarter
The service above forecasts a peak-hour average of 52 req/s next quarter. One-minute arrival rates in last month's logs peaked at 1.25 times their hourly average, so the design rate is 52 x 1.25 = 65 req/s. At the measured 3.6 req/s per replica that is ceil(65 / 3.6) = ceil(18.06) = 19 replicas just to serve the peak minute.
Redundancy then forks into two honest options. N+2 for node failure adds one replica for an unplanned GPU or host failure and one for a rolling deploy draining a replica: 21 replicas, 84 GPUs. Zone survival across three zones requires the two surviving zones to carry 19 replicas between them, so 10 per zone and 30 in total, 120 GPUs. The second option costs 36 GPUs more to cover an event that may happen once a year. Writing both down, with their GPU cost, turns an argument about headroom percentages into a business decision about risk.
Finally, record the conditions: engine version, tensor-parallel degree 4, BF16 weights, the length distribution file, prefix-cache hit rate during the test (11 percent), and the SLO definition. Attach the binding-resource finding: prefill first at about 4.3 req/s, KV memory at about 4.8. If the team wants more per replica, the cheapest next experiment is the chunked-prefill budget, with an expected ceiling of about 4.5.
Keeping the rating honest over time
A rating decays. Re-run the analysis when any of these change: model weights or quantisation, engine version, parallelism, maximum context length, or the request length mix by more than about 20 percent at the median or p90. Between sweeps, track production headroom: the ratio of observed peak-minute rate per replica to the rated rate. Alert at 0.8, act at 0.9. Also alert when production telemetry disagrees with the sweep, for example KV usage at 0.7 while the sweep predicted 0.45 at the same rate; that usually means prompts grew, and the rating is already stale.
Failure modes
- Rating from synthetic prompts. Fixed-length prompts hide the long tail that fills KV memory. Resample lengths from production.
- Warm prefix cache in the test, cold in production (or the reverse). Report the hit rate and match it.
- Throughput at saturation published as capacity. The 4.7 req/s plateau is not a rate the SLO can survive.
- Averages instead of attainment. A mean TTFT of 0.8 s can coexist with 15 percent of requests over 2 s.
- One binding resource reported. Without the second wall, optimisation plans overpromise.
- Redundancy as a percentage. 20 percent headroom does not say whether a zone loss is survivable.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Higher attainment target (99.9 vs 99) | Fewer SLO breaches | Lower rating, more replicas |
| Zone redundancy vs N+2 | Survives a zone outage | Roughly 40 percent more GPUs in the example |
| Larger replicas (TP 8 vs 4) | More KV room, longer contexts | Coarser scaling steps, more communication |
| Long sweeps per step | Stable tails | Hours of GPU time per analysis |
| Rating per workload class | Accurate per tenant | More sweeps to maintain |
What to do next
- Write the SLO as a per-request rule (TTFT and ITL thresholds) and pick an attainment target.
- Export a week of production prompt and output lengths, and build the load generator's length sampler from it.
- Run an open-loop sweep on one replica with at least six offered rates spanning underload to saturation.
- Compute goodput per step from the request log and find the rated rate by a fine sweep around the knee.
- Check each row with Little's law, then name the first and second binding resources from engine metrics.
- Compare the rating with a first-principles ceiling and explain any gap larger than about 40 percent.
- Compute replicas for both N+2 and zone survival, with GPU counts, and let the service owner choose.
- Publish the conditions with the rating, set headroom alerts at 0.8 and 0.9, and list the re-sweep triggers.