Someone asks how many GPUs a launch needs and states demand in queries per second. The honest answer starts with a question back: which queries? A request that sends 200 tokens and gets 50 back and one that sends 30,000 and gets 2,000 back differ by two orders of magnitude in GPU time. Even with the token mix pinned down, average throughput is the wrong target: a fleet sized for the mean queues badly at the peak, and LLM users notice queueing directly as time to first token.
This article turns measured engine behaviour into a requests-per-second capacity and a replica count. It uses Little's law to connect concurrency and duration, a queueing model to price the tail, and explicit adders for bursts and failures, all on one worked example. Deriving throughput from memory bandwidth and FLOPs is covered in GPU Serving Capacity Estimation; here every input comes from a load test on your own stack.
Why requests per second is a derived unit
An LLM request has two phases. Prefill processes the whole prompt in parallel and sets time to first token (TTFT). Decode generates one token per step for every running sequence and sets time per output token (TPOT). With continuous batching, a replica runs many sequences at once, so the natural unit of capacity is a slot: one sequence in flight. How many slots a replica can hold while meeting the TPOT target is the first number you need; how long each request occupies a slot is the second.
So QPS is a derived quantity: slots divided by the time each request holds a slot. Change the output length distribution and QPS changes even though the hardware did not. When a product team gives you a QPS number, ask for the token distribution behind it, and when you quote a QPS capacity, quote the mix it assumes. Token-denominated limits, such as tokens per minute, carry the mix inside them and are the better contract with callers.
Four measured numbers
Run a concurrency sweep with a load generator that replays your real prompt and output length distribution, on the exact model, engine version, quantisation and tensor-parallel layout you will ship. Hold concurrency fixed at each level long enough to reach steady state, and record p50 and p95 TTFT and TPOT. An illustrative sweep for one replica serving a chat workload averaging 1,500 input and 300 output tokens:
| Concurrent sequences | p95 TPOT | p95 TTFT (no queue) | Meets TPOT SLO of 50 ms? |
|---|---|---|---|
| 8 | 22 ms | 0.15 s | Yes |
| 16 | 28 ms | 0.19 s | Yes |
| 32 | 40 ms | 0.25 s | Yes, with margin |
| 48 | 58 ms | 0.42 s | No |
Choose S = 32 slots: the highest level that meets the TPOT target with margin for noisy neighbours and long prompts. The per-request service time at that level is prefill plus decode: T = 0.25 s + 300 tokens x 0.040 s = 12.25 s. These numbers are for illustration; yours will differ by model and engine, which is precisely why you measure them.
Concurrency, duration and per-replica capacity
Little's law says that in any stable system, average number in the system equals arrival rate times average time in the system: L = lambda x W. Apply it to one replica at its slot limit: 32 slots, each occupied for 12.25 s, sustain lambda = 32 / 12.25 = 2.61 requests per second. Cross-check in tokens: 2.61 requests x 300 output tokens is about 784 output tokens per second, close to the 800 that 32 streams at 25 tokens per second each would produce; the gap is the prefill time inside T.
Apply it to the whole demand: a peak of 40 QPS with 12.25 s per request means 40 x 12.25 = 490 requests in flight on average, which is 15.3 replicas worth of slots. That is the floor, the size at which the fleet is 100 percent busy on average. Run there and the queue grows without bound at every random burst.
Mixed workloads use the weighted service time. If 30 percent of traffic is summarisation with 8,000 input tokens and 200 output tokens, its prefill takes about 1.33 s at the same measured prefill rate, so T is 1.33 + 200 x 0.040 = 9.33 s. The blended T is 0.7 x 12.25 + 0.3 x 9.33 = 11.38 s and slot capacity rises to 2.81 QPS. Do not celebrate: long prefills stall decode steps for everyone and their KV cache can lower S, so re-run the sweep with the blended mix before trusting a blended number.
Queueing: utilisation against the tail
Arrivals are random, so some requests find every slot busy and wait. That wait lands directly in TTFT. A good first model treats each replica as an M/M/c queue with c = S servers: Poisson arrivals, exponentially distributed service times. Erlang's C formula gives the probability that an arrival waits, and the wait beyond t seconds has a simple tail.
import math
def erlang_c(c, a):
# P(arrival waits) in M/M/c, offered load a = lam / mu, valid for a < c
logs = [k * math.log(a) - math.lgamma(k + 1) for k in range(c)]
top = c * math.log(a) - math.lgamma(c + 1) + math.log(c / (c - a))
m = max(logs + [top])
s, t = sum(math.exp(l - m) for l in logs), math.exp(top - m)
return t / (s + t)
def p_wait_over(c, lam, service_s, t):
mu = 1.0 / service_s
a = lam / mu
if a >= c:
return 1.0
return erlang_c(c, a) * math.exp(-(c * mu - lam) * t)
def max_qps(c, service_s, t, target):
lo, hi = 0.0, c / service_s
for _ in range(60):
mid = (lo + hi) / 2
lo, hi = (mid, hi) if p_wait_over(c, mid, service_s, t) < target else (lo, mid)
return lo
def replicas_pooled(peak, slots, service_s, t, target):
n = math.ceil(peak * service_s / slots)
while p_wait_over(n * slots, peak, service_s, t) >= target:
n += 1
return nSet the queueing budget from the TTFT SLO. With a 1 s p99 TTFT target and 0.25 s of prefill, allow queue waits above 0.5 s for at most 1 percent of requests. For one replica (c = 32, T = 12.25 s):
| Arrival rate | Utilisation | P(wait at all) | P(wait over 0.5 s) |
|---|---|---|---|
| 1.5 QPS | 57% | 0.27% | 0.15% |
| 1.70 QPS | 65% | 1.6% | 1.0% (the limit) |
| 2.0 QPS | 77% | 10.3% | 7.6% |
| 2.3 QPS | 88% | 38% | 33% |
| 2.5 QPS | 96% | 73% | 69% |
The usable capacity of an isolated replica is 1.70 QPS, not 2.61: it must idle 35 percent of its slots to keep the tail inside budget. The curve is steep: the step from 77 to 88 percent utilisation multiplies the late-request share by four. This is the knee every capacity conversation should show.
M/M/c is a starting point, not the truth. Real arrivals are burstier than Poisson, and output lengths often have a long tail. The Allen-Cunneen approximation scales the M/M/c wait by (Ca2 + Cs2) / 2, where Ca and Cs are the coefficients of variation of inter-arrival and service times; measure both from logs. Also note that TPOT was pinned at its full-slot value, which is conservative when the replica is partly empty.
Routing changes the answer
How requests reach replicas changes the answer. Three ways to serve the 40 QPS peak:
- Divide by the mean: 40 / 2.61 rounds up to 16 replicas. At 96 percent utilisation, 9.5 percent of requests wait more than 0.5 s, so it misses the SLO by an order of magnitude.
- Random or round-robin routing: each replica is its own queue, so each can take 1.70 QPS, and 40 / 1.70 needs 24 replicas.
- Load-aware routing (a shared queue, or join-shortest-queue on free slots): the fleet behaves like one large M/M/c with c = 32n.
replicas_pooled(40, 32, 12.25, 0.5, 0.01)returns 17 replicas at 90 percent utilisation, with 0.1 percent of requests waiting over 0.5 s.
Pooling saved seven replicas, about 30 percent of the fleet, for the price of a router that knows each replica's free slots. That is usually the single highest-value change in an LLM serving stack, and it is why engine metrics such as running and waiting sequence counts should feed the load balancer. Prefix-cache-aware routing pulls the other way, concentrating requests that share a prompt; weigh cache hits against queue balance with measurements.
Burst headroom, failure spares and token limits
The queue model sizes for a steady peak. Two adders cover what it ignores.
Burst and scale-up lag. A new replica needs a node, an image pull and a model load, often several minutes for large weights. Size the headroom as demand ramp rate times scale-up lag. If traffic can climb 0.5 QPS per minute and a replica takes 8 minutes to become ready, the fleet must absorb 4 QPS above its current target, about 2 replicas at the pooled operating point of 40 / 17 = 2.35 QPS each. Pre-warmed standby nodes and cached weights cut the lag and therefore the headroom.
Failure spare. Add at least one replica so that losing one still meets the SLO; if replicas share a zone, size for the loss of the largest zone. In the example this brings the plan to 17 + 2 + 1 = 20 replicas for a 40 QPS peak, against the 16 a mean-based spreadsheet would have bought.
Finally, translate the plan into the units callers see. 40 QPS at 1,800 total tokens per request is 72,000 tokens per second, or 4.32 million tokens per minute. Publish per-tenant limits in tokens per minute that sum to less than this, so one tenant's long prompts cannot consume the queueing budget everyone shares.
Goodput and admission control
Raw throughput counts every request served. Goodput counts only requests that met both the TTFT and TPOT targets. Past the knee, throughput keeps rising slightly while goodput falls, because more requests finish late. Track goodput per replica and per tenant, and use it as the autoscaling and alerting signal instead of GPU utilisation, which reads near 100 percent on a healthy decode-bound replica and tells you nothing about queues.
When demand exceeds the plan, protect goodput by refusing early rather than serving late: admission control that returns 429 with a retry hint once the waiting queue exceeds the budgeted wait, priority classes for interactive traffic over batch, and output-length caps. Re-run the load test at the chosen fleet size with the measured arrival trace replayed, and confirm p99 TTFT and goodput before launch.
Failure modes
- QPS without a mix. A capacity number quoted without its token distribution is meaningless the day the product adds a document upload feature.
- Sizing to the mean. 100 percent average utilisation means unbounded queues.
- Benchmarks at fixed concurrency only. Closed-loop load tests hide queueing; replay an open-loop arrival trace for the final check.
- Ignoring the output tail. A few 4,000-token responses hold slots for minutes; cap max tokens per tier and measure Cs.
- Utilisation as the scaling signal. Scale on waiting requests, TTFT and goodput.
- Stale measurements. An engine upgrade, a new quantisation or a longer system prompt moves S and T. Re-sweep on every change and keep results versioned.
- Retry storms. Clients retrying on timeout multiply arrivals exactly when the fleet is saturated; require jittered backoff and honour Retry-After.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Higher slot count S | More QPS per replica | Higher TPOT, less headroom for long prompts |
| Load-aware routing | About 30% fewer replicas in the example | Router complexity, engine metrics required |
| Tighter TTFT budget | Better UX | Lower utilisation, more replicas |
| Pre-warmed standby | Smaller burst headroom | Paying for idle GPUs |
| Admission control | Protects goodput | Visible refusals at peaks |
Related reading
Related reading: GPU Inference Latency, in depth for where TTFT and TPOT come from; TTFT, in depth for the prefill side of the service time; continuous batching for why slots are the unit; GPU Capacity Planning for fleet-level budgeting; and LLM SLO Burn Rate Alerts for alerting on the targets used here.
What to do next
- Extract the input and output token distribution of your real traffic, including p95 and max.
- Run a concurrency sweep on the exact shipping stack and pick S from the TPOT target.
- Compute T and the Little's law floor; then run the Erlang C code with your TTFT wait budget.
- Check whether your router balances on free slots; if not, estimate the pooled saving.
- Measure scale-up lag end to end and size burst headroom as ramp rate times lag.
- Add a failure spare, publish per-tenant token-per-minute limits, and wire admission control.
- Replay a recorded arrival trace at the planned size and confirm p99 TTFT and goodput.