Someone asks how many GPUs a launch needs and states demand in queries per second. The honest answer starts with a question back: which queries? A request that sends 200 tokens and gets 50 back and one that sends 30,000 and gets 2,000 back differ by two orders of magnitude in GPU time. Even with the token mix pinned down, average throughput is the wrong target: a fleet sized for the mean queues badly at the peak, and LLM users notice queueing directly as time to first token.

This article turns measured engine behaviour into a requests-per-second capacity and a replica count. It uses Little's law to connect concurrency and duration, a queueing model to price the tail, and explicit adders for bursts and failures, all on one worked example. Deriving throughput from memory bandwidth and FLOPs is covered in GPU Serving Capacity Estimation; here every input comes from a load test on your own stack.

Why requests per second is a derived unit

An LLM request has two phases. Prefill processes the whole prompt in parallel and sets time to first token (TTFT). Decode generates one token per step for every running sequence and sets time per output token (TPOT). With continuous batching, a replica runs many sequences at once, so the natural unit of capacity is a slot: one sequence in flight. How many slots a replica can hold while meeting the TPOT target is the first number you need; how long each request occupies a slot is the second.

So QPS is a derived quantity: slots divided by the time each request holds a slot. Change the output length distribution and QPS changes even though the hardware did not. When a product team gives you a QPS number, ask for the token distribution behind it, and when you quote a QPS capacity, quote the mix it assumes. Token-denominated limits, such as tokens per minute, carry the mix inside them and are the better contract with callers.

Four measured numbers

Run a concurrency sweep with a load generator that replays your real prompt and output length distribution, on the exact model, engine version, quantisation and tensor-parallel layout you will ship. Hold concurrency fixed at each level long enough to reach steady state, and record p50 and p95 TTFT and TPOT. An illustrative sweep for one replica serving a chat workload averaging 1,500 input and 300 output tokens:

Concurrent sequencesp95 TPOTp95 TTFT (no queue)Meets TPOT SLO of 50 ms?
822 ms0.15 sYes
1628 ms0.19 sYes
3240 ms0.25 sYes, with margin
4858 ms0.42 sNo

Choose S = 32 slots: the highest level that meets the TPOT target with margin for noisy neighbours and long prompts. The per-request service time at that level is prefill plus decode: T = 0.25 s + 300 tokens x 0.040 s = 12.25 s. These numbers are for illustration; yours will differ by model and engine, which is precisely why you measure them.

Concurrency, duration and per-replica capacity

Little's law says that in any stable system, average number in the system equals arrival rate times average time in the system: L = lambda x W. Apply it to one replica at its slot limit: 32 slots, each occupied for 12.25 s, sustain lambda = 32 / 12.25 = 2.61 requests per second. Cross-check in tokens: 2.61 requests x 300 output tokens is about 784 output tokens per second, close to the 800 that 32 streams at 25 tokens per second each would produce; the gap is the prefill time inside T.

Apply it to the whole demand: a peak of 40 QPS with 12.25 s per request means 40 x 12.25 = 490 requests in flight on average, which is 15.3 replicas worth of slots. That is the floor, the size at which the fleet is 100 percent busy on average. Run there and the queue grows without bound at every random burst.

Mixed workloads use the weighted service time. If 30 percent of traffic is summarisation with 8,000 input tokens and 200 output tokens, its prefill takes about 1.33 s at the same measured prefill rate, so T is 1.33 + 200 x 0.040 = 9.33 s. The blended T is 0.7 x 12.25 + 0.3 x 9.33 = 11.38 s and slot capacity rises to 2.81 QPS. Do not celebrate: long prefills stall decode steps for everyone and their KV cache can lower S, so re-run the sweep with the blended mix before trusting a blended number.

Queueing: utilisation against the tail

Arrivals are random, so some requests find every slot busy and wait. That wait lands directly in TTFT. A good first model treats each replica as an M/M/c queue with c = S servers: Poisson arrivals, exponentially distributed service times. Erlang's C formula gives the probability that an arrival waits, and the wait beyond t seconds has a simple tail.

import math

def erlang_c(c, a):
    # P(arrival waits) in M/M/c, offered load a = lam / mu, valid for a < c
    logs = [k * math.log(a) - math.lgamma(k + 1) for k in range(c)]
    top = c * math.log(a) - math.lgamma(c + 1) + math.log(c / (c - a))
    m = max(logs + [top])
    s, t = sum(math.exp(l - m) for l in logs), math.exp(top - m)
    return t / (s + t)

def p_wait_over(c, lam, service_s, t):
    mu = 1.0 / service_s
    a = lam / mu
    if a >= c:
        return 1.0
    return erlang_c(c, a) * math.exp(-(c * mu - lam) * t)

def max_qps(c, service_s, t, target):
    lo, hi = 0.0, c / service_s
    for _ in range(60):
        mid = (lo + hi) / 2
        lo, hi = (mid, hi) if p_wait_over(c, mid, service_s, t) < target else (lo, mid)
    return lo

def replicas_pooled(peak, slots, service_s, t, target):
    n = math.ceil(peak * service_s / slots)
    while p_wait_over(n * slots, peak, service_s, t) >= target:
        n += 1
    return n

Set the queueing budget from the TTFT SLO. With a 1 s p99 TTFT target and 0.25 s of prefill, allow queue waits above 0.5 s for at most 1 percent of requests. For one replica (c = 32, T = 12.25 s):

Arrival rateUtilisationP(wait at all)P(wait over 0.5 s)
1.5 QPS57%0.27%0.15%
1.70 QPS65%1.6%1.0% (the limit)
2.0 QPS77%10.3%7.6%
2.3 QPS88%38%33%
2.5 QPS96%73%69%

The usable capacity of an isolated replica is 1.70 QPS, not 2.61: it must idle 35 percent of its slots to keep the tail inside budget. The curve is steep: the step from 77 to 88 percent utilisation multiplies the late-request share by four. This is the knee every capacity conversation should show.

M/M/c is a starting point, not the truth. Real arrivals are burstier than Poisson, and output lengths often have a long tail. The Allen-Cunneen approximation scales the M/M/c wait by (Ca2 + Cs2) / 2, where Ca and Cs are the coefficients of variation of inter-arrival and service times; measure both from logs. Also note that TPOT was pinned at its full-slot value, which is conservative when the replica is partly empty.

Routing changes the answer

From a load test to a fleet size: four measured numbers, one queue model, three addersConcurrency sweepTTFT and TPOT per levelTraffic logsinput/output token mixArrival tracepeak rate, burstinessPer-replica modelslots S, service time TQueue modelErlang C, tail waitPeak QPSReplicasmeet wait SLO+ burst headroomramp x scale lag+ failure spareN+1 or zone lossGoodput checkre-test at chosen sizeEvery input is measured on your model, engine version and traffic mix; none comes from a spec sheet.
The sizing pipeline. The queue model runs on the pooled fleet when the router balances on load, and per replica when it does not.

How requests reach replicas changes the answer. Three ways to serve the 40 QPS peak:

  • Divide by the mean: 40 / 2.61 rounds up to 16 replicas. At 96 percent utilisation, 9.5 percent of requests wait more than 0.5 s, so it misses the SLO by an order of magnitude.
  • Random or round-robin routing: each replica is its own queue, so each can take 1.70 QPS, and 40 / 1.70 needs 24 replicas.
  • Load-aware routing (a shared queue, or join-shortest-queue on free slots): the fleet behaves like one large M/M/c with c = 32n. replicas_pooled(40, 32, 12.25, 0.5, 0.01) returns 17 replicas at 90 percent utilisation, with 0.1 percent of requests waiting over 0.5 s.

Pooling saved seven replicas, about 30 percent of the fleet, for the price of a router that knows each replica's free slots. That is usually the single highest-value change in an LLM serving stack, and it is why engine metrics such as running and waiting sequence counts should feed the load balancer. Prefix-cache-aware routing pulls the other way, concentrating requests that share a prompt; weigh cache hits against queue balance with measurements.

Burst headroom, failure spares and token limits

The queue model sizes for a steady peak. Two adders cover what it ignores.

Burst and scale-up lag. A new replica needs a node, an image pull and a model load, often several minutes for large weights. Size the headroom as demand ramp rate times scale-up lag. If traffic can climb 0.5 QPS per minute and a replica takes 8 minutes to become ready, the fleet must absorb 4 QPS above its current target, about 2 replicas at the pooled operating point of 40 / 17 = 2.35 QPS each. Pre-warmed standby nodes and cached weights cut the lag and therefore the headroom.

Failure spare. Add at least one replica so that losing one still meets the SLO; if replicas share a zone, size for the loss of the largest zone. In the example this brings the plan to 17 + 2 + 1 = 20 replicas for a 40 QPS peak, against the 16 a mean-based spreadsheet would have bought.

Finally, translate the plan into the units callers see. 40 QPS at 1,800 total tokens per request is 72,000 tokens per second, or 4.32 million tokens per minute. Publish per-tenant limits in tokens per minute that sum to less than this, so one tenant's long prompts cannot consume the queueing budget everyone shares.

Goodput and admission control

Raw throughput counts every request served. Goodput counts only requests that met both the TTFT and TPOT targets. Past the knee, throughput keeps rising slightly while goodput falls, because more requests finish late. Track goodput per replica and per tenant, and use it as the autoscaling and alerting signal instead of GPU utilisation, which reads near 100 percent on a healthy decode-bound replica and tells you nothing about queues.

When demand exceeds the plan, protect goodput by refusing early rather than serving late: admission control that returns 429 with a retry hint once the waiting queue exceeds the budgeted wait, priority classes for interactive traffic over batch, and output-length caps. Re-run the load test at the chosen fleet size with the measured arrival trace replayed, and confirm p99 TTFT and goodput before launch.

Failure modes

  • QPS without a mix. A capacity number quoted without its token distribution is meaningless the day the product adds a document upload feature.
  • Sizing to the mean. 100 percent average utilisation means unbounded queues.
  • Benchmarks at fixed concurrency only. Closed-loop load tests hide queueing; replay an open-loop arrival trace for the final check.
  • Ignoring the output tail. A few 4,000-token responses hold slots for minutes; cap max tokens per tier and measure Cs.
  • Utilisation as the scaling signal. Scale on waiting requests, TTFT and goodput.
  • Stale measurements. An engine upgrade, a new quantisation or a longer system prompt moves S and T. Re-sweep on every change and keep results versioned.
  • Retry storms. Clients retrying on timeout multiply arrivals exactly when the fleet is saturated; require jittered backoff and honour Retry-After.

Trade-offs

ChoiceGainCost
Higher slot count SMore QPS per replicaHigher TPOT, less headroom for long prompts
Load-aware routingAbout 30% fewer replicas in the exampleRouter complexity, engine metrics required
Tighter TTFT budgetBetter UXLower utilisation, more replicas
Pre-warmed standbySmaller burst headroomPaying for idle GPUs
Admission controlProtects goodputVisible refusals at peaks

Related reading

Related reading: GPU Inference Latency, in depth for where TTFT and TPOT come from; TTFT, in depth for the prefill side of the service time; continuous batching for why slots are the unit; GPU Capacity Planning for fleet-level budgeting; and LLM SLO Burn Rate Alerts for alerting on the targets used here.

What to do next

  1. Extract the input and output token distribution of your real traffic, including p95 and max.
  2. Run a concurrency sweep on the exact shipping stack and pick S from the TPOT target.
  3. Compute T and the Little's law floor; then run the Erlang C code with your TTFT wait budget.
  4. Check whether your router balances on free slots; if not, estimate the pooled saving.
  5. Measure scale-up lag end to end and size burst headroom as ramp rate times lag.
  6. Add a failure spare, publish per-tenant token-per-minute limits, and wire admission control.
  7. Replay a recorded arrival trace at the planned size and confirm p99 TTFT and goodput.
Key takeaway: LLM QPS is slots divided by service time, both measured on your stack and mix. Little's law gives the floor; a queueing model prices the tail and shows that a replica must idle a third of its slots to keep waits short. Load-aware routing recovers much of that, burst and failure adders finish the plan, and goodput, not utilisation, tells you whether it worked.