Before buying GPUs for a language model service, someone has to answer a simple-sounding question: how many requests per second can one replica handle while meeting the latency target? Vendor throughput charts do not answer it, because they are measured with batch sizes, prompt lengths and latency budgets that are not yours. Load tests answer it precisely but need hardware you have not bought yet. Between the two sits estimation: a model built from first principles that is right to within a factor you can state, and that tells you which limit you will hit first.
This article builds that estimator for a single serving replica. It derives the memory budget, the time of one decode step from memory bandwidth, the cost of prefill from compute, and combines them into a maximum request rate under a time-per-output-token target. A runnable calculator produces the worked example. Fleet sizing on top of a per-replica number is covered in GPU capacity planning, and finding the real knee on hardware in LLM load testing.
Define the workload and the two phases
Capacity means nothing without a workload and a latency target. Fix four numbers first: the distribution of prompt lengths, the distribution of output lengths, the time-to-first-token target (TTFT) and the time-per-output-token target (TPOT, also called inter-token latency). Use percentiles from logs if you have them; averages hide the long prompts that dominate cost.
Then separate the two phases of a request. Prefill processes the whole prompt in one pass. It does about 2 floating-point operations per parameter per prompt token, in large matrix multiplies, so it is compute-bound. Decode produces one token per step for every request in the batch. Each step reads all the weights and every active request's KV cache from HBM but does little arithmetic per byte, so at practical batch sizes it is memory-bandwidth-bound. The decode compute ledger shows the per-operator detail; for estimation the bandwidth view is enough.
The memory budget
Start with memory, because it caps concurrency. Weights take parameters times bytes per parameter: a 70.6-billion-parameter model in BF16 needs about 141 GB. What remains of the memory the engine may claim, minus a reserve for activations, CUDA graphs and fragmentation, is the KV cache pool.
Each token in the cache costs 2 (keys and values) times layers times KV heads times head dimension times bytes per element. For Llama 3.1 70B, with 80 layers, 8 KV heads under grouped-query attention and head dimension 128, that is 2 x 80 x 8 x 128 x 2 = 327,680 bytes, 320 KiB per token, the same figure derived in KV cache sizing. Divide the pool by bytes per token to get pool tokens, and by the tokens a request holds at its end (prompt plus output) to get the number of requests the cache can hold at once. Paged allocators waste little, so this is close to the real ceiling.
Decode is a bandwidth problem
One decode step must stream the weights once and each active request's cache once. Its lower-bound time is bytes moved divided by achieved bandwidth:
step_time(B) = (weight_bytes + B * avg_context * kv_bytes_per_token) / (gpus * hbm_bw * bw_eff)Two consequences follow. At batch 1 the step costs roughly weight bytes over bandwidth, so a single user gets the fastest tokens but the GPU is mostly idle in arithmetic terms. As B grows, the weight read is shared across more tokens and throughput rises almost linearly, until the KV term dominates and each extra request adds as much time as it adds tokens. Every decode step also produces exactly one token per request, so the step time is the TPOT each user sees, plus any prefill work squeezed into the same iteration.
The bw_eff factor is the share of peak bandwidth a real engine achieves in decode, including tensor-parallel all-reduce and kernel launch gaps. Start at 0.6 to 0.75 and replace it with a measured value as soon as you have one.
Prefill steals decode time
Prefill costs about 2 x parameters x prompt tokens FLOPs, divided by peak FLOPs times an achieved fraction, the model FLOP utilisation or MFU. Engines with continuous batching and chunked prefill interleave prefill chunks into decode iterations, which keeps decode stalls short and TPOT smooth but means prefill time comes straight out of decode time. If prefill occupies a fraction p of the GPU, every decode step stretches by 1 / (1 - p). This is the term most back-of-envelope estimates forget, and for long prompts it is the biggest one.
Prefill also sets TTFT. A request's first token cannot appear before its own prompt has been processed, so TTFT is at least queue wait plus prompt prefill time, and with chunking it is spread over several iterations shared with decode. Larger chunks finish prefill sooner and improve TTFT, but each iteration that carries a large chunk is a long iteration for every decoding user, so TPOT spikes. Smaller chunks smooth TPOT and stretch TTFT. The calculator below models the average share of time prefill takes, not the chunk schedule, so check the TTFT target separately: if prompt prefill alone is close to the TTFT budget at your planned rate, no amount of decode tuning will rescue it, and the fixes are shorter prompts, prompt caching or dedicated prefill capacity.
A calculator you can run
Put the pieces together. At arrival rate lambda, prefill takes the share p = lambda x prefill_time. Decode concurrency B satisfies B = lambda x output_tokens x iteration_time, because each request stays in the batch for one iteration per output token. The iteration time is step_time(B) / (1 - p). Solve that fixed point, then search for the largest lambda whose iteration time meets TPOT and whose B fits the cache:
from dataclasses import dataclass
@dataclass
class Model:
params: float; layers: int; kv_heads: int; head_dim: int
w_bytes: float = 2.0; kv_bytes: float = 2.0
def kv_per_token(self):
return 2 * self.layers * self.kv_heads * self.head_dim * self.kv_bytes
@dataclass
class Replica:
gpus: int; hbm_gb: float; bw_tbs: float; tflops: float
mem_util: float = 0.90; reserve_gb: float = 4.0
bw_eff: float = 0.70; mfu: float = 0.40
def estimate(m, r, prompt, output, tpot_ms):
weights = m.params * m.w_bytes
kvt = m.kv_per_token()
pool = r.gpus * (r.hbm_gb * r.mem_util - r.reserve_gb) * 1e9 - weights
b_mem = int(pool / kvt // (prompt + output))
bw = r.gpus * r.bw_tbs * 1e12 * r.bw_eff
step = lambda b: (weights + b * (prompt + output / 2) * kvt) / bw
prefill = 2 * m.params * prompt / (r.gpus * r.tflops * 1e12 * r.mfu)
def steady(lam):
busy = lam * prefill
if busy >= 1:
return None
b = 1.0
for _ in range(200): # fixed point b = lam * output * t_iter
t_iter = step(b) / (1 - busy)
b = lam * output * t_iter
return b, t_iter
lo, hi = 0.0, 1 / prefill
for _ in range(60): # bisection on the arrival rate
mid = (lo + hi) / 2
s = steady(mid)
if s and s[0] <= b_mem and s[1] * 1000 <= tpot_ms:
lo = mid
else:
hi = mid
b, t_iter = steady(lo)
return dict(b_mem=b_mem, prefill_ms=prefill * 1e3, max_rps=lo,
batch=b, prefill_share=lo * prefill, out_tok_s=lo * output)The model assumes uniform request shapes and steady arrivals, and it ignores sampling, networking and tokenisation overheads. It is meant to find the binding limit and a ceiling, not a contract.
Worked example: Llama 3.1 70B on four H100s
Take Llama 3.1 70B in BF16 on one replica of four H100 SXM GPUs with tensor parallelism. Each GPU has 80 GB of HBM3 at 3.35 TB/s and about 989 TFLOPS of dense BF16. The workload is 2,000 prompt tokens and 400 output tokens per request, with a 40 ms TPOT target. The calculator gives:
| Quantity | Value | Comment |
|---|---|---|
| Weights | 141 GB | 70.6B parameters x 2 bytes |
| KV pool | 130.8 GB | 4 x (72 - 4) GB minus weights |
| Requests the cache holds | 166 | 2,400 tokens x 320 KiB each |
| Decode step at batch 1 | 15.1 ms | weights over achieved bandwidth |
| Prefill per request | 178 ms | 2.8e14 FLOPs at 40% MFU |
| Maximum rate at 40 ms TPOT | 2.98 req/s | about 1,190 output tokens/s |
| Decode batch at that rate | 48 | far below the 166 the cache allows |
| Share of time in prefill | 53% | the hidden majority of the work |
The striking result is that memory is not the limit. The cache could hold 166 requests, but the latency target caps the steady batch at about 48, because more than half of every second goes to prefill and stretches each decode step. Buying more memory would not help. What does help is visible from varying the inputs:
| Change | Max req/s | Binding effect |
|---|---|---|
| Baseline: 2,000 in, 400 out, 40 ms | 2.98 | TPOT, with prefill taking 53% |
| Relax TPOT to 60 ms | 3.58 | bigger batch, still TPOT-bound |
| Tighten TPOT to 25 ms | 1.90 | batch shrinks to 19 |
| Short prompts: 500 in | 11.47 | prefill cost drops fourfold |
| Long prompts: 8,000 in | 0.75 | cache now holds only 47 requests |
| FP8 weights, same prompts | 3.88 | faster decode steps, larger pool |
The FP8 row only halves weight bytes; it does not model FP8 tensor-core speedups in prefill, which would raise the number further. Every row shows the same lesson: prompt length and the TPOT target move capacity far more than the choice between two neighbouring GPU models.
Headroom and calibration
The maximum rate is the point where average inter-token latency equals the target. Run there and the p99 will miss it, queues will form at the first burst, and TTFT will climb without bound. Plan to operate at 60 to 75 percent of the estimated maximum, then refine the number with a load test that sweeps arrival rate and records TTFT and TPOT as distributions. In the example, 70 percent of 2.98 is about 2.1 requests per second per replica.
Calibrate before trusting the estimate. Measure one decode step at batch 1 and at a large batch on the real engine to fit bw_eff, and time a long prefill to fit mfu. With those two fitted, any remaining gap against the load test points to a specific term, such as sampling overhead or uneven request lengths, rather than to guesswork.
Failure modes
- Sizing from memory alone. Dividing the KV pool by context length gives the cache ceiling, not capacity. In the example it overstates the batch by more than three times.
- Ignoring prefill. Estimating from decode throughput charts alone overstates capacity whenever prompts are long, which in retrieval and agent workloads they usually are.
- Using means for lengths. A few 30,000-token prompts can dominate prefill time and KV occupancy. Estimate with percentiles, or with the full length distribution.
- Peak numbers as achieved numbers. Datasheet bandwidth and FLOPs are ceilings. Without efficiency factors, using the defaults here, the estimate is about 40 percent optimistic.
- Forgetting what else lives in HBM. LoRA adapters, speculative-decoding draft models and prefix caches all come out of the KV pool.
- Treating the result as fleet size. One replica's rate still needs headroom, failure spares and diurnal peaks applied before it becomes a GPU count.
Trade-offs
| Lever | Gain | Cost |
|---|---|---|
| Looser TPOT target | Larger batch, more throughput | Slower streaming for every user |
| Quantised weights | Faster decode, bigger KV pool | Accuracy validation per model |
| Quantised KV cache | More concurrent requests | Quality risk at long context |
| Prefill-decode disaggregation | Prefill stops stretching decode | More GPUs and KV transfer |
| Prompt caching | Skips repeated prefill | Memory for the cache, hit rate uncertain |
| Higher tensor parallelism | Lower step time, more memory | All-reduce overhead, fewer replicas |
What to do next
- Pull prompt and output length percentiles and the TTFT and TPOT targets from your product owners and logs.
- Run the calculator for each candidate GPU and parallelism layout and note which limit binds.
- If prefill share exceeds about 40%, evaluate prompt caching, chunk sizes and disaggregated prefill before buying.
- On one real replica, measure batch-1 and large-batch step times and a long prefill, then refit bw_eff and mfu.
- Load test across arrival rates to find the knee and compare it with the corrected estimate.
- Set the planning rate at 60 to 75 percent of the measured knee and hand it to fleet capacity planning.
- Re-run the estimate whenever the model, quantisation, engine version or traffic mix changes.