Every request to a large language model runs in two phases. Prefill pushes the whole prompt through the model in one forward pass and writes the key and value cache. Decode then produces the answer one token at a time, each step reading the weights and the cache again. The two phases stress the GPU in opposite ways, so every serving system has to decide how to split the hardware between them: in time on the same GPU, or in space across separate pools.
This article is about making that decision from evidence. Other pages on this site explain the mechanisms; here you will build a small harness that measures each phase on your own deployment, fit a two-number model of decode cost, and turn the result into a GPU count for each phase. A worked example shows the same model and hardware needing a 3:3 prefill-to-decode ratio for one workload and 1:4 for another, which is why no single split is right for everyone.
Why the two phases want different hardware
Start from arithmetic. A dense model with P parameters spends about 2P floating-point operations per token in its matrix multiplies. For an 8-billion-parameter model that is about 16 GFLOP per token. Prefill of a 2,000-token prompt is therefore about 32 TFLOP of work done in large matrix multiplies, which keeps the tensor cores busy: prefill is compute-bound and its cost grows with prompt length.
Decode is different. Each step produces one token per sequence, but it must read every weight from high-bandwidth memory (HBM) and every cached key and value for every sequence in the batch. In BF16 the weights of an 8B model are about 16 GB. On an H100 SXM with 3.35 TB/s of peak HBM bandwidth, just reading them takes about 4.8 ms, whether the batch holds one sequence or a hundred. Decode is memory-bound, and the way to make it efficient is to batch many sequences so the weight read is shared.
On one GPU they interfere: a decode step waiting behind a 2,000-token prefill stalls every sequence in the batch. Separate them and you pay for moving the KV cache and for two pools that can each sit idle. The prefill/decode disaggregation architecture page covers the rooflines and the transfer in detail. The question here is narrower: for your traffic, how much hardware does each phase need, and does splitting pay?
Metrics that separate the phases
Three client-side numbers map onto the phases. Time to first token (TTFT) covers queueing, prefill and the first decode step. Inter-token latency (ITL) is the gap between consecutive tokens, and time per output token (TPOT) is the mean ITL after the first token. TTFT is governed by prefill capacity; TPOT is governed by decode batch size and context length; and interference shows up as ITL spikes, not as a worse mean.
Set targets per product: a coding assistant may accept a 1.5 s TTFT but need a steady 20 ms TPOT, while a classifier returning ten tokens cares only about TTFT. Score with goodput, the requests per second that meet both targets at p99.
| Metric | Phase it measures | Main lever |
|---|---|---|
| TTFT p99 | queue + prefill | prefill capacity, prompt length, prefix caching |
| TPOT p50 | steady decode | decode batch size, resident context |
| ITL p99 | decode stalls | chunk size, prefill priority, splitting pools |
| Goodput | both | the split itself |
A measurement harness for streaming endpoints
Measure from the client, against the streaming endpoint, with request sizes drawn from production logs. The script below sends requests at a fixed arrival rate (open loop) and records, per request, the prompt and output token counts, TTFT and every token gap. It talks to any OpenAI-compatible server, which vLLM, SGLang and TensorRT-LLM's server all expose, and asks for usage in the final chunk so token counts come from the server's tokenizer rather than a guess.
import asyncio, json, random, time
import httpx
URL = "http://localhost:8000/v1/completions"
async def one(client, prompt, max_tokens, out):
body = {"model": "m", "prompt": prompt, "max_tokens": max_tokens,
"stream": True, "stream_options": {"include_usage": True}}
t_send = time.perf_counter()
stamps, usage = [], None
async with client.stream("POST", URL, json=body) as r:
async for line in r.aiter_lines():
if not line.startswith("data: ") or line == "data: [DONE]":
continue
chunk = json.loads(line[6:])
if chunk.get("usage"):
usage = chunk["usage"]
if chunk.get("choices") and chunk["choices"][0].get("text"):
stamps.append(time.perf_counter())
if not stamps:
return
out.append({"ttft": stamps[0] - t_send,
"gaps": [b - a for a, b in zip(stamps, stamps[1:])],
"prompt_tokens": usage and usage["prompt_tokens"],
"output_tokens": usage and usage["completion_tokens"],
"events": len(stamps)})
async def run(requests, rate_per_s):
out = []
async with httpx.AsyncClient(timeout=None) as client:
tasks = []
for prompt, max_tokens in requests:
tasks.append(asyncio.create_task(one(client, prompt, max_tokens, out)))
await asyncio.sleep(random.expovariate(rate_per_s)) # Poisson arrivals
await asyncio.gather(*tasks)
return outTwo details matter. The client sends on a schedule, because a closed-loop client slows down with the server and hides the queueing you want to see. And it counts events separately from tokens, because a server may pack several tokens into one streamed chunk; if events is well below output_tokens, divide gaps accordingly or your ITL is overstated.
Fitting a cost model for each phase
Client traces tell you whether you meet the targets. To plan capacity you also need the cost of each phase in isolation. Run two microbenchmarks on one GPU with nothing else loaded.
Prefill throughput. Send single requests with prompts of 512, 2,000 and 8,000 tokens and max_tokens=1, with prefix caching off or with randomised prompts. Prompt tokens divided by TTFT gives prefill tokens per second. It falls slowly as prompts grow because attention cost is quadratic in length.
Decode step time. Hold B sequences of a fixed context length in flight and record the steady ITL. Repeat over a grid of batch sizes and context lengths. Then fit the simplest model that respects the physics: a fixed cost t0 for reading the weights, plus a cost k for every resident token the attention kernels must read.
import numpy as np
# (batch, mean context tokens, measured step ms) - illustrative, measure your own
samples = [(1, 2000, 6.2), (16, 2000, 9.4), (32, 2000, 12.9), (64, 2000, 19.6),
(32, 4000, 19.8), (64, 1000, 12.8), (8, 8000, 14.3)]
X = np.array([[1.0, b * ctx / 1e3] for b, ctx, _ in samples])
y = np.array([ms for *_, ms in samples])
(t0, k), *_ = np.linalg.lstsq(X, y, rcond=None)
print(f"t0 = {t0:.2f} ms, k = {k:.4f} ms per 1k resident tokens")
# t0 = 6.23 ms, k = 0.1066 ms per 1k resident tokensThe fitted t0 of about 6.2 ms is close to the 4.8 ms weight-read floor plus kernel launch and sampling overhead, which is a useful sanity check. The k term implies the attention kernels read the cache at roughly 1.2 TB/s (128 KiB per token for an 8B model with grouped-query attention, divided by 0.107 µs). That is well below the 3.35 TB/s peak, which is normal: paged attention reads scattered blocks and rarely runs at full bandwidth. The model exists to rank options, not to predict to the millisecond.
Worked example: two workloads, two ratios
Take Llama-class 8B on H100 SXM GPUs with the fitted numbers above, illustrative prefill throughput of 25,000 tokens per second per GPU, a 20 ms TPOT target, and about 58 GB of HBM left for cache after weights and activations, which at 128 KiB per token holds roughly 442,000 tokens. Size each phase separately.
import math
def plan(rate, mean_in, mean_out, prefill_tps=25_000, tpot_ms=20,
kv_tokens=442_504, util=0.7, t0=6.23, k=0.1066):
prefill_gpus = rate * mean_in / (prefill_tps * util) # keep queues short
ctx = mean_in + mean_out / 2 # mean resident context
b_slo = (tpot_ms - t0) / (k * ctx / 1e3) # batch that meets TPOT
b_mem = kv_tokens / (mean_in + mean_out) # batch that fits in HBM
concurrency = rate * mean_out * tpot_ms / 1e3 # Little's law
decode_gpus = concurrency / min(b_slo, b_mem)
return math.ceil(prefill_gpus), round(decode_gpus, 2)
print(plan(20, 2000, 300)) # RAG-style: (3, 2.0) b_slo = 60, concurrency = 120
print(plan(20, 300, 1200)) # chat-style: (1, 3.34) b_slo = 143, concurrency = 480The RAG-style workload (20 requests per second, 2,000 tokens in, 300 out) needs 40,000 prefill tokens per second, which is 2.3 GPUs at 70% utilisation, so 3. Decode holds about 120 sequences in flight; with 2,150 tokens of resident context each, only 60 fit per GPU under the 20 ms target, so decode needs exactly 2.0 GPUs. A result sitting on the boundary means provision 3. The ratio is 3:3.
The chat-style workload has the same arrival rate but 300 tokens in and 1,200 out. Prefill needs a third of a GPU. Decode holds 480 sequences, 143 per GPU, so 3.34 GPUs, provisioned as 4. The ratio is 1:4. Same model, same hardware, opposite shape. Notice also what binds: in both cases the TPOT target limits the batch long before memory does (192 and 295 sequences would fit), so buying GPUs with more HBM would not help; a looser TPOT target or a smaller model would.
Choosing the split
With the phase models in hand, compare three ways to split.
| Split | How it divides the GPU | Wins when | Costs |
|---|---|---|---|
| Colocated, no chunking | whole prefills between decode steps | short prompts, loose ITL targets | ITL p99 spikes equal to the longest prefill |
| Colocated, chunked prefill | a token budget per iteration shared by both | mixed traffic, one pool, small scale | ITL rises by the chunk's cost every step; prefill efficiency drops |
| Disaggregated | separate prefill and decode GPUs | long prompts with tight TPOT, large fleets | KV transfer, two pools to scale, more failure paths |
Use the model to test the middle option. In the RAG example, a decode step at 60 sequences takes about 20 ms. If a 512-token prefill chunk rides along, it adds roughly 512 / 25,000 s, about 20 ms, so the step doubles and TPOT misses its target. Even a 128-token chunk adds about 5 ms, so each GPU must hold fewer sequences to stay under 20 ms, which means more decode GPUs, while prefill runs in thin, less efficient slices that raise TTFT under load. When the arithmetic says no chunk size satisfies both targets, that is the signal to split pools. The chunked prefill article shows how to derive the budget precisely, and disaggregated serving covers the proxy, routing and launch flags once you decide to split.
For the chat workload the picture flips. Prompts are short, prefill is a small fraction of GPU time, and a chunk budget of a few hundred tokens rarely fills. Colocated serving with continuous batching and chunking is usually the cheaper answer, and a dedicated prefill pool would sit mostly idle.
Proving interference before you split
Before you commit, confirm interference is real on your stack with a two-step experiment. First, run decode-only load: many requests with short prompts and long outputs, and record ITL p50 and p99. Second, keep that load and inject one long-prompt request every few seconds. If ITL p99 jumps by roughly the prefill time of the injected prompt while p50 barely moves, interference is the problem and splitting, or a smaller chunk budget, will fix it. If p99 barely moves, the scheduler is already protecting decode and a split would buy little. Compute both percentiles from the gaps lists the harness already records.
Failure modes in measuring the phases
Most bad split decisions come from bad measurements.
- Prefix cache hits. Repeating the same benchmark prompt lets the server reuse cached KV, so prefill looks nearly free. Randomise prompts or disable prefix caching for the microbenchmark, then measure your real hit rate separately and discount prefill demand by it.
- Synthetic lengths. Fixed-length tests hide the long prompts that cause stalls. Sample lengths from production logs.
- Closed-loop load. Clients that wait for each reply cut their own arrival rate when latency rises, so the overloaded case never appears in the data.
- Chunked streaming. If the server emits several tokens per event, per-event gaps overstate ITL and per-token averages hide stalls. Count both.
- Averages. The mean ITL can improve while p99 doubles. Always report p50 and p99 per phase.
Operating the split
Once a split is live, run it from the same two numbers. Export TTFT and ITL histograms per pool, plus queue depth on the prefill side and running batch size on the decode side. Scale prefill on queue depth or TTFT p99, and decode on batch size against the b_slo you computed, not on GPU utilisation, which reads near 100% for a memory-bound decode GPU that still has room. Re-run the microbenchmarks after every model, driver or engine upgrade, because t0 and k move. If you split pools, also watch KV transfer time between them; peer-to-peer KV transfer explains what limits it.
The input-to-output ratio drifts as products change, so re-run the plan monthly; a ratio sized for last quarter leaves one pool saturated and the other idle. Disaggregation lets each phase use its own parallelism, but every extra pool adds a scaling loop and a failure path, so it pays only when interference measurably costs goodput.
What to do next
- Pull a week of production request logs and record the p50 and p99 of prompt and output length per product.
- Write down TTFT and TPOT targets for each product, at p99, and define goodput from them.
- Run the open-loop harness against your current deployment at today's peak rate and record goodput.
- Run the prefill and decode microbenchmarks on one idle GPU and fit t0 and k.
- Feed the plan function your real rates and lengths, and note which limit binds: TPOT or memory.
- Run the interference experiment; split pools only if ITL p99 tracks injected prefill time.
- Schedule the microbenchmarks to re-run after every engine, driver or model change.