Ask how fast a model is and you will get a single number, usually tokens per second at some batch size nobody wrote down. That number is close to useless for a product. A person waiting on a chat reply experiences two different latencies: how long until the first word appears, and how fast the words arrive after that. A batch job cares about neither and counts tokens per GPU-hour, and the knob that helps throughput hurts per-request speed.
This article builds latency up from physics: why prefill is limited by arithmetic and decode by memory bandwidth, a short Python model you can calibrate, the contention that model leaves out, which optimisation moves which term, and a measurement harness. The worked numbers use Llama 3 8B in BF16 on one H100 SXM, because its shape and the GPU's datasheet figures are public.
Three numbers, not one
Time to first token (TTFT) runs from the moment the client sends the request to the moment the first generated token reaches it. It includes the network, any queue, the whole prompt's prefill and one decode step. Time per output token (TPOT), also called inter-token latency, is the average gap between tokens once streaming has started. End-to-end latency is the total, and for a streamed response of N tokens it decomposes exactly:
E2E = TTFT + (N - 1) x TPOTThe decomposition tells you which number to optimise. A 300-token answer with a 7.6 ms TPOT spends about 2.3 seconds streaming, so shaving 20 ms off TTFT is invisible in E2E but very visible to a user watching a blank screen. A classification call that emits 5 tokens is almost all TTFT.
Always state latency as a distribution at a stated load. The objective you will eventually write looks like p95 TTFT under 500 ms and p95 TPOT under 40 ms at 20 requests per second with prompts of 2,000 tokens.
Where the milliseconds go
| Stage | Typical scale | What sets it |
|---|---|---|
| Network and TLS | 1-50 ms | distance, connection reuse, proxies and gateways |
| Tokenisation | under 1 ms to a few ms | CPU speed and prompt length; rarely the problem |
| Queue and admission | 0 ms to seconds | load relative to capacity, KV-cache space, scheduling policy |
| Prefill | tens to hundreds of ms | prompt tokens times model FLOPs, divided by achieved FLOP rate |
| Each decode step | a few to tens of ms | bytes read per step divided by achieved memory bandwidth |
| Sample, detokenise, flush | sub-ms per step | sampling kernels, response buffering in proxies |
The two GPU stages can be reasoned about from first principles. The queue is zero in a benchmark and dominant in production.
Prefill is a compute problem
A dense transformer forward pass costs about 2 FLOPs per parameter per token: one multiply and one add for every weight the token touches. Prefill pushes every prompt token through the model at once, so the work is 2 x P x T for P parameters and T prompt tokens, done as large matrix multiplications that keep the tensor cores busy. Attention adds a term growing with T squared that only matters at long context.
For Llama 3 8B and a 2,000-token prompt that is 2 x 8e9 x 2,000 = 3.2e13 FLOPs. An H100 SXM's datasheet peak for dense BF16 is 989 TFLOPS. No real kernel sustains peak, so the model multiplies it by an achieved fraction; 0.55 is an assumption, not a measurement, and you should replace it with yours. The result is about 59 ms of prefill. Double the prompt and prefill doubles. Long retrieval contexts show up first as TTFT, and caching a shared prefix, covered in prefix caching, removes TTFT rather than TPOT.
Decode is a bandwidth problem
Decode produces one token per sequence per step. Each step still has to read every weight from HBM, but it only does 2 FLOPs per weight per sequence in the batch. At batch 1 an 8B model in BF16 reads 16 GB to do 16 GFLOPs of work, one FLOP per byte, while an H100 can do roughly 300 FLOPs in the time it takes to read one byte. The tensor cores sit idle and the step time is simply bytes moved divided by bandwidth.
Bytes per step are the weights plus the KV cache of every sequence in the batch. Llama 3 8B has 32 layers, 8 KV heads of dimension 128 and stores keys and values in 2 bytes, so each token of context costs 2 x 32 x 8 x 128 x 2 = 131,072 bytes, exactly 128 KiB. At 3.35 TB/s and an assumed 80 percent achieved bandwidth, reading 16 GB of weights takes about 6 ms, and a single sequence's 2,000-token KV cache adds about 0.1 ms. That 6 ms is a hard floor on TPOT for this model on this GPU. Only fewer bytes (quantisation, a smaller model), more bandwidth (a newer GPU, tensor parallelism across GPUs) or more tokens per read (batching, speculative decoding) can move it.
A runnable latency model
The two floors fit in twenty lines. The model below adds a fixed per-step overhead for kernel launches, sampling and the scheduler, again an assumption to measure, and computes TTFT, TPOT, E2E and throughput for a 2,000-token prompt and a 300-token answer across batch sizes.
# First-order latency model for one decoder-only model on one GPU.
PEAK_FLOPS = 989e12 # H100 SXM dense BF16 (datasheet)
HBM_BW = 3.35e12 # H100 SXM HBM3, bytes/s (datasheet)
MFU_PREFILL = 0.55 # ASSUMPTION: fraction of peak your prefill GEMMs reach
BW_EFF = 0.80 # ASSUMPTION: fraction of HBM bandwidth decode reaches
STEP_OVERHEAD = 0.0015 # ASSUMPTION: s per step (launch, sampling, scheduler)
PARAMS = 8.0e9 # Llama 3 8B
W_BYTES = PARAMS * 2 # BF16 weights
KV_PER_TOKEN = 2 * 32 * 8 * 128 * 2 # K+V x 32 layers x 8 KV heads x 128 x 2 bytes
def prefill_s(prompt_tokens):
return 2 * PARAMS * prompt_tokens / (PEAK_FLOPS * MFU_PREFILL)
def decode_step_s(batch, ctx_tokens):
bytes_moved = W_BYTES + batch * ctx_tokens * KV_PER_TOKEN
mem_s = bytes_moved / (HBM_BW * BW_EFF)
flop_s = 2 * PARAMS * batch / (PEAK_FLOPS * MFU_PREFILL)
return max(mem_s, flop_s) + STEP_OVERHEAD
def request(prompt=2000, out=300, batch=1, queue_s=0.0, net_s=0.02):
ttft = net_s + queue_s + prefill_s(prompt) + decode_step_s(batch, prompt)
tpot = decode_step_s(batch, prompt + out // 2)
return ttft, tpot, ttft + (out - 1) * tpot
for b in (1, 8, 32, 64, 128):
t, s, e = request(batch=b)
print(f"batch {b:>3}: TTFT {t*1e3:6.1f} ms TPOT {s*1e3:5.2f} ms "
f"E2E {e:5.2f} s tokens/s/GPU {b/s:7.0f}")Running it prints:
batch 1: TTFT 86.4 ms TPOT 7.58 ms E2E 2.35 s tokens/s/GPU 132
batch 8: TTFT 87.1 ms TPOT 8.31 ms E2E 2.57 s tokens/s/GPU 963
batch 32: TTFT 89.4 ms TPOT 10.83 ms E2E 3.33 s tokens/s/GPU 2953
batch 64: TTFT 92.6 ms TPOT 14.20 ms E2E 4.34 s tokens/s/GPU 4507
batch 128: TTFT 98.8 ms TPOT 20.93 ms E2E 6.36 s tokens/s/GPU 6116Check memory before believing the last row: 128 sequences at about 2,150 tokens of context is 36.1 GB of KV cache on top of 16 GB of weights, which fits in 80 GB with room for activations. At 8,000-token contexts the same batch would need about 134 GB and could not run; the batch size you can reach is set by KV capacity long before compute. Calibrate the three assumptions by running your engine at batch 1 and batch 64 with a fixed prompt and solving for them.
Reading the batch curve
The output is the central trade-off of LLM serving in five lines. Going from batch 1 to batch 64 multiplies throughput by 34 and only doubles TPOT, because the weights are read once per step whichever batch size you run. Past that the KV-cache term grows until it rivals the weights, and TPOT climbs roughly linearly while throughput gains shrink. Somewhere on this curve is a knee where your TPOT objective is still met and cost per token is near its minimum; it moves with context length.
Notice what batching does not do in this model: TTFT barely moves. That is only true because the model runs one prefill at a time with nothing else competing, which is exactly what production does not look like.
The term the model leaves out: contention
Real servers run continuous batching: new requests join the running batch between steps. A newly admitted prompt's prefill shares the GPU with everyone else's decode. If eight 2,000-token prompts arrive together and the engine prefills them back to back, each waits for the ones ahead of it. Extending the model with that serial queue gives about 87 ms TTFT for the first, 264 ms for the fourth and 499 ms for the eighth; meanwhile every decoding sequence stalls for the half second those prefills occupy, which shows up as one enormous inter-token gap in the middle of somebody else's answer.
Three mechanisms tame it. Chunked prefill caps the prompt tokens per step so decodes keep flowing, trading a slightly longer TTFT for a bounded worst gap. Prefill-decode disaggregation, compared in the disaggregation article, runs the two phases on separate GPU pools and pays for it by moving the KV cache between them. Deadline-aware scheduling, covered in SLO scheduling, decides whose prefill goes first. Beyond those, queueing theory applies: as utilisation approaches 100 percent, waiting time grows without bound, so leave headroom.
Levers, mapped to the term they move
| Lever | Moves | Cost or catch |
|---|---|---|
| Prefix caching | TTFT (skips cached prefill) | only helps repeated prefixes; eats KV memory |
| Chunked prefill | worst inter-token gap | slightly higher TTFT and some throughput |
| Weight quantisation (FP8, INT4) | TPOT (fewer bytes per step) | accuracy must be re-evaluated |
| KV-cache quantisation | TPOT at long context, batch capacity | accuracy at long context |
| Tensor parallelism | TPOT and prefill (more bandwidth and FLOPs) | all-reduce per layer; poor scaling over slow links |
| Speculative decoding | TPOT at low batch | gains shrink as batch grows and the GPU becomes busy |
| Bigger batch | throughput | TPOT rises; KV capacity caps it |
| Lower utilisation target | queue time, all tails | more GPUs per request per second |
The table also works in reverse as a diagnosis. If TTFT is bad and TPOT is fine, look at queueing, prompt length and prefix reuse. If TPOT is bad at low load, look at bytes per step. If TPOT is fine on average and awful at p99, look for prefill interference.
Measuring latency honestly
Measure from the client, stream the response and record a timestamp per chunk. Server-side histograms cannot see the load balancer, TLS or a proxy that buffers the response. The harness below works against any OpenAI-compatible streaming endpoint.
import json, time, httpx # any OpenAI-compatible streaming server (vLLM, SGLang)
def timed_request(url, model, prompt, max_tokens=300):
body = {"model": model, "prompt": prompt, "max_tokens": max_tokens, "stream": True}
t0 = time.perf_counter()
stamps = []
with httpx.stream("POST", f"{url}/v1/completions", json=body, timeout=120) as r:
for line in r.iter_lines():
if not line.startswith("data: ") or line == "data: [DONE]":
continue
chunk = json.loads(line[6:])
if chunk["choices"][0].get("text"):
stamps.append(time.perf_counter())
ttft = stamps[0] - t0
gaps = [b - a for a, b in zip(stamps, stamps[1:])] # inter-chunk latency
return {"ttft": ttft, "e2e": stamps[-1] - t0,
"tpot": (stamps[-1] - stamps[0]) / max(len(stamps) - 1, 1),
"worst_gap": max(gaps, default=0.0)}Servers may coalesce several tokens into one chunk, so per-chunk gaps are an upper bound on inter-token latency. Drive it with an open-loop arrival process at a fixed rate, not a fixed number of concurrent users, or you will never see the queue; the load-testing article explains why and how. Report p50, p95 and p99 of TTFT, TPOT and worst gap separately, at each request rate, with prompt and output length distributions written next to them.
Failure modes
- Benchmarking the wrong shape. Measuring 128-token prompts and deploying for 8,000-token retrieval contexts; prefill and KV capacity both scale with length.
- Buffered streaming. An API gateway or ingress holds SSE chunks, so the server's TTFT is 90 ms and the user's is the full E2E.
- Cold starts mixed into steady state. Model load takes seconds to minutes; measure warm-up and scale-out separately.
- KV-cache exhaustion. When the cache is full the engine queues or preempts sequences, and TTFT jumps from milliseconds to seconds with no change in GPU utilisation.
- Running hot. Utilisation near capacity makes the tail explode; a small traffic spike turns a healthy p99 into timeouts.
What to do next
- Write your latency objective as a sentence: percentile, TTFT and TPOT targets, request rate and prompt and output length distribution.
- Compute the floors for your model and GPU: 2 x params x prompt tokens over FLOP rate, and weight plus KV bytes over bandwidth.
- Run the model above, then calibrate its three assumptions against your engine at batch 1 and batch 64.
- Measure from the client with streaming, open-loop load and per-chunk timestamps; confirm nothing between client and server buffers.
- Find the knee: the highest request rate at which p95 TTFT and TPOT still meet the objective, and run production at a margin below it.
- Pick levers from the table by the term that misses its target, and re-measure after each change rather than stacking them blind.