Ask how fast a model is and you will get a single number, usually tokens per second at some batch size nobody wrote down. That number is close to useless for a product. A person waiting on a chat reply experiences two different latencies: how long until the first word appears, and how fast the words arrive after that. A batch job cares about neither and counts tokens per GPU-hour, and the knob that helps throughput hurts per-request speed.

This article builds latency up from physics: why prefill is limited by arithmetic and decode by memory bandwidth, a short Python model you can calibrate, the contention that model leaves out, which optimisation moves which term, and a measurement harness. The worked numbers use Llama 3 8B in BF16 on one H100 SXM, because its shape and the GPU's datasheet figures are public.

Advertisement

Three numbers, not one

Time to first token (TTFT) runs from the moment the client sends the request to the moment the first generated token reaches it. It includes the network, any queue, the whole prompt's prefill and one decode step. Time per output token (TPOT), also called inter-token latency, is the average gap between tokens once streaming has started. End-to-end latency is the total, and for a streamed response of N tokens it decomposes exactly:

E2E = TTFT + (N - 1) x TPOT

The decomposition tells you which number to optimise. A 300-token answer with a 7.6 ms TPOT spends about 2.3 seconds streaming, so shaving 20 ms off TTFT is invisible in E2E but very visible to a user watching a blank screen. A classification call that emits 5 tokens is almost all TTFT.

Always state latency as a distribution at a stated load. The objective you will eventually write looks like p95 TTFT under 500 ms and p95 TPOT under 40 ms at 20 requests per second with prompts of 2,000 tokens.

Where the milliseconds go

One streamed request, left to rightNetworkTLS, LB, proxyTokenizeCPU, sub-msQueueadmissionPrefillcompute-boundFirst decodetoken 1 sampledStream flushSSE chunk outTTFT = everything above, measured at the clientDecode stepread weights + KVSamplelogits to tokenDetokenize + sendone chunk per steprepeat N-1 times: each loop is one TPOTCompute floor (prefill)2 x params x prompt tokens / FLOP rateBandwidth floor (decode)(weights + batch x KV) / HBM rateE2E = TTFT + (N - 1) x TPOT, and every term has a physical floor
The path of one streamed request. Everything before the first token reaches the client is TTFT; each trip round the decode loop is one TPOT. Prefill has a compute floor, decode a bandwidth floor.
StageTypical scaleWhat sets it
Network and TLS1-50 msdistance, connection reuse, proxies and gateways
Tokenisationunder 1 ms to a few msCPU speed and prompt length; rarely the problem
Queue and admission0 ms to secondsload relative to capacity, KV-cache space, scheduling policy
Prefilltens to hundreds of msprompt tokens times model FLOPs, divided by achieved FLOP rate
Each decode stepa few to tens of msbytes read per step divided by achieved memory bandwidth
Sample, detokenise, flushsub-ms per stepsampling kernels, response buffering in proxies

The two GPU stages can be reasoned about from first principles. The queue is zero in a benchmark and dominant in production.

Advertisement

Prefill is a compute problem

A dense transformer forward pass costs about 2 FLOPs per parameter per token: one multiply and one add for every weight the token touches. Prefill pushes every prompt token through the model at once, so the work is 2 x P x T for P parameters and T prompt tokens, done as large matrix multiplications that keep the tensor cores busy. Attention adds a term growing with T squared that only matters at long context.

For Llama 3 8B and a 2,000-token prompt that is 2 x 8e9 x 2,000 = 3.2e13 FLOPs. An H100 SXM's datasheet peak for dense BF16 is 989 TFLOPS. No real kernel sustains peak, so the model multiplies it by an achieved fraction; 0.55 is an assumption, not a measurement, and you should replace it with yours. The result is about 59 ms of prefill. Double the prompt and prefill doubles. Long retrieval contexts show up first as TTFT, and caching a shared prefix, covered in prefix caching, removes TTFT rather than TPOT.

Decode is a bandwidth problem

Decode produces one token per sequence per step. Each step still has to read every weight from HBM, but it only does 2 FLOPs per weight per sequence in the batch. At batch 1 an 8B model in BF16 reads 16 GB to do 16 GFLOPs of work, one FLOP per byte, while an H100 can do roughly 300 FLOPs in the time it takes to read one byte. The tensor cores sit idle and the step time is simply bytes moved divided by bandwidth.

Bytes per step are the weights plus the KV cache of every sequence in the batch. Llama 3 8B has 32 layers, 8 KV heads of dimension 128 and stores keys and values in 2 bytes, so each token of context costs 2 x 32 x 8 x 128 x 2 = 131,072 bytes, exactly 128 KiB. At 3.35 TB/s and an assumed 80 percent achieved bandwidth, reading 16 GB of weights takes about 6 ms, and a single sequence's 2,000-token KV cache adds about 0.1 ms. That 6 ms is a hard floor on TPOT for this model on this GPU. Only fewer bytes (quantisation, a smaller model), more bandwidth (a newer GPU, tensor parallelism across GPUs) or more tokens per read (batching, speculative decoding) can move it.

A runnable latency model

The two floors fit in twenty lines. The model below adds a fixed per-step overhead for kernel launches, sampling and the scheduler, again an assumption to measure, and computes TTFT, TPOT, E2E and throughput for a 2,000-token prompt and a 300-token answer across batch sizes.

# First-order latency model for one decoder-only model on one GPU.
PEAK_FLOPS = 989e12      # H100 SXM dense BF16 (datasheet)
HBM_BW     = 3.35e12     # H100 SXM HBM3, bytes/s (datasheet)
MFU_PREFILL = 0.55       # ASSUMPTION: fraction of peak your prefill GEMMs reach
BW_EFF      = 0.80       # ASSUMPTION: fraction of HBM bandwidth decode reaches
STEP_OVERHEAD = 0.0015   # ASSUMPTION: s per step (launch, sampling, scheduler)

PARAMS = 8.0e9                          # Llama 3 8B
W_BYTES = PARAMS * 2                    # BF16 weights
KV_PER_TOKEN = 2 * 32 * 8 * 128 * 2     # K+V x 32 layers x 8 KV heads x 128 x 2 bytes

def prefill_s(prompt_tokens):
    return 2 * PARAMS * prompt_tokens / (PEAK_FLOPS * MFU_PREFILL)

def decode_step_s(batch, ctx_tokens):
    bytes_moved = W_BYTES + batch * ctx_tokens * KV_PER_TOKEN
    mem_s = bytes_moved / (HBM_BW * BW_EFF)
    flop_s = 2 * PARAMS * batch / (PEAK_FLOPS * MFU_PREFILL)
    return max(mem_s, flop_s) + STEP_OVERHEAD

def request(prompt=2000, out=300, batch=1, queue_s=0.0, net_s=0.02):
    ttft = net_s + queue_s + prefill_s(prompt) + decode_step_s(batch, prompt)
    tpot = decode_step_s(batch, prompt + out // 2)
    return ttft, tpot, ttft + (out - 1) * tpot

for b in (1, 8, 32, 64, 128):
    t, s, e = request(batch=b)
    print(f"batch {b:>3}: TTFT {t*1e3:6.1f} ms  TPOT {s*1e3:5.2f} ms  "
          f"E2E {e:5.2f} s  tokens/s/GPU {b/s:7.0f}")

Running it prints:

batch   1: TTFT   86.4 ms  TPOT  7.58 ms  E2E  2.35 s  tokens/s/GPU     132
batch   8: TTFT   87.1 ms  TPOT  8.31 ms  E2E  2.57 s  tokens/s/GPU     963
batch  32: TTFT   89.4 ms  TPOT 10.83 ms  E2E  3.33 s  tokens/s/GPU    2953
batch  64: TTFT   92.6 ms  TPOT 14.20 ms  E2E  4.34 s  tokens/s/GPU    4507
batch 128: TTFT   98.8 ms  TPOT 20.93 ms  E2E  6.36 s  tokens/s/GPU    6116

Check memory before believing the last row: 128 sequences at about 2,150 tokens of context is 36.1 GB of KV cache on top of 16 GB of weights, which fits in 80 GB with room for activations. At 8,000-token contexts the same batch would need about 134 GB and could not run; the batch size you can reach is set by KV capacity long before compute. Calibrate the three assumptions by running your engine at batch 1 and batch 64 with a fixed prompt and solving for them.

Reading the batch curve

The output is the central trade-off of LLM serving in five lines. Going from batch 1 to batch 64 multiplies throughput by 34 and only doubles TPOT, because the weights are read once per step whichever batch size you run. Past that the KV-cache term grows until it rivals the weights, and TPOT climbs roughly linearly while throughput gains shrink. Somewhere on this curve is a knee where your TPOT objective is still met and cost per token is near its minimum; it moves with context length.

Notice what batching does not do in this model: TTFT barely moves. That is only true because the model runs one prefill at a time with nothing else competing, which is exactly what production does not look like.

The term the model leaves out: contention

Real servers run continuous batching: new requests join the running batch between steps. A newly admitted prompt's prefill shares the GPU with everyone else's decode. If eight 2,000-token prompts arrive together and the engine prefills them back to back, each waits for the ones ahead of it. Extending the model with that serial queue gives about 87 ms TTFT for the first, 264 ms for the fourth and 499 ms for the eighth; meanwhile every decoding sequence stalls for the half second those prefills occupy, which shows up as one enormous inter-token gap in the middle of somebody else's answer.

Three mechanisms tame it. Chunked prefill caps the prompt tokens per step so decodes keep flowing, trading a slightly longer TTFT for a bounded worst gap. Prefill-decode disaggregation, compared in the disaggregation article, runs the two phases on separate GPU pools and pays for it by moving the KV cache between them. Deadline-aware scheduling, covered in SLO scheduling, decides whose prefill goes first. Beyond those, queueing theory applies: as utilisation approaches 100 percent, waiting time grows without bound, so leave headroom.

Levers, mapped to the term they move

LeverMovesCost or catch
Prefix cachingTTFT (skips cached prefill)only helps repeated prefixes; eats KV memory
Chunked prefillworst inter-token gapslightly higher TTFT and some throughput
Weight quantisation (FP8, INT4)TPOT (fewer bytes per step)accuracy must be re-evaluated
KV-cache quantisationTPOT at long context, batch capacityaccuracy at long context
Tensor parallelismTPOT and prefill (more bandwidth and FLOPs)all-reduce per layer; poor scaling over slow links
Speculative decodingTPOT at low batchgains shrink as batch grows and the GPU becomes busy
Bigger batchthroughputTPOT rises; KV capacity caps it
Lower utilisation targetqueue time, all tailsmore GPUs per request per second

The table also works in reverse as a diagnosis. If TTFT is bad and TPOT is fine, look at queueing, prompt length and prefix reuse. If TPOT is bad at low load, look at bytes per step. If TPOT is fine on average and awful at p99, look for prefill interference.

Measuring latency honestly

Measure from the client, stream the response and record a timestamp per chunk. Server-side histograms cannot see the load balancer, TLS or a proxy that buffers the response. The harness below works against any OpenAI-compatible streaming endpoint.

import json, time, httpx   # any OpenAI-compatible streaming server (vLLM, SGLang)

def timed_request(url, model, prompt, max_tokens=300):
    body = {"model": model, "prompt": prompt, "max_tokens": max_tokens, "stream": True}
    t0 = time.perf_counter()
    stamps = []
    with httpx.stream("POST", f"{url}/v1/completions", json=body, timeout=120) as r:
        for line in r.iter_lines():
            if not line.startswith("data: ") or line == "data: [DONE]":
                continue
            chunk = json.loads(line[6:])
            if chunk["choices"][0].get("text"):
                stamps.append(time.perf_counter())
    ttft = stamps[0] - t0
    gaps = [b - a for a, b in zip(stamps, stamps[1:])]   # inter-chunk latency
    return {"ttft": ttft, "e2e": stamps[-1] - t0,
            "tpot": (stamps[-1] - stamps[0]) / max(len(stamps) - 1, 1),
            "worst_gap": max(gaps, default=0.0)}

Servers may coalesce several tokens into one chunk, so per-chunk gaps are an upper bound on inter-token latency. Drive it with an open-loop arrival process at a fixed rate, not a fixed number of concurrent users, or you will never see the queue; the load-testing article explains why and how. Report p50, p95 and p99 of TTFT, TPOT and worst gap separately, at each request rate, with prompt and output length distributions written next to them.

Failure modes

  • Benchmarking the wrong shape. Measuring 128-token prompts and deploying for 8,000-token retrieval contexts; prefill and KV capacity both scale with length.
  • Buffered streaming. An API gateway or ingress holds SSE chunks, so the server's TTFT is 90 ms and the user's is the full E2E.
  • Cold starts mixed into steady state. Model load takes seconds to minutes; measure warm-up and scale-out separately.
  • KV-cache exhaustion. When the cache is full the engine queues or preempts sequences, and TTFT jumps from milliseconds to seconds with no change in GPU utilisation.
  • Running hot. Utilisation near capacity makes the tail explode; a small traffic spike turns a healthy p99 into timeouts.

What to do next

  1. Write your latency objective as a sentence: percentile, TTFT and TPOT targets, request rate and prompt and output length distribution.
  2. Compute the floors for your model and GPU: 2 x params x prompt tokens over FLOP rate, and weight plus KV bytes over bandwidth.
  3. Run the model above, then calibrate its three assumptions against your engine at batch 1 and batch 64.
  4. Measure from the client with streaming, open-loop load and per-chunk timestamps; confirm nothing between client and server buffers.
  5. Find the knee: the highest request rate at which p95 TTFT and TPOT still meet the objective, and run production at a margin below it.
  6. Pick levers from the table by the term that misses its target, and re-measure after each change rather than stacking them blind.
Key takeaway: LLM latency is three numbers: TTFT, TPOT and E2E = TTFT + (N - 1) x TPOT. Prefill is bounded by FLOPs, about 2 x params x prompt tokens; decode is bounded by bytes, the weights plus every sequence's KV cache per step, which puts a floor of about 6 ms per token under Llama 3 8B on an H100. Batching buys throughput almost for free until the KV term grows, contention and queueing set the tail, and every optimisation moves one specific term. Model the floors, measure from the client under open-loop load, and operate below the knee.