Time to first token, TTFT, is the delay between a request leaving the client and the first generated token arriving back. It is what users feel as responsiveness, and it is the latency number that long prompts, retrieval contexts and busy servers damage most. It is also one you can predict with arithmetic before you buy hardware or tune a scheduler.
GPU inference latency introduces TTFT alongside per-token latency with the rule of thumb that prefill costs 2 FLOPs per parameter per token. This article builds the full TTFT ledger: which parameters actually count, the attention term that grows with the square of the prompt, the memory floor for short prompts, prefix caching, tensor-parallel communication, and the queueing term that dominates at load. We finish with a calculator and a sizing example. Hardware figures are H100 SXM datasheet values; efficiency factors are labelled assumptions to replace with your own measurements.
The TTFT ledger
Write TTFT as a sum: network and TLS time, queue wait, tokenisation, prefill, sampling the first token, and flushing it through the streaming path. Two points are often misunderstood. First, there is no decode step inside TTFT. Prefill computes logits for the final prompt position, and the first output token is sampled from them; decode starts with the second token. Second, the queue term is not a constant. On an idle server it is zero; near saturation it is larger than every other term combined.
Prefill FLOPs, counted carefully
The 2 x N x P rule counts every parameter for every token. A more careful count separates three parts.
- Linear layers. Each prompt token passes through every attention projection and MLP weight: 2 FLOPs per non-embedding parameter per token. The embedding lookup is a gather, not a matmul, so it does not count.
- The output head. To sample the first token you need logits for the last position only, so the vocabulary projection costs 2 x d x V FLOPs once, not once per token. Engines that return logits for every prompt position pay it per token.
- Attention scores. QK-transpose and the weighted sum over V each cost about 2 x d FLOPs per query-key pair, where d is heads times head dimension. A causal mask halves the pairs. For P tokens that gives 2 x L x d x P squared FLOPs across L layers.
For Llama 3.1 8B (32 layers, d = 4096, MLP width 14,336, 8 KV heads of 128, vocabulary 128,256), the non-embedding weights are about 6.98 billion; the embedding table and untied output head add about 0.53 billion each. The attention term equals the linear term when P = N / (L x d), about 53,000 tokens for this model. Below that it is a correction; above it, attention dominates.
Worked numbers for Llama 3.1 8B
Assume an H100 SXM at 989 TFLOPS dense BF16 and an achieved 55 percent in prefill, about 544 TFLOPS. The calculator below gives:
| Prompt tokens | Linear FLOPs | Attention FLOPs | Attention share | Prefill time |
|---|---|---|---|---|
| 128 | 1.79e12 | 4.3e9 | 0.2% | 5.6 ms (memory floor) |
| 2,000 | 2.79e13 | 1.05e12 | 3.6% | 53 ms |
| 8,000 | 1.12e14 | 1.68e13 | 13% | 236 ms |
| 32,000 | 4.47e14 | 2.68e14 | 38% | 1,315 ms |
Two things stand out. Prefill time is roughly linear up to a few thousand tokens and then bends upward as attention grows. And very short prompts do not get faster than about 5.6 ms, which is the next term.
The memory floor for short prompts
Whatever the prompt length, prefill must read every layer weight and the output head from HBM at least once; the embedding table is only gathered row by row. In BF16 that is about 15 GB for the 8B model; at 3.35 TB/s and an assumed 80 percent achieved bandwidth it takes about 5.6 ms. Prefill time is the larger of the compute time and this read time. The crossover is where the two are equal, around 220 tokens for this model and GPU under these assumptions. Below it, a short chat message costs the same prefill time as a slightly longer one, and the only levers are fewer bytes (quantisation) or more bandwidth.
Prefix caching in the ledger
Prefix caching reuses the KV cache for a prompt prefix the server has already processed, such as a shared system prompt or the earlier turns of a conversation; prefix caching covers the mechanism. In the ledger it removes linear work for the cached tokens, but not all attention work: each new token still attends to every cached key. For c cached and s new tokens the attention term becomes 4 x L x d x (s x c + s squared / 2).
For a conversation turn with 7,000 cached tokens and 1,000 new ones, the 8B model's prefill drops from 236 ms to about 33 ms. Most of what remains is the linear term for the 1,000 new tokens, with about a fifth coming from attention to the cached prefix. That is why cache hit rate is often the single biggest TTFT lever in multi-turn products, and why a cache miss, after eviction, shows up as a sudden TTFT spike for one user.
Tensor parallelism: less compute, more communication
Tensor parallelism across T GPUs divides prefill FLOPs and weight bytes by T, but adds communication. In the common layout each layer needs two all-reduces of the activations, each P x d x 2 bytes in BF16. A ring all-reduce sends about 2 x (T - 1) / T times that size from each GPU.
Take Llama 3.1 70B (80 layers, d = 8192, about 68.5 billion non-embedding parameters) on four H100s with an 8,000-token prompt. Total prefill work is about 1.18e15 FLOPs, or 542 ms of compute per GPU at the assumed efficiency. Each all-reduce moves 131 MB of activations, so each GPU sends about 197 MB per all-reduce and 31.5 GB across 160 of them. NVLink on H100 is 900 GB/s bidirectional, about 450 GB/s each way; at an assumed 70 percent that is roughly 100 ms if none of it overlaps compute. Real engines hide part of it, so budget 0 to 100 ms on top of 542 ms and measure. Across nodes with slower links, the same arithmetic is why tensor parallelism stays inside one NVLink domain.
Queueing: the term that dominates under load
On a busy server, the request usually waits before its prefill starts. If you treat prefill as one shared resource serving requests one at a time, which is a simplification that ignores decode work competing for the same GPU, the Pollaczek-Khinchine formula for an M/G/1 queue gives the mean wait: W = lambda x E[S squared] / (2 x (1 - rho)), where lambda is the arrival rate, S the prefill service time and rho = lambda x E[S] the utilisation.
The E[S squared] term is the important one: it is dominated by long prompts. Suppose 80 percent of requests have 2,000-token prompts (53 ms) and 20 percent have 16,000 tokens (534 ms). The mean service time is 149 ms, so one GPU saturates at about 6.7 requests per second. A simulation of this queue gives the wait plus prefill seen by the short requests, before network and flush time:
| Arrivals/s | Utilisation | Mean wait (formula) | Short request wait + prefill, p50 | p95 |
|---|---|---|---|---|
| 2 | 0.30 | 85 ms | 53 ms | 552 ms |
| 4 | 0.60 | 295 ms | 112 ms | 1,236 ms |
| 5 | 0.75 | 586 ms | 417 ms | 2,104 ms |
| 6 | 0.90 | 1,716 ms | 1,173 ms | 5,383 ms |
At 30 percent utilisation the median short request sees only its own 53 ms, yet its p95 is ten times that, because one in five times it lands behind a 534 ms long prompt. The tail is set by the longest prompts, not the average one, and it explodes as utilisation passes about 70 percent. Real servers batch and interleave work, so treat these numbers as shape, not prediction.
Chunked prefill and TTFT
Chunked prefill splits long prompts into pieces that share each scheduler iteration with decode tokens, under a per-iteration token budget; see chunked prefill. In the ledger, it shortens the head-of-line blocking that produced the p95 column above, because a short prompt can start after one chunk instead of after a whole 16,000-token prefill. The cost lands on the long request: with a 2,048-token budget and 64 decode sequences in each iteration, a 16,000-token prompt needs ceil(16,000 / 1,984) = 9 iterations, each carrying decode work and per-iteration overhead, so its own TTFT rises somewhat. That is usually the right trade, because long-prompt users expect a wait and short-prompt users do not.
A TTFT calculator
The ledger fits in a short script. Replace the assumptions with measured values from your engine and hardware.
PEAK, HBM = 989e12, 3.35e12 # H100 SXM datasheet: dense BF16 FLOP/s, HBM bytes/s
MFU, BW_EFF = 0.55, 0.80 # ASSUMPTIONS: measure these on your stack
LLAMA_8B = dict(layers=32, d=4096, ffn=14336, heads=32, kv_heads=8, head_dim=128, vocab=128256)
def non_embedding_params(m):
q_o = 2 * m["d"] * m["heads"] * m["head_dim"]
k_v = 2 * m["d"] * m["kv_heads"] * m["head_dim"]
return m["layers"] * (q_o + k_v + 3 * m["d"] * m["ffn"])
def prefill_seconds(m, new, cached=0, tp=1):
linear = 2 * non_embedding_params(m) * new
attn = 4 * m["layers"] * m["heads"] * m["head_dim"] * (new * cached + new * new / 2)
head = 2 * m["d"] * m["vocab"] # last position only
compute = (linear + attn + head) / tp / (PEAK * MFU)
weight_bytes = 2 * (non_embedding_params(m) + m["d"] * m["vocab"]) / tp # embeddings are only gathered
return max(compute, weight_bytes / (HBM * BW_EFF)) # never faster than one weight read
def mg1_mean_wait(lam, services, probs):
es = sum(p * s for p, s in zip(probs, services))
es2 = sum(p * s * s for p, s in zip(probs, services))
rho = lam * es
return float("inf") if rho >= 1 else lam * es2 / (2 * (1 - rho))
def ttft(m, new, cached=0, tp=1, wait=0.0, network=0.03, tokenize=0.002, flush=0.002):
return network + wait + tokenize + prefill_seconds(m, new, cached, tp) + flush
svc = [prefill_seconds(LLAMA_8B, 2000), prefill_seconds(LLAMA_8B, 16000)]
w = mg1_mean_wait(4.0, svc, [0.8, 0.2])
print(f"mean TTFT for a 2k prompt at 4 req/s: {ttft(LLAMA_8B, 2000, wait=w) * 1e3:.0f} ms")The network, tokenisation and flush defaults are placeholders. Measure them; a cross-region hop or a proxy that buffers server-sent events can exceed the whole prefill.
Worked example: sizing for a TTFT SLO
Suppose the target is TTFT p95 under 800 ms for the 80/20 mix above on one H100 serving the 8B model. The table shows p95 is already 1.2 seconds at 4 requests per second, and the same simulation puts it at about 815 ms at 3 per second and 640 ms at 2.5, so either cap each GPU near 2.5 requests per second or change the mix. Routing the 16,000-token requests to separate replicas makes the short-prompt pool's service time a constant 53 ms; simulated, its wait-plus-prefill p95 is about 200 ms at 11 requests per second (59 percent utilisation) and about 410 ms at 15 (80 percent). Chunked prefill gets part of the same benefit without separate pools. Prefix caching shrinks the long prompts themselves when they share context. Run the numbers for each option before buying more GPUs.
Measuring TTFT so it matches the math
Measure TTFT at the client, from sending the request to receiving the first chunk that contains generated text. Some OpenAI-compatible servers send an initial chunk carrying only the role, so do not stop the clock on it. Record prompt length and cached-token count with every sample, because TTFT without prompt length cannot be compared. Report p50, p95 and p99 by prompt-length bucket, and load test at several arrival rates, as LLM load testing describes, to find the utilisation where p95 bends. Compare server-side timestamps with client-side ones to separate queue and prefill from network and proxy buffering.
Failure modes
- Averages hide the tail. Mean TTFT looks fine while p95 is dominated by short requests stuck behind long prompts.
- Prefix-cache eviction. KV memory pressure evicts cached prefixes and TTFT jumps for returning users; see KV cache sizing.
- Buffered streaming. A proxy or load balancer buffers SSE, adding seconds no GPU metric shows.
- Benchmarks with one prompt length. A fixed 512-token benchmark says nothing about a production mix with a long tail.
What to do next
- Collect your production prompt-length distribution and cache hit rate.
- Measure achieved prefill FLOP/s and bandwidth on your engine at two prompt lengths; replace MFU and BW_EFF.
- Compute the ledger for your model, including attention and the memory floor.
- Estimate queueing with the M/G/1 formula, then confirm with a load test at three arrival rates.
- Set a utilisation ceiling per replica from the rate where TTFT p95 crosses your SLO.
- If long prompts dominate the tail, evaluate chunked prefill, a separate long-prompt pool and better prefix reuse, and check KV headroom with KV cache sizing.
- Instrument client-side TTFT by prompt-length bucket and alert on p95.