Inter-token latency (ITL) is the time between consecutive tokens of a streamed response. Users feel it as smoothness. A reply that streams at a steady 30 ms per token reads well. The same average with occasional half-second freezes feels broken. Time to first token gets most of the attention, and it is covered in the TTFT math article. ITL is what decides whether the rest of the response is pleasant to read and whether an agent loop waiting on a long completion finishes on time.

The cost of one decode step, which sets the floor under ITL, is derived in the decode math article. This article is about everything above that floor. It treats ITL as a distribution, separates it from the averaged metric TPOT, traces the tail to its sources (prefill interference, preemption, speculative bursts, host overhead and network buffering), and shows how to measure, aggregate and budget it. The numbers are computed from a simple model with the assumptions stated, so you can rerun it for your own hardware.

ITL, TPOT and TBT

For a request that streams n tokens at times t_1 to t_n, the ITL samples are the n − 1 gaps t_k - t_(k-1). Two summary metrics are often confused with them.

  • TPOT, time per output token, is (t_n - t_1) / (n - 1), which is the mean of that request's gaps. Some tools compute it from the end-to-end time minus TTFT, which is the same thing.
  • ITL percentiles are computed over the gap list itself, either per request or pooled over all gaps from all requests. Some papers and serving schedulers call a per-gap deadline TBT, time between tokens.

A worked example shows why the difference matters. A request streams 201 tokens, so it has 200 gaps. 190 of them are 10 ms and 10 are 60 ms stalls. TPOT is (190 × 10 + 10 × 60) / 200 = 12.5 ms, which looks healthy. The median gap is 10 ms and the p90 is 10 ms, but the p99 is 60 ms. A user saw ten visible hitches that TPOT hides completely. Any SLO about smoothness must be written on the gap distribution, not on TPOT.

One streamed request: inter-token gaps are steps of a shared batch, not of this requesttimequeue + prefillTTFTprefill chunkpreempt / swapsteady gaps: one decode stepstalllong stallsteady againServer viewper-step time, per-token histogramsClient viewper-chunk arrival gaps, after proxiesnetworkTPOT = (end-to-end − TTFT) / (tokens − 1) averages the stalls away; the gap list keeps them.
Gaps on one request's stream. Most gaps equal one decode step of the shared batch. Stalls come from work done for other requests, or from the scheduler evicting this one.

The floor: one step of a shared batch

Every token a request receives is produced by one forward step of the whole running batch. So the smallest possible ITL is the step time, and the step time is shared: your gap depends on how many other sequences are in the batch and how long their contexts are. For Llama 3 8B in BF16 on an H100 SXM (3.35 TB/s of HBM3), each step reads about 16.06 GB of weights plus 128 KiB of KV cache per cached token per sequence. Assuming 80% of peak bandwidth is achieved, the memory-bound step time is:

BatchContext per sequenceBytes per stepStep time (model)
12,04816.3 GB6.1 ms
162,04820.4 GB7.6 ms
642,04833.2 GB12.4 ms
648,19284.8 GB31.6 ms

Two lessons follow. First, ITL rises with batch size and context even when nothing goes wrong, so the floor is a property of the load, not of the GPU. Second, at long context the KV term dominates, so memory-saving changes such as KV quantisation or grouped-query attention lower ITL directly. The model ignores kernel launch gaps and sampling, which typically add a little more. Measure the real step time from server logs before trusting any floor.

Where the tail comes from

The tail of the ITL distribution comes from steps that do more than decode, or from gaps that are not steps at all.

  • Prefill interference. With continuous batching, a new request's prompt is processed inside a step that also decodes everyone else, so that step takes longer for every running sequence. Prefill is compute bound at about 2 × parameters FLOPs per prompt token, ignoring attention. At 60% of the H100's 989 dense BF16 TFLOP/s, a 512-token chunk adds about 13.9 ms, a 2,048-token chunk about 55 ms, and an unchunked 8,192-token prompt about 222 ms, before attention cost. See continuous batching for the scheduler.
  • Preemption. When the KV cache fills, schedulers evict sequences and later recompute or swap them back. The evicted request sees a gap that covers the eviction, the wait and the rebuild, often hundreds of milliseconds or more.
  • Speculative decoding bursts. Speculation emits several accepted tokens per step. The client sees near-zero gaps within a burst and a longer gap between bursts. Mean ITL improves, but the gap distribution becomes bimodal; see speculative decoding.
  • Host overhead. Scheduling, sampling, detokenisation and Python garbage collection run on the CPU between steps. In a busy server they show up as occasional extra milliseconds on every sequence at once.
  • Network and proxies. Reverse proxies that buffer responses, compression and load balancers can batch several server-sent events into one delivery. The client then sees bursts and gaps that do not exist on the server.

The p99 cliff

Here is the most useful result of this article. Run the step model at batch 64 and 2k context, where the base step is 12.4 ms. Let a fraction f of steps also carry a prefill chunk, and read the percentiles over 200,000 simulated steps.

Stall fraction fChunkp50p99Mean (≈ TPOT)
0.5%2,04812.4 ms12.4 ms12.7 ms
2%51212.4 ms26.3 ms12.7 ms
2%2,04812.4 ms67.8 ms13.5 ms
10%2,04812.4 ms67.8 ms18.0 ms
30%51212.4 ms26.3 ms16.6 ms

The p99 is a step function. While fewer than 1% of steps stall, it equals the base step. Once more than 1% stall, it equals the base step plus the full stall length, and adding more stalls barely moves it. So two different levers act on two different metrics. The chunk size sets the height of the p99, and this is what chunked prefill is for: cutting the chunk from 2,048 to 512 tokens drops it from 67.8 ms to 26.3 ms here. The stall frequency moves the mean and TPOT, and decides whether you are on the low or the high side of the cliff. Smaller chunks cost more steps per prompt, and so they lengthen TTFT. That trade-off between ITL and TTFT is the central tuning decision for a mixed-traffic server.

def step_ms(batch, ctx, chunk_tokens=0,
            weights=16.06e9, kv_per_tok=131072, bw=3.35e12, bw_eff=0.8,
            params=8.03e9, peak=989e12, mfu=0.6):
    # memory-bound decode part: weights once + KV for every cached token
    mem = (weights + batch * ctx * kv_per_tok) / (bw * bw_eff)
    # compute-bound prefill chunk riding in the same step (attention ignored)
    pre = 2 * params * chunk_tokens / (peak * mfu)
    return 1e3 * (mem + pre)        # additive: a pessimistic but simple overlap model

print(step_ms(64, 2048), step_ms(64, 2048, 512), step_ms(64, 2048, 2048))
# about 12.4, 26.3, 67.8

Measuring ITL at the client

Measure ITL at the client, because that is what users see, and log server step times so you can explain it. The harness below streams from an OpenAI-compatible endpoint and records a timestamp per server-sent event. Each event may carry more than one token, for example under speculative decoding or proxy batching, so it records the token count per chunk too.

import json, time, requests

def stream_gaps(url, payload, headers=None):
    """Return (ttft_s, gaps_s, tokens_per_chunk) for one streamed completion."""
    t0 = time.perf_counter()
    stamps, ntoks = [], []
    with requests.post(url, json={**payload, "stream": True},
                       headers=headers, stream=True, timeout=300) as r:
        r.raise_for_status()
        for line in r.iter_lines():
            if not line.startswith(b"data: ") or line == b"data: [DONE]":
                continue
            ev = json.loads(line[6:])
            text = ev["choices"][0].get("text") or \
                   (ev["choices"][0].get("delta") or {}).get("content") or ""
            if not text:
                continue
            stamps.append(time.perf_counter())
            ntoks.append(max(1, len(text.split())))   # rough; prefer a tokenizer
    gaps = [b - a for a, b in zip(stamps, stamps[1:])]
    return stamps[0] - t0, gaps, ntoks

Three rules keep the numbers honest. Count tokens with the model's tokenizer, not by splitting on spaces as the sketch does, and record per-chunk gaps separately from per-token estimates. Run the client close to the server and also from a real user location, so you can separate network effects from serving effects. Drive realistic arrival rates and prompt-length mixes, because ITL depends on what else is in the batch. A single-request benchmark measures the floor and nothing else.

Aggregating gaps without fooling yourself

There are two legitimate ways to aggregate gaps, and they answer different questions. Pool every gap from every request and take percentiles, and you get the experience per token. Compute a percentile per request and then look at the distribution of those values, and you get the experience per user.

Take the earlier request, with 200 gaps including ten 60 ms stalls, alongside a second request with 1,000 smooth 10 ms gaps. The pooled p99 is 10 ms, because only 0.83% of all gaps are stalls. Averaging each request's p99 gives 35 ms. Neither is wrong. But one user had a visibly hitchy response, and the pooled metric cannot see it because long smooth responses dilute it. Never average percentiles, since the mean of p99s is not a p99 of anything. For SLOs, the clearest form is per request: the share of requests whose worst gap, or whose p99 gap, stays under a threshold.

Designing an ITL SLO

Write the ITL objective as a per-request condition with a target share, for example: 99% of requests have every gap under 100 ms, and 95% have a p95 gap under 40 ms. Then map each lever to the term it moves.

LeverMovesCost
Smaller prefill chunk or token budgetheight of the stall taillonger TTFT, more steps per prompt
Cap batch size or KV tokensbase step timelower throughput per GPU
Disaggregated prefill and decoderemoves prefill stalls from decode GPUsKV transfer, more complex fleet
More KV memory or KV quantisationpreemption frequency, base step at long contexthardware or small quality risk
Speculative decodingmean ITL down, gaps burstydraft compute, worse at high batch
Turn off proxy buffering for streamsclient-side burstsnone for SSE routes

Budget the SLO from the floor upward. Take the base step at your target batch and context, add the chunk stall, and check the sum against the per-gap threshold. If the sum is too high, reduce the chunk size before you cut the batch, since that is usually cheaper. If the headroom is negative even with no stall, the batch or context target itself is incompatible with the SLO.

Failure modes

  • Reporting TPOT as ITL. A mean hides the stalls users notice. Report gap percentiles.
  • Pooled percentiles only. Long responses dilute stalls; add a per-request metric.
  • Unchunked long prompts. One 8k-token prompt can freeze every stream for a fifth of a second.
  • Preemption storms. Running the KV cache near full turns occasional evictions into cascades of long gaps. Alert on preemption counts.
  • Proxy buffering. A buffering proxy can turn smooth server output into bursts; disable buffering for streaming routes and verify at the client.
  • Benchmarking at batch 1. It measures the floor, not production ITL.

What to do next

  1. Compute your base step time from weights, KV bytes and bandwidth, then compare it with the server's measured step time.
  2. Instrument a client with per-chunk timestamps and tokenizer counts, and record gaps, not just TPOT.
  3. Report both pooled gap percentiles and the per-request share meeting a worst-gap threshold.
  4. Replay a production-like mix and sweep the prefill chunk size; plot p99 gap against TTFT.
  5. Alert on preemptions and on the fraction of steps that carry prefill, which tells you which side of the p99 cliff you are on.
  6. Disable response buffering on every proxy in the streaming path and re-measure from a user location.
Key takeaway: ITL is a distribution of gaps, and its mean, TPOT, hides the stalls users feel. The floor is the decode step of a shared batch, set by weight and KV bytes over bandwidth. The tail comes from prefill chunks, preemption, speculation bursts, host overhead and proxies. The p99 jumps by the full stall length once more than 1% of steps stall, so chunk size sets its height. Measure at the client, never average percentiles, and write SLOs per request.