Tail latency is the latency of your slowest requests: the 99th or 99.9th percentile rather than the median. In LLM serving on GPUs, the tail is routinely ten or more times the median, and it is what users remember: the chat answer that sat on a spinner for eight seconds, the voice agent that paused mid-sentence. For agent workloads it matters even more, because one user task fans out into many model calls and the task is as slow as its slowest call.
This article explains where tail latency comes from in GPU inference, how to measure it without fooling yourself, and which levers move it, with the cost of each. It assumes you know the basics of prefill, decode and continuous batching; if not, read the GPU inference latency article first, which builds the median model this one extends to the tail.
Why the tail dominates
Start with why the tail dominates. If each call independently has a 1 percent chance of exceeding its p99, then a task that makes n calls has a probability of 1 - 0.99n that at least one is slow. With 10 calls that is about 9.6 percent; with 30 calls, about 26 percent; with 100 calls, about 63 percent. An agent that makes 70 or more model calls per task hits at least one p99-slow call in most tasks, so your p99 becomes its typical case. The independence assumption is optimistic, too: tail events cluster in time (a burst or a hot replica slows everything for a while), so correlated slow calls stretch end-to-end time further.
The second reason is queueing. For a single server with random arrivals and exponential service times (the M/M/1 model), mean time in system is S / (1 - rho), where S is the mean service time and rho the utilisation, and the 99th percentile is about 4.6 times that. At 50 percent utilisation p99 is roughly 9 service times; at 90 percent, roughly 46. A batching GPU server is not an M/M/1 queue, since batching raises capacity as load grows, but the shape holds: as you approach saturation, the tail grows much faster than the median. Headroom is a tail latency control, not waste.
Where slow requests come from
On a serving replica, tail events come from a short list of mechanisms. Each leaves a different signature, which is what makes attribution possible.
| Cause | Metric hit | Signature | Main lever |
|---|---|---|---|
| Queueing under bursts | TTFT | Queue time dominates; correlates with arrival rate | Headroom, admission control, autoscaling |
| Long prefill in the same step | Inter-token gaps of running requests, TTFT of others | Gap spikes coincide with a large prompt being admitted | Chunked prefill, prefill/decode separation |
| KV cache exhaustion | End-to-end time, gaps | Preemption counters rise; requests recompute or swap | Memory headroom, output caps, admission on KV |
| Output-length variance | End-to-end time | Slow requests simply generated more tokens | max_tokens policy, per-token SLOs |
| Context-length variance | TPOT | Attention cost grows with context; long-context batches step slower | Length-aware routing, separate pools |
| Cold paths | TTFT | First request after scale-up, model load, or adapter swap | Warm pools, preloading, pinned adapters |
| Host-side stalls | Gaps | Periodic spikes unrelated to load; CPU saturation | Isolate CPU cores, move tokenisation off the hot loop |
| Multi-GPU stragglers | All metrics on a replica | One rank consistently slower in tensor-parallel steps | Find and drain the slow GPU or link |
| Clock reduction under power or heat | TPOT drift | Step time rises with temperature or power draw | Cooling, consistent power caps, placement |
Prefill interference in detail
The second row deserves a closer look because it is specific to LLM serving and is often the largest single contributor to inter-token tail. With continuous batching, each engine step processes decode tokens for all running requests plus prefill work for newly admitted ones. If a 30,000-token prompt is admitted whole, that step takes as long as the prefill, and every running request waits for it: their next token arrives hundreds of milliseconds late. The user streaming an answer sees a freeze that has nothing to do with their own request.
Chunked prefill bounds the number of prefill tokens processed per step, spreading a long prompt across several steps alongside decode work. That caps the step time and therefore the inter-token gap, at the cost of a somewhat longer time to first token for the long prompt itself and some throughput lost to smaller prefill kernels. The chunked prefill article covers how to size the per-step token budget against an inter-token SLO. Disaggregating prefill and decode onto separate GPUs removes the interference entirely, at the price of moving the KV cache between them.
Measuring the tail without fooling yourself
Most tail latency numbers are wrong in a flattering direction, for three reasons. First, coordinated omission: a closed-loop load generator that waits for a response before sending the next request slows down exactly when the server does, so it never sends the requests that would have queued, and the tail vanishes from the measurement. Use open-loop arrivals at a fixed rate and measure from the scheduled send time. Second, client-side queueing: HTTP client pools cap concurrent connections, and requests waiting for a connection are invisible to the server. Third, averaging percentiles: the mean of per-minute p99 values is not the p99 of the hour. Keep histograms and merge them.
The generator below sends Poisson arrivals to an OpenAI-compatible streaming endpoint and records time to first chunk, the largest gap between chunks, and end-to-end time, all measured from the scheduled time. Note that a streamed chunk is not always exactly one token; check how your server batches the stream.
import asyncio, random, time, aiohttp
async def one(session, url, body, scheduled, out):
first = last = None
max_gap = 0.0
async with session.post(url, json=body) as resp:
async for line in resp.content:
if not line.startswith(b"data: ") or line.strip() == b"data: [DONE]":
continue
now = time.perf_counter()
if first is None:
first = now
else:
max_gap = max(max_gap, now - last)
last = now
end = time.perf_counter()
out.append({"ttft": (first or end) - scheduled,
"max_gap": max_gap,
"e2e": end - scheduled})
async def open_loop(url, model, prompts, rate, seconds):
out, tasks = [], []
conn = aiohttp.TCPConnector(limit=0) # no client-side connection cap
async with aiohttp.ClientSession(connector=conn) as s:
start, t = time.perf_counter(), 0.0
while t < seconds:
t += random.expovariate(rate) # Poisson arrivals
scheduled = start + t
await asyncio.sleep(max(0.0, scheduled - time.perf_counter()))
prompt, max_tokens = random.choice(prompts)
body = {"model": model, "prompt": prompt,
"max_tokens": max_tokens, "stream": True}
tasks.append(asyncio.create_task(one(s, url, body, scheduled, out)))
await asyncio.gather(*tasks)
return outDraw prompts from a sample of real traffic, including its long tail of prompt and output lengths. A benchmark with uniform 512-token prompts has no long prefills and will show a tail that production never sees.
Attributing the tail to causes
Measurement tells you the tail is bad; attribution tells you why. Have the server log, per request, the time spent queued, the prefill time, the number of preemptions, the largest prefill co-scheduled in any step while the request was running, and whether a cold path was hit. Then compare how often each condition holds among tail requests versus all requests. A condition with high lift is a suspect; one with equal rates is not.
def attribute(rows, metric="ttft_ms", q=0.99):
cut = sorted(r[metric] for r in rows)[int(q * (len(rows) - 1))]
tail = [r for r in rows if r[metric] >= cut]
conditions = {
"queued over 1 s": lambda r: r["queue_ms"] > 1000,
"preempted": lambda r: r["preemptions"] > 0,
"big co-scheduled prefill": lambda r: r["max_cosched_prefill"] > 8192,
"cold adapter": lambda r: r["adapter_cold"],
}
for name, cond in conditions.items():
in_tail = sum(map(cond, tail)) / len(tail)
overall = sum(map(cond, rows)) / len(rows)
print(f"{name:26s} tail {in_tail:6.1%} all {overall:6.1%} "
f"lift {in_tail / max(overall, 1e-9):5.1f}x")
Levers that move the tail
Once you know the cause, pick the lever that addresses it. The important ones:
- Headroom and admission control. Run replicas below the knee of the latency curve and reject or defer work above it, preferably on a signal that predicts trouble, such as queue depth or free KV blocks, rather than CPU. Shedding a few requests early beats serving all of them late. See admission control for LLM serving.
- Bound per-step work. Chunked prefill or prefill/decode separation, as above.
- Keep KV memory headroom. Preemption means a request loses its cache and either recomputes it or waits for a swap. Reserve memory, cap
max_tokensto what the product needs, and admit on projected KV use rather than request count. - Route by length. Send long-context requests to a separate pool so their slow attention steps do not set the pace for short chats.
- Prioritise. Interactive traffic ahead of batch or background work, with preemption of the low class rather than the high one.
- Hedge sparingly. Hedging, sending a duplicate request to a second replica when the first has not responded in time, is a classic tail cure for cheap reads. For LLM calls the duplicate costs real GPU time, so hedge only on the first token, only after a delay near the p95 time to first token, cap hedges at a few percent of traffic, and cancel the loser.
- Remove cold paths. Keep a warm pool ahead of autoscaling, preload weights and frequent adapters, and send synthetic warm-up requests before a replica takes traffic.
Hedging on the first token
A hedge on first token, in asyncio. call(replica) is assumed to return once the first token has arrived, handing back the open stream. Cancelling the losing task must close its connection; most OpenAI-compatible servers abort a request when the client disconnects, but verify that yours does, or the hedge doubles load instead of trimming tail.
import asyncio
async def hedged_first_token(call, primary, backup, hedge_after_s, budget):
first = asyncio.create_task(call(primary))
done, _ = await asyncio.wait({first}, timeout=hedge_after_s)
if done or not budget.try_spend(): # budget caps hedge rate
return await first
second = asyncio.create_task(call(backup))
done, pending = await asyncio.wait({first, second},
return_when=asyncio.FIRST_COMPLETED)
for t in pending:
t.cancel() # closes the slower stream
return done.pop().result()
Worked example
Take an illustrative chat service on eight replicas with a median time to first token of 350 ms and a p99 of 4.1 s, and a p99 inter-token gap of 900 ms against a 150 ms target. The open-loop benchmark, replaying a day of real prompt lengths at peak rate, reproduces both numbers, which already rules out a measurement artefact.
Attribution on the inter-token tail shows a co-scheduled prefill above 8,192 tokens in 71 percent of tail requests versus 6 percent overall, a lift of about 12. The time-to-first-token tail shows queueing over one second in most tail requests, and those cluster in the minutes after traffic bursts, while preemption appears in under 2 percent of either. The fixes follow the attribution: enable chunked prefill with a per-step budget sized from the inter-token target, route prompts above 16,000 tokens to two dedicated replicas, and add a ninth replica plus queue-depth admission so peak utilisation drops below the knee. Re-running the same benchmark is the only acceptable evidence that the changes worked, and the median should be checked too, since chunking and routing both cost some throughput.
Failure modes
- Optimising the median. Larger batches raise throughput and often the median looks fine while the tail grows; always report both.
- Hiding the tail with timeouts. Requests that time out drop out of latency histograms. Count them as infinitely slow, or track them as a separate SLI.
- Hedging without cancellation. Doubles load at the worst moment and makes the tail worse.
- Per-replica blind spots. A single bad GPU or link raises the fleet p99 while fleet averages look healthy. Break latency down by replica and by rank.
- Benchmarks without long prompts. They cannot reproduce interference or preemption.
Trade-offs
| Lever | Tail benefit | Cost |
|---|---|---|
| More headroom | Large; addresses queueing directly | GPU cost per request |
| Chunked prefill | Caps inter-token gaps | Slower first token for long prompts, some throughput |
| Prefill/decode separation | Removes interference | KV transfer, more complex operations |
| Length-aware pools | Isolates long-context cost | Fragmented capacity, lower utilisation |
| Hedging | Trims first-token tail | Extra GPU work; needs cancellation |
| Admission control | Protects accepted requests | Some requests rejected or delayed |
What to do next
- Build an open-loop benchmark from real prompt and output length distributions, and measure from scheduled send time with no client connection cap.
- Log per-request queue time, prefill time, preemptions, largest co-scheduled prefill and cold-path flags on the server.
- Run tail attribution and rank causes by lift before changing anything.
- Set SLOs on p99 time to first token and p99 inter-token gap, and alert on burn rate as in LLM SLO burn rate alerts.
- Apply the lever that matches the top cause, then re-run the same benchmark and check both the tail and the median.
- Review per-replica and per-rank latency weekly to catch stragglers that fleet averages hide.