Tail latency is the latency of your slowest requests: the 99th or 99.9th percentile rather than the median. In LLM serving on GPUs, the tail is routinely ten or more times the median, and it is what users remember: the chat answer that sat on a spinner for eight seconds, the voice agent that paused mid-sentence. For agent workloads it matters even more, because one user task fans out into many model calls and the task is as slow as its slowest call.

This article explains where tail latency comes from in GPU inference, how to measure it without fooling yourself, and which levers move it, with the cost of each. It assumes you know the basics of prefill, decode and continuous batching; if not, read the GPU inference latency article first, which builds the median model this one extends to the tail.

Why the tail dominates

Start with why the tail dominates. If each call independently has a 1 percent chance of exceeding its p99, then a task that makes n calls has a probability of 1 - 0.99n that at least one is slow. With 10 calls that is about 9.6 percent; with 30 calls, about 26 percent; with 100 calls, about 63 percent. An agent that makes 70 or more model calls per task hits at least one p99-slow call in most tasks, so your p99 becomes its typical case. The independence assumption is optimistic, too: tail events cluster in time (a burst or a hot replica slows everything for a while), so correlated slow calls stretch end-to-end time further.

The second reason is queueing. For a single server with random arrivals and exponential service times (the M/M/1 model), mean time in system is S / (1 - rho), where S is the mean service time and rho the utilisation, and the 99th percentile is about 4.6 times that. At 50 percent utilisation p99 is roughly 9 service times; at 90 percent, roughly 46. A batching GPU server is not an M/M/1 queue, since batching raises capacity as load grows, but the shape holds: as you approach saturation, the tail grows much faster than the median. Headroom is a tail latency control, not waste.

Where slow requests come from

On a serving replica, tail events come from a short list of mechanisms. Each leaves a different signature, which is what makes attribution possible.

CauseMetric hitSignatureMain lever
Queueing under burstsTTFTQueue time dominates; correlates with arrival rateHeadroom, admission control, autoscaling
Long prefill in the same stepInter-token gaps of running requests, TTFT of othersGap spikes coincide with a large prompt being admittedChunked prefill, prefill/decode separation
KV cache exhaustionEnd-to-end time, gapsPreemption counters rise; requests recompute or swapMemory headroom, output caps, admission on KV
Output-length varianceEnd-to-end timeSlow requests simply generated more tokensmax_tokens policy, per-token SLOs
Context-length varianceTPOTAttention cost grows with context; long-context batches step slowerLength-aware routing, separate pools
Cold pathsTTFTFirst request after scale-up, model load, or adapter swapWarm pools, preloading, pinned adapters
Host-side stallsGapsPeriodic spikes unrelated to load; CPU saturationIsolate CPU cores, move tokenisation off the hot loop
Multi-GPU stragglersAll metrics on a replicaOne rank consistently slower in tensor-parallel stepsFind and drain the slow GPU or link
Clock reduction under power or heatTPOT driftStep time rises with temperature or power drawCooling, consistent power caps, placement
Where a slow request's time goes on one serving replicaRequest A (typical)qprefilldecode: steady token gapsRequest B (p99)queue: replica busyprefilldecodestallevictrecompute + decodeLong prompt arriveslarge prefill stepKV cache fullB preemptedArrival burstutilisation near 1Typical requests see the service time; tail requests see queueing, interference and preemption stacked together.
A median request pays roughly the service time. A tail request pays queueing, interference from another request's prefill and a preemption, stacked.

Prefill interference in detail

The second row deserves a closer look because it is specific to LLM serving and is often the largest single contributor to inter-token tail. With continuous batching, each engine step processes decode tokens for all running requests plus prefill work for newly admitted ones. If a 30,000-token prompt is admitted whole, that step takes as long as the prefill, and every running request waits for it: their next token arrives hundreds of milliseconds late. The user streaming an answer sees a freeze that has nothing to do with their own request.

Chunked prefill bounds the number of prefill tokens processed per step, spreading a long prompt across several steps alongside decode work. That caps the step time and therefore the inter-token gap, at the cost of a somewhat longer time to first token for the long prompt itself and some throughput lost to smaller prefill kernels. The chunked prefill article covers how to size the per-step token budget against an inter-token SLO. Disaggregating prefill and decode onto separate GPUs removes the interference entirely, at the price of moving the KV cache between them.

Measuring the tail without fooling yourself

Most tail latency numbers are wrong in a flattering direction, for three reasons. First, coordinated omission: a closed-loop load generator that waits for a response before sending the next request slows down exactly when the server does, so it never sends the requests that would have queued, and the tail vanishes from the measurement. Use open-loop arrivals at a fixed rate and measure from the scheduled send time. Second, client-side queueing: HTTP client pools cap concurrent connections, and requests waiting for a connection are invisible to the server. Third, averaging percentiles: the mean of per-minute p99 values is not the p99 of the hour. Keep histograms and merge them.

The generator below sends Poisson arrivals to an OpenAI-compatible streaming endpoint and records time to first chunk, the largest gap between chunks, and end-to-end time, all measured from the scheduled time. Note that a streamed chunk is not always exactly one token; check how your server batches the stream.

import asyncio, random, time, aiohttp

async def one(session, url, body, scheduled, out):
    first = last = None
    max_gap = 0.0
    async with session.post(url, json=body) as resp:
        async for line in resp.content:
            if not line.startswith(b"data: ") or line.strip() == b"data: [DONE]":
                continue
            now = time.perf_counter()
            if first is None:
                first = now
            else:
                max_gap = max(max_gap, now - last)
            last = now
    end = time.perf_counter()
    out.append({"ttft": (first or end) - scheduled,
                "max_gap": max_gap,
                "e2e": end - scheduled})

async def open_loop(url, model, prompts, rate, seconds):
    out, tasks = [], []
    conn = aiohttp.TCPConnector(limit=0)        # no client-side connection cap
    async with aiohttp.ClientSession(connector=conn) as s:
        start, t = time.perf_counter(), 0.0
        while t < seconds:
            t += random.expovariate(rate)        # Poisson arrivals
            scheduled = start + t
            await asyncio.sleep(max(0.0, scheduled - time.perf_counter()))
            prompt, max_tokens = random.choice(prompts)
            body = {"model": model, "prompt": prompt,
                    "max_tokens": max_tokens, "stream": True}
            tasks.append(asyncio.create_task(one(s, url, body, scheduled, out)))
        await asyncio.gather(*tasks)
    return out

Draw prompts from a sample of real traffic, including its long tail of prompt and output lengths. A benchmark with uniform 512-token prompts has no long prefills and will show a tail that production never sees.

Attributing the tail to causes

Measurement tells you the tail is bad; attribution tells you why. Have the server log, per request, the time spent queued, the prefill time, the number of preemptions, the largest prefill co-scheduled in any step while the request was running, and whether a cold path was hit. Then compare how often each condition holds among tail requests versus all requests. A condition with high lift is a suspect; one with equal rates is not.

def attribute(rows, metric="ttft_ms", q=0.99):
    cut = sorted(r[metric] for r in rows)[int(q * (len(rows) - 1))]
    tail = [r for r in rows if r[metric] >= cut]
    conditions = {
        "queued over 1 s":        lambda r: r["queue_ms"] > 1000,
        "preempted":              lambda r: r["preemptions"] > 0,
        "big co-scheduled prefill": lambda r: r["max_cosched_prefill"] > 8192,
        "cold adapter":           lambda r: r["adapter_cold"],
    }
    for name, cond in conditions.items():
        in_tail = sum(map(cond, tail)) / len(tail)
        overall = sum(map(cond, rows)) / len(rows)
        print(f"{name:26s} tail {in_tail:6.1%}  all {overall:6.1%}  "
              f"lift {in_tail / max(overall, 1e-9):5.1f}x")

Levers that move the tail

Once you know the cause, pick the lever that addresses it. The important ones:

  • Headroom and admission control. Run replicas below the knee of the latency curve and reject or defer work above it, preferably on a signal that predicts trouble, such as queue depth or free KV blocks, rather than CPU. Shedding a few requests early beats serving all of them late. See admission control for LLM serving.
  • Bound per-step work. Chunked prefill or prefill/decode separation, as above.
  • Keep KV memory headroom. Preemption means a request loses its cache and either recomputes it or waits for a swap. Reserve memory, cap max_tokens to what the product needs, and admit on projected KV use rather than request count.
  • Route by length. Send long-context requests to a separate pool so their slow attention steps do not set the pace for short chats.
  • Prioritise. Interactive traffic ahead of batch or background work, with preemption of the low class rather than the high one.
  • Hedge sparingly. Hedging, sending a duplicate request to a second replica when the first has not responded in time, is a classic tail cure for cheap reads. For LLM calls the duplicate costs real GPU time, so hedge only on the first token, only after a delay near the p95 time to first token, cap hedges at a few percent of traffic, and cancel the loser.
  • Remove cold paths. Keep a warm pool ahead of autoscaling, preload weights and frequent adapters, and send synthetic warm-up requests before a replica takes traffic.

Hedging on the first token

A hedge on first token, in asyncio. call(replica) is assumed to return once the first token has arrived, handing back the open stream. Cancelling the losing task must close its connection; most OpenAI-compatible servers abort a request when the client disconnects, but verify that yours does, or the hedge doubles load instead of trimming tail.

import asyncio

async def hedged_first_token(call, primary, backup, hedge_after_s, budget):
    first = asyncio.create_task(call(primary))
    done, _ = await asyncio.wait({first}, timeout=hedge_after_s)
    if done or not budget.try_spend():          # budget caps hedge rate
        return await first
    second = asyncio.create_task(call(backup))
    done, pending = await asyncio.wait({first, second},
                                       return_when=asyncio.FIRST_COMPLETED)
    for t in pending:
        t.cancel()                               # closes the slower stream
    return done.pop().result()

Worked example

Take an illustrative chat service on eight replicas with a median time to first token of 350 ms and a p99 of 4.1 s, and a p99 inter-token gap of 900 ms against a 150 ms target. The open-loop benchmark, replaying a day of real prompt lengths at peak rate, reproduces both numbers, which already rules out a measurement artefact.

Attribution on the inter-token tail shows a co-scheduled prefill above 8,192 tokens in 71 percent of tail requests versus 6 percent overall, a lift of about 12. The time-to-first-token tail shows queueing over one second in most tail requests, and those cluster in the minutes after traffic bursts, while preemption appears in under 2 percent of either. The fixes follow the attribution: enable chunked prefill with a per-step budget sized from the inter-token target, route prompts above 16,000 tokens to two dedicated replicas, and add a ninth replica plus queue-depth admission so peak utilisation drops below the knee. Re-running the same benchmark is the only acceptable evidence that the changes worked, and the median should be checked too, since chunking and routing both cost some throughput.

Failure modes

  • Optimising the median. Larger batches raise throughput and often the median looks fine while the tail grows; always report both.
  • Hiding the tail with timeouts. Requests that time out drop out of latency histograms. Count them as infinitely slow, or track them as a separate SLI.
  • Hedging without cancellation. Doubles load at the worst moment and makes the tail worse.
  • Per-replica blind spots. A single bad GPU or link raises the fleet p99 while fleet averages look healthy. Break latency down by replica and by rank.
  • Benchmarks without long prompts. They cannot reproduce interference or preemption.

Trade-offs

LeverTail benefitCost
More headroomLarge; addresses queueing directlyGPU cost per request
Chunked prefillCaps inter-token gapsSlower first token for long prompts, some throughput
Prefill/decode separationRemoves interferenceKV transfer, more complex operations
Length-aware poolsIsolates long-context costFragmented capacity, lower utilisation
HedgingTrims first-token tailExtra GPU work; needs cancellation
Admission controlProtects accepted requestsSome requests rejected or delayed

What to do next

  1. Build an open-loop benchmark from real prompt and output length distributions, and measure from scheduled send time with no client connection cap.
  2. Log per-request queue time, prefill time, preemptions, largest co-scheduled prefill and cold-path flags on the server.
  3. Run tail attribution and rank causes by lift before changing anything.
  4. Set SLOs on p99 time to first token and p99 inter-token gap, and alert on burn rate as in LLM SLO burn rate alerts.
  5. Apply the lever that matches the top cause, then re-run the same benchmark and check both the tail and the median.
  6. Review per-replica and per-rank latency weekly to catch stragglers that fleet averages hide.
Key takeaway: Tail latency in GPU serving comes from a short list of mechanisms: queueing near saturation, prefill interference, KV cache preemption, length variance, cold paths, host stalls and slow hardware. Measure it open-loop from scheduled time, attribute it by comparing tail requests with all requests, and apply the lever that matches the cause, checking the median as well as the tail after every change.