Disaggregated serving runs the phases of LLM inference on separate pools of GPUs, so each pool can use the batch size, parallelism and hardware that suit its phase. Why that helps, when chunked prefill on one pool is the better choice, and how to size the two pools are covered in Prefill/Decode Disaggregation Architecture in Depth; what crosses the wire and how fast it must move is in P2P KV Cache Transfer.
This article is the operator's view. It follows one request through a real proxy, shows how vLLM and SGLang are launched in disaggregated mode, covers routing and failure handling, which the papers rarely discuss, explains the two newer forms of disaggregation, vision encoders and attention versus experts, and ends with how to prove the change paid off. Commands and parameter names were checked against the projects' documentation on 2026-10-04; both projects move fast, so re-check them against the version you deploy.
Three ways to disaggregate
Disaggregation is a general move: find two parts of inference with different resource profiles that interfere when they share a GPU, put them on separate workers, and pay for a transfer between them. Three splits are in use.
| Split | What moves between workers | Why the profiles differ | Where it appears |
|---|---|---|---|
| Prefill / decode | KV cache blocks | Prefill is compute-bound; decode is memory-bandwidth-bound | DistServe, Splitwise, Mooncake; vLLM, SGLang, NVIDIA Dynamo |
| Encoder / LLM | Image or audio embeddings | A vision encoder is small and bursty; the LLM is large and steady | vLLM encoder disaggregation; Dynamo E/PD and E/P/D |
| Attention / FFN | Activations, every layer | In MoE decode, attention holds KV state while experts are sparsely hit | MegaScale-Infer |
The first is in production at several large providers and in every major open-source engine. The second is newer and useful mainly for multimodal traffic. The third moves data every layer, so it needs fast interconnect and careful pipelining, and it is the least mature of the three.
The request path, step by step
In vLLM's reference design, the prefill and decode instances are ordinary OpenAI-compatible servers, and a proxy in front of them does the choreography. The six steps in the diagram are the whole protocol.
- The client sends a normal completion request to the proxy.
- The proxy sends the same request to a prefill instance with
max_tokensset to 1, streaming off, andkv_transfer_paramsset to{"do_remote_decode": true, "do_remote_prefill": false}with the remote fields empty. Prefill runs the prompt and produces one token. - The prefill response carries
kv_transfer_paramsdescribing where the KV blocks are: the engine id, block ids, host and port. - The proxy sends the original request to a decode instance with those parameters attached.
- The decode instance pulls the blocks from the prefill instance through the connector, NIXL in this case, directly between GPU memories where the hardware allows.
- Decode streams tokens back through the proxy to the client.
Prefill holds the blocks until decode has read them, so the cost of a decode that never arrives is KV memory pinned on the prefill side. That is the root of several failure modes below.
A minimal proxy
The proxy is small enough to read in full. The version below follows the field names in vLLM's toy proxy for the NIXL connector, adds least-loaded selection and keeps the decode response streaming. It is a teaching sketch: production routers add health checks, retries, metrics and prefix-aware routing.
import httpx, random
from fastapi import FastAPI, Request
from fastapi.responses import StreamingResponse
app = FastAPI()
PREFILL = ["http://p0:8100", "http://p1:8100"]
DECODE = ["http://d0:8200", "http://d1:8200", "http://d2:8200"]
client = httpx.AsyncClient(timeout=httpx.Timeout(30.0, connect=2.0))
def pick(pool, load): # least outstanding requests; ties broken randomly
return min(pool, key=lambda u: (load.get(u, 0), random.random()))
load: dict[str, int] = {}
@app.post("/v1/completions")
async def completions(request: Request):
req = await request.json()
p, d = pick(PREFILL, load), pick(DECODE, load)
pre = dict(req, stream=False, max_tokens=1, kv_transfer_params={
"do_remote_decode": True, "do_remote_prefill": False,
"remote_engine_id": None, "remote_block_ids": None,
"remote_host": None, "remote_port": None})
load[p] = load.get(p, 0) + 1
try:
r = await client.post(f"{p}/v1/completions", json=pre)
r.raise_for_status()
params = r.json().get("kv_transfer_params")
finally:
load[p] -= 1
dec = dict(req)
if params:
dec["kv_transfer_params"] = params # without it, decode recomputes the prompt itself
async def relay():
load[d] = load.get(d, 0) + 1
try:
async with client.stream("POST", f"{d}/v1/completions", json=dec) as resp:
async for chunk in resp.aiter_bytes():
yield chunk
finally:
load[d] -= 1
return StreamingResponse(relay(), media_type="text/event-stream")Note the fallback in the decode request: if prefill returned no transfer parameters, the decode instance receives a plain request and prefills locally. That is slower but correct, and it is the simplest degraded mode you can have.
Launching vLLM and SGLang in disaggregated mode
Both engines expose disaggregation as launch flags. In vLLM the KV connector is configured with --kv-transfer-config; with the NIXL connector the prefill side is kv_producer and the decode side kv_consumer. Each worker needs its own side-channel port on a host, which is set with VLLM_NIXL_SIDE_CHANNEL_PORT (default 5600). In SGLang, each server takes a --disaggregation-mode and a transfer backend, Mooncake, NIXL or Ascend, and the SGLang router plays the proxy's role.
# vLLM, one prefill and one decode instance with the NIXL connector
UCX_NET_DEVICES=all VLLM_NIXL_SIDE_CHANNEL_PORT=5600 \
vllm serve <model> --port 8100 \
--kv-transfer-config '{"kv_connector":"NixlConnector","kv_role":"kv_producer","kv_load_failure_policy":"fail"}'
UCX_NET_DEVICES=all VLLM_NIXL_SIDE_CHANNEL_PORT=5601 \
vllm serve <model> --port 8200 \
--kv-transfer-config '{"kv_connector":"NixlConnector","kv_role":"kv_consumer","kv_load_failure_policy":"fail"}'
# On separate hosts, also set VLLM_NIXL_SIDE_CHANNEL_HOST to an address the peer can reach.
# SGLang, with its router in front
python -m sglang.launch_server --model-path <model> --disaggregation-mode prefill \
--disaggregation-transfer-backend mooncake --disaggregation-ib-device <rdma-device> --port 30000
python -m sglang.launch_server --model-path <model> --disaggregation-mode decode \
--disaggregation-transfer-backend mooncake --disaggregation-ib-device <rdma-device> --port 30001
python -m sglang_router.launch_router --pd-disaggregation \
--prefill http://<prefill-host>:30000 --decode http://<decode-host>:30001 --host 0.0.0.0 --port 8000The kv_load_failure_policy setting is worth deciding deliberately. With fail, the default, a decode instance that cannot load the blocks fails the request; with recompute it recomputes the missing blocks locally, which keeps the request alive at the cost of decode-pool compute, exactly the interference disaggregation was meant to remove. Use recompute while you are stabilising the transfer path, and alert on how often it fires. SGLang's documentation also requires prefill and decode to run the same revision; treat a mixed-version pair as unsupported in any engine.
Routing prefill and decode
Routing in a disaggregated cluster is two decisions with different inputs.
- Choosing prefill. Load is measured in queued prompt tokens, not requests, because one 30,000-token prompt costs as much as a hundred short ones. If prefill instances keep a prefix cache, route requests that share a long system prompt or conversation history to the instance that already holds it; a hit removes most of the prefill work.
- Choosing decode. Decode capacity is KV memory. Route to the instance with enough free blocks for the prompt plus the expected output, then by running batch size. An instance that accepts a request it cannot hold will pre-empt something else later.
Multi-turn traffic raises a third question: whether a conversation's earlier KV can be reused on its next turn instead of prefilled again. Support differs by engine and connector version, so check yours, and measure how long users pause between turns before counting on it.
Failure handling
| Failure | What the user sees | What to do |
|---|---|---|
| Prefill instance dies before responding | Slow first token, then an error | Retry on another prefill instance; the prompt is idempotent |
| Decode cannot load the blocks | Error, or slow first token with recompute | fail: retry the pair; recompute: accept and alert on the rate |
| Decode instance dies mid-stream | Stream stops partway | Re-prefill prompt plus tokens already sent on a new pair; resume |
| Decode never arrives for prefilled blocks | Nothing; prefill memory leaks | Release timeout on the producer; alert on pinned blocks |
| Transfer path saturates | TTFT rises for everyone | Watch transfer time per request; add links or reduce cross-rack pairing |
| One pool scaled, the other not | Queueing in the starved pool | Autoscale each pool on its own signal: prompt-token queue vs free KV |
Resuming after a decode failure is the case worth building early. The proxy has already relayed some tokens; it can send the prompt with those tokens appended as a new prefill, then continue decoding. The user sees a pause instead of a failure, at the cost of one extra prefill. With sampling, the continuation is still a valid answer, though not the one the dead instance would have produced.
Beyond prefill and decode
Encoder disaggregation. A multimodal model first runs images through a vision encoder, whose load arrives in bursts and has nothing in common with text decode. vLLM and NVIDIA Dynamo describe running the encoder as its own worker that sends embeddings over NIXL, either to a combined prefill-and-decode worker (E/PD) or to a prefill worker that then hands KV to decode (E/P/D). Text-only requests skip the encoder pool entirely, and encoder workers scale with image traffic alone.
Attention and FFN disaggregation. In a mixture-of-experts model at decode time, attention needs the KV cache of every request, while each expert sees only the tokens routed to it, so expert GPUs sit under-used. MegaScale-Infer separates the two: attention runs data-parallel on one set of GPUs, experts run expert-parallel on another, and a ping-pong pipeline splits the batch into micro-batches so that attention computes one while experts compute another, hiding the per-layer transfer. Its authors report up to 1.90 times higher per-GPU throughput than the baselines they compared. Because activations cross the network in every layer, this split depends on fast interconnect and a purpose-built communication library; read Mixture-of-Experts Serving Architecture before considering it.
Measuring whether it paid off
Disaggregation is justified by one number: goodput, the rate of requests that meet both the time to first token and the time per output token objectives, a metric popularised by the DistServe paper. Raw throughput can rise while goodput falls, if the extra tokens come from requests that missed their objectives. GPU Inference Latency, in depth defines both latencies.
def goodput(requests, ttft_slo_s=1.0, tpot_slo_s=0.05, window_s=3600):
"""Requests per second that met BOTH latency SLOs. Each request: dict with
ttft (s), e2e (s) and out_tokens, taken from proxy logs, not engine logs."""
ok = 0
for r in requests:
tpot = (r["e2e"] - r["ttft"]) / max(r["out_tokens"] - 1, 1)
if r["ttft"] <= ttft_slo_s and tpot <= tpot_slo_s:
ok += 1
return ok / window_sMeasure at the proxy, because only the proxy sees the full path including the transfer. Compare against the strongest colocated baseline, which means a single pool with chunked prefill tuned as in Chunked Prefill in Serving, on the same number of GPUs and the same replayed traffic. If disaggregation does not beat that baseline on goodput per GPU, its operational cost is not paid for.
Worked example: migrating a sixteen-GPU deployment
Consider a team serving a chat model on sixteen GPUs as colocated replicas with chunked prefill. Their traffic mixes short chats with retrieval-augmented requests that carry 20,000-token prompts. During peaks, time per output token for chat users degrades whenever a long prompt is admitted. The figures in this plan are illustrative, not benchmarks.
- Record a day of production requests with prompt and output lengths and replay it against the current deployment; measure goodput at the proxy as the baseline.
- Stand up one prefill and one decode instance with the vLLM commands above on two GPUs, and replay a slice of traffic through the proxy to validate correctness and transfer time.
- Split the sixteen GPUs into pools using the sizing method from the pool-sizing article, for example six prefill and ten decode, and replay the full day.
- Compare goodput per GPU. Suppose the time per output token objective is now met through the peaks while time to first token rises by the transfer time; if goodput per GPU improves, keep it.
- Shift production traffic gradually, keep the colocated deployment warm for rollback, and alert on recompute rate, pinned prefill blocks and transfer time.
Trade-offs
Disaggregation buys isolation between phases and independent scaling, and it costs a second pool to operate, a transfer path that must be engineered and monitored, a proxy that becomes critical, and new failure modes. It pays most when prompts are long, output latency objectives are strict, and traffic is large enough that each pool stays busy. For small deployments, short prompts or relaxed objectives, a well-tuned colocated pool usually wins. Prefill/decode disaggregation architecture has a complementary walk through the components.
What to do next
- Measure goodput at the proxy on your current deployment with replayed production traffic.
- Check that your interconnect and engine versions support the transfer backend you plan to use.
- Run a two-GPU prefill and decode pair with the documented flags and validate output equality on a fixed prompt set.
- Choose a load-failure policy and a release timeout for prefilled blocks, and alert on both.
- Build routing on prompt-token queues for prefill and free KV blocks for decode.
- Implement resume-after-decode-failure in the proxy before you go to production.
- Compare goodput per GPU against tuned chunked prefill; only then shift traffic, gradually, with rollback ready.
- For multimodal or MoE fleets, evaluate encoder or attention-FFN splits only after phase disaggregation is stable.