Serving a text LLM is a story about the KV cache and the batch scheduler. Serving a multimodal model, one that accepts images, video or audio alongside text, adds three stages in front of that story: downloading and decoding media on the CPU, running a vision or audio encoder on the GPU, and splicing its output into the token sequence. Each stage has its own bottleneck, cache and way of failing, and a single high-resolution image can cost more prefill than a long document.
This article is about the serving system rather than the model. How encoders, projectors and cross-attention are built is in Vision-Language Model Architectures, and encoder FLOPs by resolution are in Vision Encoders, in depth. Here we follow a request end to end, count the tokens and bytes it costs, then cover scheduling, caching, disaggregation and the failure modes that take multimodal endpoints down.
The request path
A request carrying an image arrives as a URL or base64 bytes inside a chat message. The API server parses it, enforces per-request limits and hands media to CPU workers that fetch, decode, convert colour, resize to the model's grid and normalise. The resulting pixel tensor goes to the encoder, typically a ViT, which produces one feature per patch; a projector maps those to the LLM's hidden size, often merging neighbouring patches. Finally the engine replaces placeholder tokens in the prompt with these embeddings and runs ordinary prefill and decode.
Two properties shape everything after. First, the number of LLM tokens an image becomes is decided by preprocessing, before any GPU work, so the server can know the cost of a request up front if it does the arithmetic. Second, encoders generally process an image as a whole: even when the decoder's prefill is chunked across scheduler steps, the encoder output for an item must exist before its first token can be prefilled.
Counting media tokens
Token cost depends on the model family. Fixed-resolution encoders such as a CLIP ViT-L/14 at 336 pixels emit a fixed (336/14)2 = 576 patch features per image. Dynamic resolution models scale with the pixels. Qwen2-VL, for example, uses 14 pixel patches and merges each 2×2 group into one LLM token, so each token covers a 28×28 pixel area, and for video it also merges pairs of frames. Audio follows the same logic in time: Whisper's encoder turns 30 seconds into 1,500 frames, and audio LLMs usually pool those further before the decoder sees them.
import math
def image_tokens(w, h, patch=14, merge=2, max_pixels=1280 * 28 * 28):
unit = patch * merge # 28 px per LLM token side
# Resize so both sides are multiples of unit and area <= max_pixels.
scale = min(1.0, math.sqrt(max_pixels / (w * h)))
w2 = max(unit, math.floor(w * scale / unit) * unit)
h2 = max(unit, math.floor(h * scale / unit) * unit)
return (w2 // unit) * (h2 // unit), (w2, h2)
def request_tokens(text_tokens, images, max_output):
img = sum(image_tokens(w, h)[0] for w, h in images)
return text_tokens + img + max_output
print(image_tokens(1344, 1008, max_pixels=10**9)) # (1728, (1344, 1008))
print(image_tokens(1344, 1008)) # (1230, (1148, 840))The two calls at the end make a point worth internalising: the cap on pixels is a cost control. Uncapped, a 1344×1008 photo is 48×36 = 1,728 tokens; with a cap of 1,280 tokens it is resized to 1148×840 and costs 1,230. The exact resize rule lives in each model's processor, so compute it with the model's own processor in the API server rather than a reimplementation, and treat the snippet above as illustrative. What matters is that the server computes the number before admitting the request.
CPU preprocessing and its attack surface
Media preprocessing is CPU work, and it is where many multimodal endpoints first saturate. JPEG decode and resize of a 12-megapixel phone photo takes tens of milliseconds of a core; video needs demuxing and decoding of every sampled frame. If this runs in the same Python event loop as the API server, a burst of images stalls request parsing for everyone. Run it in a process pool sized to the cores you have, with a bounded queue so overload turns into fast rejections instead of growing latency.
It is also the attack surface. Image URLs make your server fetch arbitrary addresses, which is a server-side request forgery risk against internal metadata endpoints. Compressed formats allow decompression bombs: a small file that decodes to billions of pixels. Pillow raises DecompressionBombError above a pixel threshold, but you should enforce your own limits before decode.
import io, ipaddress, socket
from urllib.parse import urlparse
from PIL import Image
MAX_BYTES = 20 * 1024 * 1024
MAX_PIXELS = 40_000_000
Image.MAX_IMAGE_PIXELS = MAX_PIXELS # Pillow's own guard, tightened
def safe_fetch(url, http):
u = urlparse(url)
if u.scheme != "https":
raise ValueError("https only")
addr = ipaddress.ip_address(socket.gethostbyname(u.hostname))
if addr.is_private or addr.is_loopback or addr.is_link_local:
raise ValueError("internal address") # pin addr for the request itself
r = http.get(url, timeout=(3, 10), stream=True, allow_redirects=False)
data = r.raw.read(MAX_BYTES + 1)
if len(data) > MAX_BYTES:
raise ValueError("too large")
return data
def safe_decode(data):
img = Image.open(io.BytesIO(data)) # reads header only
if img.width * img.height > MAX_PIXELS:
raise ValueError("too many pixels")
img.draft("RGB", (2048, 2048)) # JPEG: decode at reduced scale
return img.convert("RGB")The DNS check must be bound to the connection (resolve once, connect to that address), or an attacker can rebind the name between check and fetch. Many teams avoid the whole class by refusing URLs and accepting only uploaded bytes or pre-signed object-store references.
Scheduling the encoder
On the GPU, the encoder and the LLM compete for the same step. A scheduler that admits eight new image requests in one step runs eight encoder passes and then a prefill of perhaps 14,000 image tokens, and every request already decoding waits for that step to finish. That is the multimodal version of the prefill-decode interference that continuous batching is designed to manage, with an extra term.
Engines handle it with an encoder budget per step: only so many encoder input tokens are scheduled per iteration, and items whose encoder output is not yet computed wait. vLLM also exposes --limit-mm-per-prompt to cap how many images, videos or audio clips one prompt may carry, for example '{"image": 4, "video": 1}'. Use both levers: the per-prompt limit stops a single request from monopolising the encoder, and the token budget keeps step times predictable so decode latency stays flat.
Three caches
Multimodal traffic repeats more than you might expect: the same product photo in many questions, the same PDF page in a multi-turn chat where every turn resends the whole history. Three caches catch this, one per stage.
- Preprocessing cache. Maps raw bytes plus processor settings to the pixel tensor. Saves CPU, and is mainly worth it for video.
- Encoder output cache. Maps a content hash of the media item to its embeddings, so a repeated image skips the encoder. vLLM keys its encoder cache by a per-item hash it calls
mm_hash. - Prefix KV cache. Standard prefix caching hashes token ids per block. Image placeholder tokens are identical for every image, so the engine must fold the image's content hash into the block key, or two different images with the same surrounding text would share KV, which is a correctness bug and, across tenants, a data leak.
import hashlib, json
def mm_cache_key(media_bytes: bytes, model_rev: str, processor_cfg: dict) -> str:
h = hashlib.sha256()
h.update(media_bytes) # content, never the URL
h.update(model_rev.encode()) # encoder weights
h.update(json.dumps(processor_cfg, sort_keys=True).encode()) # resize, max_pixels
return h.hexdigest()Key on content, never on URL: the same URL can serve different bytes tomorrow, and two URLs can serve the same bytes. Include the model revision and every processor setting that changes the output, because a cache that survives a model upgrade serves embeddings from the old encoder to the new decoder, and the outputs degrade quietly without errors.
Worked example: memory and compute
Work through the memory for a 7B-class model with grouped-query attention: 28 layers, 4 KV heads, head dimension 128, 16-bit KV. Each token costs 2 (K and V) × 28 × 4 × 128 × 2 bytes = 57,344 bytes, 56 KiB. An uncapped 1,728-token photo therefore occupies about 94.5 MiB of KV cache, the same as a 1,728-token document. A prompt with eight such photos is about 756 MiB before a single word of text, and a short video at 64 sampled frames multiplies that again. Admission control by request count is meaningless here; admit by projected token count against free KV blocks, as paged KV cache allocators make possible.
Compute follows the same shape. A rough encoder estimate is 2 × parameters × patches: a 675-million-parameter vision tower over the 6,912 pre-merge patches of that photo is about 9 TFLOPs plus attention, and the LLM prefill over its 1,728 tokens at 7.6 billion parameters is about 26 TFLOPs. At a few hundred TFLOPs of effective throughput, those are tens of milliseconds each; treat these as estimates and measure on your hardware, because attention at high patch counts grows faster than linearly.
Disaggregating the encoder
Once encoder load is significant, running it in a separate pool, encode then prefill and decode (sometimes written EPD), becomes attractive. vLLM documents a disaggregated encoder mode in which the encoder runs as its own instance. The arithmetic for the hand-off is favourable: what crosses the network is the projected embeddings, 1,728 tokens × 3,584 hidden × 2 bytes ≈ 12.4 MB for the photo above, versus about 99 MB of KV cache if you disaggregated after prefill instead.
Separation lets you scale encoders on cheaper or smaller GPUs, keep encoder bursts out of decode steps, and share one encoder cache across many LLM replicas. It costs a network hop on every new item, a second fleet to operate and a new failure mode: the LLM replica waiting on an encoder that is overloaded or gone. Use it when encoder time is a large fraction of prefill and traffic has many media items per request; for light image traffic a co-located encoder with a per-step budget is simpler.
Failure modes
- Head-of-line blocking by big media. One video request schedules thousands of encoder tokens and every decoding stream stutters. Cap items per prompt and encoder tokens per step.
- CPU starvation. Decode and resize in the API process saturate cores while GPUs sit idle. Profile CPU per request and run media work in a separate pool.
- Decompression bombs and SSRF. Small inputs that expand to huge pixel counts, and URLs that point inward. Check headers before decode and bind fetches to vetted addresses.
- Cache keys on URL or without model revision. Stale or wrong embeddings after an upgrade, and cross-tenant reuse if image hashes are missing from prefix blocks.
- Token count mismatch. Client-side and server-side resize rules disagree, so context-length checks pass in the client and fail in the engine. Count with the server's processor and return the count to clients.
- Silent quality loss from aggressive caps. Downscaling to save tokens makes small text in screenshots unreadable. Evaluate caps on your own document and screenshot traffic, not only on photos.
Trade-offs
| Decision | Option A | Option B | Choose A when |
|---|---|---|---|
| Resolution | Capped pixels | Native resolution | Cost and latency matter more than fine detail |
| Encoder placement | Co-located | Disaggregated | Media is a small share of prefill time |
| Media input | Uploaded bytes only | URLs fetched server-side | You cannot afford an SSRF review |
| Encoder cache | Per replica | Shared tier | Repeats are mostly within one session |
What to do next
- Measure your traffic: images, video seconds and audio seconds per request, and their pixel and duration distributions.
- Compute media tokens with the model's own processor in the API server, and admit requests by projected tokens against free KV memory.
- Set per-prompt media limits (
--limit-mm-per-promptin vLLM) and a pixel cap, and evaluate the cap on your hardest inputs. - Move fetch, decode and resize into a bounded process pool with byte, pixel and timeout limits; prefer uploads over URLs.
- Key caches on content hash plus model revision plus processor settings, and confirm that prefix cache blocks include media hashes.
- Track encoder time, CPU preprocess time and prefill time separately; consider encoder disaggregation only when encoder time is a large, bursty share.
- Load-test with mixed text and media traffic and watch decode inter-token latency during media bursts.