A vision-language model (VLM) takes images and text and produces text, and sometimes images. Underneath the many model names there are only a few architectural choices, and they decide most of what you pay at inference time: how many tokens an image becomes, whether those tokens occupy the KV cache, how much prefill compute each image needs, and whether the language model's text skills survive.
This article explains the four main families from first principles, with the published designs that define them: LLaVA, BLIP-2, Flamingo, Llama 3.2 Vision, Qwen2-VL and Chameleon. It then works through the GPU cost of each with real arithmetic and covers serving, failure modes and how to choose. Training recipes, freezing schedules and data packing are covered in multimodal LLM training; here the focus is the architecture and its runtime cost.
Encoder, connector, decoder
Most VLMs have three parts. A vision encoder, usually a Vision Transformer (ViT), cuts the image into square patches, such as 14 by 14 pixels, embeds each patch and runs transformer layers over them. The encoder is often initialised from a contrastive model such as CLIP or SigLIP, whose training is described in CLIP training on GPU. A connector turns encoder features into something the language model can consume. A language model decoder generates the answer.
The families differ in the connector and in where image information enters the decoder.
Projecting patches into the sequence
Family A: project patches into the token stream. LLaVA takes CLIP ViT-L/14 patch features and maps each one into the language model's embedding space. The first version used a linear layer; LLaVA-1.5 uses a two-layer MLP at 336 by 336 pixels. That image gives a 24 by 24 grid, so 576 visual tokens are inserted where the prompt has an image placeholder. The decoder then treats them like any other tokens.
import torch
from torch import nn
class Projector(nn.Module):
# LLaVA-1.5 style: vision width -> LM width, two layers with GELU.
def __init__(self, vision_dim=1024, lm_dim=4096):
super().__init__()
self.net = nn.Sequential(nn.Linear(vision_dim, lm_dim), nn.GELU(),
nn.Linear(lm_dim, lm_dim))
def forward(self, patch_feats): # [batch, 576, 1024]
return self.net(patch_feats) # [batch, 576, 4096]
def splice(text_embeds, image_embeds, image_pos):
# Replace one placeholder at image_pos with the projected patch tokens.
return torch.cat([text_embeds[:, :image_pos], image_embeds,
text_embeds[:, image_pos + 1:]], dim=1)The design is simple and strong, because the decoder's full attention can relate any word to any patch. Its cost grows with resolution. To read small text, LLaVA-NeXT's AnyRes splits a large image into tiles plus a downscaled overview, which multiplies the tokens. Qwen2-VL takes the native-resolution route: the ViT accepts a variable patch grid, an MLP merges each 2 by 2 group of neighbouring tokens into one, and M-RoPE gives every token temporal, height and width positions instead of a flat index. A 224 by 224 image becomes 66 tokens, which is 64 merged patches plus start and end markers.
Learned queries, cross-attention and early fusion
Family B: compress with learned queries. BLIP-2's Q-Former is a small transformer with 32 learned query vectors that cross-attend to the frozen image encoder's features. Only the 32 outputs reach the language model, whatever the image size. Flamingo's Perceiver Resampler does the same with 64 latents. The token budget is fixed and small, which is cheap at inference, but fine detail such as document text or small objects can be lost in the compression.
Family C: cross-attention inside the decoder. Flamingo inserts new cross-attention layers between the frozen language model's layers. Text queries attend to image features, and the result passes through a tanh gate whose parameter starts at zero, so at initialisation the model behaves exactly like the original language model. Llama 3.2 Vision (11B and 90B) follows this pattern: an image encoder feeds cross-attention adapter layers, and the language model's weights were kept frozen while the adapter was trained, which preserves text-only behaviour. Image features never enter the self-attention sequence, so they do not lengthen the text context.
Family D: early fusion. Chameleon trains an image tokenizer that maps a 512 by 512 image to 1,024 discrete codes from an 8,192-entry codebook. Those codes share a vocabulary with text, and one transformer models the mixed sequence from scratch. The model can output images as well as read them. The price is an expensive from-scratch training run and training instability that the authors had to address with changes such as query-key normalisation.
Positions, aspect ratio and video
Once image tokens sit in the sequence, the decoder needs to know where each one came from. The simplest designs give patch tokens ordinary consecutive positions, row after row. That works at one fixed resolution, but it tells the model nothing about two-dimensional layout: the patch directly below another is 24 positions away at 336 pixels and 96 positions away at 1,344 pixels. Models trained one way then see unfamiliar distances when the grid changes.
M-RoPE in Qwen2-VL addresses this by splitting the rotary embedding into three parts that rotate by temporal, height and width index. Text tokens get the same value in all three, so they behave as in a plain language model. Patches in the same row share a height index, whatever the image width. The same scheme extends to video, where the temporal index advances per frame.
Video makes token arithmetic harsher. A one-minute clip sampled at two frames per second gives 120 frames. At 256 tokens per frame after merging, that is 30,720 tokens, or 3.75 GiB of KV cache in the 7B example below, for a single request. Video models therefore compress across time as well as space, sample fewer frames, or lower the per-frame resolution, and each choice trades away something: motion detail, short events or small text.
Aspect ratio is the other trap. Squashing a tall receipt into a square distorts characters, and padding wastes tokens on blank space. Tiling and native-resolution encoders avoid both, at the price of a token count that varies with every input. Schedulers must then plan for the worst case, not the average one.
What each design costs on a GPU
The families differ most in GPU memory and prefill compute. Take a decoder with 7 billion parameters, 32 layers, grouped-query attention with 8 KV heads of dimension 128, in bf16. Each token in the KV cache costs 2 (keys and values) times 32 times 8 times 128 times 2 bytes, which is 128 KiB. Prefill costs about 2 FLOPs per parameter per token. The calculator below applies those rules; KV cache sizing explains the formula.
def kv_bytes(tokens, layers=32, kv_heads=8, head_dim=128, dtype_bytes=2):
return 2 * layers * kv_heads * head_dim * dtype_bytes * tokens
def image_tokens(h, w, patch=14, merge=1, special=0):
return (h // patch) * (w // patch) // (merge * merge) + special
designs = {
"A: projector, 336 px": image_tokens(336, 336),
"A: AnyRes, 4 tiles + overview": 5 * image_tokens(336, 336),
"A: native 1344 px, 2x2 merge": image_tokens(1344, 1344, merge=2, special=2),
"B: Q-Former queries": 32,
}
for name, n in designs.items():
print(f"{name:34s} {n:5d} tokens {kv_bytes(n) / 2**20:6.1f} MiB KV"
f" {2 * 7e9 * n / 1e12:5.1f} TFLOP prefill")
# Family C: image K/V live only in the cross-attention layers (here 8 of them).
print(kv_bytes(576, layers=8) / 2**20, "MiB for 576 image features")| Design | Image tokens | KV per image | Decoder prefill |
|---|---|---|---|
| A: projector, 336 px | 576 | 72 MiB | 8.1 TFLOP |
| A: AnyRes, 4 tiles plus overview | 2,880 | 360 MiB | 40.3 TFLOP |
| A: native 1344 px, 2x2 merge | 2,306 | 288 MiB | 32.3 TFLOP |
| B: 32 learned queries | 32 | 4 MiB | 0.4 TFLOP |
| C: 576 features, 8 cross-attention layers | 0 in sequence | 18 MiB | small |
Worked example. A document assistant on an 80 GB GPU has about 50 GB left for KV cache after the weights. With 1,344-pixel pages, every page holds 288 MiB, so the cache fits about 175 page images across all concurrent requests. Thirty users with 8 pages each need 240 pages, about 68 GiB, so they would not fit, and the scheduler would queue or preempt. The same pages at 336 pixels fit four times as many, but the text on them becomes unreadable. Resolution is a product decision with a direct memory price.
The vision encoder is usually the smaller cost. ViT-L has about 300 million parameters, so encoding 577 patches costs roughly 0.35 TFLOP, about 4 percent of the 576-token decoder prefill. Larger encoders and native-resolution inputs raise that share, and the encoder runs once per image whatever the answer length.
Serving a VLM
Serving engines treat a VLM request as encode, then prefill, then decode. Several practices follow from the arithmetic above:
- Cache encoder outputs by content. Key the cache on a hash of the pixels plus the preprocessing settings and the encoder version. A key built from the URL alone returns stale or wrong features when the image changes.
- Make prefix caching image-aware. Two prompts share a KV prefix only if the image tokens are identical, so the image hash must be part of the prefix key.
- Cap images and pixels per request. One request with fifty high-resolution images can take the whole cache. Enforce a pixel budget at the API layer, before the encoder runs.
- Schedule prefill separately. Image-heavy prompts produce long prefills that stall decoding for everyone else; chunked prefill and prefill and decode separation limit that.
- Use fused attention. Thousands of image tokens make attention memory matter, which is what FlashAttention-style kernels address.
Failure modes
| Failure | Cause | Fix |
|---|---|---|
| Small text read wrongly | Image downscaled to a fixed square | Native resolution or tiling for documents |
| Distorted shapes | Aspect ratio squashed during resize | Pad or tile; keep the processor's own resize |
| Token count mismatch error | Custom preprocessing differs from the model's processor | Use the published processor and test it |
| Out of memory under load | Unbounded pixels per request | Pixel budget and admission control |
| Wrong answer from cache | Encoder cache keyed on URL | Content hash plus preprocessing version |
| Objects described that are not there | Language prior overrides weak visual evidence | Grounding checks, higher resolution, refusal training |
| Text skills degrade after fine-tuning | Decoder weights updated on narrow image data | Mix text data in, or use a cross-attention design |
Choosing an architecture
Choose by workload. Documents, charts and screenshots need resolution, so use family A with native resolution or tiling and budget the memory. Captioning or classification at high volume suits family B, whose fixed budget keeps cost predictable. If you must keep an existing text model's behaviour exactly, or serve the same weights for text-only traffic, family C's gated cross-attention is the safer choice. Pick family D only when you need image output from the same model and can fund pretraining from scratch.
Measure before you decide: run your own images through the candidate processors, count the tokens, and multiply by the KV cost of your decoder. That one table usually settles the choice.
What to do next
- Collect 200 representative images from your product and record their sizes and aspect ratios.
- Run each candidate model's processor on them and log the token counts.
- Compute KV memory and prefill FLOPs per request with the calculator above, using your decoder's layer and head numbers.
- Set a pixel budget and an images-per-request limit at the API layer.
- Add a content-hashed encoder cache and include image hashes in prefix-cache keys.
- Test small-text and counting questions at two resolutions to see where accuracy falls.
- Read multimodal LLM training before you fine-tune any of these families.