Every vision-language model, image search system and multimodal guard model starts with a vision encoder that turns pixels into a sequence of feature vectors. Most are Vision Transformers (ViTs) from a few families: CLIP, SigLIP and SigLIP 2, DINOv2, and the encoders trained inside VLMs such as Qwen2-VL. In production the encoder is mostly a GPU-budget question: how many tokens does an image become, what does each token cost, and how do you avoid paying for the same image twice?

This article treats the encoder as a workload. It derives the token count and FLOP formulas from first principles, works through ViT-L/14 at four resolutions, explains tiling, native-resolution packing and token merging, covers which features you should actually extract, and ends with serving patterns, failure modes and a checklist. How the encoder is wired into a language model is covered in the VLM architectures article, and how contrastive encoders are trained in the CLIP training article; here the focus is the encoder itself.

What a vision encoder does

A vision encoder as a GPU workload: pixels in, a sequence of patch features outDecode + resizeJPEG to tensorPatchifyp x p conv = matmul+ positionslearned or 2D RoPEL transformer blocksattention + MLP, N tokensFeature tappenultimate layer, patch tokensPool or projectCLS / mean / MLP connectorConsumerLLM, retrieval, headCost lives in the blocks: linear layers scale with N, attention with N squared.N = (H/p) x (W/p), so doubling image side length quadruples the token count.
The encoder pipeline. Preprocessing runs once per image; the transformer blocks dominate compute; the feature tap decides what downstream models see.

A ViT does four things. It decodes and resizes the image to a fixed or bounded resolution and normalises the channels with the mean and standard deviation the model was trained with. It cuts the image into non-overlapping p by p patches and projects each one to a d-dimensional vector. It adds position information, because attention by itself is order-blind. Then it runs L identical transformer blocks over the resulting token sequence and hands some of the outputs downstream.

The families differ in training, not in that skeleton. CLIP and SigLIP match images to captions, so their features align with language; DINOv2 uses self-distillation without text, which gives stronger dense, spatial features. For a GPU engineer the same four numbers (patch size p, width d, depth L and input resolution) determine cost whichever family you pick.

Patchify is a matmul

The patch embedding is usually written as a convolution with kernel size and stride both equal to p. Because the windows never overlap, that convolution is exactly a matrix multiply: flatten each p x p x 3 patch into a vector of 3p² numbers and multiply by a 3p² x d weight matrix. Seeing it as a matmul matters for two reasons. It is cheap (for p = 14 and d = 1024 it is 588 x 1024 per token, a small fraction of one transformer layer), and it is the place where variable-resolution encoders cope with arbitrary image sizes, because a list of patches from any number of images can be concatenated into one long matrix.

import torch, torch.nn as nn

class PatchEmbed(nn.Module):
    def __init__(self, patch=14, dim=1024, in_ch=3):
        super().__init__()
        self.proj = nn.Conv2d(in_ch, dim, kernel_size=patch, stride=patch)

    def forward(self, x):                       # x: [B, 3, H, W]
        x = self.proj(x)                        # [B, dim, H/p, W/p]
        return x.flatten(2).transpose(1, 2)     # [B, N, dim]

def patchify_as_matmul(x, weight, bias, p):
    """Same output as the conv above, written as unfold + matmul."""
    B, C, H, W = x.shape
    cols = x.unfold(2, p, p).unfold(3, p, p)               # [B, C, H/p, W/p, p, p]
    cols = cols.permute(0, 2, 3, 1, 4, 5).reshape(B, -1, C * p * p)
    return cols @ weight.reshape(weight.shape[0], -1).T + bias

Note the integer division hidden in H/p. If the side length is not a multiple of p, the strided convolution silently drops the remainder pixels. A SigLIP so400m model with patch size 14 at 384 pixels produces 27 x 27 = 729 tokens, because 384/14 is 27.4 and the last few pixel rows and columns are never seen. Preprocessors that pick resolutions should round to multiples of p (or of 2p when tokens are later merged in 2 x 2 groups).

Counting tokens and FLOPs

With the patch embedding understood, the cost model follows. An H x W image becomes N = (H/p)(W/p) tokens, plus one if the model keeps a CLS token. Each transformer layer of width d with an MLP ratio of 4 does about 12d² multiply-accumulates (MACs) per token in its linear layers (4d² for the query, key, value and output projections, 8d² for the MLP) and about 2Nd MACs per token in attention (scores, then the weighted sum of values). So per image:

def vit_cost(h, w, patch=14, dim=1024, layers=24, mlp_ratio=4, cls=1):
    n = (h // patch) * (w // patch) + cls
    linear_macs = layers * n * (4 * dim * dim + 2 * mlp_ratio * dim * dim)
    attn_macs = layers * 2 * n * n * dim
    gflops = 2 * (linear_macs + attn_macs) / 1e9      # 1 MAC = 2 FLOPs
    return n, gflops, attn_macs / (linear_macs + attn_macs)

The linear term grows with N and the attention term with N². At the resolutions most encoders were trained at, the linear term dominates, which is why tokens per image is a good first-order proxy for cost. As resolution climbs the quadratic term catches up, and that is when fused attention kernels and native-resolution packing start to matter.

Worked example: ViT-L/14 from 224 to 672 pixels

ViT-L/14 forward cost per image (GFLOPs) and the attention share as resolution grows224 px257 tokens, 162 GFLOPs, attention 4.0%336 px577 tokens, 381 GFLOPs, attention 8.6%448 px1025 tokens, 722 GFLOPs, attention 14.3%672 px2305 tokens, 1915 GFLOPs, attention 27.3%Blue: linear layers (QKV, output, MLP). Red: attention scores and weighted sums.
Computed with the cost function above for ViT-L/14 (d = 1024, L = 24, one CLS token). Attention is a small share at 224 pixels and over a quarter at 672.

Take CLIP's ViT-L/14, the encoder LLaVA-1.5 uses: 24 layers, width 1024, about 300 million parameters, patch size 14. Plugging the four resolutions into the function gives the numbers in the figure.

InputTokens (incl. CLS)GFLOPs per imageAttention shareRelative cost
224 x 2242571624%1x
336 x 3365773819%2.4x
448 x 4481,02572214%4.5x
672 x 6722,3051,91527%11.8x

Three conclusions fall out. First, going from 224 to 336 pixels, as LLaVA-1.5 did to read small text, costs 2.4 times the compute per image. Second, the weights (about 0.6 GB in BF16) are small, so the batched encoder is compute-bound: at an assumed 400 TFLOP/s achieved, a 336-pixel image costs about one millisecond of GPU time. Measure your own number. Third, memory surprises come from attention if it is ever materialised: 16 heads x 577 x 577 scores in BF16 is about 10.7 MB per layer per image, which at batch 256 is 2.7 GB for one layer. Fused kernels such as FlashAttention never build that matrix, so make sure your encoder actually uses them.

Variable resolution: tiling, packing and merging

Fixed square inputs waste pixels on letterboxing and destroy detail in tall documents or wide screenshots. Three strategies handle real images, and each has a different GPU profile.

  • Tiling (any-resolution grids). LLaVA-NeXT splits a large image into 336-pixel tiles from a set of grid shapes and adds one downscaled global view. A 2 x 2 grid plus the global view is 5 x 576 = 2,880 tokens. Tiles are uniform, so they batch perfectly with a fixed-resolution encoder, but the LLM pays for every token.
  • Native resolution with packing. NaViT showed that a ViT can be trained on patches from images of different sizes packed into one sequence, with attention masked so that patches only attend within their own image. SigLIP 2 ships NaFlex variants built on this idea that keep the native aspect ratio, and Qwen2-VL's encoder processes images at their native resolution, bounded by configurable minimum and maximum pixel counts in its processor.
  • Token merging. After the encoder, neighbouring tokens are combined before the LLM sees them. Qwen2-VL merges each 2 x 2 group of 14-pixel patches with a small MLP, so each LLM token covers 28 x 28 pixels: a 1344 x 1344 image is 96 x 96 = 9,216 patches inside the encoder but 2,304 tokens for the language model.

Packing is implemented on the GPU with variable-length attention. Instead of padding every image to the longest one, you concatenate all patches into one sequence and pass cumulative sequence boundaries to the kernel:

from flash_attn import flash_attn_varlen_func

# three images of 576, 1024 and 256 patches packed into one sequence of 1856 tokens
cu = torch.tensor([0, 576, 1600, 1856], dtype=torch.int32, device="cuda")
# q, k, v: [1856, heads, head_dim]; attention never crosses an image boundary
out = flash_attn_varlen_func(q, k, v, cu, cu, max_seqlen_q=1024, max_seqlen_k=1024)

Position information has to follow. Encoders with learned absolute position embeddings interpolate them (usually bicubically) to the new grid, which works for modest changes and degrades for large ones; encoders built for native resolution use 2D rotary embeddings or similar schemes that are defined for any grid.

Which features leave the encoder

What leaves the encoder matters as much as what goes in. Contrastive models were trained to produce one pooled vector (a CLS token or an attention-pooling head) that matches a caption embedding, and that vector is what you want for retrieval and zero-shot classification. VLMs instead take the full grid of patch tokens, and LLaVA famously takes them from the penultimate layer rather than the last, because the final layer is specialised for the contrastive objective and loses spatial detail. Dense tasks such as segmentation or grounding favour DINOv2-style features.

Never assume the feature tap: read the consuming model's configuration for the layer index, CLS handling and pooling, and cache exactly that tensor. Features from the wrong layer produce plausible but worse outputs, not errors.

Running it fast

Getting close to the hardware's throughput takes a few habits. Run in BF16 in inference mode; make sure attention dispatches to a fused kernel (PyTorch's scaled_dot_product_attention backends or FlashAttention); compile the model so layer norms, GELUs and residual adds fuse. Batch images: one 336-pixel image is 577 tokens, too few to fill a large GPU. Move decode and resize onto the GPU (nvJPEG via torchvision, or DALI) once CPU preprocessing becomes the bottleneck, which often happens first.

@torch.inference_mode()
def images_per_second(model, batch, res, iters=50):
    x = torch.randn(batch, 3, res, res, device="cuda", dtype=torch.bfloat16)
    for _ in range(5):
        model(x)                                  # warm-up: compile, autotune, allocator
    start = torch.cuda.Event(enable_timing=True)
    end = torch.cuda.Event(enable_timing=True)
    torch.cuda.synchronize(); start.record()
    for _ in range(iters):
        model(x)
    end.record(); torch.cuda.synchronize()
    return batch * iters / (start.elapsed_time(end) / 1000)

Sweep batch size and resolution with this harness and plot images per second; the knee of the curve is your serving batch size.

Serving: scheduling and caching encoder output

In a VLM server the encoder runs once per image, before the prefill that consumes it. vLLM's V1 engine keeps an encoder cache so outputs are not recomputed across scheduling steps, and its limit_mm_per_prompt setting caps how many images a single request may carry, which bounds worst-case encoder work. Across requests, the same image often appears many times (a product photo, a document in a multi-turn chat), so an application-level feature cache keyed on content pays off:

import collections, hashlib

class EncoderCache:
    def __init__(self, max_bytes):
        self.max_bytes, self.used = max_bytes, 0
        self.entries = collections.OrderedDict()

    @staticmethod
    def key(image_bytes, preprocess_cfg, model_rev):
        h = hashlib.sha256(image_bytes)
        h.update(repr(preprocess_cfg).encode())   # resolution, tiling, normalisation
        h.update(model_rev.encode())              # encoder weights and feature tap
        return h.hexdigest()

    def get(self, k):
        if k in self.entries:
            self.entries.move_to_end(k)
            return self.entries[k]

    def put(self, k, feats):
        size = feats.numel() * feats.element_size()
        while self.used + size > self.max_bytes and self.entries:
            _, old = self.entries.popitem(last=False)
            self.used -= old.numel() * old.element_size()
        self.entries[k] = feats
        self.used += size

Size it with arithmetic: 576 patch tokens at width 1024 in BF16 is about 1.2 MB per image before the connector, and 4.7 MB after projection into a 4096-wide LLM, so cache pre-projection features if the connector is cheap. Under heavy image traffic, encoders can run on separate GPUs so large images do not stall token generation.

Failure modes

  • Preprocessing skew. Wrong mean and standard deviation, RGB versus BGR, or a different resize filter than training. Outputs degrade quietly. Pin the processor configuration with the weights and test with a golden image and expected features.
  • Silent cropping. Center-crop preprocessing cuts text off wide screenshots. Inspect what the encoder actually sees.
  • Token blow-up. Tiling or native resolution turns one large upload into thousands of tokens and evicts other requests' KV cache. Enforce pixel and image-count caps at the gateway.
  • Unfused attention. A fallback to the math path materialises N² scores and runs out of memory at batch sizes that worked yesterday. Log which attention backend is active.
  • Stale caches. A cache key missing the model revision or preprocessing config serves features from the previous encoder after an upgrade.

Trade-offs

ChoiceGainCost
Higher fixed resolutionSmall text and fine detailTokens and FLOPs grow with the square of the side length
TilingUniform batches, works with any fixed encoderMany LLM tokens; seams between tiles
Native resolution with packingNo distortion, fewer wasted tokensVarlen kernels, harder batching and caching
2 x 2 token merge4x fewer LLM tokensCoarser spatial detail per token
Feature cacheRepeat images nearly freeMemory, invalidation discipline

What to do next

  1. Write down p, d, L and the input resolution policy of the encoder you run, and compute tokens and GFLOPs per image with the cost function above.
  2. Benchmark images per second across batch sizes and resolutions with CUDA events, and confirm a fused attention backend is active.
  3. Check the feature tap your consumer expects (layer index, CLS handling, pooling) against the code that produces features.
  4. Add a golden-image test that compares preprocessed tensors and encoder outputs against stored references on every deploy.
  5. Put caps on pixels and images per request, and add a content-hash feature cache whose key includes the model revision and preprocessing config.
  6. If images dominate traffic, measure whether moving the encoder to separate GPUs improves time to first token; see the video LLM article for how the same budget math scales to frames.
Key takeaway: A vision encoder's cost is set by tokens per image, N = (H/p)(W/p), times about 24d squared FLOPs per token per layer, plus an attention term that grows with N squared and matters at high resolution. Control resolution deliberately, pack or merge tokens instead of padding, extract the features your consumer was trained on, use fused attention, and cache features by content hash so no image is encoded twice.