A video LLM answers questions about a clip: what happened, when, in what order. Architecturally it is an image-text model with a time axis bolted on, but operationally it is a different beast. A single image costs a few hundred to a few thousand tokens. A ten-minute clip, sampled naively, costs hundreds of thousands. Almost every engineering decision in a video LLM system is a decision about which of those tokens to keep and where to spend the GPU time.

This article follows a clip through the pipeline, from compressed bytes to the language model's KV cache, and puts numbers on each stage using a real, published model shape. It covers how frames become tokens, how time is encoded, what dominates latency and memory, how serving and training pipelines are built, and where they break. For the general image-text architecture, read vision language models first; this page assumes it and focuses on what video changes.

Advertisement

The pipeline, stage by stage

ContainerMP4 / H.264 / HEVCDecodeNVDEC or CPUSample + resizefps, max pixelsPatchify2 x 14 x 14 cubesVision encoderViT, 32 layersMerge 2x2MLP projectorInterleave with texttime-aware positionsLLM prefillbuilds KV cache for all tokensLLM decodeone token per stepEmbedding cachekey: video hash + fps + sizestore / reuseCost centres: decode (CPU or NVDEC), encoder FLOPs per frame, prefill FLOPs over all visual tokens, KV memory per token.
A clip passes through decode, sampling, patchification, a vision encoder and a token-merging projector before the LLM ever sees it. Each stage is a separate budget.

Decode. Video arrives compressed (H.264, HEVC, AV1). The model needs raw RGB frames, so something has to decode. On a CPU, decoding high-resolution video is expensive enough to starve a GPU; data-centre GPUs include fixed-function decoders (NVDEC on NVIDIA parts) that decode without using CUDA cores. Decoding also has a structural trap: you can only seek to keyframes, so pulling 2 frames per second out of a stream with a keyframe every 10 seconds still decodes most frames in between.

Sample and resize. The model does not consume every frame. A sampler picks frames (fixed fps, a fixed count spread uniformly, or content-aware selection) and resizes them to a pixel budget. These two knobs, frames and pixels per frame, are the main controls on everything that follows.

Encode. A vision transformer turns each frame, or each short group of frames, into patch embeddings. Some models patchify in 3D: a patch spans two consecutive frames as well as a 14 by 14 pixel square, which halves the token count for video at once.

Compress and project. Neighbouring patch embeddings are merged (for example 2 by 2 into one) by a small MLP that also maps them into the LLM's embedding space. Other designs use pooling or a learned resampler that emits a fixed number of tokens per frame.

Language model. Visual tokens are interleaved with the text prompt, the LLM runs prefill over the whole sequence, building a KV cache, then decodes the answer one token at a time.

Worked example: tokens, FLOPs and memory for one minute of video

Numbers make the trade-offs concrete. Take the published configuration of Qwen2-VL-7B-Instruct. Its vision encoder has 32 layers of width 1280, patch size 14, temporal patch size 2 and spatial merge size 2. Its language model has 28 layers, 28 attention heads and 4 key-value heads over a hidden size of 3584, so the head dimension is 128. Everything below is arithmetic on those numbers; throughput figures are assumptions you should replace with your own measurements.

Sample one minute at 2 fps and resize to 448 by 448. That is 120 frames, grouped in pairs into 60 temporal patches. Each pair yields a 32 by 32 grid (448 / 14 = 32), so 1,024 encoder tokens per pair and 61,440 encoder tokens in total. The 2 by 2 merge divides by four: 256 LLM tokens per pair, 15,360 visual tokens for the minute.

KV cache per token is 2 (keys and values) x 28 layers x 4 KV heads x 128 dims x 2 bytes in BF16 = 57,344 bytes, 56 KiB. The minute of video therefore costs about 0.82 GiB of KV cache per sequence. Ten minutes at the same settings is 153,600 tokens and roughly 8.2 GiB, per request, before the batch dimension. Without grouped-query attention (28 KV heads instead of 4) it would be seven times larger.

Compute follows the rule of thumb of about 2 FLOPs per parameter per token for a forward pass. The roughly 7.6B-parameter language model spends about 2 x 7.6e9 x 15,360, or 2.3e14 FLOPs, on prefill. A 32-layer, 1280-wide encoder has on the order of 0.6 to 0.7B parameters, so its 61,440 tokens cost about 8e13 FLOPs plus attention. If a GPU sustains, say, 400 TFLOP/s on these kernels, prefill is around 0.6 s and encoding about 0.2 s. Decoding a 100-token answer is cheap in FLOPs but memory-bound, since each step reads the weights and the whole KV cache.

KnobEffect on visual tokensWhat you lose
fps 2 to 1HalvesEvents shorter than about a second
448 to 336 pixels squareDrops to about 56 percentSmall text, distant objects
Merge 2x2 to 4x4QuartersFine spatial detail
Fixed frame count (e.g. 32)Constant per clipTemporal density on long clips

The lesson is that video cost is quadratic in nothing and linear in everything: frames, pixels per frame and merge ratio multiply. Attention adds a quadratic term over sequence length, which is why fused attention kernels matter at these lengths and why KV-cache sizing decides how many concurrent requests fit.

Advertisement

Encoding time: positions that know about frames

Text positions are one-dimensional. A visual token has a frame index, a row and a column. If the model flattens them into a single counter, it can still learn ordering, but it confuses spatial adjacency with temporal adjacency and degrades on long clips. Two families of fixes exist.

Multi-axis rotary embeddings. Qwen2-VL's configuration uses rope type "mrope" with sections [16, 24, 24]: the rotary frequency pairs of each head are split into a temporal block and two spatial blocks, so a token's rotation depends on (time, row, column). Text tokens use the same index on all three axes, which reduces to ordinary RoPE. Qwen2.5-VL's report describes aligning the temporal axis to absolute time rather than frame index, so the model can reason about seconds even when the sampling rate varies.

Explicit timestamps. Simpler models insert text such as "<t=12.5s>" before each frame's tokens. This costs a few tokens per frame and works with any LLM, but the model must learn to read numbers as time.

Either way, record the real timestamp of every sampled frame. Variable-frame-rate phone video breaks the assumption that frame k sits at k / fps seconds, and a model asked "when" will answer confidently and wrongly.

Serving: a pipeline, not a forward pass

A production video LLM service is mostly plumbing around two GPU stages. The sketch below shows the shape; the model calls are placeholders for whatever runtime you use.

import hashlib, av, numpy as np

def sample_frames(path, fps=2.0, max_frames=256):
    """Decode once, keep frames on a time grid, return (frames, timestamps)."""
    out, ts, next_t, step = [], [], None, 1.0 / fps
    with av.open(path) as box:
        stream = box.streams.video[0]
        stream.thread_type = "AUTO"
        for frame in box.decode(stream):
            t = float(frame.pts * stream.time_base)     # real time, survives VFR
            if next_t is None:
                next_t = t                              # first pts need not be 0
            if t + 1e-6 >= next_t:
                out.append(frame.to_ndarray(format="rgb24"))
                ts.append(t)
                while next_t <= t:                      # skip past gaps, no bursts
                    next_t += step
                if len(out) == max_frames:
                    break
    if len(out) % 2:                                    # temporal patch of 2
        out.append(out[-1]); ts.append(ts[-1])
    return np.stack(out), ts

def video_key(path, fps, size):
    h = hashlib.sha256(open(path, "rb").read()).hexdigest()
    return f"{h}:{fps}:{size}"

def answer(path, question, cache, encoder, llm, fps=2.0, size=448):
    key = video_key(path, fps, size)
    vis = cache.get(key)
    if vis is None:
        frames, ts = sample_frames(path, fps)
        vis = encoder.encode(resize(frames, size), timestamps=ts)  # GPU stage 1
        cache.put(key, vis)                                        # reuse across questions
    return llm.generate(visual=vis, prompt=question)               # GPU stage 2

Three design points carry most of the value. First, cache visual embeddings keyed by content hash and sampling settings; users ask several questions about the same clip, and re-encoding is pure waste. Second, use prefix caching in the LLM runtime with the video tokens placed before the question, so follow-up questions reuse the KV cache instead of repeating prefill. Third, separate the stages: decoding and encoding can run on a different pool from LLM decode, because encoding is compute-bound and bursty while decode is memory-bound and steady.

Admission control must count tokens, not requests. A scheduler that admits sixteen requests per GPU will run out of KV memory the first time sixteen users upload ten-minute videos. Estimate tokens from duration, fps and resolution before admitting, and reject or downsample clips that exceed the budget.

Training: where the time actually goes

Video LLMs are usually trained in stages: align the projector with the encoder and LLM frozen, then unfreeze progressively on image and video instruction data, then extend to longer clips. The modelling is familiar. The systems problems are not.

  • The data loader is the bottleneck. Decoding video on CPU workers often cannot keep a GPU busy. Pre-extract frames at the training fps into a sharded store, or decode on the GPU, and measure GPU utilisation before blaming the model.
  • Sequence length varies wildly. A 5-second clip and a 3-minute clip in one batch waste most of the padding. Pack several samples into one sequence with attention masks that stop them seeing each other, and bucket by token count.
  • Activation memory scales with tokens. Long clips need activation checkpointing and, past a point, sequence parallelism across GPUs. See long-context strategies for the options and their costs.
  • Mask the loss. Compute loss only on answer tokens. Visual tokens are inputs, and training the LLM to predict them wastes capacity.
  • Freeze carefully. Unfreezing the vision encoder on a small video set often degrades image skills. Use a lower learning rate for the encoder and keep image data in the mix.

Failure modes you will see

SymptomLikely causeFix
Misses a brief eventSampling interval longer than the eventRaise fps on short clips; shot-boundary or motion-aware sampling
Wrong "when" answersTimestamps assumed from frame indexUse container timestamps; feed absolute time
OOM on some requests onlyLong uploads blow the KV budgetToken-based admission, max frames, downsampling
GPU idle at 30 to 50 percentCPU decode or resize bottleneckNVDEC, pre-extraction, more workers, cache
Describes a typical scene, not this oneToo few tokens per frame; language prior dominatesMore pixels per frame for detail questions; grounded evaluation
Order of events swappedWeak temporal encoding or shuffled packingCheck positions per frame; test with reversed-clip probes

A cheap and effective evaluation trick: play the clip backwards and ask the same ordering question. A model that gives the same answer is not using time.

Trade-offs to decide explicitly

Uniform versus adaptive sampling. Uniform sampling is predictable and easy to cache; adaptive sampling (more frames where motion or scene changes are high) spends tokens better but makes cost per clip harder to forecast.

Many small frames versus few large ones. Activity recognition wants temporal density; reading a whiteboard wants pixels. Some systems run two passes: a cheap low-resolution pass over the whole clip to find the interesting window, then a high-resolution pass over that window.

One long context versus retrieval. For hour-long video, splitting into segments, captioning or embedding each and retrieving the relevant ones before asking the LLM is usually cheaper and more accurate than one enormous context. The cost is that questions spanning the whole video, such as "how many times did X happen", need aggregation logic outside the model.

What to do next

  1. Write a token calculator for your model: frames x tokens per frame after merging, then KV bytes per token from its config. Run it on your real clip-length distribution.
  2. Measure decode throughput separately from GPU stages; switch to NVDEC or pre-extraction if the GPU waits.
  3. Store real timestamps with every sampled frame and pass them to the model.
  4. Add a content-hash embedding cache and enable prefix caching for multi-question sessions.
  5. Admit requests by estimated tokens, with a hard cap on frames per request.
  6. Build an evaluation set with short events, ordering questions and reversed-clip probes before tuning fps or resolution.
Key takeaway: A video LLM is an image-text model whose cost is set by three multiplied knobs: frames, pixels per frame and the merge ratio. Do the token and KV arithmetic for your model before choosing them, keep real timestamps, keep the GPU fed with hardware decode and caching, and admit work by tokens. Most failures are sampling and plumbing problems, not model problems.