A thumbnail is the most-viewed frame of any video: it appears on every browse row, search result and share card, usually long before anyone presses play. At catalogue scale, choosing it by hand stops working, and choosing it naively (the first frame, the midpoint) produces black frames, motion blur and closed eyes. A thumbnail pipeline is the system that turns every uploaded video into good still images automatically, in several shapes and sizes, cheaply enough to run on every asset and every re-encode.
This article builds that pipeline from first principles: what to extract, how to extract it without decoding the whole file when you do not need to, how to reject and score candidates, how to crop for different screens, how to generate the scrubbing previews players show, and how to run all of it as an idempotent job farm. ffmpeg filter names and defaults were checked against the ffmpeg-filters documentation on 2026-10-02.
Three products that share a name
"Thumbnail" means three different outputs with different requirements, and conflating them is the first design mistake.
| Output | Count per video | Quality bar | Consumer |
|---|---|---|---|
| Poster (hero) image | 1 to 5 candidates | Highest: chosen frame, sharp, well cropped | Browse rows, search, social cards |
| Aspect-ratio crops | One per surface (16:9, 1:1, 9:16, 2:3) | Subject must survive the crop | Feeds, mobile, TV tiles |
| Trick-play previews | One every few seconds | Low resolution, consistent spacing | Player seek bar |
Posters need selection: many candidates, a few winners. Trick-play needs coverage: a regular grid of small frames, no selection at all. Both should come from one pipeline run so that the expensive part, decoding, is paid once.
The architecture
The pipeline usually hangs off the same upload event that starts transcoding, but it should not depend on transcode output: run it against the mezzanine (the high-quality source) so poster quality does not inherit encoding artefacts, and so thumbnails are ready before the full encode ladder finishes. The transcoding farm itself is described in Video transcoding architecture.
Extracting candidates without decoding everything
Decoding is the dominant cost. A 90-minute 1080p title at 24 fps is about 130,000 frames; decoding all of them to examine a few hundred is wasteful. Compressed video stores occasional keyframes (independently decodable) and many predicted frames that need their references decoded first. That structure gives you three cheap candidate strategies and one expensive one.
# 1. Keyframes only: decoder skips everything else, so this is fast even on long titles.
ffmpeg -skip_frame nokey -i in.mp4 -vf "scale=640:-2" -fps_mode vfr -q:v 3 kf_%05d.jpg
# 2. Scene cuts: frames whose scene score exceeds 0.4 (docs suggest 0.3-0.5).
ffmpeg -i in.mp4 -vf "select='gt(scene,0.4)',scale=640:-2" -fps_mode vfr sc_%05d.jpg
# 3. "Most representative" frame per batch of 100 frames (filter default n=100).
ffmpeg -i in.mp4 -vf "thumbnail,scale=640:-2" -fps_mode vfr th_%05d.jpg
# 4. One frame at a known timestamp: -ss before -i seeks the input instead of decoding to it.
ffmpeg -ss 00:12:31.500 -i in.mp4 -frames:v 1 -q:v 2 poster_candidate.jpg
# 5. Trick-play sprite: one frame every 10 s, 160x90, 10x10 grid per sheet.
ffmpeg -i in.mp4 -vf "fps=1/10,scale=160:90,tile=10x10" -q:v 5 sprite_%03d.jpg
# 6. Black-frame report (defaults: amount=98 percent of pixels below threshold=32).
ffmpeg -i in.mp4 -vf "blackframe" -f null -- Keyframes only (
-skip_frame nokey): the decoder discards non-key frames, so cost scales with keyframe count. Encoders often place keyframes at scene cuts, which makes them reasonable candidates, but with a fixed GOP they are just evenly spaced. - Scene-change select:
select='gt(scene,0.4)'keeps frames whose scene-change score exceeds the threshold. It decodes everything but yields frames at the start of shots, which tend to be compositionally deliberate. Take a frame a little after the cut rather than the cut itself, to avoid transition blends. - The
thumbnailfilter picks the frame closest to the average of each batch ofnframes (default 100). It is a cheap way to avoid outliers like flashes; largerncosts memory. - Seek to timestamps:
-ssplaced before-iseeks in the input, so grabbing a frame at minute 40 does not decode minutes 0 to 40. Use this to fetch full-resolution versions of candidates you picked from a low-resolution pass.
A good default is a two-pass design: a cheap low-resolution pass (keyframes or scene cuts, scaled to 640 pixels wide) to find candidates, then exact seeks to extract the winners at full resolution. Skip the first and last few percent of duration, where logos, slates and credits live. On NVIDIA hardware, ffmpeg can decode with NVDEC and offers a thumbnail_cuda filter, which matters when the farm processes thousands of hours a day.
Rejecting bad frames
Most raw candidates are unusable: fades to black, motion blur, title cards, frames with burned-in subtitles, letterbox bars. Rejection is cheap and should run before any ML scoring. ffmpeg's blackframe filter flags frames where at least 98 percent of pixels fall below luma 32 by default, and blurdetect reports a no-reference blur metric per frame. In Python, three numbers catch most problems:
import numpy as np, cv2
def frame_quality(img):
"""Cheap rejection metrics on a decoded BGR candidate."""
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
return {
"mean": float(gray.mean()), # near 0 or 255: black / white
"sharpness": float(cv2.Laplacian(gray, cv2.CV_64F).var()), # low: blur or fade
"contrast": float(gray.std()), # low: flat title card
}
def keep(q, cfg):
return (cfg.min_mean < q["mean"] < cfg.max_mean
and q["sharpness"] > cfg.min_sharpness
and q["contrast"] > cfg.min_contrast)
def choose(cands, k=3, min_gap_s=30.0):
"""Highest score first, but never two picks closer than min_gap_s (diversity)."""
picks = []
for cnd in sorted(cands, key=lambda x: x.score, reverse=True):
if all(abs(cnd.t - p.t) >= min_gap_s for p in picks):
picks.append(cnd)
if len(picks) == k:
break
return picks
def crop_box(w, h, target_ratio, subject):
"""Largest window of target_ratio inside (w, h), centred on subject box, clamped."""
if w / h > target_ratio:
cw, ch = int(h * target_ratio), h
else:
cw, ch = w, int(w / target_ratio)
sx = (subject[0] + subject[2]) / 2
sy = (subject[1] + subject[3]) / 2
x = int(min(max(sx - cw / 2, 0), w - cw))
y = int(min(max(sy - ch / 2, 0), h - ch))
return x, y, cw, chThe variance of the Laplacian is a standard sharpness proxy: edges produce large second derivatives, blur flattens them. Calibrate thresholds per content type on a few hundred labelled frames; animation is sharp and flat-coloured, film grain inflates sharpness, and night scenes are legitimately dark. Detect letterboxing (ffmpeg's cropdetect helps) and crop the bars off before scoring, or every downstream crop will be off-centre.
Scoring, choosing and diversity
Surviving candidates are scored. Common signals: face presence, size and whether eyes are open; an aesthetic-quality model; text-overlay detection (penalise frames dominated by captions); and colourfulness. Combine them with a weighted sum whose weights you fit against a label you care about, such as editor picks or click-through on past titles. Keep the model and weight version with every pick, so you can explain and re-run any decision.
Choose with diversity. The choose function above enforces a minimum time gap, so the top three are not three frames of the same shot. Offer those candidates to an editor where one exists, and record overrides: they are your best training data.
Click-through testing between candidates is powerful and easy to get wrong. A thumbnail that over-promises raises clicks and lowers watch time, so measure a downstream metric such as completed views, not clicks alone. Policy also matters: avoid frames that misrepresent content or contain sensitive imagery, and keep a human review path for flagged titles.
Smart crops for every aspect ratio
A 16:9 frame cropped to 9:16 keeps barely a third of its width, so the crop position decides whether the subject survives. The approach is: find a subject box (faces first, then a saliency model or object detector), compute the largest window of the target ratio, centre it on the subject, and clamp it inside the frame. That is crop_box above.
Refinements that matter in practice: when there are several faces, centre on the union box if it fits, otherwise on the largest face; respect UI safe zones where your apps overlay titles or badges; and set a minimum crop size, falling back to a different candidate rather than upscaling a tiny region. Generate crops from the full-resolution extract, never from a smaller rendition.
Trick-play sprites
Seek-bar previews need a frame every few seconds at small size. Fetching hundreds of tiny images is slow, so frames are packed into sprite sheets with ffmpeg's tile filter (default grid 6x5; the example uses 10x10), and the player is told where each tile is.
There are three common ways to tell it. A WebVTT file whose cues point to a sprite region with a #xywh= media fragment is widely supported by web players, but it is a player convention rather than part of the HLS or DASH specifications:
WEBVTT
00:00:00.000 --> 00:00:10.000
sprite_001.jpg#xywh=0,0,160,90
00:00:10.000 --> 00:00:20.000
sprite_001.jpg#xywh=160,0,160,90DASH has a standard mechanism from the DASH-IF interoperability guidelines: an adaptation set with contentType="image" and an EssentialProperty with scheme http://dashif.org/thumbnail_tile whose value gives the grid (for example 10x10); each sheet is a segment and tile duration follows from segment duration. HLS offers I-frame playlists (EXT-X-I-FRAMES-ONLY), which let a player decode keyframes from the video itself instead of shipping images. Check which your target players support before choosing; many services ship both a VTT sprite track and the manifest-native form. Container-level details are in fragmented MP4 internals.
Encoding and delivery
Encode posters at a few widths (for example 320, 640, 1280, 1920) in JPEG plus WebP or AVIF, chosen by the client's Accept header or picture element. Name files by content: a hash of source, recipe version and parameters. Content-addressed names can be cached as immutable forever, and a change produces a new URL rather than a cache purge. Thumbnails are small but extremely numerous and hot, so they benefit from the same tiered caching as segments; see Video CDN architecture.
Orchestration at scale
Make each job idempotent with a key of (source content hash, recipe version). Re-running a finished job writes nothing new; bumping the recipe version (new scorer, new crop sizes) regenerates exactly the affected outputs. Write outputs to a staging prefix and publish the catalogue entry last, so a half-finished job never exposes broken images.
Shard long titles by time range for the scene-detection pass, with a few seconds of overlap so a cut at the boundary is not lost, then merge candidates. Keep ML scoring on a separate GPU worker pool fed by a queue of small candidate images, so decode workers and model workers scale independently. Track per-asset cost and duration, and alert on zero-candidate outcomes, which usually mean a corrupt source or an all-dark film needing different thresholds.
Worked example: a 90-minute film
Pass 1 decodes keyframes only at 640 pixels: with a 2-second GOP that is about 2,700 frames, plus a scene-cut pass yielding roughly 1,100 shot starts. After trimming the first and last 3 percent and deduplicating near-identical frames, about 1,500 candidates remain. Rejection removes about a third (dark scenes, blur, text-heavy frames). Scoring ranks the rest; diversity picks five, at least 30 seconds apart. Pass 2 seeks to those five timestamps and extracts full-resolution frames, which are cropped to 16:9, 1:1 and 9:16 and encoded at four widths and two formats: 120 files. In parallel the sprite pass produces 540 frames (one per 10 s) on six 10x10 sheets plus a VTT file. All counts here are illustrative arithmetic, not benchmarks; measure your own throughput.
Failure modes and trade-offs
- Spoilers: the best-scoring frame is the climax. Restrict poster candidates to the first half or let editors exclude ranges.
- Variable frame rate and odd timestamps: sprite timing drifts from the seek bar. Use
fpsfilter output timestamps, not frame counts. - HDR sources: extracting from HDR without tone mapping gives washed-out posters. Tone-map to SDR in the filter graph.
- Rotation metadata: phone video displays rotated; make sure extraction honours it.
- Trade-off: decoding everything (scene detection) finds better shots; keyframes-only is several times cheaper. Most catalogues use keyframes for the long tail and full analysis for titles that get promoted.
Per-title encoding decisions interact with keyframe placement and therefore with candidate quality; per-title encoding covers that side.
What to do next
- Separate your requirements into posters, crops and trick-play, and decide which surfaces need which ratios.
- Run the two-pass extraction on 50 representative titles and label usable versus unusable candidates.
- Calibrate the rejection thresholds on those labels per content type.
- Add face and saliency boxes, implement diversity-aware selection and the crop window, and review the results with an editor.
- Generate sprite sheets with a VTT track and, for DASH, a thumbnail_tile adaptation set; test seeking on each target player.
- Key every job by source hash and recipe version, publish atomically, and track cost per asset.