Transcoding is the step between the file a studio or user uploads and the dozen files a player actually streams. For video on demand, the input is a mezzanine: a high-quality master, often large, sometimes oddly made. The output is a ladder of renditions at different resolutions and bitrates, each with keyframes in exactly the same places so a player can switch between them, plus audio tracks, ready for packaging into HLS and DASH.
Running one ffmpeg command per rendition works for a demo and fails at scale: a feature film at a slow preset takes hours per rendition, one bad input stalls the queue, and a crash at 90 percent wastes everything. This article describes the architecture that replaces that command: probing and validation, a plan that splits work into GOP-aligned chunks, many workers encoding in parallel, stitching and verification, quality control, and an orchestration layer whose tasks can be retried safely. Codec internals and ladder design are covered elsewhere and linked where they matter.
The pipeline at a glance
Ingest stores the upload and records a SHA-256 of the bytes, which becomes the identity of the source for everything downstream. Probe reads the container and stream metadata and validation applies policy. Plan decides the ladder and cuts the timeline into chunks. Encode workers each produce one chunk of one rendition. Stitch joins chunks; verify checks frame counts and keyframe positions. QC measures quality and looks for defects. Packaging turns renditions into segments and manifests, and publish makes them visible. Audio follows its own path, as a whole-file job, for reasons explained below.
Probe and validate the mezzanine
Most transcoding incidents start with an input that the pipeline silently mishandled: variable frame rate from a phone, interlaced broadcast material, a rotation flag, HDR metadata, an extra audio stream with commentary, or a video track shorter than the audio. Decide policy for each case explicitly and apply it before any encoding. Rejecting is acceptable; guessing is not.
import json, subprocess
def probe(path: str) -> dict:
out = subprocess.run(
["ffprobe", "-v", "error", "-print_format", "json",
"-show_format", "-show_streams", path],
check=True, capture_output=True, text=True).stdout
return json.loads(out)
def validate(info: dict) -> list[str]:
"""Return policy violations; an empty list means the mezzanine is accepted."""
problems = []
video = [s for s in info["streams"] if s["codec_type"] == "video"]
audio = [s for s in info["streams"] if s["codec_type"] == "audio"]
if len(video) != 1:
problems.append(f"expected 1 video stream, found {len(video)}")
return problems
v = video[0]
if v.get("r_frame_rate") != v.get("avg_frame_rate"):
problems.append("variable frame rate: normalise to CFR before chunking")
if v.get("field_order") not in (None, "progressive", "unknown"):
problems.append(f"interlaced ({v['field_order']}): deinterlace in the plan")
if not audio:
problems.append("no audio stream")
rotation = [d for d in v.get("side_data_list", []) if "rotation" in d]
if rotation:
problems.append("rotation metadata present: apply it explicitly")
if v.get("color_transfer") in ("smpte2084", "arib-std-b67"):
problems.append("HDR transfer: route to the HDR ladder, never tone-map silently")
return problemsStore the probe output with the job, and write normalisation decisions such as constant frame rate or deinterlacing into the plan as filter settings, not hidden pre-processing.
Why chunk, and where the cuts must go
Encoding time scales with duration, so the only way to make a two-hour title finish in minutes is to split the timeline and encode pieces in parallel. The constraint is that the stitched result must be indistinguishable, to a player, from a continuous encode. That means two things. Each chunk must start with an IDR frame and contain only closed GOPs, so no frame refers to a frame in another chunk. And chunk boundaries must fall exactly where the continuous encode would have placed keyframes, so that segment boundaries line up across every rendition.
The simplest rule that satisfies both is a fixed GOP in seconds, for example 2 seconds, and chunks that are a whole number of GOPs. Scene-cut keyframes are disabled so the encoder never inserts an extra keyframe that one rendition has and another does not. Packaging later cuts segments of a whole number of GOPs, for example 4 or 6 seconds, and every rendition's segments then cover the same time range, which is what lets a player switch rendition at any segment boundary. CMAF architecture explains why the packager depends on this alignment.
from dataclasses import dataclass
import hashlib, json
@dataclass(frozen=True)
class Chunk:
index: int
start_frame: int
frames: int
def plan_chunks(total_frames: int, fps_num: int, fps_den: int,
gop_seconds: int = 2, chunk_gops: int = 30) -> list[Chunk]:
"""Chunk boundaries fall on GOP boundaries, so every chunk starts with an IDR
at exactly the position a continuous encode would have placed one."""
gop = gop_seconds * fps_num // fps_den # assumes an integer-frame GOP
assert gop * fps_den == gop_seconds * fps_num, "GOP must be a whole number of frames"
size = gop * chunk_gops
return [Chunk(i, s, min(size, total_frames - s))
for i, s in enumerate(range(0, total_frames, size))]
def task_key(source_sha256: str, rendition: dict, chunk: Chunk, encoder: str) -> str:
"""Idempotency key: same inputs, same output object, so retries are free."""
blob = json.dumps([source_sha256, rendition, chunk.start_frame, chunk.frames, encoder],
sort_keys=True)
return hashlib.sha256(blob.encode()).hexdigest()
Encoding a chunk
A worker receives a source reference, a start frame, a frame count and a rendition. It seeks into the mezzanine, decodes from the nearest earlier keyframe, discards frames up to the exact start, and encodes exactly the requested number of frames. Putting -ss before -i makes the seek fast, and because the stream is being re-encoded ffmpeg still lands on the exact time. Converting the start frame to a time requires an exact rational frame rate, which is another reason to normalise variable-frame-rate sources first.
# One chunk of the 720p rendition, 24 fps source, 2 s GOP = 48 frames, chunk = 1440 frames.
# -ss before -i seeks fast; when transcoding, ffmpeg decodes and discards up to the exact time.
ffmpeg -ss 120.000 -i mezzanine.mov -frames:v 1440 -an \
-vf "scale=1280:720:flags=lanczos,format=yuv420p" \
-c:v libx264 -preset slow -profile:v high \
-crf 21 -maxrate 3000k -bufsize 6000k \
-x264-params "keyint=48:min-keyint=48:scenecut=0:open-gop=0" \
-force_key_frames "expr:gte(t,n_forced*2)" \
chunk_0002_720p.mp4
# Audio once, over the whole programme, never per chunk.
ffmpeg -i mezzanine.mov -vn -c:a aac -b:a 128k -ac 2 audio_stereo.m4a
# Stitch: the concat demuxer with stream copy, valid only because every chunk
# used identical encoder settings. list.txt holds lines like: file 'chunk_0002_720p.mp4'
ffmpeg -f concat -safe 0 -i list.txt -c copy video_720p.mp4The x264 parameters pin the GOP to 48 frames, disable scene-cut keyframes and forbid open GOPs; -force_key_frames adds a keyframe every 2 seconds as a belt-and-braces measure. The rate control is capped CRF: constant quality, with a VBV ceiling so the stream respects a maximum bitrate that the ABR algorithm and CDN can rely on. The choice of codec, preset and tuning belongs with video encoder architecture; the transcoding system's job is to apply the same settings identically on every chunk.
One consequence of chunking is that rate control restarts at every chunk boundary. The encoder in chunk 3 knows nothing about the buffer state at the end of chunk 2, so VBV compliance across the joins is not guaranteed by construction, and quality can dip slightly at the start of each chunk. Longer chunks reduce the number of joins; verify VBV compliance on the stitched output, not per chunk.
Audio is not chunked
Compressed audio codecs such as AAC prepend encoder delay, known as priming samples, and work in fixed-size frames that do not line up with video frame boundaries. Encode audio per chunk and every join gets a small gap or overlap, which listeners hear as clicks and which accumulates into lip-sync drift. Audio encoding is cheap compared with video, so encode each audio track once over the full programme on a single worker. Do loudness normalisation in the same job if your platform applies it.
Stitch and verify
Because every chunk of a rendition was encoded with identical parameters, the concat demuxer can join them with stream copy, which takes seconds and introduces no further loss. Then verify, because a stitched file that plays is not the same as a correct one. Check that the frame count equals the planned total, that keyframes sit exactly at multiples of the GOP, that timestamps increase monotonically with no gaps, and that the video duration matches the audio duration within a frame. A worker that encoded one frame too few produces a file that looks fine and drifts out of sync by the end.
Quality control
QC has two halves. The first is a quality score: VMAF compares a rendition with the mezzanine. The libvmaf filter takes the distorted stream first and the reference second, and both must have the same resolution, so scale the rendition up to the reference size before comparing. Swapping the inputs does not crash; it produces a plausible, wrong number. Scoring every frame of every rendition is expensive, so many systems score a sample of segments per title and fully score only flagged ones.
# VMAF: distorted FIRST, reference SECOND, both at the same resolution.
ffmpeg -i video_720p.mp4 -i mezzanine.mov -lavfi \
"[0:v]scale=1920:1080:flags=bicubic,setpts=PTS-STARTPTS[dist];\
[1:v]setpts=PTS-STARTPTS[ref];[dist][ref]libvmaf=log_fmt=json:log_path=vmaf.json" \
-f null -
# Defects a metric average hides: black runs, frozen runs, silent runs.
ffmpeg -i video_720p.mp4 -vf "blackdetect=d=2,freezedetect=d=5" -an -f null -
ffmpeg -i audio_stereo.m4a -af "silencedetect=d=3" -f null -The second half is defect detection, because an average score hides a two-second black hole or a frozen scene. The blackdetect, freezedetect and silencedetect filters report runs, which you compare against the source: black in the output where the source has picture is a defect; black in both is the director's choice. Per-title ladder decisions that use these quality scores are described in per-title encoding.
Orchestration: idempotent tasks and cost
Every task writes to an object key derived from its inputs: the source hash, the rendition settings, the chunk range and the encoder version. If a worker dies, the task is simply re-run; if two workers run the same task, they write the same bytes to the same key. A task whose output already exists and passes a checksum is skipped. That makes spot or preemptible capacity safe, because losing a machine costs only its in-flight chunks.
Use priority queues, cap chunks in flight per title so one film cannot starve the rest, track cost as encoder-seconds per rendition, and quarantine titles that fail policy or QC instead of retrying forever.
Worked example: a 90-minute film
The source is 90 minutes at 24 frames per second: 129,600 frames. With a 2-second GOP of 48 frames and chunks of 30 GOPs, each chunk is 1,440 frames or 60 seconds, giving 90 chunks. A ladder of six renditions makes 540 video tasks, plus one audio task per track. If a slow-preset 1080p chunk takes ten minutes on one worker, serial encoding of that rendition takes fifteen hours; with 90 workers it takes about ten minutes plus scheduling overhead. These timings are illustrative; measure your own encoder-seconds per chunk.
When the worker for chunk 37 of the 1080p rendition is preempted, only that task re-runs. When the team later changes the 480p bitrate cap, the task keys for 480p change and only the 90 tasks for 480p run again; the other 450 outputs are reused. Packaging then cuts 6-second segments, three GOPs each, identically for all renditions. Deferring packaging until request time is an option covered in just-in-time packaging.
Failure modes
- Misaligned keyframes. Scene-cut detection left on in one rendition adds keyframes the others lack; segments drift apart and players stall when switching.
- Open GOPs at chunk joins. Leading frames reference the previous chunk, which does not exist in this file; the join shows a corrupt frame.
- Frame count off by one. Rounding a fractional frame rate when converting frames to seconds drops or duplicates a frame per chunk; audio drifts by the end.
- Per-chunk audio. Priming gaps at every join cause clicks and drift.
- Silent normalisation. HDR tone-mapped, or interlaced material encoded without deinterlacing, because no one decided policy.
- Swapped VMAF inputs. A plausible score that measures the wrong thing.
- Non-idempotent outputs. Tasks writing to random or timestamped keys make retries double the storage and make a partial output look complete.
Trade-offs
Smaller chunks give more parallelism and faster turnaround but more joins, more rate-control restarts and more scheduling overhead. A fixed GOP gives perfect alignment but places keyframes in the middle of scenes, slightly hurting efficiency compared with scene-aware placement. Capped CRF gives consistent quality with a hard ceiling, while two-pass VBR hits a target size more exactly at the cost of a second pass. Sampling VMAF is cheap but can miss a localised failure, which is why defect detectors run on everything. Live transcoding inverts most of these choices because latency, not throughput, dominates; see live streaming architecture.
What to do next
- Write your mezzanine policy: what you reject, what you normalise, and how, and run the probe and validate code on your last hundred uploads.
- Fix a GOP length in seconds and confirm it is a whole number of frames for every frame rate you accept.
- Disable scene-cut keyframes and open GOPs in every rendition, then check keyframe positions across renditions on one title.
- Move audio to a single whole-file job per track.
- Derive every output key from the task inputs and make workers skip tasks whose outputs already verify.
- Add frame-count, keyframe and duration checks after stitching, plus blackdetect, freezedetect and silencedetect on every rendition.
- Audit your VMAF command for input order and resolution.