AI video editing sounds like one feature, but in production it is a pipeline of quite ordinary components with models plugged into a few of them. A user deletes words from a transcript and the video cuts itself; a request such as "make a 60-second highlight of the product demo" becomes a rough cut; a microphone stand vanishes from a shot. Behind each of those sits the same architecture: analyse the footage once into a searchable index, express every edit as an operation on a timeline, and only touch pixels at render time.
This article explains that architecture from first principles. It covers why time representation is the hard part, what the analysis layer produces, why the timeline rather than the video file is the source of truth, how text-based editing turns word timestamps into frame-accurate cuts, how to let a language model plan edits without letting it invent timestamps, where generative effects belong, and how rendering decides between copying and re-encoding. It ends with failure modes and a checklist.
The core idea: models write operations, renderers write pixels
A traditional editor is non-destructive: the source media never changes, and the project is a list of decisions, which clip, from which frame to which frame, on which track, with which effect. Rendering applies that list to the sources. AI editing keeps this design and adds two things: analysis, which turns footage into data a program can reason about, and planning, which turns a goal into a list of operations.
Keeping the models on the operation side has large practical benefits. Every AI suggestion is reviewable and reversible, because it is a diff to the timeline. A preview can render from cheap proxies in seconds while the final render uses full-quality sources. And expensive generative work runs only for the frames that survive the edit.
Time is the hard part
Every component must agree on what a moment in the video is, and that is harder than it looks. Video is a sequence of frames at a rate such as 30000/1001 (the 29.97 fps used in North American broadcast), not 29.97 exactly. Audio has its own sample rate, typically 48,000 per second. Containers stamp each packet with a presentation timestamp in a timebase, and phone footage is often variable frame rate, so frame number and time are not interchangeable.
Represent every time in the system as a rational number tied to a rate, never as a float in seconds that drifts with rounding. A worked example: a word ends at 12.345 s in a 30000/1001 video. That is 12.345 x 30000 / 1001 = 369.98 frames, so the cut belongs at frame 370, which starts at 370 x 1001 / 30000 = 12.3457 s. Snap to frames once, at the edge of the system, and store integers after that. Normalise variable-frame-rate sources to a constant rate at ingest, or every downstream tool will disagree about where frame 370 is.
Ingest and proxies
Ingest probes each upload with ffprobe for codec, rate, duration, rotation and audio layout, and rejects or normalises the surprises. It then writes two derivatives. The mezzanine is a high-quality, constant-frame-rate copy used for the final render. The proxy is a small, low-resolution file encoded with every frame as a keyframe, so scrubbing and preview renders can seek anywhere instantly. Long-GOP delivery codecs are slow to seek because decoding a frame requires the frames it depends on; GOP structure explains why.
The analysis layer
Analysis runs once per asset and writes time intervals with IDs into an index. The useful set for editing is small.
| Signal | Typical tool | Used for |
|---|---|---|
| Shot boundaries | PySceneDetect content detector | Cut points that look natural; B-roll selection |
| Words with start and end times | Speech recognition with word timestamps | Text-based editing, captions, search |
| Speaker turns | Diarization | Interview edits, per-speaker cuts |
| Silence and loudness | ffmpeg silencedetect, loudness measurement | Removing dead air, level matching |
| Faces, objects, text on screen | Detection and tracking models | Reframing, masks for effects, search |
| Embeddings per shot | Image or video-text encoders | "Find the shot where..." queries |
from scenedetect import detect, ContentDetector
import whisper
shots = detect("talk_proxy.mp4", ContentDetector(threshold=27.0))
shot_rows = [(i, s.get_frames(), e.get_frames()) for i, (s, e) in enumerate(shots)]
asr = whisper.load_model("small").transcribe("talk.wav", word_timestamps=True)
words = [(w["word"].strip(), w["start"], w["end"])
for seg in asr["segments"] for w in seg["words"]]
# ffmpeg -i talk.wav -af silencedetect=noise=-35dB:d=0.5 -f null -
# parse "silence_start" / "silence_end" lines from stderr into intervalsStore word times as reported, then convert to frames with the asset's rate. Recognised words are probabilistic: timestamps can be off by a fraction of a second, especially at word boundaries after silence, which is why the cutting step below pads and snaps instead of trusting them exactly. The same transcript feeds captions, so build it once.
The timeline as source of truth
The edit is a data structure, not a file. OpenTimelineIO (OTIO), an open interchange format from the Academy Software Foundation, models it well: a timeline holds tracks, tracks hold clips, gaps and transitions, and each clip points at media with a source range in rational time. Using it also lets a rough cut leave your system for a professional editing application through OTIO adapters.
import opentimelineio as otio
RATE = 30000 / 1001
def clip(name, url, start_frame, n_frames):
return otio.schema.Clip(
name=name,
media_reference=otio.schema.ExternalReference(target_url=url),
source_range=otio.opentime.TimeRange(
start_time=otio.opentime.RationalTime(start_frame, RATE),
duration=otio.opentime.RationalTime(n_frames, RATE)))
tl = otio.schema.Timeline(name="rough_cut_v3")
v1 = otio.schema.Track(name="V1", kind=otio.schema.TrackKind.Video)
tl.tracks.append(v1)
for i, (s, e) in enumerate(kept_ranges): # frame ranges from the planner
v1.append(clip(f"keep_{i:03d}", "file:///media/talk_mezz.mov", s, e - s))
otio.adapters.write_to_file(tl, "rough_cut_v3.otio")Version every timeline. A model suggestion becomes a new version with a diff the user can accept or reject, and undo is free. Keep audio on its own tracks so audio edits such as J-cuts and L-cuts, where sound leads or trails the picture, remain possible.
Text-based editing, worked through
The user deletes "um, so, basically" and a repeated sentence from a transcript. The system must turn kept words into kept frame ranges that sound natural. The algorithm:
def kept_ranges(words, deleted_ids, rate, pad_s=0.06, min_gap_s=0.25, min_keep_s=0.4):
spans = []
for i, (text, start, end) in enumerate(words):
if i in deleted_ids:
continue
s, e = start - pad_s, end + pad_s
if spans and s - spans[-1][1] < min_gap_s: # join words separated by tiny gaps
spans[-1][1] = e
else:
spans.append([s, e])
spans = [sp for sp in spans if sp[1] - sp[0] >= min_keep_s]
return [(round(s * rate), round(e * rate)) for s, e in spans] # snap once, to framesEach parameter exists because of a failure seen in practice. Padding avoids clipping consonants that the recogniser timed slightly late. Joining short gaps avoids a stutter of micro-cuts inside a sentence. Dropping very short keeps avoids flash frames. At render time, apply a short audio crossfade of a few milliseconds at every cut, because joining two waveforms at arbitrary samples produces an audible click. If a cut lands mid-shot on a talking head, the visual jump is unavoidable; offer to cover it with B-roll or a slight zoom rather than pretending it is seamless.
Planning edits with a language model
For goal-level requests such as "cut a 60-second highlight", a language model plans. The safe design gives the model the index, not the video, and constrains its output to operations that reference IDs the index contains.
{"ops": [
{"op": "keep_words", "from_word": 812, "to_word": 870, "reason": "demo of export"},
{"op": "keep_shot", "shot": 41, "trim_head_frames": 12},
{"op": "insert_broll", "shot": 77, "over_words": [845, 860]},
{"op": "title_card", "text": "Exporting in one click", "at_word": 812}
]}A deterministic validator then checks every reference exists, ranges are ordered, the total duration meets the target within tolerance and no op cuts through a word. Only validated ops become a timeline version. The rule that matters: the model never emits raw timestamps, because it will invent plausible ones that do not match the footage. Referencing word and shot IDs makes hallucinated edits fail validation instead of producing a wrong cut, and makes every decision explainable by pointing at the words it kept.
Generative effects belong at render time
Removing an object, replacing a background, extending a shot by a second or upscaling footage changes pixels, so these are effects attached to clips on the timeline, with their parameters: a mask track, a prompt, a model version and a seed. The renderer computes them only for the frames that survive the edit and caches results by a hash of inputs and parameters, so re-rendering after an unrelated change does not pay again.
The engineering problem is temporal consistency. Per-frame image models flicker; video models or approaches that propagate results along motion are needed, and long shots are processed in overlapping windows that are blended. Budget them explicitly: generative frames cost orders of magnitude more compute than decoding. Video super-resolution covers the motion-aligned techniques that also apply here. Record that frames were generated, for example with C2PA content credentials, so downstream viewers and platforms can tell.
Rendering: copy or re-encode
Compressed video can only be cut cleanly at keyframes without decoding. ffmpeg -ss 12.3457 -i in.mp4 -t 30 -c copy out.mp4 is fast because it copies packets, but the output can only begin at a keyframe at or before the requested time, so the cut is inexact and players may show frozen or garbled frames until the next keyframe. Re-encoding is frame-accurate but slow and costs a generation of quality.
Production renderers combine the two. Preview renders re-encode from all-intra proxies, which is fast because every frame is a keyframe. Final renders re-encode from the mezzanine, split into segments at kept-range boundaries and encoded in parallel on many workers, then concatenated; segment boundaries must be closed GOPs with matching encoder settings so concatenation is clean. A smart render re-encodes only the GOPs touched by cuts and copies the rest, which is fast but only works when the copied stream already matches the target codec and settings exactly. Encoder settings and the quality cost of re-encoding are covered in video encoder architecture.
Failure modes
- Audio drifts out of sync over a long render. Usually a variable-frame-rate source or float-seconds arithmetic; normalise at ingest and use rational time everywhere.
- Clipped words and clicks. Cuts placed exactly on recogniser timestamps without padding or crossfades.
- Flash frames. Kept ranges of one or two frames from over-eager deletion; enforce a minimum keep length.
- Hallucinated edits. A planner allowed to output timestamps; force ID references and validate.
- Flicker in generative effects applied frame by frame; use temporally consistent methods and overlapping windows.
- Inexact cuts from stream copy on long-GOP sources; re-encode cut points.
- Stale caches after a model upgrade; include the model version in every effect's cache key.
Operational guidance and trade-offs
Analysis is a one-time cost per asset and should run asynchronously after upload, with a progress state the editor can show; editing features unlock as each signal lands. Store proxies close to the editing clients and mezzanines close to the render farm. Measure render time per output minute and cost per generative second, and gate the expensive effects behind an explicit user action. For QC, check duration against the timeline, decode every output frame once, measure integrated loudness and compare a few frames at cut points with the timeline's expectation.
The central trade-off is automation versus control. Fully automatic edits are fast but generic; the timeline-of-operations design lets you start automatic and let a human finish, which is what most users actually want.
What to do next
- Pick one rate representation (rational, per asset) and convert every timestamp in your system to it; normalise variable-frame-rate uploads at ingest.
- Generate an all-intra proxy and a mezzanine for each upload, and build the analysis index with shots, word-timed transcript and silence intervals.
- Model edits as an OpenTimelineIO timeline with versioning; implement text-based cutting with padding, gap joining, minimum keep length and audio crossfades.
- If you add an LLM planner, define an ID-referencing op schema and a validator before writing the prompt.
- Implement preview renders from proxies and segment-parallel final renders from the mezzanine; test cut accuracy frame by frame.
- Add generative effects last, as cached, versioned clip effects with provenance labels.