Serving a text-to-video model looks like serving an image diffusion model, scaled up, until you put it in front of users. A single 5-second 720p clip from an open model in the Wan or HunyuanVideo class can occupy several data-center GPUs for minutes. Requests differ in cost by an order of magnitude depending on resolution and duration. Users wait long enough to cancel, retry or close the tab. None of the assumptions behind token-streaming LLM serving hold: there is no useful partial output to stream, batching across requests rarely pays, and one request can need more than one GPU.
This page is about the serving system rather than the model. The video diffusion article covers the 3D VAE, spacetime tokens and the diffusion transformer. Here we cover the job API, the shape-based cost model, scheduling onto sequence-parallel GPU groups, splitting VAE decode and encoding from denoising, progress, cancellation and preemption, capacity planning, and the failure modes that show up in production.
Why video generation breaks request-response serving
Three properties drive the whole design.
- Requests are long. Tens of denoising steps, each a full transformer forward pass over a long token sequence, often doubled by classifier-free guidance. Wall-clock time per clip is measured in tens of seconds to minutes, which rules out holding an HTTP connection open and makes every failure expensive.
- Cost depends on shape, superlinearly. Attention is quadratic in sequence length, and sequence length grows with height, width and frame count. Two requests that look similar to a user can differ in cost by five times or more.
- One request can span GPUs. Large models at 720p are often run with sequence parallelism, splitting the token sequence across GPUs within one forward pass, using approaches such as Ulysses-style all-to-all or ring attention. The unit of scheduling becomes a group of GPUs, not a single device.
The architecture
The request path has four parts. An API accepts a job, validates it, estimates its cost and returns 202 with a job ID. A durable job store holds state transitions (queued, running, decoding, encoding, succeeded, failed, cancelled) and progress. A scheduler matches queued jobs to free GPU groups by shape and priority. Worker pools run the stages. The denoising GPUs are the bottleneck and the expensive resource, so every other component exists to keep them busy on work that will be delivered.
The job API
The contract is asynchronous and idempotent:
POST /v1/video-jobs
Idempotency-Key: 6f1c2d8e-...
{
"prompt": "a red fox running through fresh snow, golden hour",
"width": 1280, "height": 720, "frames": 81, "fps": 16,
"steps": 40, "guidance": 5.0, "seed": 1234,
"callback_url": "https://client.example/hooks/video"
}
202 Accepted
{"id": "vj_01J9...", "status": "queued", "estimated_gpu_seconds": 1900,
"queue_position": 7}
GET /v1/video-jobs/vj_01J9...
{"id": "vj_01J9...", "status": "running", "step": 22, "steps": 40}
GET /v1/video-jobs/vj_01J9... (later)
{"id": "vj_01J9...", "status": "succeeded",
"video_url": "<signed URL, expires in 1 h>", "seed": 1234}The idempotency key matters more here than in most APIs. A client that times out and retries a synchronous call would otherwise pay for two multi-minute generations. Record the seed and all sampling parameters with the result so a clip can be reproduced or debugged. Only accept resolutions and frame counts from an allowed list, for the reasons in the next section.
Shapes, buckets and the cost model
Estimate cost from the token count. For a VAE that compresses time by 4 and each spatial axis by 8, followed by 2 by 2 patchification, which is the layout Wan2.1 uses, an 81-frame clip becomes 21 latent frames. At 1280 by 720 each latent frame is 160 by 90, or 80 by 45 = 3,600 tokens after patchifying, so the clip is 75,600 tokens. The same computation at other shapes:
| Shape | Latent frames | Tokens | Attention cost vs 480p 81f |
|---|---|---|---|
| 832 x 480, 81 frames | 21 | 32,760 | 1.0x |
| 1280 x 720, 81 frames | 21 | 75,600 | about 5.3x |
| 832 x 480, 161 frames | 41 | 63,960 | about 3.8x |
| 1280 x 720, 161 frames | 41 | 147,600 | about 20x |
Linear layers scale with tokens and attention with tokens squared, so measure the real ratio on your hardware. The planning consequence holds regardless: restrict requests to a few buckets of fixed shape. Buckets give you predictable cost estimates for admission and pricing, compiled graphs and attention kernels that do not recompile per request, and a clean way to map work to group sizes. Arbitrary sizes defeat all three.
Scheduling onto sequence-parallel groups
Partition the denoising fleet into fixed groups, for example eight-GPU groups for 720p and two-GPU groups for 480p. Sequence parallelism scales less well on short sequences, because communication stops being hidden behind compute, so small shapes run more efficiently on small groups. Keep groups inside one NVLink domain: the all-to-all traffic of sequence parallelism over a slower interconnect erases the speed-up. A minimal scheduler:
import time, heapq
from dataclasses import dataclass, field
GROUP_SIZE_FOR = {"480p_81": 2, "480p_161": 4, "720p_81": 8, "720p_161": 8}
TIER_WEIGHT = {"interactive": 0, "standard": 1, "batch": 2}
AGING_S = 120 # every 2 min waiting improves priority by one tier
@dataclass(order=True)
class Job:
sort_key: float
id: str = field(compare=False)
bucket: str = field(compare=False)
tier: str = field(compare=False)
submitted: float = field(compare=False)
def priority(tier, submitted, now):
return TIER_WEIGHT[tier] - (now - submitted) / AGING_S
class Scheduler:
def __init__(self, free_groups): # {group_size: [group_id, ...]}
self.free = free_groups
self.queues = {b: [] for b in GROUP_SIZE_FOR}
def submit(self, job_id, bucket, tier):
now = time.time()
heapq.heappush(self.queues[bucket],
Job(TIER_WEIGHT[tier] + now / AGING_S, job_id, bucket, tier, now))
def dispatch(self):
"""Assign the most urgent job that has a free group of the right size."""
now, assigned = time.time(), []
while True:
best = None
for bucket, q in self.queues.items():
if q and self.free.get(GROUP_SIZE_FOR[bucket]):
j = q[0]
p = priority(j.tier, j.submitted, now)
if best is None or p < best[0]:
best = (p, bucket)
if best is None:
return assigned
job = heapq.heappop(self.queues[best[1]])
group = self.free[GROUP_SIZE_FOR[best[1]]].pop()
assigned.append((job.id, group))The heap key is the tier weight plus submission time divided by AGING_S, so priority at any moment is the key minus now / AGING_S. Every job ages at the same rate, the heap order never goes stale, and an old batch job eventually overtakes fresh interactive work. Production versions add per-tenant fair share and let an idle eight-GPU group split into four two-GPU groups when the 480p queue is long, merging back when 720p work arrives.
Splitting decode and encode from denoising
Run VAE decode and video encoding outside the denoising group. The VAE decoder turns the final latent into full-resolution frames. Its memory peak is dominated by activations at pixel resolution, often higher than the transformer's during sampling, and it runs once per clip. If it runs inside the denoising group, eight expensive GPUs sit idle while one decodes. Instead, ship the final latent to a decode pool. The latent is small: 16 channels by 21 by 90 by 160 values is about 4.8 million numbers, under 20 MB in fp32, so the transfer is cheap. Tiled or chunked decoding bounds peak memory if your VAE implementation supports it.
Encoding is the third stage. A GPU with NVENC can encode on the decode node, or frames can be piped to CPU encoders:
ffmpeg -f rawvideo -pix_fmt rgb24 -s 1280x720 -r 16 -i - \
-c:v libx264 -pix_fmt yuv420p -crf 18 -movflags +faststart out.mp4This is the same idea as disaggregated prefill and decode for LLMs: put each phase on the resource it needs, so the bottleneck resource only does bottleneck work.
Progress, cancellation and preemption
Because jobs are long, the worker loop has to be interruptible. Check for cancellation between denoising steps, report progress, and save a checkpoint of the latent periodically so that preemption or a node failure does not discard minutes of GPU time:
def run_denoise(job, model, store, ckpt_every=10):
state = store.load_checkpoint(job.id) or model.init_latent(job.seed, job.shape)
for step in range(state.step, job.steps):
if store.is_cancelled(job.id):
store.mark(job.id, "cancelled")
return None
if store.preempt_requested(job.group):
store.save_checkpoint(job.id, state) # latent, step, scheduler state
store.requeue(job.id)
return None
state = model.denoise_step(state, job.cond, step) # one transformer forward (x2 with CFG)
store.set_progress(job.id, step + 1, job.steps)
if (step + 1) % ckpt_every == 0:
store.save_checkpoint(job.id, state)
return state.latentThis is pseudocode around your own model wrapper, not a library API. A checkpoint must include the sampler's state and the random generator state as well as the latent, or a resumed job will not match an uninterrupted one. With checkpoints, interactive jobs can preempt batch jobs at a step boundary, and spot or preemptible capacity becomes usable for the batch tier.
Worked example: capacity planning
Plan from measurements, not FLOP estimates. Suppose load testing shows that a 720p, 81-frame clip at 40 steps takes 240 seconds on an eight-GPU group, and a 480p clip takes 60 seconds on a two-GPU group. These are illustrative numbers, so substitute your own. Traffic is 100 clips an hour, 30 percent of them 720p.
- 720p demand: 30 clips × 240 s = 7,200 group-seconds per hour, or 2.0 eight-GPU groups busy.
- 480p demand: 70 clips × 60 s = 4,200 group-seconds per hour, or about 1.2 two-GPU groups busy.
- At a target utilization of 70 percent, to keep queueing delay bounded, you need 3 eight-GPU groups and 2 two-GPU groups: 28 GPUs.
- GPU-seconds per clip are 1,920 for 720p and 120 for 480p, a 16-fold difference. Price and quota by bucket, not per request.
Queueing delay grows sharply as utilization approaches 100 percent, and video jobs are long, so a small queue means a long wait. Show the estimated start time at submission. The biggest lever is usually the model rather than the fleet: step-distilled or guidance-distilled variants and faster attention kernels such as FlashAttention change the per-clip cost directly, but each needs a quality evaluation before it ships.
Failure modes
- Out of memory at decode on the largest bucket after a dependency upgrade. Load-test every bucket on every release, and decode in tiles.
- One slow GPU in a group slows every step, because sequence parallelism synchronizes per layer. Track per-group step time and drain groups that run slower than their peers.
- Retry storms: clients retry long jobs without idempotency keys and double the load during an incident.
- Orphaned work: a user cancels or the result is never fetched. Honor cancellation between steps, and expire results with object-store lifecycle rules.
- Head-of-line blocking: a strict FIFO lets a burst of 720p batch jobs starve interactive 480p requests. Separate queues per bucket and tier, with aging.
- Content safety ordering: screen the prompt before admission, so you do not spend GPU minutes on a clip you will refuse to deliver, and screen the output before releasing the URL.
Trade-offs
| Decision | Option A | Option B |
|---|---|---|
| Group size | Large groups: lowest latency per clip | Small groups: higher throughput per GPU |
| Shapes | Fixed buckets: predictable, compilable | Free-form: flexible, unpredictable cost |
| Decode placement | In the denoise group: simpler | Separate pool: keeps the bottleneck busy |
| Checkpointing | Frequent: cheap preemption, more I/O | Rare: less I/O, more lost work on failure |
| Cross-request batching | Same-bucket batches: some throughput gain | Batch of one: simpler, lower latency |
What to do next
- Define three or four buckets and measure seconds per step for each at several group sizes on your hardware.
- Build the asynchronous job API with idempotency keys, recorded seeds and signed result URLs.
- Split the decode and encode stages out of the denoising group, and confirm denoising GPU utilization rises.
- Add step-boundary cancellation, progress and latent checkpoints, and verify that a resumed job reproduces an uninterrupted one bit for bit, or document why it cannot.
- Size the fleet from measured group-seconds per bucket at 70 percent utilization, and price by bucket.
- Keep learning: image diffusion on GPUs for the single-image case, and the model and attention articles linked above.