Serving a diffusion model looks like serving an LLM until you measure it. There is no token stream and no KV cache. Each request runs the same denoising network tens of times over a latent tensor whose size is fixed by the output resolution, then decodes the latent into pixels. Cost is set before the first step runs: resolution times step count times guidance mode. That makes diffusion serving more predictable than LLM serving, and it means the levers are different: bucketing shapes, batching at step granularity, caching everything that does not depend on the noise, and moving decoding off the critical GPU.

This article builds the request path for text-to-image models such as Stable Diffusion XL, using Hugging Face Diffusers for code. For the model internals and where inference time goes, read Stable Diffusion Training + Inference; for transformer-based FLUX see FLUX on the GPU; and for minute-long video jobs see Video Generation Serving.

A cost model before anything else

Start with a cost model, because every scheduling decision follows from it. One image costs roughly:

gpu_time = encode_text
         + steps * denoise_step(batch_shape, cfg_multiplier)
         + vae_decode(resolution)
         + safety_check

Text encoding runs once and is small. Denoising dominates: with classifier-free guidance (CFG) the network runs on a conditional and an unconditional copy of each latent, so a batch of B images is a network batch of 2B. Some distilled models, such as FLUX.1-dev with its guidance distillation, take the guidance scale as an input instead and need one pass. The step cost grows faster than linearly with resolution for attention-heavy models, because the number of latent tokens grows with area and self-attention grows with the square of tokens. VAE decoding runs once per image but at full pixel resolution, which makes it memory-hungry rather than slow.

Worked example, with assumed numbers you should replace with your own measurements. Suppose an SDXL UNet step at 1024x1024 with CFG (network batch 2) takes 60 ms on your GPU, and a step at network batch 8 takes 200 ms. Thirty steps at batch 1 cost 1.8 s of GPU time per image; at batch 4 (network batch 8) thirty steps cost 6 s for four images, 1.5 s each, about 20 percent more throughput at the price of 6 s latency instead of 1.8 s. Add an assumed 150 ms decode. If the SLO is p95 under 8 s including queueing, batch 4 fits only when the queue is short, so the scheduler must choose the batch size from the queue depth and the deadline, not a fixed setting.

The request path

Image diffusion serving: stages with different shapes of workAPI gatewayauth, quotasAdmission + queueper bucket, deadlineText encodersprompt cacheStep schedulergroups by shape, stepsDenoise pool1024x1024, CFG batchDenoise pool832x1216, CFG batchDenoise pool768x768, CFG batchVAE decodelatents to pixelsSafety + watermarkclassifier, policyObject store + CDNsigned URL to clientDenoising dominates GPU time; decode and safety are separate pools so they never stall it
The request path. Admission and the step scheduler decide what runs together; denoising runs in pools keyed by shape; decode and safety are separate stages.

Split the pipeline into stages that can be scaled separately. The gateway authenticates and applies quotas. Admission validates the request, maps it to a shape bucket and a step count, estimates its cost from the model above and rejects it early if the deadline cannot be met, which is far better than timing out after spending the GPU. Text encoders run once per request and are cached. Denoise workers are grouped into pools by resolution bucket, because tensors of different shapes cannot share a batch. VAE decode and safety classification are separate stages so that a large decode never blocks the next denoising step. The result goes to object storage and the client receives a signed URL; returning megabytes of PNG through the inference process wastes its time.

Buckets and batching

Clients ask for arbitrary sizes. Do not serve them directly. Define a small set of buckets with the same pixel area but different aspect ratios, for example 1024x1024, 1152x896, 896x1152, 1216x832 and 832x1216 for SDXL, which was trained at roughly one megapixel across aspect ratios. Map each request to the nearest bucket and crop or pad on output. Buckets make batching possible, make compiled graphs reusable and make the cost model accurate.

Step count and scheduler belong in the bucket key too. Requests at 20 steps and 40 steps can share a batch only if the scheduler handles per-request timesteps. The simple, robust approach is cohort batching: start requests with the same bucket, step count and sampler together and run them to completion. The more efficient approach is step-level batching, the diffusion equivalent of continuous batching: the network is called on whatever latents are ready, each with its own timestep, and new requests join at step boundaries. The UNet accepts a timestep per sample, but stock Diffusers schedulers keep state per pipeline call, so step-level batching means keeping one scheduler object per request and stepping each slice separately.

# Step-level batching loop for one shape bucket (sketch; assumes CFG and an SDXL-style UNet)
active = []                                     # requests mid-denoise
while True:
    # admit() gives each request its own scheduler (set_timesteps(steps)) and
    # r.latent = randn(seed) * r.sched.init_noise_sigma
    active += admit(queue, bucket, max_batch - len(active), deadline_aware=True)
    if not active:
        wait_for_work(); continue
    lat  = torch.cat([r.sched.scale_model_input(r.latent, r.sched.timesteps[r.i])
                      for r in active])
    t    = torch.cat([r.sched.timesteps[r.i].reshape(1) for r in active])
    emb  = torch.cat([r.cond for r in active] + [r.uncond for r in active])
    eps  = unet(torch.cat([lat, lat]), torch.cat([t, t]), encoder_hidden_states=emb,
                added_cond_kwargs=stack_added(active)).sample
    cond, uncond = eps.chunk(2)
    for j, r in enumerate(active):
        guided = uncond[j:j+1] + r.cfg * (cond[j:j+1] - uncond[j:j+1])
        r.latent = r.sched.step(guided, r.sched.timesteps[r.i], r.latent).prev_sample
        r.i += 1
    done = [r for r in active if r.i == r.steps]
    for r in done:
        decode_queue.put(r)                     # VAE runs in another stage
    active = [r for r in active if r.i < r.steps]

This loop gives a request that arrives mid-batch a start within one step time instead of a whole generation time. Its costs are a recompile or a padded graph for each batch size, and per-sample Python work in the scheduler step, which matters at small step times. Measure both before preferring it over cohorts. GPU Batching Strategies compares the general patterns.

Caching what does not depend on noise

Anything that does not depend on the noise can be cached. Prompt embeddings are the obvious case: products with templates, retries and "generate four variations" buttons see the same prompt many times, and the negative prompt is often a constant. Diffusers pipelines expose encode_prompt and accept precomputed embeddings, so the cache sits outside the pipeline:

import hashlib, torch
from diffusers import StableDiffusionXLPipeline

pipe = StableDiffusionXLPipeline.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0", torch_dtype=torch.float16).to("cuda")
cache = {}                                      # use an LRU with a byte budget in production

def embeddings(prompt, negative):
    key = hashlib.sha256(f"{prompt}\x00{negative}".encode()).hexdigest()
    if key not in cache:
        cache[key] = pipe.encode_prompt(prompt=prompt, device="cuda",
                                        num_images_per_prompt=1,
                                        do_classifier_free_guidance=True,
                                        negative_prompt=negative)
    return cache[key]

pe, npe, ppe, nppe = embeddings("a red bicycle, studio photo", "blurry")
image = pipe(prompt_embeds=pe, negative_prompt_embeds=npe,
             pooled_prompt_embeds=ppe, negative_pooled_prompt_embeds=nppe,
             num_inference_steps=30, guidance_scale=6.0).images[0]

Key the cache on the model version and LoRA set as well as the text, or an adapter change will serve stale embeddings. Never cache images by prompt unless the seed is part of the key; users expect variation.

Compilation and warmup

Diffusion is a good fit for compilation because the same network runs many times on the same shapes. Diffusers documents compiling the UNet with torch.compile(pipe.unet, mode="max-autotune", fullgraph=True); the max-autotune mode also uses CUDA graphs, which remove per-kernel launch overhead that matters at small batch sizes. Compilation is per shape: every new resolution or batch size triggers a recompile that can take minutes. So compile on startup for every bucket and batch size you will serve, run a warmup image through each, and only then mark the worker ready. A worker that recompiles on live traffic shows up as rare multi-minute latency outliers that look like a hang.

Precision is the other lever: fp16 or bf16 is standard; quantized weights reduce memory and can raise throughput, but check image quality against a fixed evaluation set of prompts and seeds before shipping. If you are on Triton, the Triton Inference Server ensemble pattern maps cleanly onto the encode, denoise and decode stages.

Serving many LoRA adapters

Products often serve many styles as LoRA adapters on one base model. Diffusers supports pipe.load_lora_weights(path, adapter_name="ink"), pipe.set_adapters(["ink"], adapter_weights=[0.8]), pipe.fuse_lora() and pipe.unload_lora_weights(). Fusing merges the adapter into the base weights, which removes its per-step overhead but makes switching expensive and invalidates compiled graphs. Unfused adapters switch quickly but add work to every step. A practical policy: fuse for the few adapters that carry most traffic and dedicate workers to them; serve the long tail unfused, routed so that requests for the same adapter land on the same worker; and include the adapter in the batching key, since a batch can only share one adapter configuration.

VAE decode and safety as their own stages

The VAE decoder turns a 128x128x4 latent into a 1024x1024x3 image and its activations are large; decoding a batch of high-resolution images can need more memory than the denoising itself. pipe.vae.enable_tiling() decodes in tiles to cap memory, at some cost in time and occasional seam artefacts. Running decode in its own stage, on the same GPU in a separate stream or on cheaper GPUs, keeps the denoise pools full. Latents are small (128 x 128 x 4 values in fp16 is 128 KB), so handing them between processes is cheap compared with handing over pixels. Then run safety classification. SD 1.x pipelines in Diffusers ship an optional safety checker; SDXL pipelines do not, so you must supply your own classifier and policy, plus a prompt filter before admission so that blocked requests never spend GPU time.

Few-step models

Distilled few-step models, such as SDXL-Turbo (adversarial diffusion distillation) or LCM-LoRA adapters, produce usable images in one to eight steps instead of twenty to fifty, which cuts denoise cost by roughly the same factor. The trade-offs are real: lower diversity, different or no use of guidance, sometimes lower resolution, and licences that differ from the base model. A common product split is a few-step model for interactive previews and a full-step model for the final render, with the preview's seed and prompt carried forward.

Observability and capacity

Because cost is known at admission, the most useful metrics compare prediction with reality. Record, per request: bucket, steps, adapter, batch size at each step, queue time, denoise time, decode time and predicted cost. Alert when actual denoise time per step drifts from the cost model for a bucket, which catches thermal throttling, a noisy neighbour or a silent recompile. Track GPU utilisation per pool, not per node: a decode pool at 30 percent is fine, a denoise pool at 30 percent with a non-empty queue means the scheduler is starving it. Capacity planning then becomes arithmetic: expected requests per second per bucket times GPU-seconds per request, divided by a target utilisation of around 70 to 80 percent to leave headroom for bursts.

Failure modes

  • Recompile on live traffic. A new size or batch size triggers torch.compile. Fix: fixed buckets, compile and warm all of them at startup.
  • Out of memory at decode. Large batches decode at once. Fix: separate decode stage, tiling, decode batch limit.
  • Head-of-line blocking. A 50-step 1536 px job delays 4-step previews. Fix: separate pools or priority by deadline.
  • Doomed work. Requests that will miss their deadline still run. Fix: cost-based admission and early rejection.
  • Stale embeddings. Prompt cache keyed only on text after a model or adapter change. Fix: include versions in the key.
  • Silent quality regressions. A new precision or scheduler shifts outputs. Fix: a fixed prompt-and-seed evaluation set compared on every release.
  • Large payloads through the GPU worker. PNG encoding and upload stall the loop. Fix: hand off to a CPU stage and object storage.

Trade-offs

Bigger batches raise throughput and latency together; choose batch size per bucket from the queue and deadline. Step-level batching cuts queueing delay but costs engineering and per-step overhead. Fused LoRAs are fast but sticky. Few-step models are cheap but change the product. Fewer buckets mean better batching but more cropping. Keeping all stages on one GPU is simpler; splitting them improves utilisation at the price of moving latents between processes, which is cheap because latents are small.

What to do next

  1. Measure denoise step time per bucket and batch size, decode time and encode time on your GPU; fill in the cost model.
  2. Define five or fewer resolution buckets and map every request to one.
  3. Add cost-based admission that rejects requests that cannot meet their deadline.
  4. Cache prompt embeddings keyed by text, model version and adapters.
  5. Compile and warm every bucket and batch size before marking a worker ready.
  6. Move VAE decode, safety and upload into separate stages.
  7. Build a fixed prompt-and-seed evaluation set and run it on every model, precision or scheduler change.
  8. Try cohort batching first; move to step-level batching only if queueing delay dominates.
Key takeaway: Diffusion serving cost is known at admission: resolution, steps and guidance. Use that to reject doomed work, bucket shapes so batches and compiled graphs can be shared, batch at cohort or step level by deadline, cache prompt embeddings, warm every compiled shape before taking traffic, and move decode, safety and upload off the denoising GPU loop.