FLUX.1 is a family of text-to-image models from Black Forest Labs, released in August 2024 and built around a transformer of about 12 billion parameters trained with rectified flow. There are three variants. [pro] is served through an API. [dev] has open weights under a non-commercial licence and is guidance-distilled. [schnell] has open weights under Apache 2.0 and is timestep-distilled, so it produces images in one to four steps. In November 2025 the company released FLUX.2, whose [dev] model is a 32 billion parameter flow transformer paired with a Mistral Small 3.2 24B text encoder. The arithmetic below uses FLUX.1 because its configuration is public and widely deployed, and the method carries over to FLUX.2 with the larger numbers.
The aim here is the GPU view: what runs, how many tokens it processes, where the FLOPs and the bytes go, and which knobs trade quality, latency and memory. The underlying maths of rectified flow is covered in rectified flow maths and flow matching.
The pipeline: what runs once and what runs every step
A FLUX.1 request runs four networks. CLIP-L turns the prompt into one pooled vector that conditions every block. The encoder half of T5-XXL turns the prompt into a sequence of up to 512 token embeddings, or 256 for [schnell], which take part in attention. The transformer then starts from Gaussian noise in the latent space and runs once per sampling step, predicting a velocity that moves the latent towards an image. Finally the VAE decoder turns the latent into pixels.
This split decides the optimisation strategy. The text encoders cost little per image, but T5-XXL is large in memory. The transformer runs N times, so it dominates time. The VAE decoder runs once, but at high resolution its activations can be the peak of the whole request.
From pixels to tokens
The VAE compresses each spatial dimension by 8 and produces 16 channels. The transformer then packs each 2 by 2 patch of latent pixels into one token of 64 values. A 1024 by 1024 image therefore becomes a 128 by 128 by 16 latent and 64 by 64 = 4,096 image tokens. Adding the text tokens gives the sequence length that every attention layer sees.
def flux_tokens(height, width, text_tokens=512):
assert height % 16 == 0 and width % 16 == 0 # 8x VAE, then 2x2 patches
image_tokens = (height // 16) * (width // 16)
return image_tokens, image_tokens + text_tokens
for h, w in [(512, 512), (768, 1360), (1024, 1024), (1536, 1536), (2048, 2048)]:
img, total = flux_tokens(h, w)
print(f"{h}x{w}: {img} image tokens, {total} in attention")| Resolution | Image tokens | Sequence with 512 text tokens |
|---|---|---|
| 512 x 512 | 1,024 | 1,536 |
| 768 x 1360 | 4,080 | 4,592 |
| 1024 x 1024 | 4,096 | 4,608 |
| 1536 x 1536 | 9,216 | 9,728 |
| 2048 x 2048 | 16,384 | 16,896 |
Doubling both sides quadruples the image tokens. Linear layers scale with tokens, and attention scales with tokens squared, so high resolution costs much more than proportionally.
Inside the transformer: double-stream and single-stream blocks
The FLUX.1 transformer has a hidden size of 3,072 and 24 attention heads, which gives 128 dimensions per head, a size that fused attention kernels handle efficiently. It runs 19 double-stream blocks followed by 38 single-stream blocks.
In a double-stream block, text tokens and image tokens keep separate weights for their projections and MLPs, but attention runs over the concatenated sequence, so each stream can attend to the other. This lets the two modalities keep different statistics early in the network. In a single-stream block, the two sequences are concatenated and processed by one set of weights, with the attention and MLP computed in parallel from the same input. Sharing one set of weights makes each single-stream block cheaper in parameters than a double-stream block.
Every block is conditioned through modulation: a small network turns the timestep and the pooled CLIP vector into scale, shift and gate values that are applied around the normalisation layers. In [dev], the guidance strength is also fed in as an embedding. Positions use rotary embeddings over separate axes for the latent grid, so the model can generate at resolutions and aspect ratios other than its training size. On the GPU nearly all of this is large matrix multiplies plus attention, so the tensor cores do the work and the usual transformer tooling applies, including FlashAttention through PyTorch scaled dot-product attention.
The sampling loop
Rectified flow trains the network to predict a velocity that carries a noisy latent along a nearly straight path towards data. Sampling integrates that velocity from pure noise at t = 1 to the image at t = 0, usually with simple Euler steps. Straighter paths need fewer steps, and distillation straightens them further. That is how [schnell] works in one to four steps, while [dev] normally uses around 28 to 50.
# Euler sampling for a velocity-prediction flow model (simplified; no batching or caching).
x = torch.randn(1, image_tokens, 64, device="cuda", dtype=torch.bfloat16)
sigmas = shifted_schedule(num_steps, image_tokens) # 1.0 -> 0.0, shifted for larger images
for s, s_next in zip(sigmas[:-1], sigmas[1:]):
v = transformer(x, t=s, text=t5_tokens, pooled=clip_vec, guidance=g) # one forward pass
x = x + (s_next - s) * v # dt is negative
latents = unpack_2x2(x, height, width)
image = vae.decode(latents / vae_scale + vae_shift)Two points matter for cost. First, the schedule is shifted towards the noisy end for larger images, so do not reuse a step count tuned at 512 pixels for 2,048 pixels without checking quality. Second, [dev] is guidance-distilled. The guidance strength is an input, so each step is one forward pass. Classic classifier-free guidance needs two passes per step, one with the prompt and one without, so this halves the work compared with an undistilled model of the same size. The guidance_scale argument for [dev] therefore drives that embedding rather than the usual two-pass combination, and [schnell] ignores it and expects 0.
FLOP and latency arithmetic
A useful estimate counts about 2 FLOPs per parameter per token for the linear layers, plus about 4 x L squared x hidden FLOPs per attention layer for the score and value products, where L is the sequence length.
PARAMS, HIDDEN, BLOCKS = 12e9, 3072, 19 + 38
def step_flops(seq_len):
linear = 2 * PARAMS * seq_len
attention = BLOCKS * 4 * seq_len**2 * HIDDEN
return linear, attention
for L in (4608, 16896): # 1024x1024 and 2048x2048 with 512 text tokens
lin, att = step_flops(L)
print(f"L={L}: {lin/1e12:.0f} TFLOP linear + {att/1e12:.0f} TFLOP attention per step")At 1024 by 1024 that gives about 111 TFLOP of linear work and 15 TFLOP of attention per step, about 126 TFLOP in all, or about 6.3 PFLOP for 50 steps. At 2048 by 2048 it is about 405 plus 200 TFLOP per step, and attention has grown from about an eighth of the work to about a third. To turn FLOPs into time, divide by the throughput your GPU actually sustains in BF16, not its peak. A GPU sustaining 400 TFLOP/s would need roughly 0.3 seconds per step at 1 megapixel, so about 16 seconds for 50 steps and well under a second for four [schnell] steps. These are estimates for planning. Measure your own steps per second, because kernel choice, compilation and offload change the result considerably.
Memory budget, offload and quantization
In BF16 each parameter takes 2 bytes. The 12B transformer needs about 24 GB, the T5-XXL encoder, about 4.7 billion parameters, roughly 9.5 GB, and CLIP-L and the VAE together under 1 GB. That is about 34 GB of weights before activations, which at 1 megapixel are small by comparison, except during VAE decoding. The diffusers documentation says that loading all components needs about 50 GB of RAM or VRAM. A single 80 GB or 48 GB GPU holds everything. A 24 GB consumer card does not.
import torch
from diffusers import FluxPipeline
pipe = FluxPipeline.from_pretrained("black-forest-labs/FLUX.1-dev", torch_dtype=torch.bfloat16)
pipe.enable_model_cpu_offload() # moves whole components to the GPU only while they run
pipe.vae.enable_tiling() # decode large images in tiles to cap the VAE peak
image = pipe("a lighthouse at dusk, oil painting", height=1024, width=1024,
guidance_scale=3.5, num_inference_steps=50,
generator=torch.Generator("cpu").manual_seed(0)).images[0]| Option | GPU memory | Speed cost | Quality risk |
|---|---|---|---|
| Everything resident, BF16 | About 34 GB plus activations | None | None |
enable_model_cpu_offload() | Largest single component | Small: T5 and transformer swap once per image | None |
enable_sequential_cpu_offload() | A few GB | Very large: weights stream every step | None |
| Group offloading with CUDA streams | Configurable | Moderate, overlaps copies with compute | None |
| FP8 weights (transformer, T5) | About half of BF16 | Depends on kernels and GPU support | Small, check with a fixed seed set |
| 8-bit or 4-bit bitsandbytes | About half or a quarter | Usually slower than BF16 | Grows at 4-bit |
Two warnings from the diffusers documentation: FP8 inference can be brittle depending on GPU, CUDA and PyTorch versions, and when running in FP16 the text encoders should be kept in FP32 to avoid activation overflow. BF16 avoids that problem, as explained in mixed precision, in depth.
Serving FLUX in production
Image requests take seconds, so serve them through a queue with workers that each own a GPU, rather than as synchronous HTTP calls. Batch requests that share a resolution and step count, because a batch of four at 1 megapixel uses the GPU much better than four sequential runs. Bucket resolutions to a small set of sizes. That keeps batches full and stops torch.compile from recompiling for every new shape. Cache T5 embeddings for repeated prompts and templates, since that skips a 9.5 GB model on a cache hit.
Record the seed, model revision, scheduler settings, step count and guidance with each output. Reproducing an image or investigating a complaint then becomes a single replay. Check licences per variant: [dev] weights are non-commercial unless you have a commercial licence, while [schnell] is Apache 2.0.
Worked example: planning capacity for 1-megapixel images
A product needs 1024 by 1024 images from [dev] at 28 steps, with a peak of 1,200 images per hour. On the target GPU the team measures 3.1 steps per second at batch 1 and 2.0 steps per second per image at batch 4, which is 8 image-steps per second. With batch 4, one GPU produces 8 / 28 = 0.29 images per second, about 1,030 per hour, after allowing a few percent for text encoding and VAE decoding. Two GPUs cover the peak with headroom. The latency of one image is now 28 / 2.0 = 14 seconds instead of 9 at batch 1, so the team caps batch size at 2 for the interactive tier and uses 4 for bulk jobs. These measured numbers are specific to this example. Repeat the measurement on your own hardware.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| Out of memory only at large sizes | VAE decode activations | VAE tiling, or decode on a separate pass |
| Black or NaN images in FP16 | Text-encoder overflow | Use BF16, or keep text encoders in FP32 |
| Prompt details ignored with [schnell] | Prompt truncated at 256 T5 tokens | Shorten the prompt or use [dev] |
| Over-saturated [dev] images | Guidance treated like classic CFG values | Start near 3.5 and tune down |
| Slow first requests after deploy | Compilation for new shapes | Warm up every resolution bucket |
| Throughput far below estimate | Sequential offload left enabled | Use model offload or keep weights resident |
What to do next
- Decide the variant by licence and latency: [schnell] for fast and commercial use, [dev] for quality under its licence terms, or the API.
- Compute tokens per image for your resolutions and estimate FLOPs per image with the script above.
- Choose a memory plan: resident BF16 on 48 GB or larger GPUs, model offload or FP8 on smaller ones, and validate quality on a fixed seed set.
- Measure steps per second at batch sizes 1, 2 and 4 on your GPU and derive images per hour.
- Bucket resolutions, enable VAE tiling for large sizes and warm up compiled shapes before taking traffic.
- Log seed, revision, steps and guidance with every output.
- If you plan to move to FLUX.2, redo the memory budget for its 32B transformer and 24B text encoder before choosing hardware.