FLUX.1 is a family of text-to-image models from Black Forest Labs, released in August 2024 and built around a transformer of about 12 billion parameters trained with rectified flow. There are three variants. [pro] is served through an API. [dev] has open weights under a non-commercial licence and is guidance-distilled. [schnell] has open weights under Apache 2.0 and is timestep-distilled, so it produces images in one to four steps. In November 2025 the company released FLUX.2, whose [dev] model is a 32 billion parameter flow transformer paired with a Mistral Small 3.2 24B text encoder. The arithmetic below uses FLUX.1 because its configuration is public and widely deployed, and the method carries over to FLUX.2 with the larger numbers.

The aim here is the GPU view: what runs, how many tokens it processes, where the FLOPs and the bytes go, and which knobs trade quality, latency and memory. The underlying maths of rectified flow is covered in rectified flow maths and flow matching.

Advertisement

The pipeline: what runs once and what runs every step

FLUX.1 inference: what runs once, what runs every stepPrompttextCLIP-Lpooled vector, onceT5-XXL encoderup to 512 tokens, onceGaussian noiselatent tokensTransformer, run N times (N = 1 to 4 schnell, about 50 dev)19 double-streamown weights per stream38 single-streamshared weightsjoint attention over text plus image tokensconditioning: timestep, pooled CLIP, guidance (dev)text tokenspooledEuler updatex = x + dt times vnext stepVAE decoderonce, 8x upsampleImagepixelsafter N stepsCost is dominated by the transformer loop; memory is dominated by transformer plus T5 weights.
Figure 1. FLUX.1 inference. The two text encoders and the VAE decoder run once per image; the 12B transformer runs once per sampling step.

A FLUX.1 request runs four networks. CLIP-L turns the prompt into one pooled vector that conditions every block. The encoder half of T5-XXL turns the prompt into a sequence of up to 512 token embeddings, or 256 for [schnell], which take part in attention. The transformer then starts from Gaussian noise in the latent space and runs once per sampling step, predicting a velocity that moves the latent towards an image. Finally the VAE decoder turns the latent into pixels.

This split decides the optimisation strategy. The text encoders cost little per image, but T5-XXL is large in memory. The transformer runs N times, so it dominates time. The VAE decoder runs once, but at high resolution its activations can be the peak of the whole request.

From pixels to tokens

The VAE compresses each spatial dimension by 8 and produces 16 channels. The transformer then packs each 2 by 2 patch of latent pixels into one token of 64 values. A 1024 by 1024 image therefore becomes a 128 by 128 by 16 latent and 64 by 64 = 4,096 image tokens. Adding the text tokens gives the sequence length that every attention layer sees.

def flux_tokens(height, width, text_tokens=512):
    assert height % 16 == 0 and width % 16 == 0     # 8x VAE, then 2x2 patches
    image_tokens = (height // 16) * (width // 16)
    return image_tokens, image_tokens + text_tokens

for h, w in [(512, 512), (768, 1360), (1024, 1024), (1536, 1536), (2048, 2048)]:
    img, total = flux_tokens(h, w)
    print(f"{h}x{w}: {img} image tokens, {total} in attention")
ResolutionImage tokensSequence with 512 text tokens
512 x 5121,0241,536
768 x 13604,0804,592
1024 x 10244,0964,608
1536 x 15369,2169,728
2048 x 204816,38416,896

Doubling both sides quadruples the image tokens. Linear layers scale with tokens, and attention scales with tokens squared, so high resolution costs much more than proportionally.

Advertisement

Inside the transformer: double-stream and single-stream blocks

The FLUX.1 transformer has a hidden size of 3,072 and 24 attention heads, which gives 128 dimensions per head, a size that fused attention kernels handle efficiently. It runs 19 double-stream blocks followed by 38 single-stream blocks.

In a double-stream block, text tokens and image tokens keep separate weights for their projections and MLPs, but attention runs over the concatenated sequence, so each stream can attend to the other. This lets the two modalities keep different statistics early in the network. In a single-stream block, the two sequences are concatenated and processed by one set of weights, with the attention and MLP computed in parallel from the same input. Sharing one set of weights makes each single-stream block cheaper in parameters than a double-stream block.

Every block is conditioned through modulation: a small network turns the timestep and the pooled CLIP vector into scale, shift and gate values that are applied around the normalisation layers. In [dev], the guidance strength is also fed in as an embedding. Positions use rotary embeddings over separate axes for the latent grid, so the model can generate at resolutions and aspect ratios other than its training size. On the GPU nearly all of this is large matrix multiplies plus attention, so the tensor cores do the work and the usual transformer tooling applies, including FlashAttention through PyTorch scaled dot-product attention.

The sampling loop

Rectified flow trains the network to predict a velocity that carries a noisy latent along a nearly straight path towards data. Sampling integrates that velocity from pure noise at t = 1 to the image at t = 0, usually with simple Euler steps. Straighter paths need fewer steps, and distillation straightens them further. That is how [schnell] works in one to four steps, while [dev] normally uses around 28 to 50.

# Euler sampling for a velocity-prediction flow model (simplified; no batching or caching).
x = torch.randn(1, image_tokens, 64, device="cuda", dtype=torch.bfloat16)
sigmas = shifted_schedule(num_steps, image_tokens)      # 1.0 -> 0.0, shifted for larger images
for s, s_next in zip(sigmas[:-1], sigmas[1:]):
    v = transformer(x, t=s, text=t5_tokens, pooled=clip_vec, guidance=g)   # one forward pass
    x = x + (s_next - s) * v                                               # dt is negative
latents = unpack_2x2(x, height, width)
image = vae.decode(latents / vae_scale + vae_shift)

Two points matter for cost. First, the schedule is shifted towards the noisy end for larger images, so do not reuse a step count tuned at 512 pixels for 2,048 pixels without checking quality. Second, [dev] is guidance-distilled. The guidance strength is an input, so each step is one forward pass. Classic classifier-free guidance needs two passes per step, one with the prompt and one without, so this halves the work compared with an undistilled model of the same size. The guidance_scale argument for [dev] therefore drives that embedding rather than the usual two-pass combination, and [schnell] ignores it and expects 0.

FLOP and latency arithmetic

A useful estimate counts about 2 FLOPs per parameter per token for the linear layers, plus about 4 x L squared x hidden FLOPs per attention layer for the score and value products, where L is the sequence length.

PARAMS, HIDDEN, BLOCKS = 12e9, 3072, 19 + 38

def step_flops(seq_len):
    linear = 2 * PARAMS * seq_len
    attention = BLOCKS * 4 * seq_len**2 * HIDDEN
    return linear, attention

for L in (4608, 16896):          # 1024x1024 and 2048x2048 with 512 text tokens
    lin, att = step_flops(L)
    print(f"L={L}: {lin/1e12:.0f} TFLOP linear + {att/1e12:.0f} TFLOP attention per step")

At 1024 by 1024 that gives about 111 TFLOP of linear work and 15 TFLOP of attention per step, about 126 TFLOP in all, or about 6.3 PFLOP for 50 steps. At 2048 by 2048 it is about 405 plus 200 TFLOP per step, and attention has grown from about an eighth of the work to about a third. To turn FLOPs into time, divide by the throughput your GPU actually sustains in BF16, not its peak. A GPU sustaining 400 TFLOP/s would need roughly 0.3 seconds per step at 1 megapixel, so about 16 seconds for 50 steps and well under a second for four [schnell] steps. These are estimates for planning. Measure your own steps per second, because kernel choice, compilation and offload change the result considerably.

Memory budget, offload and quantization

In BF16 each parameter takes 2 bytes. The 12B transformer needs about 24 GB, the T5-XXL encoder, about 4.7 billion parameters, roughly 9.5 GB, and CLIP-L and the VAE together under 1 GB. That is about 34 GB of weights before activations, which at 1 megapixel are small by comparison, except during VAE decoding. The diffusers documentation says that loading all components needs about 50 GB of RAM or VRAM. A single 80 GB or 48 GB GPU holds everything. A 24 GB consumer card does not.

import torch
from diffusers import FluxPipeline

pipe = FluxPipeline.from_pretrained("black-forest-labs/FLUX.1-dev", torch_dtype=torch.bfloat16)
pipe.enable_model_cpu_offload()     # moves whole components to the GPU only while they run
pipe.vae.enable_tiling()            # decode large images in tiles to cap the VAE peak

image = pipe("a lighthouse at dusk, oil painting", height=1024, width=1024,
             guidance_scale=3.5, num_inference_steps=50,
             generator=torch.Generator("cpu").manual_seed(0)).images[0]
OptionGPU memorySpeed costQuality risk
Everything resident, BF16About 34 GB plus activationsNoneNone
enable_model_cpu_offload()Largest single componentSmall: T5 and transformer swap once per imageNone
enable_sequential_cpu_offload()A few GBVery large: weights stream every stepNone
Group offloading with CUDA streamsConfigurableModerate, overlaps copies with computeNone
FP8 weights (transformer, T5)About half of BF16Depends on kernels and GPU supportSmall, check with a fixed seed set
8-bit or 4-bit bitsandbytesAbout half or a quarterUsually slower than BF16Grows at 4-bit

Two warnings from the diffusers documentation: FP8 inference can be brittle depending on GPU, CUDA and PyTorch versions, and when running in FP16 the text encoders should be kept in FP32 to avoid activation overflow. BF16 avoids that problem, as explained in mixed precision, in depth.

Serving FLUX in production

Image requests take seconds, so serve them through a queue with workers that each own a GPU, rather than as synchronous HTTP calls. Batch requests that share a resolution and step count, because a batch of four at 1 megapixel uses the GPU much better than four sequential runs. Bucket resolutions to a small set of sizes. That keeps batches full and stops torch.compile from recompiling for every new shape. Cache T5 embeddings for repeated prompts and templates, since that skips a 9.5 GB model on a cache hit.

Record the seed, model revision, scheduler settings, step count and guidance with each output. Reproducing an image or investigating a complaint then becomes a single replay. Check licences per variant: [dev] weights are non-commercial unless you have a commercial licence, while [schnell] is Apache 2.0.

Worked example: planning capacity for 1-megapixel images

A product needs 1024 by 1024 images from [dev] at 28 steps, with a peak of 1,200 images per hour. On the target GPU the team measures 3.1 steps per second at batch 1 and 2.0 steps per second per image at batch 4, which is 8 image-steps per second. With batch 4, one GPU produces 8 / 28 = 0.29 images per second, about 1,030 per hour, after allowing a few percent for text encoding and VAE decoding. Two GPUs cover the peak with headroom. The latency of one image is now 28 / 2.0 = 14 seconds instead of 9 at batch 1, so the team caps batch size at 2 for the interactive tier and uses 4 for bulk jobs. These measured numbers are specific to this example. Repeat the measurement on your own hardware.

Failure modes

SymptomLikely causeFix
Out of memory only at large sizesVAE decode activationsVAE tiling, or decode on a separate pass
Black or NaN images in FP16Text-encoder overflowUse BF16, or keep text encoders in FP32
Prompt details ignored with [schnell]Prompt truncated at 256 T5 tokensShorten the prompt or use [dev]
Over-saturated [dev] imagesGuidance treated like classic CFG valuesStart near 3.5 and tune down
Slow first requests after deployCompilation for new shapesWarm up every resolution bucket
Throughput far below estimateSequential offload left enabledUse model offload or keep weights resident

What to do next

  1. Decide the variant by licence and latency: [schnell] for fast and commercial use, [dev] for quality under its licence terms, or the API.
  2. Compute tokens per image for your resolutions and estimate FLOPs per image with the script above.
  3. Choose a memory plan: resident BF16 on 48 GB or larger GPUs, model offload or FP8 on smaller ones, and validate quality on a fixed seed set.
  4. Measure steps per second at batch sizes 1, 2 and 4 on your GPU and derive images per hour.
  5. Bucket resolutions, enable VAE tiling for large sizes and warm up compiled shapes before taking traffic.
  6. Log seed, revision, steps and guidance with every output.
  7. If you plan to move to FLUX.2, redo the memory budget for its 32B transformer and 24B text encoder before choosing hardware.
Key takeaway: FLUX.1 is a 12B rectified-flow transformer flanked by two text encoders and a VAE. Resolution sets the token count, through 8x VAE compression and 2x2 patches, so 1 megapixel means about 4,600 tokens per attention layer and attention grows quadratically beyond that. Time goes into the transformer loop, about 126 TFLOP per step at 1 megapixel, and memory goes into the transformer and T5 weights, about 34 GB in BF16. Choose the variant by licence and step count, the memory plan by GPU size, measure steps per second at several batch sizes, and serve through a queue with resolution buckets and full provenance per image.