A FLOPS budget turns “we want to train this model on that much data” into GPU-hours, calendar days and money before anyone reserves a cluster. Two terms get confused constantly. FLOPs with a lowercase s is a count of floating-point operations, the total work. FLOPS or FLOP/s is a rate, what a GPU delivers per second. The budget is a count, written C; the schedule comes from dividing it by a rate you actually achieve.

This article derives the count from first principles, shows where the standard 6ND rule is incomplete, converts it into GPU-hours with model FLOPs utilisation and goodput, and works a complete example for a 6.6-billion parameter model on 2 trillion tokens. It ends with a ledger you can reuse and the mistakes that most often make budgets wrong by a factor of two. Every number in the example comes from a short script you can rerun with your own configuration.

Counting the FLOPs

Nearly all training compute in a dense transformer is matrix multiplication. Multiplying an activation vector by a weight matrix with N entries costs N multiplies and N adds, so 2N FLOPs per token in the forward pass. The backward pass computes two products per weight matrix, one for the gradient with respect to the activations and one for the gradient with respect to the weights, so it costs about twice the forward. Forward plus backward is therefore about 6N FLOPs per token, and a run over D tokens costs about C = 6ND. That is the rule popularised by the Kaplan and Chinchilla scaling papers.

Which N? Count the parameters that take part in matrix multiplications: the attention projections, the feed-forward layers and the output projection to the vocabulary. The input embedding is a table lookup, not a multiply, so it adds memory but almost no FLOPs. Norms, activations, softmax and the optimizer step are small next to the matmuls and are usually left out.

The 6N term misses one thing: attention scores. Computing the query-key products and applying them to the values depends on sequence length, not on parameters. The PaLM paper's accounting adds 12·L·d·T FLOPs per token for L layers, model width d and sequence length T, covering forward and backward. That convention counts the full attention matrix; causal kernels such as FlashAttention skip the masked half, so they do somewhat less work than the formula says. Pick one convention, state it, and use it consistently, because MFU figures are only comparable under the same convention.

def params_for_flops(L, d, d_ff, vocab, gated_mlp=True):
    attn = 4 * d * d                       # Q, K, V, O projections (no GQA)
    mlp = (3 if gated_mlp else 2) * d * d_ff
    return L * (attn + mlp) + vocab * d    # + output head; input embedding excluded

def flops_per_token(L, d, d_ff, vocab, seq_len):
    n = params_for_flops(L, d, d_ff, vocab)
    dense = 6 * n                          # matmuls, forward + backward
    attn_scores = 12 * L * d * seq_len     # PaLM convention, full attention matrix
    return dense + attn_scores, dense, attn_scores

per_tok, dense, attn = flops_per_token(L=32, d=4096, d_ff=11008, vocab=32000, seq_len=4096)
C = per_tok * 2e12                         # 2 trillion training tokens
print(f"{per_tok:.4e} FLOPs/token, attention share {attn / per_tok:.1%}, C = {C:.4e}")
# 4.6085e+10 FLOPs/token, attention share 14.0%, C = 9.2170e+22

Adjust attn for grouped-query attention, where K and V projections are smaller, and for mixture of experts, where N means the parameters active per token, not the total.

Sequence length changes the answer

The attention share grows linearly with sequence length, which matters for long-context training. For the example configuration it is about 7.5% at 2,048 tokens, 14% at 4,096, 24.5% at 8,192 and 56.5% at 32,768. A budget built on 6ND alone underestimates a long-context phase by more than half. Budget each phase with its own sequence length.

Peak, MFU and HFU

The denominator starts from the GPU's peak dense throughput for the data type you train in. For an H100 SXM, NVIDIA's datasheet lists 1,979 teraFLOPS for BF16 tensor cores, with a footnote that the figure assumes structured sparsity. Dense training gets half of that, about 989 teraFLOPS. Using the sparsity number is the single most common way to halve a schedule on paper. Check the footnotes for every GPU you budget, and see NVIDIA H100, Hopper architecture for how the tensor cores reach that peak.

No real job reaches peak. Model FLOPs utilisation (MFU) is the model's FLOPs per second, using the count above, divided by peak. It is lowered by communication not hidden behind compute, pipeline bubbles, memory-bound kernels, data loading stalls and small matrices. The Llama 3 paper reports 38 to 43 percent BF16 MFU across its large H100 configurations, a useful reference for well-tuned dense training at scale. Untuned jobs commonly sit well below that; LLM Bottleneck Analysis shows how to find what is holding yours down.

Hardware FLOPs utilisation (HFU) counts the FLOPs the GPU actually executed, including recomputation. With full activation checkpointing, each layer's forward runs twice, so the hardware does about 8N per token instead of 6N. A job at 40% MFU with full recompute is running at about 53% HFU. Budget with MFU, because recompute is overhead the model does not benefit from, and never report HFU as MFU.

Goodput: the time you actually train

MFU describes the job while it is running. A budget also needs goodput: the fraction of reserved time spent making forward progress. It is reduced by hardware failures and the work lost since the last checkpoint, restart time, checkpoint saves that block training, evaluation pauses and scheduled maintenance. Llama 3 reports more than 90 percent effective training time over a 54-day snapshot on a very large cluster, achieved with considerable automation. A new team should plan for less until it has measured its own. Faster, asynchronous checkpoints, as covered in GPU Checkpointing Deep Dive, raise goodput directly.

From a model and a token count to GPU-hours, days and moneyModel configL, d, d_ff, vocabTokens Dand sequence length TFLOPs per token6N + 12·L·d·TBudget Cper-token × DPeak dense FLOPSper GPU, per dtypeMFUmeasured, not assumedGPU-hoursC / (peak × MFU)Wall clock, cost÷ GPUs, ÷ goodputGoodputrestarts, evals, stallsEvery box is a place where a factor of two hides: sparsity peaks, recompute, wrong N, ignored attention.
The budget pipeline. Counting gives C; measured MFU and goodput convert it into time on a given number of GPUs.

Worked example: 6.6B parameters, 2T tokens

Take a dense decoder with 32 layers, width 4,096, a gated feed-forward width of 11,008 and a 32,000-token vocabulary. Matmul parameters come to 6.607 billion. Train it on 2 trillion tokens at sequence length 4,096 on H100s in BF16, assuming 40% MFU and 90% goodput.

QuantityFormulaValue
FLOPs per token6N + 12·L·d·T4.6085 × 1010
Total budget Cper-token × D9.217 × 1022
Achieved rate per GPU989.5 TFLOPS × 0.40395.8 TFLOPS
Tokens per second per GPUrate ÷ FLOPs per tokenabout 8,590
Compute GPU-hoursC ÷ rate ÷ 3,60064,686
Reserved GPU-hours÷ 0.90 goodput71,873
Wall clock on 512 GPUs÷ 512 ÷ 245.85 days
Cost at an illustrative US$2.50 per GPU-hour× priceabout US$180,000

Sensitivity is the useful output. At 35% MFU the compute hours rise to 73,927; at 45% they fall to 57,499. A five-point MFU change moves the budget by more than 10 percent, which is why the first week of any project should be spent measuring MFU on the real configuration, not quoting someone else's.

Check the token count against compute-optimal scaling. Chinchilla's rule of thumb is about 20 tokens per parameter, which here is about 132 billion tokens. Two trillion is roughly 15 times that. This is a deliberate choice when the model will serve heavy inference traffic: a smaller model trained longer is cheaper to run, at the cost of more training compute than the loss alone would justify. The Chinchilla scaling article explains the trade.

Finally, sanity-check the method against a published run. Llama 3's 405-billion parameter model on 15.6 trillion tokens gives 6ND = 3.79 × 1025, matching the paper's stated 3.8 × 1025 FLOPs. At 40% MFU on about 16,000 H100s that is roughly 70 days of pure compute, the right order of magnitude for a run of that size.

Measuring MFU during the run

Once training starts, compute MFU from step time every few hundred steps and log it next to loss. The formula needs only the tokens per step, the per-token FLOPs you already computed and the GPU count.

PEAK_BF16_DENSE = 989.5e12      # H100 SXM, dense; change per GPU and dtype

def mfu(tokens_per_step, step_seconds, n_gpus, flops_per_tok, peak=PEAK_BF16_DENSE):
    achieved = tokens_per_step * flops_per_tok / step_seconds
    return achieved / (n_gpus * peak)

# global batch of 1,024 sequences x 4,096 tokens on 512 GPUs, measured 1.0 s per step
print(f"{mfu(1024 * 4096, 1.0, 512, 4.6085e10):.1%}")   # 38.2%

Count only real tokens. If sequences are padded rather than packed, padding inflates tokens per step and therefore MFU while doing no useful work; packing sequences fixes both. Exclude steps that include checkpoint saves or evaluation from the MFU average and account for them in goodput instead, so each number measures one thing. If the measured MFU sits well below plan, re-run the schedule with the measured value immediately; the budget is a living forecast, not a promise made once.

The full ledger

The main run is never the whole bill. A practical ledger reserves compute for the work around it. The split below is an example, not a standard; adjust it from your own history.

Line itemExample shareWhy it exists
Main pre-training run60 to 70%The budget computed above, divided by goodput
Ablations and scaling probes10 to 20%Small runs that choose data mix, LR and width
Evaluation3 to 5%Periodic benchmark passes during and after training
Long-context and annealing phases5 to 10%Budgeted with their own sequence length
Contingency10%Loss spikes, rollbacks, a lower MFU than planned

Precision changes the denominator, not the numerator. Training in FP8 on hardware that supports it raises peak throughput, but the FLOP count is the same and the achievable MFU against the higher peak is usually lower. Always state which peak an MFU figure uses. Mixed precision training covers the numerics.

Mistakes that cost a factor of two

MistakeEffect on the budgetFix
Sparsity peak used for dense trainingSchedule 2× too optimisticUse the dense figure from the footnote
Input embedding counted in NSmall overcount; large for small modelsCount matmul parameters only
Attention term ignoredUnderestimate, worse at long contextAdd 12·L·d·T per token
Recompute counted as usefulMFU inflated by up to a thirdReport MFU and HFU separately
Padding counted as tokensMFU inflated, data under-deliveredPack sequences; count real tokens
Total MoE parameters usedHuge overestimateUse active parameters per token
Goodput assumed to be 100%Overruns of weeksMeasure restarts; plan 85 to 90% at best

Trade-offs

A FLOPS budget is a model, and it trades precision for speed. The 6N rule ignores small operations and the attention convention overstates causal work, so two careful teams can disagree by a few percent; that is fine as long as each is consistent. The larger uncertainties are MFU and goodput, which is why the budget should be a range. Spending more GPUs shortens wall clock but usually lowers MFU through communication overhead. Whether a shorter calendar is worth a higher bill is a business decision; the budget's job is to make that trade visible. Turning the budget into a cluster order, with memory floors, fabric and lead times, is covered in GPU infrastructure planning.

What to do next

  1. Write down your model configuration and compute matmul parameters with embedding excluded and the output head included.
  2. Compute FLOPs per token as 6N plus the attention term, separately for each training phase's sequence length.
  3. Look up dense peak throughput for your GPU and data type, reading the sparsity footnote.
  4. Run a short job on the real configuration, measure MFU from step time, and replace any assumed value.
  5. Convert to GPU-hours, apply a measured or conservative goodput, and add ablation, evaluation and contingency lines.
  6. Log MFU and goodput throughout the run and re-forecast the end date whenever either drifts.
Key takeaway: A training budget is C = (6N + 12·L·d·T) × D FLOPs, divided by dense peak throughput, measured MFU and goodput. Count only matmul parameters, budget each sequence-length phase separately, never use sparsity peaks, keep MFU and HFU apart, measure MFU on the real configuration before committing, and reserve compute for ablations, evaluation and contingency.