A FLOPS budget turns “we want to train this model on that much data” into GPU-hours, calendar days and money before anyone reserves a cluster. Two terms get confused constantly. FLOPs with a lowercase s is a count of floating-point operations, the total work. FLOPS or FLOP/s is a rate, what a GPU delivers per second. The budget is a count, written C; the schedule comes from dividing it by a rate you actually achieve.
This article derives the count from first principles, shows where the standard 6ND rule is incomplete, converts it into GPU-hours with model FLOPs utilisation and goodput, and works a complete example for a 6.6-billion parameter model on 2 trillion tokens. It ends with a ledger you can reuse and the mistakes that most often make budgets wrong by a factor of two. Every number in the example comes from a short script you can rerun with your own configuration.
Counting the FLOPs
Nearly all training compute in a dense transformer is matrix multiplication. Multiplying an activation vector by a weight matrix with N entries costs N multiplies and N adds, so 2N FLOPs per token in the forward pass. The backward pass computes two products per weight matrix, one for the gradient with respect to the activations and one for the gradient with respect to the weights, so it costs about twice the forward. Forward plus backward is therefore about 6N FLOPs per token, and a run over D tokens costs about C = 6ND. That is the rule popularised by the Kaplan and Chinchilla scaling papers.
Which N? Count the parameters that take part in matrix multiplications: the attention projections, the feed-forward layers and the output projection to the vocabulary. The input embedding is a table lookup, not a multiply, so it adds memory but almost no FLOPs. Norms, activations, softmax and the optimizer step are small next to the matmuls and are usually left out.
The 6N term misses one thing: attention scores. Computing the query-key products and applying them to the values depends on sequence length, not on parameters. The PaLM paper's accounting adds 12·L·d·T FLOPs per token for L layers, model width d and sequence length T, covering forward and backward. That convention counts the full attention matrix; causal kernels such as FlashAttention skip the masked half, so they do somewhat less work than the formula says. Pick one convention, state it, and use it consistently, because MFU figures are only comparable under the same convention.
def params_for_flops(L, d, d_ff, vocab, gated_mlp=True):
attn = 4 * d * d # Q, K, V, O projections (no GQA)
mlp = (3 if gated_mlp else 2) * d * d_ff
return L * (attn + mlp) + vocab * d # + output head; input embedding excluded
def flops_per_token(L, d, d_ff, vocab, seq_len):
n = params_for_flops(L, d, d_ff, vocab)
dense = 6 * n # matmuls, forward + backward
attn_scores = 12 * L * d * seq_len # PaLM convention, full attention matrix
return dense + attn_scores, dense, attn_scores
per_tok, dense, attn = flops_per_token(L=32, d=4096, d_ff=11008, vocab=32000, seq_len=4096)
C = per_tok * 2e12 # 2 trillion training tokens
print(f"{per_tok:.4e} FLOPs/token, attention share {attn / per_tok:.1%}, C = {C:.4e}")
# 4.6085e+10 FLOPs/token, attention share 14.0%, C = 9.2170e+22Adjust attn for grouped-query attention, where K and V projections are smaller, and for mixture of experts, where N means the parameters active per token, not the total.
Sequence length changes the answer
The attention share grows linearly with sequence length, which matters for long-context training. For the example configuration it is about 7.5% at 2,048 tokens, 14% at 4,096, 24.5% at 8,192 and 56.5% at 32,768. A budget built on 6ND alone underestimates a long-context phase by more than half. Budget each phase with its own sequence length.
Peak, MFU and HFU
The denominator starts from the GPU's peak dense throughput for the data type you train in. For an H100 SXM, NVIDIA's datasheet lists 1,979 teraFLOPS for BF16 tensor cores, with a footnote that the figure assumes structured sparsity. Dense training gets half of that, about 989 teraFLOPS. Using the sparsity number is the single most common way to halve a schedule on paper. Check the footnotes for every GPU you budget, and see NVIDIA H100, Hopper architecture for how the tensor cores reach that peak.
No real job reaches peak. Model FLOPs utilisation (MFU) is the model's FLOPs per second, using the count above, divided by peak. It is lowered by communication not hidden behind compute, pipeline bubbles, memory-bound kernels, data loading stalls and small matrices. The Llama 3 paper reports 38 to 43 percent BF16 MFU across its large H100 configurations, a useful reference for well-tuned dense training at scale. Untuned jobs commonly sit well below that; LLM Bottleneck Analysis shows how to find what is holding yours down.
Hardware FLOPs utilisation (HFU) counts the FLOPs the GPU actually executed, including recomputation. With full activation checkpointing, each layer's forward runs twice, so the hardware does about 8N per token instead of 6N. A job at 40% MFU with full recompute is running at about 53% HFU. Budget with MFU, because recompute is overhead the model does not benefit from, and never report HFU as MFU.
Goodput: the time you actually train
MFU describes the job while it is running. A budget also needs goodput: the fraction of reserved time spent making forward progress. It is reduced by hardware failures and the work lost since the last checkpoint, restart time, checkpoint saves that block training, evaluation pauses and scheduled maintenance. Llama 3 reports more than 90 percent effective training time over a 54-day snapshot on a very large cluster, achieved with considerable automation. A new team should plan for less until it has measured its own. Faster, asynchronous checkpoints, as covered in GPU Checkpointing Deep Dive, raise goodput directly.
Worked example: 6.6B parameters, 2T tokens
Take a dense decoder with 32 layers, width 4,096, a gated feed-forward width of 11,008 and a 32,000-token vocabulary. Matmul parameters come to 6.607 billion. Train it on 2 trillion tokens at sequence length 4,096 on H100s in BF16, assuming 40% MFU and 90% goodput.
| Quantity | Formula | Value |
|---|---|---|
| FLOPs per token | 6N + 12·L·d·T | 4.6085 × 1010 |
| Total budget C | per-token × D | 9.217 × 1022 |
| Achieved rate per GPU | 989.5 TFLOPS × 0.40 | 395.8 TFLOPS |
| Tokens per second per GPU | rate ÷ FLOPs per token | about 8,590 |
| Compute GPU-hours | C ÷ rate ÷ 3,600 | 64,686 |
| Reserved GPU-hours | ÷ 0.90 goodput | 71,873 |
| Wall clock on 512 GPUs | ÷ 512 ÷ 24 | 5.85 days |
| Cost at an illustrative US$2.50 per GPU-hour | × price | about US$180,000 |
Sensitivity is the useful output. At 35% MFU the compute hours rise to 73,927; at 45% they fall to 57,499. A five-point MFU change moves the budget by more than 10 percent, which is why the first week of any project should be spent measuring MFU on the real configuration, not quoting someone else's.
Check the token count against compute-optimal scaling. Chinchilla's rule of thumb is about 20 tokens per parameter, which here is about 132 billion tokens. Two trillion is roughly 15 times that. This is a deliberate choice when the model will serve heavy inference traffic: a smaller model trained longer is cheaper to run, at the cost of more training compute than the loss alone would justify. The Chinchilla scaling article explains the trade.
Finally, sanity-check the method against a published run. Llama 3's 405-billion parameter model on 15.6 trillion tokens gives 6ND = 3.79 × 1025, matching the paper's stated 3.8 × 1025 FLOPs. At 40% MFU on about 16,000 H100s that is roughly 70 days of pure compute, the right order of magnitude for a run of that size.
Measuring MFU during the run
Once training starts, compute MFU from step time every few hundred steps and log it next to loss. The formula needs only the tokens per step, the per-token FLOPs you already computed and the GPU count.
PEAK_BF16_DENSE = 989.5e12 # H100 SXM, dense; change per GPU and dtype
def mfu(tokens_per_step, step_seconds, n_gpus, flops_per_tok, peak=PEAK_BF16_DENSE):
achieved = tokens_per_step * flops_per_tok / step_seconds
return achieved / (n_gpus * peak)
# global batch of 1,024 sequences x 4,096 tokens on 512 GPUs, measured 1.0 s per step
print(f"{mfu(1024 * 4096, 1.0, 512, 4.6085e10):.1%}") # 38.2%Count only real tokens. If sequences are padded rather than packed, padding inflates tokens per step and therefore MFU while doing no useful work; packing sequences fixes both. Exclude steps that include checkpoint saves or evaluation from the MFU average and account for them in goodput instead, so each number measures one thing. If the measured MFU sits well below plan, re-run the schedule with the measured value immediately; the budget is a living forecast, not a promise made once.
The full ledger
The main run is never the whole bill. A practical ledger reserves compute for the work around it. The split below is an example, not a standard; adjust it from your own history.
| Line item | Example share | Why it exists |
|---|---|---|
| Main pre-training run | 60 to 70% | The budget computed above, divided by goodput |
| Ablations and scaling probes | 10 to 20% | Small runs that choose data mix, LR and width |
| Evaluation | 3 to 5% | Periodic benchmark passes during and after training |
| Long-context and annealing phases | 5 to 10% | Budgeted with their own sequence length |
| Contingency | 10% | Loss spikes, rollbacks, a lower MFU than planned |
Precision changes the denominator, not the numerator. Training in FP8 on hardware that supports it raises peak throughput, but the FLOP count is the same and the achievable MFU against the higher peak is usually lower. Always state which peak an MFU figure uses. Mixed precision training covers the numerics.
Mistakes that cost a factor of two
| Mistake | Effect on the budget | Fix |
|---|---|---|
| Sparsity peak used for dense training | Schedule 2× too optimistic | Use the dense figure from the footnote |
| Input embedding counted in N | Small overcount; large for small models | Count matmul parameters only |
| Attention term ignored | Underestimate, worse at long context | Add 12·L·d·T per token |
| Recompute counted as useful | MFU inflated by up to a third | Report MFU and HFU separately |
| Padding counted as tokens | MFU inflated, data under-delivered | Pack sequences; count real tokens |
| Total MoE parameters used | Huge overestimate | Use active parameters per token |
| Goodput assumed to be 100% | Overruns of weeks | Measure restarts; plan 85 to 90% at best |
Trade-offs
A FLOPS budget is a model, and it trades precision for speed. The 6N rule ignores small operations and the attention convention overstates causal work, so two careful teams can disagree by a few percent; that is fine as long as each is consistent. The larger uncertainties are MFU and goodput, which is why the budget should be a range. Spending more GPUs shortens wall clock but usually lowers MFU through communication overhead. Whether a shorter calendar is worth a higher bill is a business decision; the budget's job is to make that trade visible. Turning the budget into a cluster order, with memory floors, fabric and lead times, is covered in GPU infrastructure planning.
What to do next
- Write down your model configuration and compute matmul parameters with embedding excluded and the output head included.
- Compute FLOPs per token as 6N plus the attention term, separately for each training phase's sequence length.
- Look up dense peak throughput for your GPU and data type, reading the sparsity footnote.
- Run a short job on the real configuration, measure MFU from step time, and replace any assumed value.
- Convert to GPU-hours, apply a measured or conservative goodput, and add ablation, evaluation and contingency lines.
- Log MFU and goodput throughout the run and re-forecast the end date whenever either drifts.