Ask what it costs to train a foundation model and you will usually hear one number: GPU hours for the final pretraining run multiplied by a rental rate. That number is real, and it is the smallest honest answer. A project also pays for the experiments that decided the architecture and data mix, the runs that crashed, post-training, evaluation, data and labels, storage, and above all the people who did the work. Budgets built from the headline number are routinely short by a multiple.
This article builds the whole-project ledger. It starts with what published figures actually measure, derives compute from first principles, puts every line into a small model you can run, and works through an illustrative 30B-parameter project to show which lines dominate and which inputs move the total. For the per-token view of a single run, see cost per token trained.
What a headline training cost measures
The DeepSeek-V3 technical report is a model of clear disclosure. Its Table 1 splits training into 2,664K H800 GPU hours for pretraining, 119K for context extension and 5K for post-training, 2,788K in total, and prices them at $5.576M assuming a rental price of $2 per GPU hour. The figures are internally consistent: the report gives 180K GPU hours per trillion tokens, and 180K x 14.8 trillion tokens is exactly 2,664K. The report then states that these costs include only the official training run, excluding prior research and ablation experiments on architectures, algorithms or data.
That sentence is the point of this article. The headline is one line of a ledger, and the authors said so. Meta's Llama 3.1 model card is another useful reference: 39.3M H100-80GB GPU hours cumulatively, of which 30.84M were for the 405B model. The rest is the 8B and 70B models, not research overhead. Read every published figure for its scope: which runs, which hardware, whose price and what was left out.
The price is a separate question from the hours. A rental rate bundles the hardware, power, facility and the provider's margin. If you own the cluster, the right rate is the amortised cost of a GPU hour: capital spread over its useful life plus power, cooling, networking and operations, divided by the hours you actually use, which is higher than it looks when utilisation is low. Deriving that figure is covered in GPU infrastructure cost. Either way, write the rate down as an assumption with a source and a date, because it is one of the inputs that moves the total most.
The project ledger
A complete project ledger has about ten lines. Compute lines are GPU hours times a rate; the rest are priced directly.
| Line | What it covers | How to estimate it |
|---|---|---|
| Final pretraining run | The run you ship | 6ND FLOPs / (peak x MFU) |
| Failures and restarts | Lost work since the last checkpoint, hangs, bad nodes | Final run x (1/goodput - 1) |
| Research and ablations | Scaling-law sweeps, data-mix and architecture tests, false starts | Ratio to final run, from your own history |
| Post-training compute | Context extension, SFT, preference tuning or RL with rollouts | Ratio, or a bottom-up rollout count |
| Evaluation and red-teaming | Benchmark inference on many checkpoints, safety testing | Checkpoints x suites x tokens |
| Data | Crawl processing, deduplication, filtering, licences | Bottom-up, per retained token |
| Human labels | SFT demonstrations, preference comparisons, expert review | Items x price per item |
| People | Researchers, infrastructure, data and evaluation engineers | FTEs x months x loaded cost |
| Storage and network | Datasets, checkpoints, logs, egress | TB-months plus transfers |
| Contingency | Re-runs, price changes, schedule slip | Percentage of the rest |
Each line has a different owner and a different way of going wrong. Ablation spend is decided by research leads and grows with uncertainty. Failure cost is set by infrastructure quality. People cost is set by calendar time, so a slipped schedule costs money even if no GPU runs. Data costs are covered in depth in the data curation cost article.
Compute from first principles
Dense transformer training costs about 6 x N x D FLOPs for N parameters and D tokens: two for the forward pass and four for the backward pass, per parameter per token. GPU hours follow from peak throughput and model FLOPs utilisation (MFU), the fraction of peak your stack actually delivers:
gpu_hours = 6 * N * D / (peak_flops_per_gpu * MFU * 3600)
30e9 params, 6e12 tokens, H100 dense BF16 989e12, MFU 0.40
FLOPs = 6 * 30e9 * 6e12 = 1.08e24
gpu_hours = 1.08e24 / (989e12*0.4*3600) = ~758,000Measure MFU on your own stack, model shape and cluster with a short run before you budget; do not borrow someone else's. As a sanity check, apply the formula to Llama 3.1 405B: 6 x 405e9 x 15e12 is about 3.6e25 FLOPs, and 30.84M H100 hours at 989 TFLOPS can deliver about 1.1e26. That implies an average of roughly a third of peak over everything those hours covered. Treat this as approximate, because the token count is given only as about 15 trillion. Sizing compute budgets in more detail is covered in FLOPs budgets for training.
The ledger as code
The model is deliberately small. Every input is a named, documented assumption that someone owns, and the output is a dictionary of lines rather than one number, so reviews argue about inputs instead of totals.
from dataclasses import dataclass, replace
@dataclass
class Plan:
params: float = 30e9 # dense parameters
tokens: float = 6e12 # pretraining tokens
peak_flops: float = 989e12 # H100 SXM dense BF16
mfu: float = 0.40 # measured on your stack, not assumed
gpu_hour_usd: float = 2.50 # illustrative rate
ablation_ratio: float = 1.0 # research compute / final-run compute
goodput: float = 0.90 # useful fraction after failures, restarts
posttrain_ratio: float = 0.10 # SFT + preference/RL compute / final run
eval_gpu_hours: float = 40_000 # benchmark and red-team inference
data_usd: float = 600_000 # curation compute, licences, filtering
labels_usd: float = 450_000 # human SFT and preference data
people: int = 12 # FTEs on the project
months: float = 9
fte_month_usd: float = 25_000 # fully loaded
storage_net_usd: float = 120_000
contingency: float = 0.15
def ledger(p: Plan) -> dict:
final_hours = 6 * p.params * p.tokens / (p.peak_flops * p.mfu * 3600)
rate = p.gpu_hour_usd
lines = {
"final pretraining run": final_hours * rate,
"failures and restarts": final_hours * (1 / p.goodput - 1) * rate,
"research and ablations": final_hours * p.ablation_ratio * rate,
"post-training compute": final_hours * p.posttrain_ratio * rate,
"evaluation and red-teaming": p.eval_gpu_hours * rate,
"data acquisition and curation": p.data_usd,
"human labels": p.labels_usd,
"people": p.people * p.months * p.fte_month_usd,
"storage and network": p.storage_net_usd,
}
lines["contingency"] = sum(lines.values()) * p.contingency
lines["TOTAL"] = sum(lines.values())
return lines
def swing(field, lo, hi):
a = ledger(replace(Plan(), **{field: lo}))["TOTAL"]
b = ledger(replace(Plan(), **{field: hi}))["TOTAL"]
return abs(b - a)
Worked example: an illustrative 30B project
Take an illustrative 30B dense model trained on 6 trillion tokens at 40 percent MFU, with research compute equal to the final run, 90 percent goodput, post-training at 10 percent of the final run, 12 people for nine months at $25,000 per fully loaded month, and a $2.50 GPU hour. None of these is a quote; replace each one with your own.
Three things stand out. The final run is about a fifth of the total. People are the largest single line. Research compute matches the final run line for line, and in a project still finding its recipe it can be several times larger. A budget that requested $1.9M for this project would have been short by about $7.5M.
Sensitivity: which inputs move the total
Change one input at a time across a plausible range and record the swing in the total:
| Input | Range | Total | Swing |
|---|---|---|---|
| Research ratio | 0.5x to 2x final run | $8.30M to $11.57M | $3.27M |
| People | 8 to 20 FTEs | $8.35M to $11.46M | $3.11M |
| GPU hour rate | $2.00 to $3.50 | $8.40M to $11.36M | $2.96M |
| Duration | 6 to 14 months | $8.35M to $11.11M | $2.76M |
| MFU | 50% to 30% | $8.42M to $10.99M | $2.57M |
| Goodput | 97% to 80% | $9.21M to $9.69M | $0.48M |
The research ratio and team size move the total as much as the GPU price does, and they are usually the least scrutinised inputs. Goodput matters less in dollars, but it matters in calendar time, which feeds back into people cost. Spend estimating effort in proportion to the swing.
Lines teams forget
- Evaluation inference. Running a full suite on every saved checkpoint, plus long-context and safety tests, is real inference spend.
- Checkpoints. With mixed-precision Adam, a saved training state (fp32 master weights and two moments, often plus bf16 weights) is roughly 12 to 16 bytes per parameter: about 0.4 to 0.5 TB for 30B parameters, per checkpoint, and many are kept.
- Idle reserved capacity. Reserved clusters bill while you debug a data loader. Bring-up and burn-in can take weeks.
- Minimum terms. Reservations often run longer than the job; the tail is either wasted or must be filled with useful work.
- Data reprocessing. A tokenizer or filter change late in the project re-runs the pipeline.
- Serving pilot. The first deployment needs capacity before revenue; see LLM serving TCO.
Tracking actuals against the plan
A forecast is useful only if actuals are compared against it. Tag every job at submission with a ledger line (ablation, final, posttrain, eval), record GPU hours from the scheduler, and produce a weekly burn chart per line against plan. Track goodput as useful training steps over paid GPU hours, so failure cost is measured rather than guessed. Re-forecast the total weekly from the current burn rate and remaining scope. Attribution across shared clusters is covered in LLM cost attribution.
The weekly review needs only three numbers per line: planned to date, actual to date and the forecast at completion. A line whose forecast exceeds plan by more than its share of contingency gets a decision that week: cut scope, move money from another line, or accept the overrun explicitly.
Put gates on the expensive transitions. Do not start the final run until the scaling-law fit and the data mix are frozen, a short run has measured MFU on the production configuration, and checkpoint restore has been tested end to end. Each gate converts a likely re-run into a small, planned cost.
Failure modes
- Budgeting the headline only. The project runs out of money after the first successful run, with post-training and evaluation unfunded.
- Assumed MFU. A budget at 50 percent on a stack that delivers 30 percent is two-thirds over on every compute line.
- Untagged jobs. Without line tags, ablation creep is invisible until the cluster invoice arrives.
- No restart drill. The first real failure costs days because restore was never tested.
- Contingency spent early. It gets used on research and is gone before the final run.
- Scope drift without re-forecast. Adding a modality, a longer context or a second model size mid-project multiplies several lines at once; re-run the ledger the day the scope changes, not at the quarterly review.
- Double counting shared capacity. When one cluster serves several projects, charging each project for the whole reservation inflates every budget and hides idle time.
Trade-offs
Spending more on ablations lowers the risk of a failed final run but delays it. Reserved capacity is cheaper per hour but bills through idle time; on-demand is flexible but may not be available at the scale you need. A smaller model trained longer costs more to train and less to serve. Buying labelled data is fast; building an internal labelling team is cheaper at volume but adds people cost and management time. None of these has a universal answer, which is why the ledger exposes them as inputs.
What to do next
- Copy the ledger model and replace every default with a number you can source or defend.
- Run a short job on the production configuration to measure MFU, and update the model.
- Pull last year's GPU hours by job type to estimate your real research-to-final ratio.
- Add a ledger-line tag to every job submission and build the weekly burn chart.
- Run the sensitivity table and spend estimating effort on the largest swings.
- Define the gates before the final run, including a tested checkpoint restore.