Ask what it costs to train a foundation model and you will usually hear one number: GPU hours for the final pretraining run multiplied by a rental rate. That number is real, and it is the smallest honest answer. A project also pays for the experiments that decided the architecture and data mix, the runs that crashed, post-training, evaluation, data and labels, storage, and above all the people who did the work. Budgets built from the headline number are routinely short by a multiple.

This article builds the whole-project ledger. It starts with what published figures actually measure, derives compute from first principles, puts every line into a small model you can run, and works through an illustrative 30B-parameter project to show which lines dominate and which inputs move the total. For the per-token view of a single run, see cost per token trained.

What a headline training cost measures

The DeepSeek-V3 technical report is a model of clear disclosure. Its Table 1 splits training into 2,664K H800 GPU hours for pretraining, 119K for context extension and 5K for post-training, 2,788K in total, and prices them at $5.576M assuming a rental price of $2 per GPU hour. The figures are internally consistent: the report gives 180K GPU hours per trillion tokens, and 180K x 14.8 trillion tokens is exactly 2,664K. The report then states that these costs include only the official training run, excluding prior research and ablation experiments on architectures, algorithms or data.

That sentence is the point of this article. The headline is one line of a ledger, and the authors said so. Meta's Llama 3.1 model card is another useful reference: 39.3M H100-80GB GPU hours cumulatively, of which 30.84M were for the 405B model. The rest is the 8B and 70B models, not research overhead. Read every published figure for its scope: which runs, which hardware, whose price and what was left out.

The price is a separate question from the hours. A rental rate bundles the hardware, power, facility and the provider's margin. If you own the cluster, the right rate is the amortised cost of a GPU hour: capital spread over its useful life plus power, cooling, networking and operations, divided by the hours you actually use, which is higher than it looks when utilisation is low. Deriving that figure is covered in GPU infrastructure cost. Either way, write the rate down as an assumption with a source and a date, because it is one of the inputs that moves the total most.

The project ledger

A complete project ledger has about ten lines. Compute lines are GPU hours times a rate; the rest are priced directly.

LineWhat it coversHow to estimate it
Final pretraining runThe run you ship6ND FLOPs / (peak x MFU)
Failures and restartsLost work since the last checkpoint, hangs, bad nodesFinal run x (1/goodput - 1)
Research and ablationsScaling-law sweeps, data-mix and architecture tests, false startsRatio to final run, from your own history
Post-training computeContext extension, SFT, preference tuning or RL with rolloutsRatio, or a bottom-up rollout count
Evaluation and red-teamingBenchmark inference on many checkpoints, safety testingCheckpoints x suites x tokens
DataCrawl processing, deduplication, filtering, licencesBottom-up, per retained token
Human labelsSFT demonstrations, preference comparisons, expert reviewItems x price per item
PeopleResearchers, infrastructure, data and evaluation engineersFTEs x months x loaded cost
Storage and networkDatasets, checkpoints, logs, egressTB-months plus transfers
ContingencyRe-runs, price changes, schedule slipPercentage of the rest

Each line has a different owner and a different way of going wrong. Ablation spend is decided by research leads and grows with uncertainty. Failure cost is set by infrastructure quality. People cost is set by calendar time, so a slipped schedule costs money even if no GPU runs. Data costs are covered in depth in the data curation cost article.

Compute from first principles

Dense transformer training costs about 6 x N x D FLOPs for N parameters and D tokens: two for the forward pass and four for the backward pass, per parameter per token. GPU hours follow from peak throughput and model FLOPs utilisation (MFU), the fraction of peak your stack actually delivers:

gpu_hours = 6 * N * D / (peak_flops_per_gpu * MFU * 3600)

30e9 params, 6e12 tokens, H100 dense BF16 989e12, MFU 0.40
  FLOPs     = 6 * 30e9 * 6e12            = 1.08e24
  gpu_hours = 1.08e24 / (989e12*0.4*3600) = ~758,000

Measure MFU on your own stack, model shape and cluster with a short run before you budget; do not borrow someone else's. As a sanity check, apply the formula to Llama 3.1 405B: 6 x 405e9 x 15e12 is about 3.6e25 FLOPs, and 30.84M H100 hours at 989 TFLOPS can deliver about 1.1e26. That implies an average of roughly a third of peak over everything those hours covered. Treat this as approximate, because the token count is given only as about 15 trillion. Sizing compute budgets in more detail is covered in FLOPs budgets for training.

The ledger as code

The model is deliberately small. Every input is a named, documented assumption that someone owns, and the output is a dictionary of lines rather than one number, so reviews argue about inputs instead of totals.

from dataclasses import dataclass, replace

@dataclass
class Plan:
    params: float = 30e9            # dense parameters
    tokens: float = 6e12            # pretraining tokens
    peak_flops: float = 989e12      # H100 SXM dense BF16
    mfu: float = 0.40               # measured on your stack, not assumed
    gpu_hour_usd: float = 2.50      # illustrative rate
    ablation_ratio: float = 1.0     # research compute / final-run compute
    goodput: float = 0.90           # useful fraction after failures, restarts
    posttrain_ratio: float = 0.10   # SFT + preference/RL compute / final run
    eval_gpu_hours: float = 40_000  # benchmark and red-team inference
    data_usd: float = 600_000       # curation compute, licences, filtering
    labels_usd: float = 450_000     # human SFT and preference data
    people: int = 12                # FTEs on the project
    months: float = 9
    fte_month_usd: float = 25_000   # fully loaded
    storage_net_usd: float = 120_000
    contingency: float = 0.15

def ledger(p: Plan) -> dict:
    final_hours = 6 * p.params * p.tokens / (p.peak_flops * p.mfu * 3600)
    rate = p.gpu_hour_usd
    lines = {
        "final pretraining run": final_hours * rate,
        "failures and restarts": final_hours * (1 / p.goodput - 1) * rate,
        "research and ablations": final_hours * p.ablation_ratio * rate,
        "post-training compute": final_hours * p.posttrain_ratio * rate,
        "evaluation and red-teaming": p.eval_gpu_hours * rate,
        "data acquisition and curation": p.data_usd,
        "human labels": p.labels_usd,
        "people": p.people * p.months * p.fte_month_usd,
        "storage and network": p.storage_net_usd,
    }
    lines["contingency"] = sum(lines.values()) * p.contingency
    lines["TOTAL"] = sum(lines.values())
    return lines

def swing(field, lo, hi):
    a = ledger(replace(Plan(), **{field: lo}))["TOTAL"]
    b = ledger(replace(Plan(), **{field: hi}))["TOTAL"]
    return abs(b - a)

Worked example: an illustrative 30B project

Take an illustrative 30B dense model trained on 6 trillion tokens at 40 percent MFU, with research compute equal to the final run, 90 percent goodput, post-training at 10 percent of the final run, 12 people for nine months at $25,000 per fully loaded month, and a $2.50 GPU hour. None of these is a quote; replace each one with your own.

Illustrative 30B project ledger: $9.39M total; the final run is about 20%final pretraining run$1.90Mfailures and restarts$0.21Mresearch and ablations$1.90Mpost-training compute$0.19Mevaluation and red-teaming$0.10Mdata acquisition and curation$0.60Mhuman labels$0.45Mpeople$2.70Mstorage and network$0.12Mcontingency (15%)$1.22MBlue: compute. Yellow: data. Red: people. Grey: storage and contingency.
Output of the ledger with the default Plan. Bar lengths are to scale.

Three things stand out. The final run is about a fifth of the total. People are the largest single line. Research compute matches the final run line for line, and in a project still finding its recipe it can be several times larger. A budget that requested $1.9M for this project would have been short by about $7.5M.

Sensitivity: which inputs move the total

Change one input at a time across a plausible range and record the swing in the total:

InputRangeTotalSwing
Research ratio0.5x to 2x final run$8.30M to $11.57M$3.27M
People8 to 20 FTEs$8.35M to $11.46M$3.11M
GPU hour rate$2.00 to $3.50$8.40M to $11.36M$2.96M
Duration6 to 14 months$8.35M to $11.11M$2.76M
MFU50% to 30%$8.42M to $10.99M$2.57M
Goodput97% to 80%$9.21M to $9.69M$0.48M

The research ratio and team size move the total as much as the GPU price does, and they are usually the least scrutinised inputs. Goodput matters less in dollars, but it matters in calendar time, which feeds back into people cost. Spend estimating effort in proportion to the swing.

Lines teams forget

  • Evaluation inference. Running a full suite on every saved checkpoint, plus long-context and safety tests, is real inference spend.
  • Checkpoints. With mixed-precision Adam, a saved training state (fp32 master weights and two moments, often plus bf16 weights) is roughly 12 to 16 bytes per parameter: about 0.4 to 0.5 TB for 30B parameters, per checkpoint, and many are kept.
  • Idle reserved capacity. Reserved clusters bill while you debug a data loader. Bring-up and burn-in can take weeks.
  • Minimum terms. Reservations often run longer than the job; the tail is either wasted or must be filled with useful work.
  • Data reprocessing. A tokenizer or filter change late in the project re-runs the pipeline.
  • Serving pilot. The first deployment needs capacity before revenue; see LLM serving TCO.

Tracking actuals against the plan

A forecast is useful only if actuals are compared against it. Tag every job at submission with a ledger line (ablation, final, posttrain, eval), record GPU hours from the scheduler, and produce a weekly burn chart per line against plan. Track goodput as useful training steps over paid GPU hours, so failure cost is measured rather than guessed. Re-forecast the total weekly from the current burn rate and remaining scope. Attribution across shared clusters is covered in LLM cost attribution.

The weekly review needs only three numbers per line: planned to date, actual to date and the forecast at completion. A line whose forecast exceeds plan by more than its share of contingency gets a decision that week: cut scope, move money from another line, or accept the overrun explicitly.

Put gates on the expensive transitions. Do not start the final run until the scaling-law fit and the data mix are frozen, a short run has measured MFU on the production configuration, and checkpoint restore has been tested end to end. Each gate converts a likely re-run into a small, planned cost.

Failure modes

  • Budgeting the headline only. The project runs out of money after the first successful run, with post-training and evaluation unfunded.
  • Assumed MFU. A budget at 50 percent on a stack that delivers 30 percent is two-thirds over on every compute line.
  • Untagged jobs. Without line tags, ablation creep is invisible until the cluster invoice arrives.
  • No restart drill. The first real failure costs days because restore was never tested.
  • Contingency spent early. It gets used on research and is gone before the final run.
  • Scope drift without re-forecast. Adding a modality, a longer context or a second model size mid-project multiplies several lines at once; re-run the ledger the day the scope changes, not at the quarterly review.
  • Double counting shared capacity. When one cluster serves several projects, charging each project for the whole reservation inflates every budget and hides idle time.

Trade-offs

Spending more on ablations lowers the risk of a failed final run but delays it. Reserved capacity is cheaper per hour but bills through idle time; on-demand is flexible but may not be available at the scale you need. A smaller model trained longer costs more to train and less to serve. Buying labelled data is fast; building an internal labelling team is cheaper at volume but adds people cost and management time. None of these has a universal answer, which is why the ledger exposes them as inputs.

What to do next

  1. Copy the ledger model and replace every default with a number you can source or defend.
  2. Run a short job on the production configuration to measure MFU, and update the model.
  3. Pull last year's GPU hours by job type to estimate your real research-to-final ratio.
  4. Add a ledger-line tag to every job submission and build the weekly burn chart.
  5. Run the sensitivity table and spend estimating effort on the largest swings.
  6. Define the gates before the final run, including a tested checkpoint restore.
Key takeaway: The final pretraining run is one line of a foundation-model budget, often around a fifth of it. Build the whole ledger: research, failures, post-training, evaluation, data, labels, people, storage and contingency. Derive compute from 6ND and measured MFU, keep every input as a named assumption, find the inputs with the largest swings, and tag jobs so actuals can be compared with the plan every week.