How much will it cost to train a frontier model in 2027? The question is usually answered with a single headline number, and those numbers rarely say what they include. This article builds the answer from parts. It starts with the best public measurement of the trend, breaks it into an identity you can reason about, calibrates the identity against two published training runs, and projects 2024 to 2027 with explicit uncertainty. You leave with a small model you can run and update yourself, plus an understanding of which lever moves which number.

Most readers do not train frontier models. The trajectory still matters to you. The same forces that raise frontier spend also make a fixed level of capability cheaper every year, and they decide whether a model you would like to train next year fits your budget. All forward-looking figures here are projections from stated assumptions, not reports.

Which cost? Three definitions

There are three costs, and they are routinely mixed up:

  • Final-run compute cost. This is the accelerators, servers, networking and energy used by the one training run that produced the released model. It is either amortised from purchase price or priced at cloud rental rates.
  • Programme compute cost. This is the final run plus every experiment, ablation, failed run and restart. It is often several times larger than the final run.
  • Full development cost. This is programme compute plus research and engineering staff, data acquisition and evaluation.

The choice of method matters too. Epoch AI's 2024 study estimated final-run costs for frontier models. It found cloud-rental estimates averaged about twice the hardware-amortisation estimates, because rental prices include the provider's margin and idle capacity. Always ask which cost and which method a number uses before comparing it with another.

The measured trend

Epoch AI (Cottier and colleagues, May 2024) estimated costs for frontier models, defined as models in the top ten by training compute at release. They found the amortised hardware and energy cost of the final run grew about 2.4x per year since 2016, with a 95% confidence interval of 2.0x to 3.1x. Cloud-rental estimates grew about 2.6x per year. For flagship models in their sample, hardware (accelerators, servers and interconnect) was 47-67% of development cost and R&D staff was 29-49%. Energy was only 2-6%. Their method added 23% on top of server costs for cluster-level networking. Extrapolating the trend, they wrote that the largest training runs would cost more than a billion dollars by 2027.

Two features of that result matter for using it. First, it is a fit to a few dozen estimates, each with wide error bars. The confidence interval is wide, and three years at 2.0x per year gives 8x while three years at 3.1x gives about 30x. Second, energy is a small share of cost even though power is a binding constraint on where clusters can be built. Cheap electricity does not make training cheap. Available power decides how big a cluster you can build at all.

The study's data end in early 2024, so everything after that date in this article is projection. A trend measured over eight years is a reasonable default for three more, but it is not a law. It continues only while investors keep funding larger runs and while the returns from scale stay visible. Treat the 2.4x figure as a prior that you update as new disclosures arrive, not as a schedule.

An identity for training cost

Training cost = FLOPs x price per FLOP, and each factor has its own trendTraining FLOPs6 x params x tokens, grows fastPeak FLOP/s per GPUnew generation, lower precisionPrice per GPU-hourcapex amortisation or rentalUtilisation (MFU, goodput)what share of peak you getPrice per useful FLOPGPU-hour price / (peak x MFU)Final-run cost= FLOPs x price per useful FLOP; total programme cost adds experiments, failures, staffFrontier spend has grown about 2.4x a year because FLOPs grew faster than price per FLOP fell.
The decomposition used in this article. Each box can be estimated separately, and the trend in the bottom line is the product of the trends above it.

Write final-run cost as an identity:

cost = FLOPs / (peak_flops_per_gpu * MFU * 3600) * price_per_gpu_hour
     = FLOPs * price_per_useful_flop

FLOPs for a dense transformer are about 6 x parameters x training tokens. The method and its corrections are covered in FLOPS budget for LLM training. Peak throughput rises with each hardware generation and again with each drop in precision. FP8 doubles the dense peak relative to BF16 on Hopper-class parts. MFU (model FLOPs utilisation) is the share of peak that becomes useful model arithmetic. Large dense runs commonly report around 35-45%. Price per GPU-hour is either your amortised cost or a rental rate.

The trend in cost is therefore the trend in FLOPs divided by the trend in useful FLOPs per dollar. Frontier compute has grown much faster than price-performance has improved, so cost has risen. Algorithmic efficiency, meaning better architectures, data and training recipes, does not show up in this identity directly. It lowers the FLOPs needed for a given capability. Frontier labs spend the saving on more capability rather than lower bills. That is why the cost to reach a fixed benchmark score falls quickly while the frontier bill rises.

Calibrating on published runs

Two published runs make good calibration points, because both disclosed GPU-hours.

Llama 3.1 405B. Meta's model card reports 30.84 million H100 GPU-hours for the 405B model, and 39.3 million for the whole Llama 3.1 family. The paper gives about 3.8 x 1025 FLOPs, which matches 6 x 405B parameters x about 15.6T tokens. Back-solve the utilisation:

flops      = 3.8e25
gpu_hours  = 30.84e6
peak_bf16  = 989e12           # H100 SXM dense BF16, FLOP/s
mfu = flops / (gpu_hours * 3600 * peak_bf16)
print(f"implied MFU = {mfu:.0%}")                 # about 35%

for price in (2.0, 3.0, 4.0):                     # $/GPU-hour, your assumption
    print(price, f"${gpu_hours * price / 1e6:,.0f}M")   # about $62M, $93M, $123M

An implied MFU of about 35% is plausible for a run that also absorbed failures and restarts. The dollar figure depends entirely on the price you assume, which is why published estimates for the same run differ by a factor of two.

DeepSeek-V3. The technical report gives 2.788 million H800 GPU-hours. At an assumed $2 per GPU-hour, that is $5.576 million. The report states that this covers the final run only and excludes prior research and ablations. The low figure comes from fewer active parameters per token (a mixture-of-experts model), FP8 training and heavy systems work. It is a final-run number, not a programme cost, so it should not be compared with full development budgets. The per-run method is worked through in H100-hours per foundation model.

Projecting 2024 to 2027 with uncertainty

Now project. The forecast below treats the growth rate as uncertain and propagates it, instead of picking one line. The base is a parameter. Set it to a 2024 frontier final-run cost you believe, with its method stated:

import random, statistics

def project(base_2024, years=(2025, 2026, 2027), trials=20000,
            g_lo=2.0, g_hi=3.1, g_mid=2.4, seed=1):
    rng = random.Random(seed)
    out = {y: [] for y in years}
    for _ in range(trials):
        g = rng.triangular(g_lo, g_hi, g_mid)       # one growth path per trial
        for y in years:
            out[y].append(base_2024 * g ** (y - 2024))
    for y in years:
        xs = sorted(out[y])
        p10, p50, p90 = (xs[int(q * trials)] for q in (0.1, 0.5, 0.9))
        print(y, f"p10 ${p10/1e6:,.0f}M  median ${p50/1e6:,.0f}M  p90 ${p90/1e6:,.0f}M")

project(base_2024=100e6)     # replace with your own base and method

With an illustrative $100M base, the simulation gives a 2027 median of about $1.5B, with p10 near $1.1B and p90 near $2.3B. The edges of the range are easier to read as fixed-rate paths:

YearLow path (2.0x/yr)Central (2.4x/yr)High path (3.1x/yr)
2025$200M$240M$310M
2026$400M$576M$961M
2027$800M$1.38B$2.98B

The spread between the fixed-rate paths by 2027 is almost 4x, and that comes from growth-rate uncertainty alone. Uncertainty in the base adds more. That is the honest form of the claim: if the trend continues, the largest final runs land somewhere around one to a few billion dollars by 2027, and a single-point forecast hides most of what is known. Things that would bend the curve: power and grid connections limiting cluster size, chip supply, a shift of spend from pre-training to post-training and inference-time compute, and a fall in rental prices as supply grows. The last of these lowers cost per FLOP without changing the hardware.

What moves each factor

Each factor in the identity has its own driver over 2024-2027. It helps to see them side by side, because a headline about one factor is often read as a statement about the total.

FactorWhat moves itDirection for a fixed modelDirection at the frontier
Training FLOPsParameters, tokens, post-training and reasoning-style RLFalls as recipes improveRises: labs spend savings on scale
Peak FLOP/s per GPUNew generations, FP8 and lower-precision formatsRisesRises
UtilisationParallelism strategy, failures, checkpoint stalls, networkRises with engineeringOften falls as clusters grow
Price per GPU-hourSupply, depreciation schedule, rental competition, powerOlder GPUs get cheaperNewest GPUs carry a premium
Programme overheadExperiments, ablations, restartsFalls with reuse of known recipesRises with novelty

Two rows deserve comment. Utilisation tends to fall as cluster size grows. More GPUs mean more frequent hardware failures, longer collective operations and more time lost to checkpointing and restarts. A recipe that reaches 45% MFU on a thousand GPUs may give less on tens of thousands. Goodput, the share of wall-clock time that advances training, is the measure to track. Price per GPU-hour also behaves differently by generation. When a new generation arrives, the rental price of the previous one tends to fall. A team that does not need the newest part can ride that curve, while the frontier pays the premium for it.

The practical lesson is that you should forecast the factors separately and multiply, rather than extrapolate a total. When a new data point arrives, such as a price change, a new chip or a measured MFU, you can update the one factor it touches and see its effect on the total.

Using the trajectory to plan your own runs

For a team planning its own run, the same identity becomes a budgeting tool:

  • Price per useful FLOP is where you have leverage. Moving MFU from 30% to 45% cuts cost by a third. Moving from BF16 to FP8 where accuracy allows can nearly halve it. Both are engineering work, not purchases.
  • Budget programme cost, not final-run cost. Multiply the final-run estimate by a factor for experiments and failures. Set the factor from your own history. If you have none, use a conservative 2-3x and revise it.
  • Price the hardware path you will really use. A rental quote, a reserved contract and owned hardware amortised over three years give very different per-hour figures. See cost per token trained for marginal versus loaded cost.
  • Re-plan yearly. Fixed capability gets cheaper every year. A run that does not fit this year's budget may fit next year's with a newer recipe and cheaper GPU-hours.

Failure modes in cost estimates

Common errors in training-cost estimates:

  • Peak instead of achieved FLOPs. Dividing FLOPs by peak throughput understates GPU-hours by 2-3x. Always include MFU.
  • Sparse peak numbers. Vendors often quote structured-sparsity peaks that are double the dense figure. Training uses dense throughput.
  • Mixing methods. Comparing one lab's amortised cost with another's rental-priced estimate builds a 2x error into the comparison.
  • Mixing scopes. Comparing a final-run figure with a full development budget.
  • Ignoring the confidence interval. Compounding a point estimate for three years turns a modest error in the rate into a large error in the answer.
  • Assuming energy dominates. At 2-6% of cost it rarely does, but available power can cap the cluster you can build.

What to do next

  1. Write down which cost you mean: final run, programme or full development, and amortised or rental.
  2. Estimate FLOPs for your planned run from parameters and tokens.
  3. Measure MFU on a short run on your real stack instead of assuming it.
  4. Price GPU-hours from quotes you can actually get, then compute final-run cost with the identity.
  5. Multiply by your experiment-and-failure factor, and add staff and data.
  6. Run the projection code with your own base and range, and report p10 to p90, not a single number.
  7. Revisit the numbers every quarter as prices, hardware and recipes change.
Key takeaway: Frontier final-run cost has grown about 2.4x a year, with wide uncertainty, because compute demand outran gains in price per useful FLOP. Decompose any estimate into FLOPs, peak, utilisation and GPU-hour price. Calibrate on disclosed runs, project with a range rather than a point, and spend your effort on utilisation and precision, which are the levers you control.