"How many H100-hours did that model take?" is the question behind every training budget, cluster reservation and build-versus-buy decision. A few model developers publish the answer, and those disclosures are the best calibration data you will find. They are also easy to misread: they mix hardware generations, count different stages, and quietly leave out the experiments that came before the final run.

This article does three things. It collects the disclosures that are specific enough to use, quoted from model cards and technical reports. It back-solves each one into delivered throughput per GPU, so you can see what utilisation the developers actually achieved. Then it turns that into a method, with code, for estimating the GPU-hours of your own run and for checking the estimate while the run is in flight. The FLOP accounting itself is covered in FLOPS budget for LLM training; this page is about the hours.

What a GPU-hour measures

A GPU-hour is one accelerator held for one hour. It is a unit of allocation, not of work: a GPU that spent the hour waiting on a straggler, reloading a checkpoint or idling in a failed node still counts. That distinction drives everything below. The chain from work to hours has four links:

model_flops   = 6 * N * D                  # dense transformer, training; N = active params
ideal_hours   = model_flops / (peak_flops_per_s * 3600)
step_hours    = ideal_hours / MFU          # kernels, communication, pipeline bubbles
billed_hours  = step_hours / goodput       # restarts, checkpoint stalls, lost work

MFU, model FLOPs utilisation, is the fraction of peak the training steps achieve, counting only the model's useful FLOPs. Goodput is the fraction of allocated time spent making forward progress. For the H100 SXM, the dense BF16 peak is about 989 TFLOP/s; the often-quoted 1,979 figure assumes structured sparsity, which training does not use. Mixing those two peaks is the single most common factor-of-two error in GPU-hour estimates.

From a FLOP budget to billed GPU-hoursModel FLOPs6 N D (+ attention)/ peakIdeal hoursat 100% of peak/ MFUStep hourskernels as measured/ goodputPre-training hoursrestarts, stalls, ckpt+ other stageslong context, anneal, post-train+ R&D overheadablations, failed runs, evalsBilled GPU-hourswhat the budget paysPublished model-card numbers usually cover the first row plus some later stages; never the R&D overhead.
Every published GPU-hour figure sits somewhere on this chain; read the source to find out where before comparing it with anything.

The published numbers

Three sources, covering five models, are precise enough to calibrate against. All figures below are as published; nothing is interpolated.

ModelHardwarePublished GPU-hoursWhat is countedTokens
Llama 3.1 405BH100-80GB30.84MModel card: total GPU time to train the model15.6T (paper)
Llama 3.1 70BH100-80GB7.0MSame~15T (card)
Llama 3.1 8BH100-80GB1.46MSame~15T (card)
DeepSeek-V3 (671B total, 37B active)H8002.788M2.664M pre-train + 0.119M context extension + 0.005M post-train14.8T
Llama 2 70BA100-80GB1.72MModel card: total GPU time2T

The Llama 3.1 card gives 39.3M GPU-hours for the three models together. The DeepSeek-V3 report prices its 2.788M hours at an assumed rental rate of $2 per GPU-hour, $5.576M, and states explicitly that the figure excludes prior research and ablation experiments. It also gives a rate: 180K H800 GPU-hours per trillion tokens, about 3.7 days on its 2,048-GPU cluster. The H800 is the export variant of the H100 with reduced NVLink bandwidth, so its hours are close to, but not interchangeable with, H100-hours.

Back-solving delivered throughput

Back-solving divides the model FLOPs by the published hours to get delivered throughput per GPU-hour, averaged over everything the figure includes. Using 6ND and ignoring attention FLOPs:

def delivered_tflops(active_params, tokens, gpu_hours):
    return 6 * active_params * tokens / (gpu_hours * 3600) / 1e12

delivered_tflops(405e9, 15.6e12, 30.84e6)   # 341  -> 34.5% of 989
delivered_tflops(70e9,  15e12,   7.0e6)     # 250  -> 25.3%
delivered_tflops(8e9,   15e12,   1.46e6)    # 137  -> 13.9%
delivered_tflops(37e9,  14.8e12, 2.664e6)   # 343  (DeepSeek-V3 pre-training, H800)
delivered_tflops(70e9,  2e12,    1.72e6)    # 136  -> 43.5% of the A100's 312

The 405B number is the one to study, because the Llama 3 paper also publishes the inside of the chain. Its 405B pre-training ran at 38 to 43 percent BF16 MFU, between 380 and 430 TFLOP/s per GPU depending on the stage, on up to 16,384 GPUs, and the team reports more than 90 percent effective training time. The paper's own budget is 3.8e25 FLOPs; at 400 TFLOP/s that is 26.4M ideal step-hours. Divide by roughly 0.9 goodput and you get about 29M, close to the card's 30.84M, and the card's figure also covers the stages after the main pre-training. The chain reconciles to within a few percent.

The smaller models do not reconcile as neatly: 137 TFLOP/s for the 8B is about 14 percent of peak. The card does not break the hours down, so the gap cannot be attributed from public data. Plausible contributors are a less efficient parallel layout at small scale, where communication and fixed overheads are a larger share, and stages that the token count does not capture. The lesson is general: small models are not proportionally cheaper per token in GPU-hours, and a back-of-envelope estimate that assumes 40 percent MFU will be too optimistic for them.

DeepSeek-V3 is the interesting contrast. Its active-parameter FLOPs per GPU-hour, 343 TFLOP/s, match the 405B's, but it trained in FP8, whose dense peak is twice BF16's, and as a mixture-of-experts it moves more data between GPUs per FLOP. Compare MoE and dense runs on active parameters, and never compare hours across precisions without saying which peak you used.

Where the hours go: goodput at scale

Goodput losses are not exotic at this scale. During a 54-day snapshot of Llama 3 405B pre-training the team logged 466 job interruptions: 47 planned, for firmware and configuration updates, and 419 unexpected. About 78 percent of the unexpected ones were attributed to confirmed or suspected hardware issues, and GPU issues alone were 58.7 percent. That is roughly eight unexpected interruptions a day. Each one costs the time to detect the failure, restart the job and redo the work since the last checkpoint.

A simple model shows how those terms combine. With F failures per day, a restart cost of R minutes and a checkpoint interval of C minutes, each failure loses about R + C/2 minutes of the whole cluster:

def goodput(failures_per_day, restart_min, ckpt_interval_min, ckpt_stall_min):
    lost_per_failure = restart_min + ckpt_interval_min / 2      # expected redo
    failure_loss = failures_per_day * lost_per_failure / (24 * 60)
    ckpt_loss = ckpt_stall_min / ckpt_interval_min               # blocking saves
    return max(0.0, 1 - failure_loss - ckpt_loss)

goodput(8, 10, 30, 0.5)   # 0.844: eight failures a day, 10 min restart, 30 min checkpoints
goodput(8, 5, 10, 0.2)    # 0.924: faster restart and async checkpoints pay for themselves

These inputs are illustrative, but the shape holds: at thousands of GPUs, restart time and checkpoint interval are worth as much as a kernel optimisation. See training checkpointing for the mechanics and DGX H100 system architecture for the node the hours are billed on.

Estimating your own run

Turn the calibration into an estimator with explicit, named assumptions. Every input is something you can measure early and correct later.

from dataclasses import dataclass

@dataclass
class RunPlan:
    active_params: float          # N, active per token for MoE
    tokens: float                 # D
    peak_tflops: float = 989.0    # H100 SXM dense BF16; set to the precision you train in
    mfu: float = 0.35             # measure on a short run at target scale; never assume 0.5
    goodput: float = 0.90         # from your cluster's failure rate and restart time
    attn_overhead: float = 0.05   # extra FLOPs beyond 6ND; grows with sequence length
    later_stages: float = 0.08    # long-context, annealing, post-training, as a fraction
    rnd_multiplier: float = 1.5   # ablations and failed runs; 1.0 if you only count the final

    def hours(self):
        flops = 6 * self.active_params * self.tokens * (1 + self.attn_overhead)
        step = flops / (self.peak_tflops * 1e12 * self.mfu) / 3600
        final_run = step / self.goodput * (1 + self.later_stages)
        return final_run, final_run * self.rnd_multiplier

final, programme = RunPlan(active_params=13e9, tokens=4e12).hours()
# final ~ 0.32M H100-hours; programme ~ 0.47M. On 1,024 GPUs the final run is ~12.8 days.

Walk through the example. A 13B dense model on 4T tokens is 3.12e23 model FLOPs, 3.28e23 with the attention allowance. At 35 percent MFU of 989 TFLOP/s each GPU delivers 346 TFLOP/s, so the steps need about 263K GPU-hours; dividing by 0.9 goodput and adding 8 percent for later stages gives about 316K for the final run, and the 1.5 multiplier makes the programme about 473K. On 1,024 GPUs the final run is about 12.8 days of wall clock, before any queueing for the reservation itself.

Two parameters deserve the most scepticism. MFU must come from a measured run at the target parallel layout, not from a paper, because it varies with model size, sequence length and interconnect. The R&D multiplier is a business decision, not a constant: published final-run figures are lower bounds on what a programme costs, and the gap can be large.

Converting between GPUs and into energy

Disclosures come in A100-hours, H800-hours and H100-hours, and quotes from vendors come in all three. Convert through delivered FLOPs, never through a ratio of spec-sheet peaks. Llama 2 70B delivered about 136 TFLOP/s per A100-hour; the Llama 3.1 405B run delivered about 341 per H100-hour. The ratio, roughly 2.5, is much lower than the ratio of dense BF16 peaks, which is about 3.2, because the larger run paid for 16,384-way parallelism and long context. If you plan to move a workload from A100s to H100s, measure both on your own layout rather than assuming the peak ratio carries over.

Hours also convert into energy, which increasingly sets the budget. The Llama 3.1 card lists a 700 W TDP per H100. Multiplying 30.84M hours by 0.7 kW gives about 21.6 GWh as a GPU-only upper bound for the 405B; the facility draws more once host CPUs, networking, storage and cooling are added, so apply your data centre's measured overhead factor before using the number in a power or carbon plan.

Tracking hours during the run

Once the run starts, replace assumptions with measurements every day:

  • Log tokens processed and GPU-hours allocated per job attempt, including failed attempts. Realised GPU-hours per trillion tokens is the number to compare with the plan, and with DeepSeek-V3's published 180K.
  • Compute MFU from step time: 6 times active parameters times tokens per step, divided by step time times GPU count times peak. A falling MFU usually means a slow node or a degraded link.
  • Compute goodput as tokens trained divided by the tokens the cluster could have trained at the measured step rate over the allocated time.
  • Re-forecast completion with the realised rates and publish the new finish date. A plan that is 10 percent behind after a week will not catch up on its own.

Mistakes that skew estimates

MistakeEffectFix
Using the sparse peak (1,979 BF16)Hours underestimated by 2xUse dense peak for the precision trained
Comparing MoE on total parametersMoE looks impossibly efficientUse active parameters per token
Treating a card figure as the programme costBudget short by the R&D shareAdd ablations, failed runs and evals explicitly
Mixing H800, A100 and H100 hoursWrong per-GPU throughputConvert via delivered FLOPs, then re-divide
Assuming small models scale down linearlyOptimistic plan for small runsMeasure MFU at the real size and layout
Ignoring sequence lengthAttention FLOPs missing at long contextAdd the attention term, or measure

What to do next

  1. Reproduce the back-solve table above from the published figures; it takes five minutes and calibrates your intuition.
  2. Run a short job at your target scale and parallel layout to measure MFU before you commit to a reservation.
  3. Estimate goodput from your cluster's actual failure log, restart time and checkpoint interval with the goodput function.
  4. Fill in RunPlan for your model and write down every assumption with its source.
  5. Track realised GPU-hours per trillion tokens daily and re-forecast weekly.
  6. Convert hours into money with cost per token trained, and review the H100 itself in the Hopper architecture article.
Key takeaway: Published GPU-hour figures are the best calibration data available, but each counts a different slice of the work and none includes the experiments before the final run. Back-solve them into delivered TFLOP/s using active parameters and the dense peak, estimate your own run as FLOPs over measured MFU and goodput plus later stages and an explicit R&D multiplier, and replace every assumption with a measurement as soon as the run starts.