Teams that train and serve large language models produce a flood of metrics, and most of them belong to the system: utilisation panels, latency histograms, error budgets. Those are covered in LLM KPI Dashboards, in depth. This article is about a smaller and more personal set: the handful of numbers an individual contributor, the engineer writing the training loop, the inference kernels or the eval harness, should own, compute and move.
A good individual KPI ties a person's work to what the GPUs actually do. It is something the engineer can change by their own decisions, it is computed from data that already exists, and it comes with a second metric that makes gaming it pointless. The sections below define the core set by role, derive model FLOPs utilisation from first principles with the attention term included, work through a training and an inference example with numbers, give a small script that computes the KPIs from logs, and explain how to keep them out of the performance-review trap.
What makes a KPI individual
A dashboard metric answers "is the system healthy?" An individual KPI answers "did my last change make our GPUs do more useful work?" That changes the requirements. It must be:
- Controllable. The engineer's own decisions move it: a kernel, a parallelism layout, a batch policy, a data loader fix. Fleet-wide utilisation fails this test because scheduling and demand dominate it.
- A ratio of useful output to resource. Tokens per GPU-second, dollars per million tokens, GPU hours per accepted experiment. Raw counts reward spending more.
- Paired with a guard. Every efficiency number can be improved by doing worse work. A guard metric, such as loss parity, a latency SLO or an eval score, catches that.
- Cheap and reproducible. Computed by a script from step logs, scheduler accounting and eval results, not assembled by hand for a meeting.
The core set by role
| Role | Primary KPI | Guard metric | What moves it |
|---|---|---|---|
| Training engineer | MFU, and MFU x goodput | Loss curve matches the reference run | Kernels, parallelism layout, overlap of communication, checkpoint and restart time |
| Inference engineer | Cost per 1M output tokens at the SLO | p95 TTFT and inter-token latency SLOs, eval score | Batching policy, quantisation, KV cache layout, speculative decoding |
| Eval or data engineer | Regressions caught before release | False alarm rate, eval wall-clock | Coverage of the eval suite, statistical power, harness efficiency |
| Every IC | GPU-hours per accepted change | Cycle time from idea to decision | Smaller pilot runs, early stopping, reuse of cached results |
Four is enough. An engineer with ten KPIs has none, because no decision can be judged against all of them. Pick the row that matches the role, add the shared GPU-hours metric, and leave the rest on the team dashboard.
MFU from first principles
Model FLOPs utilisation (MFU) is the fraction of the hardware's peak arithmetic that went into the model's own computation. It was popularised by the PaLM paper (Chowdhery et al., 2022) because it is comparable across systems: it counts the FLOPs the model mathematically requires, not the FLOPs the implementation happened to run.
For a dense decoder-only transformer with N non-embedding parameters, one token costs about 2N FLOPs in the forward pass and 4N in the backward pass, 6N in total. Attention adds a term that grows with sequence length. Following PaLM's accounting, with L layers, H heads, head dimension Q and sequence length T, training FLOPs per token are:
flops_per_token = 6 * N + 12 * L * H * Q * T
MFU = tokens_per_second * flops_per_token / (num_gpus * peak_flops_per_gpu)Hardware FLOPs utilisation (HFU) counts what the hardware actually executed. With full activation recomputation, every layer's forward pass runs twice, adding 2N + 4LHQT per token, so HFU is higher than MFU for the same run. Report MFU as the KPI, because recomputation is a cost, not useful work; report HFU beside it to show how much of the gap is recomputation rather than idle hardware.
Use the dense peak for the datatype you train in. For an H100 SXM in BF16 that is 989 TFLOPS dense; NVIDIA's headline 1,979 figure assumes 2:4 structured sparsity, which ordinary training does not use. Using the sparse figure halves every MFU you report. For inference, the forward pass alone is about 2N FLOPs per generated token plus attention over the cache, but decoding is usually memory-bound, so MFU is the wrong KPI there; use cost per token at the SLO instead.
Worked example: a training run
A worked example. A team trains an 8B-parameter dense model (L = 32, H = 32, Q = 128) at sequence length T = 8,192 on 64 H100 SXM GPUs in BF16, and the step logs show 400,000 tokens per second.
| Quantity | Calculation | Value |
|---|---|---|
| 6N | 6 x 8e9 | 4.80e10 FLOPs per token |
| Attention term | 12 x 32 x 32 x 128 x 8,192 | 1.29e10 FLOPs per token (21% of the total) |
| FLOPs per token | sum | 6.09e10 |
| Peak | 64 x 989e12 | 6.33e16 FLOPs per second |
| MFU | 400,000 x 6.09e10 / 6.33e16 | 38.5% |
| HFU with full recompute | adds 2N + 4LHQT per token | 51.3% |
Dropping the attention term would report 30.3 percent and understate the engineer's work by a fifth; at long sequence lengths the error grows. That is why the formula, not a rule of thumb, belongs in the KPI script.
MFU measures a running job. Goodput measures how much of the allocated time produced progress that survived. Suppose the job holds the 64 GPUs for a week, 168 hours. It fails three times; each failure loses on average 30 minutes of work since the last hourly checkpoint plus 20 minutes to restart, 2.5 hours in total. Data loader stalls and evaluation pauses cost another 3.4 hours. Goodput is (168 - 5.9) / 168 = 96.5 percent, and the effective number, MFU x goodput, is 37.1 percent. An engineer who halves checkpoint and restart time moves this KPI as surely as one who writes a faster kernel, which is the point: see checkpointing in depth for that lever.
The guard is loss parity. A change that raises MFU but changes the numerics, a different reduction order or a lower precision, must reproduce the reference loss curve within its run-to-run noise over a fixed number of tokens before the gain counts.
Inference: cost per token at the SLO
For an inference engineer the useful output is tokens delivered within the latency objective, and the resource is GPU time. The KPI is cost per million output tokens, measured at the operating point where the SLO still holds:
cost_per_1M = gpu_cost_per_hour / (tokens_per_second_per_gpu_at_slo * 3600) * 1_000_000With an assumed internal rate of $2.50 per GPU-hour and 2,500 output tokens per second per GPU at the point where p95 time to first token is still inside its target, cost is about $0.28 per million tokens. The rate is an illustration; use your own accounting figure and keep it fixed across comparisons so the KPI moves only when the engineering does.
The phrase "at the SLO" carries the weight. Throughput measured with an unbounded queue always looks better, because larger batches amortise weight reads, but users experience the latency. Run a load sweep, find the highest offered load at which p95 TTFT and inter-token latency meet their targets, and report throughput there. The latency definitions are in LLM Metrics Deep Dive and the budget logic in LLM SLO Burn Rate Alerts. The second guard is quality: quantisation or a speculative decoding change must hold the eval score within its confidence interval.
Experiment and eval KPIs
The shared KPI is GPU-hours per accepted change: total GPU-hours an engineer's experiments consumed in a period, divided by the number of changes that were merged or decisions that were made. It rewards pilot runs at small scale, early stopping when a run is clearly worse, caching eval results, and killing jobs that have lost their purpose. Its guard is cycle time, the days from proposal to decision, so that nobody improves it by simply running fewer experiments.
For an eval engineer the primary KPI is regressions caught before release, measured as the share of quality regressions found later in production that the pre-release suite would have flagged. Its guard is the false alarm rate, because a suite that flags everything catches everything. Harness efficiency, eval GPU-hours per release, is a secondary number; LLM Eval Frameworks, in depth covers where that compute goes.
Computing the KPIs from logs
Compute everything from the same records, in one script that lives in the repository and runs on a schedule. The sketch below assumes step logs with token counts and timestamps, and scheduler records of allocated GPU-hours; adapt the field names to your own logging.
from dataclasses import dataclass
H100_BF16_DENSE = 989e12 # FLOPs/s per GPU, dense (no 2:4 sparsity)
@dataclass
class ModelShape:
n_params: float # non-embedding parameters
layers: int
heads: int
head_dim: int
seq_len: int
def train_flops_per_token(m: ModelShape) -> float:
return 6 * m.n_params + 12 * m.layers * m.heads * m.head_dim * m.seq_len
def mfu(tokens: int, seconds: float, gpus: int, m: ModelShape, peak=H100_BF16_DENSE) -> float:
return (tokens / seconds) * train_flops_per_token(m) / (gpus * peak)
def goodput(allocated_hours: float, lost_hours: float) -> float:
return (allocated_hours - lost_hours) / allocated_hours
def cost_per_million(gpu_cost_per_hour: float, tok_per_s_per_gpu_at_slo: float) -> float:
return gpu_cost_per_hour / (tok_per_s_per_gpu_at_slo * 3600) * 1e6
def gpu_hours_per_accepted(jobs, accepted_ids) -> float:
used = sum(j["gpu_hours"] for j in jobs)
accepted = len(set(accepted_ids))
return float("inf") if accepted == 0 else used / accepted
shape = ModelShape(8e9, 32, 32, 128, 8192)
print(f"MFU {mfu(400_000 * 3600, 3600, 64, shape):.1%}") # 38.5%
print(f"goodput {goodput(168, 5.9):.1%}") # 96.5%
print(f"$/1M tokens {cost_per_million(2.50, 2500):.2f}") # 0.28Store the inputs alongside each result: the model shape, the peak figure used, the token window, the commit. A KPI whose inputs cannot be reproduced will be argued about rather than acted on.
Goodhart and the review trap
Goodhart's law applies with full force. Once a number becomes a target for a person, it stops measuring what it used to. The common forms on GPU teams are predictable: raising MFU by increasing sequence length or batch size until convergence suffers; lowering cost per token by quietly loosening the latency target; improving GPU-hours per accepted change by running only safe experiments.
Three rules keep the KPIs honest. First, every KPI has its guard, and a change only counts if the guard holds. Second, the KPIs belong to the engineer as a tool for choosing work, and are reviewed by the engineer weekly; they are inputs to a conversation with a manager, not a score in a performance review. Third, attribution is honest: a fleet driver upgrade or a new GPU generation moves every number, and the script should mark such step changes rather than credit them to whoever was on call.
Failure modes
- Sparse peak in the denominator. Using 1,979 rather than 989 TFLOPS for an H100 in BF16 halves the reported MFU and makes comparisons with published numbers meaningless.
- No attention term. 6N alone understates MFU, by about a fifth in the example and more at long context.
- Throughput without the SLO. Cost per token measured on a saturated queue rewards changes that make users wait.
- Counting warm-up. The first steps include compilation and cache warming; exclude a fixed window so runs compare fairly.
- Changing several things at once. If a run changes kernels and batch size together, the KPI cannot say which one helped.
- Hand-built numbers. A spreadsheet assembled before a meeting drifts from the logs; compute from the source records only.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| MFU over HFU as headline | Comparable across implementations | Hides recompute work the engineer did |
| MFU x goodput | Rewards reliability work too | Needs scheduler accounting to be accurate |
| Cost per token at SLO | Matches what users and finance feel | Needs a load sweep for every change |
| Few KPIs per person | Clear decisions | Some useful work is invisible; record it in notes |
| Engineer-owned review | Honest numbers | Less convenient for managers who want a score |
What to do next
- Pick the primary KPI for your role from the table above, plus GPU-hours per accepted change.
- Write down the guard metric for each and the threshold at which a gain stops counting.
- Put the formulas in a script in the repository, with the dense peak for your GPU and datatype and the attention term included.
- Backfill the last month from step logs and scheduler records to get a baseline.
- For every change you ship, record the KPI before and after, with the guard, the commit and the measurement window.
- Review the numbers yourself weekly and bring them to one-to-ones as evidence for choosing the next piece of work, not as a score.