Most experiment tracking setups answer one question well: did the loss go down? On GPU clusters that is only half of what a run record has to explain. The other half is efficiency and cost: how many GPU-hours the run consumed, how much of each step the hardware spent doing useful math, and whether a run that looks faster really is faster or simply ran on a different machine, driver or library version. Without that half, teams compare a run on one cluster with a run on another and draw conclusions about a code change that are really about hardware.
This article builds the efficiency and cost side of a tracking record: a stack fingerprint used as a comparison key, step timing that is correct on an asynchronous device, model FLOPs utilisation (MFU) and goodput, a cost ledger, and a run diff that refuses unfair comparisons. Loss curves, token axes and lineage are covered in LLM training experiment tracking, and server topology in MLflow for GPU training ops.
The four layers of a GPU run record
Think of a GPU run record as four layers, each with a different source and update rate:
- Identity and intent, written once at launch: config, git commit, dataset version, parent sweep, and the question the run is meant to answer.
- Stack fingerprint, written once per node at start: GPU model, driver, CUDA and library versions, interconnect, clocks and power limits.
- Efficiency series, written every N steps: step time and its breakdown, tokens per second, MFU, memory high-water mark.
- Cost ledger, written at start, restart and end: allocated GPU-hours, productive GPU-hours, restarts and their causes.
The data flow matters as much as the schema. Each rank measures locally, rank 0 reduces and buffers, and a background thread flushes to the tracking server so a slow network never stalls a step. Cluster telemetry from the DCGM exporter stays in your metrics system and is joined to the run by run id and time range rather than copied into the tracker.
Fingerprinting the stack
The fingerprint answers the question every surprising result raises first: was anything different underneath? Collect it on every node, not just rank 0, because mixed nodes in one job are a common cause of stragglers. The probe below uses only standard PyTorch calls and documented nvidia-smi query fields.
import hashlib, json, platform, socket, subprocess
import torch
def stack_fingerprint():
smi = subprocess.run(
["nvidia-smi",
"--query-gpu=name,driver_version,memory.total,clocks.max.sm,power.limit,"
"ecc.mode.current,mig.mode.current",
"--format=csv,noheader"],
capture_output=True, text=True, check=True).stdout.strip().splitlines()
fp = {
"host": socket.gethostname(),
"python": platform.python_version(),
"torch": torch.__version__,
"cuda_runtime": torch.version.cuda,
"cudnn": torch.backends.cudnn.version(),
"nccl": ".".join(map(str, torch.cuda.nccl.version())),
"gpus": [line.strip() for line in smi],
}
# The comparison key ignores the hostname: same stack on another node is still comparable.
comparable = {k: v for k, v in fp.items() if k != "host"}
fp["stack_hash"] = hashlib.sha256(
json.dumps(comparable, sort_keys=True).encode()).hexdigest()[:12]
return fpGather the fingerprints to rank 0, log the full set as an artifact (for MLflow, mlflow.log_dict(fps, 'fingerprint.json')), and set the distinct stack hashes as a tag. If a job reports more than one stack hash, flag the run before anyone reads its throughput. The full replay manifest, with seeds and data order, is a separate concern covered in LLM training reproducibility; the fingerprint here is deliberately narrower, a key for comparing performance.
Measuring step time on an asynchronous device
GPU kernels run asynchronously: a Python call returns as soon as the work is queued. Timing a step with time.time() around the forward and backward pass therefore measures how fast the CPU enqueues work, not how long the GPU takes, until the queue fills and the numbers suddenly jump. Measure intervals between synchronisation points instead, and separate the data wait so a slow loader is visible as a loader problem.
import time
import torch
class StepTimer:
def __init__(self, every=50):
self.every, self.n, self.data_wait, self.tokens = every, 0, 0.0, 0
torch.cuda.synchronize()
self.t0 = time.perf_counter()
def batch_ready(self, waited_s, tokens):
self.data_wait += waited_s
self.tokens += tokens
def step_done(self):
self.n += 1
if self.n % self.every:
return None
torch.cuda.synchronize() # one sync per interval, not per step
wall = time.perf_counter() - self.t0
out = {"step_time_s": wall / self.every,
"data_wait_frac": self.data_wait / wall,
"tokens_per_s": self.tokens / wall}
self.t0, self.data_wait, self.tokens = time.perf_counter(), 0.0, 0
return outA synchronisation every 50 steps costs almost nothing; one every step can cost several percent by draining the queue. Reduce the interval figures across ranks and log the maximum step time as well as the mean: in synchronous data parallel training the slowest rank sets the pace, and a rising max-to-mean ratio points to a straggler node. For a breakdown inside the step, use the profiler for a short window rather than permanent instrumentation; see LLM performance profiling.
Log memory in the same interval. torch.cuda.max_memory_allocated() returns the peak bytes held by tensors since the last torch.cuda.reset_peak_memory_stats(), so read it and reset it at each logging point to get a per-interval high-water mark. torch.cuda.max_memory_reserved() adds what the caching allocator holds but has not handed out; a large and growing gap between the two signals fragmentation. Both matter for experiment tracking because headroom is what decides whether the next run can raise its micro-batch size, and a run that sits at 98 percent of device memory is one longer sequence away from an out-of-memory crash. Record the device total from the fingerprint next to the peak so the dashboard can show headroom as a percentage rather than a raw byte count that means different things on 40 GB and 80 GB parts.
From step time to MFU and goodput
Tokens per second is not comparable across model sizes or GPU types. MFU is: the fraction of the hardware's peak arithmetic that went into the model's own math. For dense transformer training, a standard approximation is 6 times the parameter count FLOPs per token (2 for the forward pass, 4 for the backward pass). That approximation leaves out attention score FLOPs, which grow with sequence length, and it deliberately does not count activation recomputation, which is extra work rather than model work. State the formula you used in the run tags so numbers stay comparable.
def mfu(params, tokens_per_s_per_gpu, peak_flops_per_gpu):
achieved = 6 * params * tokens_per_s_per_gpu
return achieved / peak_flops_per_gpu
# Peak must match the precision you train in, dense, without structured sparsity.
PEAK_BF16_DENSE = {"H100-SXM": 989e12, "A100": 312e12}Worked example: a 7-billion-parameter model on eight H100 SXM GPUs reaches 9,000 tokens per second per GPU. That is 6 times 7e9 times 9,000, or 378 TFLOPS per GPU, and 378 divided by the BF16 dense peak of about 989 TFLOPS gives an MFU of 38.2 percent. The same model on eight A100 GPUs at 3,100 tokens per second per GPU achieves 130 TFLOPS against a 312 TFLOPS peak, an MFU of 41.7 percent. The A100 run uses its hardware slightly better and is still almost three times slower, which is why MFU and throughput must be logged side by side.
MFU measures the step. Goodput measures the job: productive GPU time, the steps that survive into the final model, divided by all GPU time you were allocated. Restarts that replay steps since the last checkpoint, slow startup, and nodes held idle while waiting for the rest of the job all lower goodput without touching MFU.
Cost per run, and cost per result
Continue the example with a 50-billion-token training budget. At 72,000 tokens per second across the eight H100s, the job needs 192.9 hours of productive time, or 1,543 GPU-hours. Suppose two node failures each replay about six hours of work, so 12.4 hours are lost; goodput is 192.9 / 205.3, about 94 percent. Prices vary widely by provider and contract, so treat the following as assumptions, not quotes: at an assumed 2.50 dollars per H100 GPU-hour the productive part costs about 3,860 dollars. The A100 run needs 560 hours and 4,480 GPU-hours; at an assumed 1.20 dollars per A100 GPU-hour that is about 5,380 dollars, more expensive despite the higher MFU and the cheaper hour.
Record cost as a ledger, not a single number:
| Field | Written when | Why |
|---|---|---|
| allocated_gpu_hours | Every heartbeat and at end | What the bill follows |
| productive_gpu_hours | At each checkpoint | Numerator of goodput |
| restarts, with cause | On each resume | Separates hardware faults from code faults |
| price_per_gpu_hour | At launch, as a tag | Makes cost reproducible when prices change |
| cost_per_eval_point | After each evaluation | What a unit of improvement cost |
The last row is the one leadership asks about. Cost per run is easy; cost per percentage point on the evaluation that matters is what decides whether a line of research continues.
Comparing runs fairly
A run comparison tool should refuse, or at least loudly warn, when the comparison is not fair. The minimal rule: efficiency metrics are comparable only when stack hash, GPU count, global batch size and sequence length match; otherwise compare MFU, not throughput, and label it.
COMPARE_KEYS = ["stack_hash", "num_gpus", "global_batch", "seq_len", "precision"]
def diff_runs(a, b, metrics=("tokens_per_s", "mfu", "goodput", "cost_per_eval_point")):
mismatched = [k for k in COMPARE_KEYS if a["tags"].get(k) != b["tags"].get(k)]
rows = []
for m in metrics:
va, vb = a["summary"].get(m), b["summary"].get(m)
if va is None or vb is None:
continue
rows.append((m, va, vb, (vb - va) / va))
if mismatched:
rows = [r for r in rows if r[0] in ("mfu", "goodput", "cost_per_eval_point")]
return {"not_comparable_on": mismatched, "rows": rows}Summaries should use the median of interval measurements after warm-up, not the mean over the whole run, because the first intervals include compilation and cache warming. For sweeps, attach this record to every child run, including runs stopped early: their GPU-hours are part of what the sweep cost. The search side is covered in hyperparameter search on GPU clusters.
Failure modes
- Unsynchronised timing. Step times measured without a device sync look excellent until the queue fills, then jump; the trend is an artefact.
- Wrong peak. Using a sparsity peak, or an FP8 peak for a BF16 run, halves or quarters the reported MFU and makes every run look broken.
- Mixed stacks in one job. One node with an older driver or lower power limit becomes a straggler; without per-node fingerprints it looks like noise.
- Logging that blocks. Synchronous calls to the tracker from inside the step turn a network blip into a GPU stall on every rank.
- Cost without restarts. Counting only the final segment's hours understates cost, sometimes badly for long jobs on unreliable capacity.
- Cross-hardware throughput charts. Plotting tokens per second for runs on different GPU types in one chart invites the wrong conclusion.
Trade-offs
Every extra field costs storage and attention. Log the fingerprint once, the efficiency series at the same interval as the loss, and leave per-kernel detail to short profiler windows. A full replay manifest is worth its cost for long and expensive runs; for small sweeps, the stack hash and config are usually enough. MFU with the 6N approximation is easy to compare across teams but understates long-sequence work; if attention dominates, log a second, attention-inclusive figure and keep both clearly named.
What to do next
- Add the fingerprint probe to your launcher, log it per node, and tag runs with the set of stack hashes.
- Replace per-step wall-clock timing with interval timing between device syncs, and log data wait separately.
- Compute MFU with a documented formula and a peak that matches your precision, and record both in tags.
- Start a cost ledger with allocated and productive GPU-hours, restarts and an explicit price assumption.
- Make your comparison dashboards check stack hash, GPU count, batch and sequence length before showing throughput deltas.
- Report cost per evaluation point for every finished sweep.