MLPerf is the benchmark suite that hardware vendors quote when they say a chip trains or serves a model faster than its rival. It is run by MLCommons, an industry consortium, and its results are peer reviewed by the other submitters before publication. That review is what separates MLPerf from a vendor slide. It is also why the results are easy to misread: every number is true under a precise set of rules, and those rules are rarely quoted next to the number.
This article explains how the two main suites work, Training and Inference. It covers what is timed, the four inference scenarios and LLM latency limits, how many runs a training result averages, what closed and open divisions permit, and how to turn a published result into a decision about your own hardware. Specifics are checked against the MLCommons rules repositories and round announcements as of October 2026. Rules change every round, so treat the numbers here as a snapshot and check the round you are reading.
What MLPerf measures
MLCommons publishes several benchmark suites. Two matter most for GPU buyers. MLPerf Training measures time to train a reference model to a fixed quality target on a fixed dataset. MLPerf Inference measures how fast a system serves a reference model under a defined traffic pattern while staying above an accuracy target and, for some scenarios, below a latency bound. Each suite publishes results roughly twice a year. Recent rounds were Inference v5.1 (September 2025), v6.0 (April 2026) and v6.1 (September 2026), and Training v5.1 (November 2025) and v6.0 (June 2026).
Workloads rotate to track the industry. Training v5.1 had seven benchmarks: RetinaNet object detection, DLRMv2 recommendation, Llama 3.1 405B pretraining, Llama 2 70B LoRA fine-tuning, Llama 3.1 8B pretraining (replacing BERT), the R-GAT graph network and Flux.1 text-to-image (replacing Stable Diffusion v2). Training v6.0 added two mixture-of-experts benchmarks, DeepSeek V3 at 671B parameters and GPT-OSS 20B. On the inference side, v5.1 added DeepSeek-R1 reasoning, Llama 3.1 8B and Whisper Large v3. v6.0 added a GPT-OSS 120B benchmark, an interactive DeepSeek-R1 scenario that permits speculative decoding, and the suite's first multimodal and video models. v6.1 added an end-to-end RAG pipeline and an edge agentic benchmark. The lesson for readers is that a result is only comparable with results from the same benchmark and version.
How an inference result is produced
Inference results are produced by LoadGen, a C++ library with Python bindings maintained by MLCommons. LoadGen generates queries according to the scenario, records when each is issued and completed, and writes the logs that the submission checker validates. The query sample library (QSL) holds the dataset samples the queries refer to. Everything else is the system under test (SUT): the submitter's pre-processing, serving engine, batching policy, quantised model and hardware. Because LoadGen owns the clock and the arrival pattern, a submitter cannot choose convenient query timings.
Each submission includes an accuracy run, which processes the full dataset and must reach the benchmark's quality target, and a performance run, which produces the reported number. Compliance tests check for shortcuts such as caching results across queries. Submissions then go through a review period in which competitors can challenge results before publication.
Scenarios and latency limits
The scenario defines the traffic pattern and the metric. These come from the current inference rules:
| Scenario | How LoadGen issues queries | Metric |
|---|---|---|
| Single stream | Next query as soon as the previous completes, one sample each, at least 600 s | 90th-percentile latency (early-stopping estimate) |
| Multistream | Next query as soon as the previous completes, 8 samples each, at least 600 s | 99th-percentile latency (early-stopping estimate) |
| Server and Interactive | Poisson arrivals, one sample each, at least 600 s, latency bound per benchmark | Highest Poisson rate that meets the bound |
| Offline | All samples in one query at the start, at least 24,576 samples | Measured throughput |
For LLM benchmarks the latency bound is two numbers, time to first token (TTFT) and time per output token (TPOT), each checked at the 99th percentile. The current rules set these limits:
| Benchmark | Server (conversational) TTFT / TPOT | Interactive TTFT / TPOT |
|---|---|---|
| Llama 2 70B | 2000 ms / 200 ms | 450 ms / 40 ms |
| Llama 3.1 8B | 2000 ms / 100 ms | 500 ms / 30 ms |
| Llama 3.1 405B | 6000 ms / 175 ms | 4500 ms / 80 ms |
| DeepSeek-R1 | 2000 ms / 80 ms | 1500 ms / 15 ms |
| GPT-OSS 120B | 3000 ms / 80 ms | 2000 ms / 20 ms |
Offline answers "how much work per second at any latency". Server answers "how much traffic while users stay happy". The ratio between them tells you how much throughput the latency target costs on that system, often more useful than either number alone. Interactive scenarios set much tighter TPOT bounds, so they reward small effective batch sizes and techniques like speculative decoding where the rules permit them.
Statistics, early stopping and a simulation
Tail percentiles from short runs are noisy, so the rules set minimum durations and use an early-stopping criterion. After the minimum duration, LoadGen counts how many queries exceeded the latency bound and computes how many queries are needed to be 99% confident that the true percentile meets the target. A system comfortably inside the bound can stop sooner; a system near the bound must run longer; one above it never passes. The rules suggest a starting point of 24,576 queries for a 90th percentile and 270,336 for a 99th.
Why the Server metric is the maximum sustainable rate, and why finite runs need care, is easy to see in a simulation. The script below models eight identical workers with log-normal service times and searches for the highest Poisson rate that keeps the 99th-percentile latency under a bound. It is a teaching model of the scenario, not LoadGen.
import heapq
import random
def p99_latency(rate, workers, n_queries=20000, seed=0):
"""Poisson arrivals at `rate` per second into `workers` identical FIFO servers."""
rng = random.Random(seed)
free_at = [0.0] * workers
t, lat = 0.0, []
for _ in range(n_queries):
t += rng.expovariate(rate)
service = rng.lognormvariate(-2.3, 0.5) # median about 0.10 s
start = max(t, heapq.heappop(free_at))
heapq.heappush(free_at, start + service)
lat.append(start + service - t)
lat.sort()
return lat[int(0.99 * len(lat))]
def max_rate(bound, workers, lo=1.0, hi=1000.0):
for _ in range(40): # bisection on the arrival rate
mid = (lo + hi) / 2
lo, hi = (mid, hi) if p99_latency(mid, workers) <= bound else (lo, mid)
return lo| p99 bound | Max Poisson rate | Share of offline ceiling (70.3 q/s) |
|---|---|---|
| 1.00 s | 68.1 q/s | 97% |
| 0.60 s | 65.3 q/s | 93% |
| 0.45 s | 61.8 q/s | 88% |
| 0.35 s | 51.0 q/s | 73% |
Mean service time was 0.114 s, so eight workers give an offline ceiling of 70.3 queries per second. The 99th-percentile service time alone is 0.324 s, so no rate can meet a 0.30 s bound. As the bound tightens towards that floor, the sustainable rate falls quickly. The result at 1.0 s is also flattering: near saturation, queues grow slowly, and a run of 20,000 queries ends before they have grown. That is exactly why LoadGen enforces a minimum duration.
How training is timed and scored
A training result is wall-clock time from the start of the clock until the model reaches the benchmark's quality target, such as an evaluation loss or log perplexity. The current rules allow up to 30 minutes of untimed model initialisation in the closed division and 4 hours in the open division. The clock must start before the system touches the dataset, includes on-the-clock preprocessing, training and evaluation, and may stop as soon as the target is reached.
Convergence time is random, so each benchmark requires multiple runs: 10 for most benchmarks and 3 for the largest LLM pretraining tasks. The score drops the fastest and slowest runs and averages the rest, and one non-converging run may be dropped as the slowest. To stop submitters winning by luckier hyperparameters, results are checked against reference convergence points (RCPs): reference runs at several batch sizes that establish how many epochs or samples convergence normally takes. A submission that converges faster than the RCP within a statistically acceptable range is normalised back to the RCP mean, so the score reflects system speed rather than a lucky schedule.
Large training results are therefore as much a test of the interconnect and software as of the GPU. Scaling to hundreds of accelerators exposes collective communication, as discussed in NCCL collectives, along with data loading and keeping every node healthy for the whole timed run.
Divisions and categories
Every result sits in a division and a category, and comparisons across them are not valid.
- Closed division. The model must be equivalent to the reference, with the same preprocessing; training must use reference hyperparameters except those the rules allow to change. Inference permits approved quantisation if the accuracy run still meets the target, typically 99% or 99.9% of the reference score depending on the benchmark. This is the division for comparing hardware.
- Open division. Different models, hyperparameters, preprocessing or data order are allowed, which makes it a showcase for techniques such as pruning or distillation rather than hardware comparisons.
- Network division (inference). The SUT receives queries over a network, which adds serving overheads the other divisions omit.
- Availability categories. Available systems can be bought or rented now; Preview systems must become available by a later round; Research, Development or Internal systems carry no availability promise.
- Power. Some inference submissions include measured system power alongside performance. Compare energy figures only with others that also measured power.
Worked example: reading two results
Suppose two published closed-division results for Llama 2 70B Offline report 98,000 tokens per second on an 8-accelerator system A and 60,000 tokens per second on a 4-accelerator system B. These figures are invented for illustration. Work through them in order.
- Check that they are comparable. Same benchmark version, same scenario, same accuracy target (99% and 99.9% are separate entries), same division and availability category.
- Normalise carefully. Per accelerator, A delivers 12,250 tokens/s and B 15,000. That division is your arithmetic, not an official MLPerf metric, and it ignores host CPUs, interconnect and power that differ between systems.
- Compare Server with Offline. If A's Server result is 85% of its Offline result and B's is 60%, B struggles with the latency target and its advantage may vanish under your SLOs.
- Read the software stack. Results name the inference engine, its version and the precision. A large jump between rounds on the same hardware is usually software; you get it only if you run that stack.
- Map to your workload. The Llama 2 70B dataset averages about 294 output tokens per sample with prompts up to 1,024 tokens. If your prompts are 8,000 tokens, prefill dominates and the ranking can change. Use capacity estimation to translate.
Reproducing a result on your hardware
The most valuable use of MLPerf is as a reproducible baseline for your own hardware. Each round's results repository holds the submitters' code, configurations and logs, and several vendors publish step-by-step reproduction guides for their submissions. A practical procedure:
- Pick the result closest to your hardware, and clone the exact repository tag for that round.
- Build the submitter's container and run the accuracy mode first; a performance number without a passing accuracy run is meaningless.
- Run the performance mode with the published settings and compare. Within a few percent means your system is healthy; a large gap usually points to clocks, power caps, NUMA placement, interconnect topology or driver versions.
- Change one thing at a time, such as your own sequence lengths, batch limits or engine version, and measure the effect with profiling rather than guessing.
- Keep the run as a regression test for driver and firmware upgrades.
Failure modes in interpretation
- Comparing across versions or scenarios. A v6.1 Server result and a v5.1 Offline result answer different questions.
- Treating Offline as capacity. Production traffic arrives randomly and has latency targets; Server or Interactive is closer to reality, and still not your traffic.
- Ignoring system size. Headline records often come from very large systems. Throughput per accelerator at that scale may not hold at your scale.
- Assuming the software is available to you. Closed-division numbers use tuned, sometimes pre-release stacks; check what your provider actually runs.
- Reading open-division wins as hardware wins. Different models make those results technique demonstrations.
- Missing benchmarks. Vendors submit where they look good. Absence from a benchmark tells you something too.
Trade-offs
| Approach | Strength | Weakness |
|---|---|---|
| MLPerf closed results | Peer-reviewed, fixed rules, comparable hardware data | Reference models and datasets, not your traffic |
| Reproducing MLPerf locally | Validates your systems against a known-good number | Effort to build vendor stacks |
| Your own benchmark | Your prompts, lengths and SLOs | No external comparability, easy to bias; see custom benchmarks |
Use MLPerf to narrow a shortlist and to validate a deployment, and your own benchmark to make the final decision.
What to do next
- Identify the MLPerf benchmark and scenario closest to your workload, usually an LLM Server or Interactive entry.
- Filter results to closed division, Available category, same version and accuracy target.
- Compute Server-to-Offline ratios and per-accelerator figures, labelled as your own arithmetic.
- Note each result's engine, version and precision, and confirm you can run that stack.
- Reproduce one result on your hardware, accuracy run first, and investigate any gap above a few percent.
- Re-run with your own sequence lengths and latency targets before committing to a purchase.
- Check the next round's rules when it lands; benchmarks and limits change every six months.