vLLM ships its own benchmark suite, and most published vLLM performance numbers, including the ones in pull requests that claim a speedup, come from it. Older releases kept the tools as scripts in the repository's benchmarks/ folder with names like benchmark_serving.py; current releases expose them as subcommands of the CLI: vllm bench serve, vllm bench throughput and vllm bench latency. They answer three different questions, attach at different points in the stack, and have defaults that silently shape the result.

This article explains what each tool measures, how its load model works, which flags change the numbers, and how to run a reproducible sweep that finds the concurrency your SLO allows. Flag names and defaults here were checked against the vLLM CLI documentation in October 2026; the tool evolves quickly, so confirm with vllm bench serve --help on the version you deploy.

Three tools, three questions

Where each vLLM benchmark attachesvllm bench serveload generator, client sideHTTP streamvllm serveOpenAI-compatible APIEnginescheduler, batchingModel runnerGPU kernels, paged KVbench throughputoffline, whole datasetbench latencyoffline, fixed batchmeasures TTFT, TPOT, ITL,E2E latency, goodputno HTTP, no tokenizer-in-API pathincludes HTTP, detokenisation,queueing and arrival process
bench serve drives a running server over HTTP; bench throughput and bench latency construct the engine in-process and skip the API layer.

The split between offline and online tools is the most important thing to understand. vllm bench latency and vllm bench throughput build the engine inside the benchmark process. There is no HTTP server, no request parsing and no network, so they isolate the engine and the GPU. vllm bench serve is a client: it sends requests to an already running vllm serve (or another OpenAI-compatible endpoint) and measures what a user would see.

ToolQuestion it answersKey defaults
bench latencyHow long does one fixed batch take end to end?input 32, output 128, batch 8, 10 warm-up and 30 measured iterations
bench throughputHow fast can the engine chew through a whole dataset?sharegpt dataset, 1000 prompts, all submitted at once
bench serveWhat latency do clients see at a given load?random dataset, 1024 in / 128 out, 1000 prompts, request rate inf

Use latency for kernel and small-batch work, such as checking a quantisation or attention backend change. Use throughput to compare engine configurations for batch jobs where nobody waits on a single response. Use serve for anything that has an SLO.

What bench serve measures

The serving benchmark streams every response and timestamps each chunk on the client, which gives per-request metrics. Time to first token (TTFT) runs from send to the first output token and includes queueing and prefill. Inter-token latency (ITL) is each gap between successive chunks. Time per output token (TPOT) is a per-request average that excludes the first token:

TPOT = (E2EL - TTFT) / (output_tokens - 1)       # per request, then aggregated
E2EL = time from request sent to last token received

Aggregate throughput is reported as requests per second and output tokens per second over the run. By default percentiles are reported for TTFT, TPOT and ITL at the 99th percentile; --percentile-metrics selects metrics (ttft, tpot, itl, e2el and client-queue variants) and --metric-percentiles 50,90,99 adds more cut points. Ask for the median and p99 at minimum: a mean hides the tail that users feel. A detailed treatment of choosing targets is in LLM serving SLOs.

Two subtleties matter when reading results. ITL and TPOT differ when the server sends several tokens per chunk, for example with speculative decoding, and TTFT includes time spent waiting in the server queue, so under overload TTFT explodes while TPOT looks fine.

The offline tools in practice

The offline tools are the quickest way to answer a narrow engine question, such as whether a quantised checkpoint or a different attention backend is faster, because they remove everything that is not the engine. Typical invocations look like this:

# One fixed batch, repeated: good for kernel and backend changes.
vllm bench latency --model meta-llama/Llama-3.1-8B-Instruct \
  --input-len 1000 --output-len 250 --batch-size 1 \
  --num-iters-warmup 10 --num-iters 30 --output-json lat_b1.json

# Whole dataset at once: good for batch-job capacity.
vllm bench throughput --model meta-llama/Llama-3.1-8B-Instruct \
  --dataset-name random --random-input-len 1000 --output-len 250 \
  --num-prompts 2000 --output-json tput.json

Run the latency tool at batch size 1 and again at the batch size you expect in production: batch 1 is dominated by reading weights from memory, while larger batches shift the bottleneck toward compute and KV-cache reads, and a change can help one and hurt the other. The throughput tool submits everything immediately and lets the scheduler pack the GPU, so it reports an upper bound on tokens per second with no latency target attached. Treat that number as a ceiling for offline batch work, never as an interactive capacity figure. Because both tools build the engine in-process, they accept the same engine arguments as vllm serve, such as tensor parallel size or quantisation, which makes A/B comparisons of engine settings a one-flag change. Keep everything else fixed between the two runs, repeat each at least three times, and only believe differences larger than the run-to-run spread.

Load shape and request content

How requests arrive determines what you measure. --request-rate sets the average arrival rate in requests per second. Its default is inf, which sends all requests at time zero: a saturation test, useful for peak throughput and meaningless for latency, because every request queues behind the batch. With a finite rate, --burstiness shapes the gaps: 1.0 gives a Poisson process (exponential gaps), values below 1 make traffic burstier and values above 1 make it more uniform, using gamma-distributed intervals.

--max-concurrency caps how many requests are in flight at once, independent of arrival rate. With rate inf and a concurrency cap of 32 you get a closed-loop test: 32 virtual users, each sending a new request as soon as the last finishes. With a finite rate and no cap you get an open-loop test, which is what real traffic looks like and the only kind that exposes queue growth when the server falls behind. Newer versions also offer --ramp-up-strategy (linear or exponential) with start and end rates, which is useful for finding a knee in a single run.

Request content matters as much as arrival. The random dataset generates prompts of --random-input-len tokens and requests --random-output-len tokens; --random-range-ratio samples lengths around those values instead of fixing them. sharegpt uses real conversation lengths from a file you supply with --dataset-path, and prefix_repetition deliberately shares prefixes to test prefix caching. Fixed lengths make runs comparable; realistic distributions make them predictive. Use both.

A reproducible sweep

Here is a reproducible sweep for a chat workload with prompts of about 1,000 tokens and answers of about 250. Start the server with the exact flags you intend to deploy, then sweep concurrency with fixed lengths.

# Server: pin everything that affects scheduling.
vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --max-num-seqs 256 --gpu-memory-utilization 0.90 --seed 0 &

# Wait until the model is loaded and the API answers.
until curl -sf localhost:8000/v1/models > /dev/null; do sleep 5; done

# Warm up once (compilation, CUDA graph capture), discard the result.
vllm bench serve --model meta-llama/Llama-3.1-8B-Instruct \
  --dataset-name random --random-input-len 1000 --random-output-len 250 \
  --num-prompts 64 --max-concurrency 16

for C in 1 4 8 16 32 64 128; do
  vllm bench serve --model meta-llama/Llama-3.1-8B-Instruct \
    --dataset-name random --random-input-len 1000 --random-output-len 250 \
    --random-range-ratio 0.2 --ignore-eos --seed 0 \
    --num-prompts $((C * 20)) --max-concurrency $C \
    --percentile-metrics ttft,tpot,itl,e2el --metric-percentiles 50,90,99 \
    --goodput ttft:1000 tpot:50 \
    --save-result --result-dir results/ --result-filename c${C}.json
done

Scaling --num-prompts with concurrency keeps each run long enough that the start and end ramps do not dominate. --ignore-eos forces every request to produce its full output length, so a model that happens to stop early does not look faster. Then parse the saved files. Key names in the result JSON can change between versions, so the parser discovers them rather than hard-coding a schema.

import json, pathlib, re

rows = []
for f in sorted(pathlib.Path("results").glob("c*.json"),
                key=lambda p: int(re.findall(r"\d+", p.stem)[0])):
    d = json.loads(f.read_text())
    pick = {k: v for k, v in d.items()
            if isinstance(v, (int, float)) and
            re.search(r"(throughput|goodput|p\d+_)", k)}
    rows.append((f.stem, pick))

for name, pick in rows:
    print(name, {k: round(v, 1) for k, v in sorted(pick.items())})

Worked example: reading the sweep

Suppose the sweep produces the following (illustrative numbers for one GPU; your hardware and model will differ):

ConcurrencyOutput tok/sp99 TTFT msp99 TPOT msGoodput req/s
1954510.50.38
864012012.82.5
321,85041017.97.1
642,70095026.09.8
1283,1003,40044.04.2

Throughput keeps rising to 128, which is the number a throughput-only benchmark would report. But goodput, the rate of requests that met both SLOs (TTFT under 1 second and TPOT under 50 ms), peaks at 64 and collapses at 128 because p99 TTFT blows past the target as requests queue for prefill. The operating point is concurrency 64, or somewhat below it for headroom. Capacity planning then divides expected peak concurrency by about 50 to get replica count, as described in serving capacity planning.

Why does TPOT rise with concurrency? Decode is memory-bandwidth bound: each step reads the weights once regardless of batch size, so batching is nearly free at first, but growing batches also read more KV cache, and new prefills interleave with decode steps. The paged KV cache is what lets vLLM pack that many sequences at all.

Before trusting the knee, repeat the two runs either side of it at least three times. If goodput at 64 varies by more than the gap between 32 and 64, the sweep has not located the knee, and you need more prompts per run or finer concurrency steps such as 48 and 80. Record the spread alongside the median, so the next person who re-runs the sweep after an upgrade can tell a real regression from noise.

Failure modes

Most bad vLLM benchmark numbers come from a short list of mistakes.

  • No warm-up. The first requests pay for compilation and CUDA graph capture. Run a discarded warm-up pass; the latency tool does this for you with --num-iters-warmup.
  • Rate inf for latency. The default sends everything at once, so TTFT measures queue depth. Set a rate or a concurrency cap.
  • Prefix-cache flattery. Repeated or templated prompts hit the prefix cache and make prefill look free. Use random prompts for baseline numbers and the prefix workload only when you mean to test caching.
  • Early stopping. Without --ignore-eos, output lengths vary by model and checkpoint, so two models are compared on different work.
  • Client bottleneck. One Python client can become the limit at high rates; if achieved request rate falls below the target, the client is saturated.
  • Mismatched server flags. --max-num-seqs, --max-num-batched-tokens and memory utilisation change scheduling. Record them with every result.
  • Version drift. Comparing results from different vLLM versions measures the release, not your change. Pin the version and record the commit.
  • Too few prompts. A p99 from 100 requests is one sample. Use enough prompts that the tail is populated, and repeat runs to see the noise; regression analysis shows how to compare runs with confidence intervals.

Trade-offs

vLLM's suite is the shortest path to numbers that vLLM maintainers and other users will recognise, it understands vLLM-specific workloads such as prefix repetition and multimodal inputs, and the offline tools isolate the engine cleanly. Its limits: it is tied to one project's release cadence, synthetic datasets are not your traffic, and client-side measurement on the same host can perturb results. LLMPerf is a neutral alternative when comparing different serving stacks, and replaying your own request traces, as in custom benchmarks, is the only way to predict production behaviour. A reasonable practice is all three: vLLM's tools for tuning, a neutral harness for cross-stack comparisons, and trace replay before launch.

What to do next

  1. Start vllm serve with your production flags and record the vLLM version, GPU and driver.
  2. Run vllm bench latency at batch sizes 1, 8 and 32 to get an engine baseline without HTTP.
  3. Run the concurrency sweep above with --ignore-eos, a warm-up pass and --goodput set to your SLOs.
  4. Plot goodput against concurrency and pick the operating point below the peak.
  5. Repeat the sweep with an open-loop finite --request-rate and burstiness below 1 to check behaviour under bursts.
  6. Re-run the same sweep on every vLLM upgrade and flag changes beyond your noise floor.
Key takeaway: vLLM's suite has two offline tools that isolate the engine and one online client that measures what users see. Defaults such as an infinite request rate and fixed random lengths shape every number, so set the arrival process, cap concurrency, force output lengths, warm up, and choose the operating point from goodput against your SLOs rather than from peak throughput.