LLMPerf is a small open-source load generator from the Ray project for measuring how fast an LLM endpoint answers: time to first token, the gap between streamed tokens, end-to-end latency and aggregate output throughput under concurrent load. It talks to any OpenAI-compatible chat endpoint, which covers vLLM, SGLang, TGI's Messages API and most hosted providers, and it was widely used to compare hosted providers in 2023 and 2024.
Two facts shape how you should use it today. First, the repository is archived: its last push was in December 2024, so it receives no fixes and you should pin the commit you run. Second, several of its metrics do not mean what their names suggest. This article reads the tool from its source, shows exactly how each number is produced, corrects the misleading ones with a short analysis script, and walks through a concurrency sweep that turns raw output into a capacity decision for a GPU deployment.
How the load test works
The load test lives in token_benchmark_ray.py. It builds a list of prompts up front. Each prompt starts with the instruction Randomly stream lines from the following text. Don't generate eos tokens: followed by shuffled lines of Shakespeare's sonnets, padded until the prompt reaches a length drawn from a gaussian with mean --mean-input-tokens and standard deviation --stddev-input-tokens. The requested output length is drawn the same way and sent as max_tokens.
It then starts --num-concurrent-requests threads. Each thread owns one client (a Ray actor), sends a streaming request, waits until the stream ends, records the result and only then sends its next request. That makes LLMPerf a closed-loop test: the offered load is a fixed number of in-flight requests, not an arrival rate. The run stops when --max-num-completed-requests have finished or --timeout seconds have passed.
Every token count, input and output, is measured with one Llama tokenizer whatever model is behind the endpoint. The authors chose this so prompts are the same size across providers. The side effect is that the output token counts it reports are Llama-token counts, which can differ noticeably from what your model's own tokenizer produces.
Running a sweep
Install from a pinned commit, point the OpenAI-compatible client at your endpoint with two environment variables, and run one concurrency level at a time. Every flag below appears in the project README; --model is sent in the request body, so it must match the name your server registered.
git clone https://github.com/ray-project/llmperf.git && cd llmperf
git checkout <pinned-commit> && pip install -e .
export OPENAI_API_BASE="http://10.0.0.12:8000/v1" # your vLLM / SGLang server
export OPENAI_API_KEY="not-used-but-required"
for N in 1 8 32 64; do
python token_benchmark_ray.py \
--model "meta-llama/Llama-3.1-8B-Instruct" \
--mean-input-tokens 550 --stddev-input-tokens 150 \
--mean-output-tokens 150 --stddev-output-tokens 10 \
--max-num-completed-requests 400 --timeout 900 \
--num-concurrent-requests $N \
--results-dir "results/c$N" \
--llm-api openai \
--additional-sampling-params '{"ignore_eos": true, "temperature": 0}'
doneThe ignore_eos parameter is a vLLM extension, so check your server accepts it. It matters because the prompt only asks the model not to stop; instruction-tuned models often stop early anyway, and then output length, and therefore every throughput number, drifts between runs and between models. Each run writes a _summary.json and an _individual_responses.json file. The second one is the one to keep.
What each metric really means
Reading the client code gives these definitions. Three of them need correcting before you put them in a report.
| Field | What the code computes | Watch out for |
|---|---|---|
ttft_s | Time from sending the request to the first chunk whose delta has content | Includes network, queueing and prefill; fine as defined |
end_to_end_latency_s | Time from send to end of stream | Fine as defined |
inter_token_latency_s | Sum of all chunk gaps including the TTFT, divided by the number of SSE chunks received | This is roughly E2E divided by chunks, not decode-only time per token |
number_output_tokens | Llama-tokenizer count of the generated text | Not your model's token count |
request_output_throughput_token_per_s | Output tokens divided by E2E for one request | Per-user speed, includes the TTFT |
mean_output_throughput_token_per_s | All successful output tokens divided by wall time of the run | Includes ramp-up and the drain at the end |
The inter-token figure is the one that misleads most. The client appends the TTFT to its list of gaps and the driver divides the sum by a chunk count, so a request with a slow prefill looks like it has slow decoding. On a long-prompt workload the reported value can be several times the true decode interval. A comment in the source even notes that the sum should equal the end-to-end latency.
Two more effects are easy to miss. The summary computes quantiles only over requests without an error code, so a server that times out its slowest requests reports a better p99 than one that serves them; check error_rate before any latency figure. And the HTTP client has a hardcoded 180-second timeout, separate from --timeout, so very long generations fail as errors rather than appearing as slow successes.
Recomputing TPOT and goodput
Recompute decode-only latency and goodput from the per-request file. Time per output token (TPOT) is the decode time divided by the gaps after the first token. Goodput is the fraction of all started requests that succeeded and met both latency targets; it is the number a capacity plan should use.
import json, sys
def pct(xs, q):
xs = sorted(xs)
return xs[min(len(xs) - 1, int(q * len(xs)))] if xs else float("nan")
rows = json.load(open(sys.argv[1])) # *_individual_responses.json
ok = [r for r in rows if r.get("error_code") is None and r["number_output_tokens"] > 1]
ttft = [r["ttft_s"] for r in ok]
tpot = [(r["end_to_end_latency_s"] - r["ttft_s"]) / (r["number_output_tokens"] - 1) for r in ok]
TTFT_SLO, TPOT_SLO = 0.5, 0.035 # seconds; set from your product
good = sum(1 for a, b in zip(ttft, tpot) if a <= TTFT_SLO and b <= TPOT_SLO)
print(f"requests={len(rows)} errors={len(rows) - len(ok)}")
print(f"ttft p50={pct(ttft, .5):.3f} p95={pct(ttft, .95):.3f}")
print(f"tpot p50={pct(tpot, .5):.4f} p95={pct(tpot, .95):.4f}")
print(f"goodput={good / len(rows):.1%}")Because output tokens are Llama-token counts, TPOT here is per Llama token. If you need figures in your model's own tokens, re-tokenize the generated text, which the file does not store; patch the client to save it, or count tokens server-side from the engine's metrics endpoint for the same window.
Worked example: from sweep to replica count
Suppose one H100 serves an 8B model with vLLM and the product needs p95 TTFT under 0.5 s and TPOT under 35 ms. You run the sweep above. The numbers below are illustrative, made self-consistent to show the reasoning, not a measurement of any particular stack.
| Concurrency | TTFT p50 / p95 (s) | TPOT p50 (ms) | E2E p50 (s) | Aggregate output (tok/s) |
|---|---|---|---|---|
| 1 | 0.08 / 0.10 | 12 | 1.87 | 80 |
| 8 | 0.12 / 0.17 | 16 | 2.50 | 479 |
| 32 | 0.35 / 0.48 | 28 | 4.52 | 1,061 |
| 64 | 1.90 / 3.40 | 41 | 8.01 | 1,199 |
Check the arithmetic the way you would check real data. At concurrency 32, E2E is 0.35 + 149 x 0.028 = 4.52 s for a 150-token answer, and Little's law gives 32 / 4.52 = 7.1 requests per second, or 7.1 x 150 = 1,061 output tokens per second. Going from 32 to 64 in-flight requests buys only 13 percent more throughput while TTFT p95 grows sevenfold: the batch is full and new requests are queueing for prefill.
LLMPerf's own inter-token figure at concurrency 64 would be about 8.01 s divided by roughly 151 chunks, around 53 ms, which overstates the true 41 ms decode interval by almost a third. Reported as is, it would wrongly suggest the GPU is decode-bound.
The decision: run at most 32 concurrent requests per replica and admit no more at the router. If you expect a peak of 15 requests per second of this shape, you need 15 / 7.1 = 2.1, so three replicas, and you should re-run the sweep with your real prompt length distribution before ordering hardware.
Cross-checking against the server
A client-side benchmark tells you what users experience; it does not tell you why. Run it with the server's own telemetry recording, on the same clock, so every knee in the sweep has an explanation. Three server signals are enough for most diagnoses: how many requests are running in the batch, how many are waiting, and how full the KV cache is. Serving engines such as vLLM expose these on a Prometheus /metrics endpoint; metric names have changed between versions, so read them from your server rather than copying them from a blog.
Interpret the combinations. If TTFT rises while the waiting count stays at zero, prefill itself is slow, usually because prompts are long or chunked prefill is competing with decode. If the waiting count grows during the run, the server is at its batch or memory limit and requests are queueing; that is the TTFT knee in the worked example. If KV-cache usage sits near 100 percent and the engine reports preemptions, requests are being evicted and recomputed, which shows up as bimodal end-to-end latency. Alongside the engine, record GPU utilisation, memory and power from DCGM or nvidia-smi; a run where the GPU is idle half the time while latency climbs points at the client, the network or the router, not the model.
Failure modes
- Comparing providers on different tokenizers. Requests-per-minute is comparable; tokens-per-second is only comparable after you fix the token unit.
- Early stopping. Without
ignore_eosor an equivalent, mean output length falls below--mean-output-tokensand throughput looks lower. Always report achieved output tokens alongside throughput. - Prefix-cache effects. Every prompt shares the same instruction prefix and draws from one small text, so engines with prefix caching get some free hits that your production traffic may not.
- Client saturation. At high concurrency the Python client can become the bottleneck. Run the generator on a separate machine and confirm its CPU is not pegged.
- Too few requests. A p99 from 100 requests is the single worst request. Collect at least several hundred completed requests per point.
- Treating closed-loop results as open-loop capacity. A closed loop slows its own arrivals when the server slows, so it hides the queue growth that real traffic causes.
Trade-offs and where it fits
LLMPerf's strengths are simplicity and neutrality: one command, any OpenAI-compatible endpoint, and the same synthetic prompt for everyone. That makes it good for a quick provider comparison and for regression checks between engine versions on the same hardware. Its weaknesses are the synthetic text, the closed loop, the tokenizer choice and the fact that it is no longer maintained.
For capacity planning, add an open-loop test that replays your real arrival pattern and prompt lengths, and set SLOs on TTFT and TPOT separately. Treat LLMPerf as one input to a benchmark plan, next to a quality check such as its own llm_correctness.py test, which asks the model to convert spelled-out numbers to digits and counts mismatches. A fast endpoint that quietly serves a degraded quantized model is not a win.
To keep learning, read the TTFT ledger and throughput math for the theory behind the sweep, tail latency for why p99 behaves as it does, MLPerf for a standardised suite with rules, and custom benchmarks for quality evaluation on your own data.
What to do next
- Pin an LLMPerf commit and record it, the engine version, the GPU and the model in every result.
- Set
ignore_eosor an equivalent and fixtemperatureso output length is controlled. - Sweep concurrency at 1, 8, 32 and 64 (or until TTFT breaks) with at least 400 completed requests per point.
- Analyse
_individual_responses.jsonwith the script above; never quote the raw inter-token field. - Report error rate first, then TTFT and TPOT percentiles, then goodput and aggregate throughput.
- Pick the highest concurrency that meets both SLOs and enforce it as an admission limit.
- Repeat with your production prompt-length distribution and an open-loop arrival test before sizing hardware.