LLM inference is slow for a small number of reasons, and profiling is how you find out which one applies to your deployment instead of guessing. The thin page that used to be here listed the tools. Tools are the easy part. The hard part is method: knowing what number you should be getting before you measure, measuring at the level where the problem is visible, and only then descending into traces and kernels.
This article is the method. It starts with a back-of-envelope budget for prefill and decode, moves through service metrics, PyTorch profiler traces and Nsight Systems timelines, and ends at kernel analysis, with a worked example on an 8B model. For tool mechanics, Nsight for LLMs and Nsight Systems and Nsight Compute go deeper; this page tells you when to reach for them and what to look for.
The profiling ladder
The ladder exists because each level of detail costs more to collect and more to read. A kernel profile of a model that is actually waiting on a request queue is a beautifully detailed answer to the wrong question. Start at the top, and write down the number you expect before every measurement, so the measurement can surprise you.
Start with a budget
Transformer inference has two phases with different limits. Prefill processes the whole prompt in parallel and is dominated by matrix multiplies: roughly 2 floating point operations per parameter per token, so it tends to be compute-bound. Decode produces one token per sequence per step and must read every weight once per step regardless of batch size, plus each sequence's KV cache, so at small batch it is bound by memory bandwidth. Decode math derives this in detail.
That gives two utilisation metrics. Model FLOPs utilisation (MFU) is achieved FLOPs divided by peak, and is the right lens for prefill. Model bandwidth utilisation (MBU) is achieved bytes per second divided by peak memory bandwidth, and is the right lens for decode. Use the peak figures from the datasheet of your exact part, and remember that dense and sparse tensor core figures differ by 2x; use dense.
def decode_floor_ms(params_b, bytes_per_param, kv_bytes_per_tok, ctx, batch, hbm_tb_s):
weights = params_b * 1e9 * bytes_per_param # read once per step
kv = kv_bytes_per_tok * ctx * batch # every sequence reads its cache
return (weights + kv) / (hbm_tb_s * 1e12) * 1e3 # ms per decode step
def prefill_floor_ms(params_b, prompt_tokens, peak_tflops):
return 2 * params_b * 1e9 * prompt_tokens / (peak_tflops * 1e12) * 1e3
# Llama-3-8B-class model in BF16: 32 layers, 8 KV heads, head_dim 128
kv_per_tok = 2 * 32 * 8 * 128 * 2 # K and V, 2 bytes -> 131,072 bytes
print(decode_floor_ms(8.0, 2, kv_per_tok, 4096, 1, 3.35)) # ~4.9 ms on 3.35 TB/s HBM
print(prefill_floor_ms(8.0, 2048, 989)) # ~33 ms at 989 dense TFLOPSThe figures 3.35 TB/s and 989 dense BF16 TFLOPS are the published H100 SXM numbers; substitute yours. These floors are not achievable, they are ceilings on speed. Real deployments typically reach a fraction of them, and the useful question is how large the fraction is and where the remainder goes. Attention FLOPs, ignored above, grow with context and matter for long prompts.
Measure the service first
Measure what users feel before anything else: time to first token (TTFT), the gap between subsequent tokens (inter-token latency, often reported as time per output token), throughput in output tokens per second, and their p50 and p99 at a stated request rate and prompt and output length distribution. Use a load generator that streams, because non-streaming clients cannot see TTFT. vLLM ships one as vllm bench serve; other servers have equivalents.
Benchmark hygiene decides whether two runs can be compared. Fix the input and output length distribution and the random seed. Warm up until compilation and CUDA graph capture are done. Report the arrival rate, not just concurrency. Record the GPU clocks and power limit, since a power-capped card runs slower without any error. Run each configuration at least three times and report spread.
Then split the latency. TTFT is queueing plus prefill. If TTFT rises sharply with load while prefill time per request is flat, the problem is scheduling or capacity, not kernels; see continuous batching and prefill and decode disaggregation. If inter-token latency degrades when long prompts arrive, prefill is stealing decode steps.
Framework traces with torch.profiler
When a phase is slower than its budget, record a framework-level trace. The PyTorch profiler shows operators, the CUDA kernels they launch, and the CPU time between launches. Use a schedule so that only a few steady-state steps are recorded; profiling everything produces unreadable multi-gigabyte traces and perturbs the run.
import torch
from torch.profiler import profile, schedule, ProfilerActivity, record_function
sched = schedule(wait=5, warmup=3, active=4, repeat=1) # skip compile and warmup steps
with profile(activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA],
schedule=sched, record_shapes=True, with_stack=False) as prof:
for step in range(12):
with record_function("decode_step"):
logits = model(next_ids, past_key_values=cache, use_cache=True).logits
next_ids = logits[:, -1].argmax(-1, keepdim=True)
prof.step()
print(prof.key_averages().table(sort_by="cuda_time_total", row_limit=15))
prof.export_chrome_trace("decode_trace.json") # open in Perfetto UIRead the table first: the share of GPU time in GEMMs, attention, normalisation, and elementwise kernels. Then open the trace and look at the GPU row between kernels. Long empty stretches with a busy CPU row mean the GPU is waiting for Python to launch work. Keep record_shapes on when you need GEMM shapes, and leave with_stack and profile_memory off unless you need them, because they add overhead.
System timelines with Nsight Systems
The PyTorch trace stops at the process boundary. Nsight Systems sees the whole machine: all CUDA streams, NCCL communication, memory copies, CPU threads and OS activity. Mark phases with NVTX ranges and limit capture to a window so the report is small.
# in the server: bracket a few steady-state steps
torch.cuda.cudart().cudaProfilerStart()
with torch.cuda.nvtx.range("decode_step"):
run_step()
torch.cuda.cudart().cudaProfilerStop()
# on the command line
nsys profile -t cuda,nvtx,osrt --capture-range=cudaProfilerApi \
--capture-range-end=stop -o decode_run python serve.py
nsys stats --report cuda_gpu_kern_sum decode_run.nsys-repIn the timeline, look for four patterns. Launch gaps: many tiny kernels with idle GPU between them, the typical decode signature at small batch, fixed by CUDA graphs or larger batches. Synchronisation: a host-side wait after each step, often from moving a tensor to the CPU to check a stop condition. Copies: host-to-device transfers inside the step, from unpinned buffers or tensors created on the CPU. Communication: in tensor-parallel runs, all-reduce kernels whose duration grows with GPU count and that do not overlap compute.
Kernel analysis with Nsight Compute
Descend to Nsight Compute only for a kernel that already accounts for a large share of step time and runs well below its budget. It replays each profiled kernel many times to collect hardware counters, so it slows the program dramatically; filter by kernel name and count.
ncu --set full -k regex:gemm -c 5 -o gemm_report python serve.pyFor a decode GEMM, the headline is DRAM throughput as a percentage of peak; a memory-bound kernel near peak is already done, and the remaining lever is reading fewer bytes through quantisation or batching. Low throughput with low occupancy points at launch configuration or tiny shapes; low throughput with high occupancy and high stall reasons on memory points at access patterns. For prefill GEMMs, the headline is tensor core utilisation. Shapes matter: a matrix dimension that is not a multiple of the tile size wastes a partial tile on every launch.
Reading the evidence
| Symptom | Likely cause | First fix to try |
|---|---|---|
| Decode MBU well under 50% at batch 1 | Kernel launch overhead, Python in the loop | CUDA graphs, fused kernels |
| TTFT grows with load, prefill time flat | Queueing, admission policy | More replicas, chunked prefill, scheduling |
| Inter-token latency spikes with long prompts | Prefill interleaved with decode | Chunked prefill or disaggregation |
| Throughput plateaus as batch rises | KV cache memory limits batch | Paged KV, KV quantisation, shorter max length |
| Step time grows with GPU count | All-reduce not overlapped | Fewer TP ranks, faster interconnect, overlap |
| Host-to-device copies in each step | Unpinned or CPU-created tensors | Preallocate on device, pin host buffers |
Worked example: a launch-bound decode loop
Here is the method applied to an illustrative case. An 8B BF16 model serves one stream on an H100 SXM with a 4,096-token context. The budget above says decode cannot be faster than about 4.9 ms per token. The service benchmark reports 9.8 ms inter-token latency at concurrency 1, which is an MBU of about 50%: half the bandwidth is unused.
The PyTorch trace shows several hundred kernels per step, most under 20 microseconds, and the GPU row is idle for a large fraction of the step while the CPU row is busy. That is a launch bound loop. Enabling CUDA graph capture for decode, which most serving engines support for a fixed set of batch sizes, removes per-kernel launch cost. A rerun shows the gaps closed and inter-token latency near 6 ms, roughly 80% MBU. The Nsight Systems timeline now shows back-to-back kernels, and the kernel summary is led by GEMMs reading weights, which is exactly what a healthy decode step should look like. Further gains now require reading fewer bytes, through weight quantisation or a larger batch that shares each weight read across more tokens. The numbers in this scenario are representative, not a benchmark result; run the same steps on your hardware.
Timing regions and catching regressions
Many questions do not need a profiler at all, just correct timing of one region. CUDA launches are asynchronous, so time on the device with events and read the result after a synchronise:
start, end = torch.cuda.Event(enable_timing=True), torch.cuda.Event(enable_timing=True)
times = []
for _ in range(50):
start.record()
run_step()
end.record()
torch.cuda.synchronize()
times.append(start.elapsed_time(end)) # milliseconds
times.sort()
print("p50", times[25], "p95", times[47])Make this a regression test. Store the floor, the measured step time and the derived MBU or MFU per model, precision and batch size in CI, and fail a change that drops utilisation by more than a few percent. Performance regressions in inference stacks usually arrive through dependency upgrades, such as a new framework, driver or attention library, and nobody notices until the cloud bill does. A ten-minute nightly run on one GPU catches most of them, and the stored traces from the last good run make the diff obvious.
Profiling pitfalls
- Profiling cold. Compilation, autotuning and graph capture dominate the first steps. Always skip them with a schedule or a capture range.
- Timing without synchronisation. CUDA is asynchronous; a Python timer around a launch measures the launch. Use CUDA events or a synchronise before reading the clock.
- Observer effect. Stack collection, memory profiling and kernel replay change timing. Confirm improvements with the unprofiled service benchmark.
- Wrong peak. Sparse TFLOPS or a different SKU's bandwidth make utilisation look poor or impossible. Use the dense figure for your exact part and clocks.
- Averaging away the tail. A mean step time hides the one step in twenty that stalls on a copy or a garbage-collection pause; look at distributions.
Trade-offs between tools
| Tool | Sees | Overhead | Use when |
|---|---|---|---|
| Service benchmark | User latency, throughput | None | Always, first and last |
| torch.profiler | Ops, kernels, CPU launch path | Low to moderate | A phase misses its budget |
| Nsight Systems | Whole system, streams, NCCL, OS | Low | Gaps, syncs, comms, multi-process |
| Nsight Compute | One kernel's hardware counters | Very high per kernel | A proven hot kernel underperforms |
What to do next
- Compute the decode and prefill floors for your model, precision, context and GPU with the snippet above.
- Run a streaming load test at fixed lengths and rates, and record TTFT, inter-token latency and throughput with p50 and p99.
- Compute MBU for decode and MFU for prefill, and pick the phase furthest from its floor.
- Capture four steady-state steps with torch.profiler and classify the time into kernels and gaps.
- If the cause crosses processes or GPUs, capture an NVTX-scoped Nsight Systems window.
- Change one thing, rerun the unprofiled benchmark, and keep the change only if the user-facing number moves.