LLM inference is slow for a small number of reasons, and profiling is how you find out which one applies to your deployment instead of guessing. The thin page that used to be here listed the tools. Tools are the easy part. The hard part is method: knowing what number you should be getting before you measure, measuring at the level where the problem is visible, and only then descending into traces and kernels.

This article is the method. It starts with a back-of-envelope budget for prefill and decode, moves through service metrics, PyTorch profiler traces and Nsight Systems timelines, and ends at kernel analysis, with a worked example on an 8B model. For tool mechanics, Nsight for LLMs and Nsight Systems and Nsight Compute go deeper; this page tells you when to reach for them and what to look for.

The profiling ladder

Profile top-down: each rung tells you which rung to open next1. budgetbytes per token, FLOPs per promptpen and paper2. serviceTTFT, inter-token latency, tokens/sload generator3. phasesprefill vs decode, queue vs computeserver metrics + NVTX4. timelinegaps, syncs, copies, commstorch.profiler, Nsight Systems5. kernelsbandwidth, occupancy, stallsNsight ComputeMost wins are found on rungs 2 to 4. Open Nsight Compute only for a kernel already proven to matter.
The profiling ladder. Descend one rung only when the rung above has localised the problem.

The ladder exists because each level of detail costs more to collect and more to read. A kernel profile of a model that is actually waiting on a request queue is a beautifully detailed answer to the wrong question. Start at the top, and write down the number you expect before every measurement, so the measurement can surprise you.

Start with a budget

Transformer inference has two phases with different limits. Prefill processes the whole prompt in parallel and is dominated by matrix multiplies: roughly 2 floating point operations per parameter per token, so it tends to be compute-bound. Decode produces one token per sequence per step and must read every weight once per step regardless of batch size, plus each sequence's KV cache, so at small batch it is bound by memory bandwidth. Decode math derives this in detail.

That gives two utilisation metrics. Model FLOPs utilisation (MFU) is achieved FLOPs divided by peak, and is the right lens for prefill. Model bandwidth utilisation (MBU) is achieved bytes per second divided by peak memory bandwidth, and is the right lens for decode. Use the peak figures from the datasheet of your exact part, and remember that dense and sparse tensor core figures differ by 2x; use dense.

def decode_floor_ms(params_b, bytes_per_param, kv_bytes_per_tok, ctx, batch, hbm_tb_s):
    weights = params_b * 1e9 * bytes_per_param           # read once per step
    kv = kv_bytes_per_tok * ctx * batch                  # every sequence reads its cache
    return (weights + kv) / (hbm_tb_s * 1e12) * 1e3      # ms per decode step

def prefill_floor_ms(params_b, prompt_tokens, peak_tflops):
    return 2 * params_b * 1e9 * prompt_tokens / (peak_tflops * 1e12) * 1e3

# Llama-3-8B-class model in BF16: 32 layers, 8 KV heads, head_dim 128
kv_per_tok = 2 * 32 * 8 * 128 * 2                        # K and V, 2 bytes -> 131,072 bytes
print(decode_floor_ms(8.0, 2, kv_per_tok, 4096, 1, 3.35))   # ~4.9 ms on 3.35 TB/s HBM
print(prefill_floor_ms(8.0, 2048, 989))                     # ~33 ms at 989 dense TFLOPS

The figures 3.35 TB/s and 989 dense BF16 TFLOPS are the published H100 SXM numbers; substitute yours. These floors are not achievable, they are ceilings on speed. Real deployments typically reach a fraction of them, and the useful question is how large the fraction is and where the remainder goes. Attention FLOPs, ignored above, grow with context and matter for long prompts.

Measure the service first

Measure what users feel before anything else: time to first token (TTFT), the gap between subsequent tokens (inter-token latency, often reported as time per output token), throughput in output tokens per second, and their p50 and p99 at a stated request rate and prompt and output length distribution. Use a load generator that streams, because non-streaming clients cannot see TTFT. vLLM ships one as vllm bench serve; other servers have equivalents.

Benchmark hygiene decides whether two runs can be compared. Fix the input and output length distribution and the random seed. Warm up until compilation and CUDA graph capture are done. Report the arrival rate, not just concurrency. Record the GPU clocks and power limit, since a power-capped card runs slower without any error. Run each configuration at least three times and report spread.

Then split the latency. TTFT is queueing plus prefill. If TTFT rises sharply with load while prefill time per request is flat, the problem is scheduling or capacity, not kernels; see continuous batching and prefill and decode disaggregation. If inter-token latency degrades when long prompts arrive, prefill is stealing decode steps.

Framework traces with torch.profiler

When a phase is slower than its budget, record a framework-level trace. The PyTorch profiler shows operators, the CUDA kernels they launch, and the CPU time between launches. Use a schedule so that only a few steady-state steps are recorded; profiling everything produces unreadable multi-gigabyte traces and perturbs the run.

import torch
from torch.profiler import profile, schedule, ProfilerActivity, record_function

sched = schedule(wait=5, warmup=3, active=4, repeat=1)    # skip compile and warmup steps

with profile(activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA],
             schedule=sched, record_shapes=True, with_stack=False) as prof:
    for step in range(12):
        with record_function("decode_step"):
            logits = model(next_ids, past_key_values=cache, use_cache=True).logits
            next_ids = logits[:, -1].argmax(-1, keepdim=True)
        prof.step()

print(prof.key_averages().table(sort_by="cuda_time_total", row_limit=15))
prof.export_chrome_trace("decode_trace.json")             # open in Perfetto UI

Read the table first: the share of GPU time in GEMMs, attention, normalisation, and elementwise kernels. Then open the trace and look at the GPU row between kernels. Long empty stretches with a busy CPU row mean the GPU is waiting for Python to launch work. Keep record_shapes on when you need GEMM shapes, and leave with_stack and profile_memory off unless you need them, because they add overhead.

System timelines with Nsight Systems

The PyTorch trace stops at the process boundary. Nsight Systems sees the whole machine: all CUDA streams, NCCL communication, memory copies, CPU threads and OS activity. Mark phases with NVTX ranges and limit capture to a window so the report is small.

# in the server: bracket a few steady-state steps
torch.cuda.cudart().cudaProfilerStart()
with torch.cuda.nvtx.range("decode_step"):
    run_step()
torch.cuda.cudart().cudaProfilerStop()

# on the command line
nsys profile -t cuda,nvtx,osrt --capture-range=cudaProfilerApi \
     --capture-range-end=stop -o decode_run python serve.py
nsys stats --report cuda_gpu_kern_sum decode_run.nsys-rep

In the timeline, look for four patterns. Launch gaps: many tiny kernels with idle GPU between them, the typical decode signature at small batch, fixed by CUDA graphs or larger batches. Synchronisation: a host-side wait after each step, often from moving a tensor to the CPU to check a stop condition. Copies: host-to-device transfers inside the step, from unpinned buffers or tensors created on the CPU. Communication: in tensor-parallel runs, all-reduce kernels whose duration grows with GPU count and that do not overlap compute.

Kernel analysis with Nsight Compute

Descend to Nsight Compute only for a kernel that already accounts for a large share of step time and runs well below its budget. It replays each profiled kernel many times to collect hardware counters, so it slows the program dramatically; filter by kernel name and count.

ncu --set full -k regex:gemm -c 5 -o gemm_report python serve.py

For a decode GEMM, the headline is DRAM throughput as a percentage of peak; a memory-bound kernel near peak is already done, and the remaining lever is reading fewer bytes through quantisation or batching. Low throughput with low occupancy points at launch configuration or tiny shapes; low throughput with high occupancy and high stall reasons on memory points at access patterns. For prefill GEMMs, the headline is tensor core utilisation. Shapes matter: a matrix dimension that is not a multiple of the tile size wastes a partial tile on every launch.

Reading the evidence

SymptomLikely causeFirst fix to try
Decode MBU well under 50% at batch 1Kernel launch overhead, Python in the loopCUDA graphs, fused kernels
TTFT grows with load, prefill time flatQueueing, admission policyMore replicas, chunked prefill, scheduling
Inter-token latency spikes with long promptsPrefill interleaved with decodeChunked prefill or disaggregation
Throughput plateaus as batch risesKV cache memory limits batchPaged KV, KV quantisation, shorter max length
Step time grows with GPU countAll-reduce not overlappedFewer TP ranks, faster interconnect, overlap
Host-to-device copies in each stepUnpinned or CPU-created tensorsPreallocate on device, pin host buffers

Worked example: a launch-bound decode loop

Here is the method applied to an illustrative case. An 8B BF16 model serves one stream on an H100 SXM with a 4,096-token context. The budget above says decode cannot be faster than about 4.9 ms per token. The service benchmark reports 9.8 ms inter-token latency at concurrency 1, which is an MBU of about 50%: half the bandwidth is unused.

The PyTorch trace shows several hundred kernels per step, most under 20 microseconds, and the GPU row is idle for a large fraction of the step while the CPU row is busy. That is a launch bound loop. Enabling CUDA graph capture for decode, which most serving engines support for a fixed set of batch sizes, removes per-kernel launch cost. A rerun shows the gaps closed and inter-token latency near 6 ms, roughly 80% MBU. The Nsight Systems timeline now shows back-to-back kernels, and the kernel summary is led by GEMMs reading weights, which is exactly what a healthy decode step should look like. Further gains now require reading fewer bytes, through weight quantisation or a larger batch that shares each weight read across more tokens. The numbers in this scenario are representative, not a benchmark result; run the same steps on your hardware.

Timing regions and catching regressions

Many questions do not need a profiler at all, just correct timing of one region. CUDA launches are asynchronous, so time on the device with events and read the result after a synchronise:

start, end = torch.cuda.Event(enable_timing=True), torch.cuda.Event(enable_timing=True)
times = []
for _ in range(50):
    start.record()
    run_step()
    end.record()
    torch.cuda.synchronize()
    times.append(start.elapsed_time(end))     # milliseconds
times.sort()
print("p50", times[25], "p95", times[47])

Make this a regression test. Store the floor, the measured step time and the derived MBU or MFU per model, precision and batch size in CI, and fail a change that drops utilisation by more than a few percent. Performance regressions in inference stacks usually arrive through dependency upgrades, such as a new framework, driver or attention library, and nobody notices until the cloud bill does. A ten-minute nightly run on one GPU catches most of them, and the stored traces from the last good run make the diff obvious.

Profiling pitfalls

  • Profiling cold. Compilation, autotuning and graph capture dominate the first steps. Always skip them with a schedule or a capture range.
  • Timing without synchronisation. CUDA is asynchronous; a Python timer around a launch measures the launch. Use CUDA events or a synchronise before reading the clock.
  • Observer effect. Stack collection, memory profiling and kernel replay change timing. Confirm improvements with the unprofiled service benchmark.
  • Wrong peak. Sparse TFLOPS or a different SKU's bandwidth make utilisation look poor or impossible. Use the dense figure for your exact part and clocks.
  • Averaging away the tail. A mean step time hides the one step in twenty that stalls on a copy or a garbage-collection pause; look at distributions.

Trade-offs between tools

ToolSeesOverheadUse when
Service benchmarkUser latency, throughputNoneAlways, first and last
torch.profilerOps, kernels, CPU launch pathLow to moderateA phase misses its budget
Nsight SystemsWhole system, streams, NCCL, OSLowGaps, syncs, comms, multi-process
Nsight ComputeOne kernel's hardware countersVery high per kernelA proven hot kernel underperforms

What to do next

  1. Compute the decode and prefill floors for your model, precision, context and GPU with the snippet above.
  2. Run a streaming load test at fixed lengths and rates, and record TTFT, inter-token latency and throughput with p50 and p99.
  3. Compute MBU for decode and MFU for prefill, and pick the phase furthest from its floor.
  4. Capture four steady-state steps with torch.profiler and classify the time into kernels and gaps.
  5. If the cause crosses processes or GPUs, capture an NVTX-scoped Nsight Systems window.
  6. Change one thing, rerun the unprofiled benchmark, and keep the change only if the user-facing number moves.
Key takeaway: Know the floor before you measure: decode is bounded by bytes read per step and prefill by FLOPs. Measure user-facing latency first, express it as MBU or MFU, then descend from framework traces to system timelines to kernel counters only as far as the evidence requires, and confirm every fix on the unprofiled benchmark.