Framework profilers tell you which Python operator was slow. NVIDIA Nsight Systems and Nsight Compute tell you what the GPU actually did: which kernels ran, on which stream, how long the GPU sat idle between them, and how close a single kernel came to the hardware's memory or math limits. For LLMs the expensive problems are rarely one slow operator; they are idle gaps in the decode loop, unoverlapped all-reduces, or GEMMs running at a fraction of memory bandwidth.
This article is a working method for profiling LLM workloads with the two tools: how to annotate a model so the timeline reads in terms of prefill, decode and layers; how to capture a short steady-state window from a long training job or a serving engine; how to classify an LLM step's kernels; and how to use arithmetic intensity to decide what a decode GEMM should achieve before Nsight Compute tells you what it did achieve. The tools' general mechanics, such as replay, overhead and timeline rows, are covered in GPU Profiling: Nsight Systems and Nsight Compute; this page applies them to models.
Two tools, in order
Nsight Systems (nsys) records a timeline. It traces CUDA API calls, kernels, memory copies, NVTX ranges and, if asked, NCCL and cuBLAS activity, with low overhead. Use it first, to find where wall-clock time goes.
Nsight Compute (ncu) profiles individual kernels. It replays a kernel to collect hardware counters: achieved DRAM bandwidth, FLOP rate, occupancy, cache hit rates and stall reasons. Use it second, on a kernel that nsys already showed to be expensive.
Start with the timeline: tuning one GEMM from 80% to 90% of bandwidth is irrelevant if a third of the step is idle time between kernels.
Two tools, in order
Nsight Systems (nsys) records a timeline. It traces CUDA API calls, kernels, memory copies, NVTX ranges and, if asked, NCCL and cuBLAS activity, with low overhead. Use it first, to find where wall-clock time goes.
Nsight Compute (ncu) profiles individual kernels. It replays a kernel to collect hardware counters: achieved DRAM bandwidth, FLOP rate, occupancy, cache hit rates and stall reasons. Use it second, on a kernel that nsys already showed to be expensive.
Start with the timeline: tuning one GEMM from 80% to 90% of bandwidth is irrelevant if a third of the step is idle time between kernels.
Make the timeline speak in model terms
A raw timeline of an LLM is thousands of kernels with templated names. NVTX ranges turn it into a story. Push a range for each training step or serving iteration, nested ranges for prefill and decode, and, during investigation, one per transformer layer. Keep layer-level ranges behind a profiling flag; they add CPU overhead.
import torch
import torch.cuda.nvtx as nvtx
def forward_step(model, batch, phase):
with nvtx.range(f"{phase} bs={batch.input_ids.shape[0]}"):
hidden = model.embed(batch.input_ids)
for i, layer in enumerate(model.layers):
with nvtx.range(f"layer {i}"):
hidden = layer(hidden, batch.kv_cache, batch.positions)
with nvtx.range("lm_head + sample"):
logits = model.lm_head(hidden[:, -1])
return torch.argmax(logits, dim=-1)NVTX ranges are recorded on the CPU timeline; Nsight Systems projects them onto the GPU by correlating the kernels launched inside each range. Because launches are asynchronous, a range can close on the CPU long before its kernels finish, so read GPU durations from the projected range, not from the CPU range. If you cannot edit the model, nsys profile --pytorch=autograd-nvtx asks PyTorch to emit an NVTX range per operator automatically, at higher overhead and with less meaningful names.
Capture a steady-state window
Profiling a whole job produces gigabyte reports dominated by start-up, weight loading and compilation. Capture a short steady-state window instead: a few training steps after warm-up, or a few seconds of serving under representative load. The cleanest control is the CUDA profiler API. The program calls torch.cuda.profiler.start() and stop(), and nsys records only between them.
# train.py: capture steps 50-54 only
for step, batch in enumerate(loader):
if step == 50:
torch.cuda.synchronize()
torch.cuda.profiler.start()
with torch.cuda.nvtx.range(f"step {step}"):
loss = model(batch).loss
loss.backward()
optimizer.step()
optimizer.zero_grad(set_to_none=True)
if step == 54:
torch.cuda.synchronize()
torch.cuda.profiler.stop()
breaknsys profile -t cuda,nvtx,osrt,cublas,cudnn \
--capture-range=cudaProfilerApi --capture-range-end=stop \
--cuda-memory-usage=true \
-o llm_train_steps50_54 python train.pyThe synchronize calls make the window start and end on a quiet GPU, so the first kernel in the report really belongs to step 50. --cuda-memory-usage=true adds a device-memory track, useful when you suspect allocator churn, at extra overhead; leave it off for timing work.
Profiling a serving engine
Serving engines add two complications: worker processes and CUDA graphs. Engines such as vLLM run model execution in child processes, and they capture the decode step as CUDA graphs to remove per-kernel launch cost. By default Nsight Systems traces a graph launch as a single item, which hides the kernels you came to see.
The vLLM profiling guide (checked on 2026-10-04 against its latest documentation; flags change between releases, so read the guide for your version) recommends --trace-fork-before-exec=true so the workers are traced, --cuda-graph-trace=node so each kernel inside a graph appears, and VLLM_WORKER_MULTIPROC_METHOD=spawn. For a server it adds a capture range driven by the engine's own profiler hooks:
export VLLM_WORKER_MULTIPROC_METHOD=spawn
nsys profile --trace-fork-before-exec=true --cuda-graph-trace=node \
--capture-range=cudaProfilerApi --capture-range-end=repeat \
vllm serve meta-llama/Llama-3.1-8B-Instruct --profiler-config.profiler cuda
# in another shell: a short load run that toggles profiling
vllm bench serve --profile --num-prompts 2Node-level graph tracing adds overhead per kernel, so take headline latency from the engine's own metrics and use the report for the structure of a step. Other engines follow the same pattern: trace child processes, expand graphs to nodes, and capture steady load rather than model loading.
Where an LLM step goes: kernel families
Every transformer step is built from a small number of kernel families. Classifying kernel time into them is the fastest way to see where a step goes. Export the kernel summary as CSV and bucket it by name:
nsys stats --report cuda_gpu_kern_sum --format csv \
--output step llm_train_steps50_54.nsys-repimport csv, glob, re
from collections import defaultdict
FAMILIES = [ # first match wins; extend the patterns for your stack
("communication", r"nccl"),
("attention", r"flash|fmha|attn|attention|paged"),
("gemm", r"gemm|cutlass|xmma|nvjet|matmul"),
("norm/activation", r"norm|silu|gelu|act_and_mul|rotary|rope|softmax"),
("optimizer", r"adam|multi_tensor|foreach"),
]
path = glob.glob("step*cuda_gpu_kern_sum*.csv")[0]
rows = list(csv.DictReader(open(path, newline="")))
time_col = next(k for k in rows[0] if k.startswith("Total Time"))
totals = defaultdict(float)
for r in rows:
name = r["Name"].lower()
fam = next((f for f, pat in FAMILIES if re.search(pat, name)), "other")
totals[fam] += float(r[time_col])
grand = sum(totals.values())
for fam, t in sorted(totals.items(), key=lambda kv: -kv[1]):
print(f"{fam:18s} {t / 1e6:9.2f} ms {100 * t / grand:5.1f}%")Kernel names differ by library and GPU generation, which is why the patterns are a list you maintain rather than a fixed rule. Then compare the kernel total with the wall-clock length of the steps in the window. The difference is time when the GPU ran nothing: launch overhead, host-side scheduling, synchronisation, or data loading. For PyTorch-level attribution of the same gaps, see LLM Trace Analysis.
Set the target before profiling a kernel
Before opening Nsight Compute, decide what a kernel should achieve. A decode step multiplies each weight matrix by a thin activation matrix with one row per sequence in the batch. Per parameter it reads two bytes of BF16 weight once and performs two floating-point operations per sequence. Arithmetic intensity is therefore about B FLOP per byte for a batch of B sequences, ignoring KV-cache and activation traffic.
Take an H100 SXM: its datasheet lists 3.35 TB/s of HBM3 bandwidth and about 989 TFLOPS of dense BF16 tensor-core throughput. The ridge point, where a kernel stops being bandwidth-bound and becomes compute-bound, is roughly 989 / 3.35, about 295 FLOP per byte. Decode at batch 16 sits at about 16 FLOP per byte, more than an order of magnitude below the ridge. Decode GEMMs are bandwidth-bound, so the right yardstick is achieved DRAM throughput, not tensor-core utilisation. Prefill, with thousands of tokens per GEMM, is compute-bound and judged by FLOP rate.
This gives a floor. An 8-billion-parameter model in BF16 is about 16 GB of weights; reading them once at 3.35 TB/s takes about 4.8 ms. No single-GPU decode step for that model can be faster, whatever the batch size, until weights are quantised or sharded. The reasoning is developed further in the decode-math article.
Nsight Compute on a decode GEMM
With a target in hand, profile a handful of instances of the expensive kernel. Kernel replay multiplies run time many times over, so point ncu at a small benchmark that runs the same shapes rather than at a live server, and skip the warm-up launches:
ncu --kernel-name regex:"gemm|nvjet" --launch-skip 200 --launch-count 3 \
--section SpeedOfLight --section MemoryWorkloadAnalysis --section Occupancy \
-o decode_gemm python decode_bench.py --batch 16 --steps 300
ncu --import decode_gemm.ncu-rep --page detailsRead the SpeedOfLight section first. It reports compute and memory throughput as percentages of the hardware peak. For a decode GEMM you want memory throughput high and compute low; a decode GEMM that is low on both is latency-bound, usually because the grid is too small to fill the GPU at that batch size or because the library picked a poor tile shape for a skinny matrix.
Two cautions apply. ncu controls GPU clocks while profiling (see --clock-control), so its durations do not match production; compare percentages of peak, or pass --clock-control none when you need comparable times. It also flushes caches between replays by default (--cache-control all). Inside CUDA graphs, --graph-profiling node (the default) profiles each kernel node as an ordinary kernel.
Tensor-parallel communication in the timeline
With tensor parallelism, each transformer layer typically ends its attention block and its MLP block with an all-reduce across the GPUs, so a 32-layer model performs about 64 all-reduces per forward pass. In decode they are small, latency-bound and cannot overlap with the next layer, which needs their result.
Add nccl to the trace list (-t cuda,nvtx,nccl) so NCCL calls are annotated, then look for three patterns. First, no overlap of compute and communication in training. Second, a straggler rank: the healthy ranks' NCCL kernels spin waiting for it, so they look slow and it looks fast. Third, a decode step where NCCL is a large share, meaning the tensor-parallel degree is too high for the batch size. nsys recipe nccl_sum summarises collective time across reports. Collective algorithms and their costs are covered in NCCL collectives.
Worked example: a decode step 60% over its floor
Here is a diagnosis on the 8B model above, serving on one H100 SXM at batch 16. The numbers are illustrative but realistic in shape. The engine reports 7.6 ms per decode step against the 4.8 ms floor.
- Capture with the serving recipe and open the projected NVTX range for one decode step: 7.6 ms on the GPU lane.
- Run the classifier over the window: GEMMs 5.3 ms, attention 0.8 ms, norm and activation 0.3 ms, other 0.1 ms. Kernels total 6.5 ms, so the GPU is idle for 1.1 ms of every step.
- Zoom into the idle time. It sits in short gaps between nearly every kernel, each a few microseconds, and the CUDA API lane shows one launch call per kernel rather than one graph launch. CUDA graphs are off; the deployment had set eager mode to work around an old bug.
- Profile three decode GEMMs with
ncu. SpeedOfLight shows memory throughput around 85% of peak with low compute, as the intensity estimate predicted. GEMMs are close to the floor; tuning them further has little to gain. - Re-enable graphs and capture again. Inter-kernel gaps mostly disappear and the step drops towards the 6.5 ms kernel total. The remaining gap to 4.8 ms is attention reading the KV cache, plus GEMM efficiency below 100%. The next lever is batch size or KV-cache quantisation, not kernel work.
The timeline found idle time no kernel profile could show, and the intensity estimate stopped work on GEMMs already near their limit. LLM bottleneck analysis generalises this into MFU and MBU checks with perturbation tests.
Failure modes
- Profiling the warm-up. The first steps include compilation, autotuning and allocator growth. A report that starts at step 0 blames kernels that never run again. Use a capture range after warm-up.
- Graphs hiding kernels. Without
--cuda-graph-trace=node, a decode step appears as one opaque launch and per-kernel analysis is impossible. - Missing worker processes. The report shows only the API server's CPU threads and no kernels, because the GPU work happened in children that were not traced.
- Comparing
ncutimes with production. Clock control, cache flushing and replay change absolute durations. Compare percentages of peak, or disable those controls deliberately.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Layer-level NVTX | Timeline reads per layer | CPU overhead; keep it behind a profiling flag |
| Graph tracing at node level | Kernels inside graphs become visible | Per-kernel overhead inflates step time |
| Profiler-API capture range | Small reports tied to program state | Requires a code or engine hook |
ncu --set full | Every section in one pass | Many replays; minutes per kernel |
| Tracing all ranks | Finds stragglers and skew | Large reports; more overhead |
What to do next
- Add NVTX ranges for step, prefill and decode to your training loop or serving engine, with layer ranges behind a flag.
- Add a profiler-API capture hook that records a few steps after warm-up, and keep the
nsyscommand in the repository. - For serving engines, trace child processes and expand CUDA graphs to nodes.
- Export the kernel summary and bucket it into GEMM, attention, norm and activation, communication and other; compare the kernel total with step time to measure idle time.
- Compute the arithmetic intensity and the weight-read floor for your model, GPU and batch size before profiling any kernel.
- Run
ncuon the two or three most expensive kernels in a microbenchmark and compare achieved throughput with the floor. - With tensor parallelism, trace NCCL, check overlap in training and the communication share of decode, and look for straggler ranks.
- Confirm every fix with an untraced benchmark before you claim the gain.