The PyTorch profiler answers one question better than any other tool on the machine: which line of your model code caused the GPU time you are paying for. Nsight Systems sees every kernel but not your module names; a wall clock around the training step sees the total but not the parts. torch.profiler sits in the middle, recording Python and ATen operators on the host, the CUDA runtime calls they make, and (through NVIDIA's CUPTI interface) the kernels and copies that actually ran on the device, then stitching them together with correlation ids.

This article is a working manual for that one tool applied to large language models. It covers what the profiler records and what each column of its tables means, how to use the step schedule so a trace captures steady state rather than warmup, how to read a trace for the patterns that matter in LLM training and inference, how to profile memory with allocator snapshots, how to handle multi-GPU jobs, and how much the profiler itself distorts what it measures. The worked example is a fine-tuning loop whose GPU sits idle for a third of every step. For the wider ladder of tools, from service metrics down to single kernels, see LLM performance profiling; this page stays inside the PyTorch profiler.

What the profiler records

Three event streams feed every profile. The first is host operators: each call into an ATen operator such as aten::mm or aten::scaled_dot_product_attention, nested inside any record_function ranges you add and, with with_stack=True, annotated with the Python source line. The second is the CUDA runtime: cudaLaunchKernel, cudaMemcpyAsync, cudaStreamSynchronize and friends. The third is device activity collected by CUPTI: the kernels, memory copies and memsets, with their start and end on the GPU's own clock and the stream they ran on.

The key structural fact is that the host and device run asynchronously. When Python calls torch.matmul, the CPU enqueues a kernel and returns in a few microseconds; the kernel may run milliseconds later. The profiler records a correlation id on the launch and on the kernel, so the trace viewer can draw an arrow between them. Every diagnosis below is a statement about that arrow: whether the CPU is far enough ahead of the GPU (healthy), or whether the GPU is waiting on the CPU (launch-bound, data-bound or sync-bound).

What torch.profiler records and where it ends upPython + ATen opsrecord_function scopesCUDA runtime callscudaLaunchKernel, memcpyDevice activitykernels via CUPTIProfiler (Kineto)events + correlation idskey_averages() tableaggregate per opTrace JSONPerfetto, HTAMemory snapshotpytorch.org/memory_vizA correlation id ties each CPU-side launch to the GPU kernel it produced; that link is what makes the timeline readable.
Host operators, runtime calls and device activity are recorded separately and joined by correlation id.

Capturing steady state with a schedule

A useful LLM profile is short, captures steady state, and lands in a file you can open later. The schedule does that. wait steps run with the profiler off, warmup steps run with tracing on but results discarded (the first traced steps carry CUPTI start-up cost), active steps are recorded, and repeat controls how many cycles happen; skip_first skips steps once at the start, which is where you hide compilation, cuDNN autotuning and the first allocator growth. After each active window the on_trace_ready callback receives the profiler and normally writes a trace.

schedule(skip_first=10, wait=1, warmup=1, active=3, repeat=2)skip x10waitwarmupactive x3waitwarmupactive x3on_trace_ready fires after each active window: two trace filesprof.step() advances one step; a step is whatever you call it after, normally one optimizer or decode iteration.
The step schedule: steps are counted by prof.step(), so one call per training or decode iteration.
import torch
from torch.profiler import (profile, schedule, ProfilerActivity,
                            record_function, tensorboard_trace_handler)

sched = schedule(skip_first=10, wait=1, warmup=1, active=3, repeat=2)

with profile(
    activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA],
    schedule=sched,
    on_trace_ready=tensorboard_trace_handler("./traces", use_gzip=True),
    record_shapes=True,      # input shapes per op: needed to group by shape
    profile_memory=True,     # allocator events per op
    with_stack=True,         # Python source lines; adds overhead
) as prof:
    for step, batch in enumerate(loader):
        with record_function("data_to_device"):
            batch = {k: v.cuda(non_blocking=True) for k, v in batch.items()}
        with record_function("forward"):
            loss = model(**batch).loss
        with record_function("backward"):
            loss.backward()
        with record_function("optimizer"):
            opt.step(); opt.zero_grad(set_to_none=True)
        prof.step()          # one schedule tick per iteration
        if step >= 20:
            break

print(prof.key_averages().table(sort_by="self_cuda_time_total", row_limit=15))

Despite the name, tensorboard_trace_handler just writes Chrome-trace JSON files, one per worker per window; they open in the Perfetto UI or in chrome://tracing with no TensorBoard involved. For a single ad-hoc capture without a schedule, call prof.export_chrome_trace('trace.json') after the context exits. Two flags deserve caution. with_modules is documented as deprecated and only does anything for TorchScript, so leave it off in eager code. with_flops estimates FLOPs for a subset of operators (matrix multiplies and convolutions) from their shapes; it is a sanity check on arithmetic intensity, not a substitute for counting your model's FLOPs.

Reading key_averages

The key_averages() table aggregates events by operator name. The columns that matter split along two axes: CPU versus CUDA, and total versus self. Total time includes child operators; self time excludes them. A high-level op such as aten::linear has large total time and almost no self time because the work is in its child aten::addmm. Sort by self_cuda_time_total to find where GPU time actually goes, and by self_cpu_time_total to find Python and dispatcher overhead. The profiler recipe also documents self_cpu_memory_usage and cpu_memory_usage for the memory columns when profile_memory=True is set.

Two grouping options turn the table from a list into a diagnosis. key_averages(group_by_input_shape=True) splits each op by input shapes, which in an LLM separates the QKV projection from the MLP up-projection and the LM head, and shows immediately when an odd sequence length or vocabulary size is producing an inefficient matmul. key_averages(group_by_stack_n=5) (with with_stack=True) groups by the top five Python frames, which attributes an op to the call site that issued it, useful when the same aten::copy_ comes from three places.

Symptom in the tableLikely meaningNext check
High self CPU on many tiny ops, low CUDA totallaunch or dispatcher boundtrace: gaps between kernels
aten::copy_ / Memcpy HtoD near the tophost copies in the hot pathpin memory, non_blocking, move work to GPU
cudaStreamSynchronize or aten::item with large CPU timehost waiting on devicefind the .item(), print or bool check
One matmul shape far slower per FLOPbad shape or dtype pathgroup_by_input_shape, pad to multiple of 8/64
Attention op dominates at long contextexpected; check which kerneltrace: is it the fused SDPA backend

Reading the trace

The table tells you what; the trace tells you when, and most LLM problems are timing problems. Open the JSON in Perfetto and look at three rows: the Python thread, the CUDA runtime row beneath it, and the GPU stream rows. Then look for these shapes.

A solid GPU stream with the CPU running ahead is the healthy state for training at reasonable batch sizes: launches happen early, kernels are back to back, and the gaps between kernels are a few microseconds. Speeding this up means faster kernels, not profiler work. A comb of short kernels with gaps between them, where each gap lines up with a launch on the CPU row, is launch-bound execution; it is typical of small-batch decode and is the case CUDA graphs and torch.compile exist for. A wide idle hole on every stream that lines up with a cudaStreamSynchronize or a long Python region is a host stall: a data loader that is not keeping up, a tokenizer running in the training process, a logging call that pulls a tensor to the CPU. Copies and kernels on separate streams that never overlap means an intended prefetch is serialized, usually because the source tensor was not pinned.

Because a trace of even three steps of a large model can contain hundreds of thousands of events, it pays to summarize it programmatically. Holistic Trace Analysis (HTA), a PyTorch-ecosystem library, reads a directory of these traces and reports where GPU time went per rank:

from hta.trace_analysis import TraceAnalysis

analyzer = TraceAnalysis(trace_dir="./traces")
time_df = analyzer.get_temporal_breakdown()      # idle / compute / non-compute per rank
idle_df = analyzer.get_idle_time_breakdown()     # why the GPU was idle
kernels = analyzer.get_gpu_kernel_breakdown()    # time by kernel type and name
overlap = analyzer.get_comm_comp_overlap()       # NCCL vs compute overlap per rank

Memory: allocator snapshots

Out-of-memory errors in LLM work are rarely about weights; they come from activations, optimizer state, the KV cache and allocator fragmentation. profile_memory=True attributes allocations to operators, which helps when one op allocates surprisingly. For the full picture use the CUDA allocator's history recorder, which logs every allocation and free with a stack trace and lets you replay the memory timeline. The profiler's own export_memory_timeline method is documented as deprecated in favour of this recorder.

import torch

torch.cuda.memory._record_memory_history(max_entries=100_000)
for step in range(3):
    train_step(model, next(it))          # run only a few steps: history grows fast
torch.cuda.memory._dump_snapshot("finetune_snapshot.pickle")
torch.cuda.memory._record_memory_history(enabled=None)   # stop recording
# Drag the pickle onto https://pytorch.org/memory_viz (runs locally in the browser)

The visualizer shows allocated memory over time as stacked blocks, each traceable to the Python line that allocated it. A training step has a characteristic sawtooth: memory climbs through the forward pass as activations are saved, peaks at the start of backward, and falls as gradients consume them. The peak's height minus the flat baseline (weights, gradients, optimizer state) is your activation footprint; if it dominates, activation checkpointing or a smaller micro-batch is the lever. If the snapshot shows reserved memory well above allocated, the problem is fragmentation, and the allocator settings in GPU memory pools apply. The underscore prefix on these functions is meaningful: they are semi-private APIs whose arguments have changed across releases, so pin them behind a helper in your codebase.

Worked example: a fine-tune that idles a third of every step

A team fine-tunes a 7B-parameter decoder with LoRA on a single 80 GB GPU. Throughput looks low, and the GPU utilization graph hovers around 65%. The numbers below are illustrative of the pattern, not a benchmark. The first profile, with the schedule above, shows a step of roughly 900 ms. In the trace, the GPU stream is busy for about 600 ms and then idle for about 300 ms on every step.

Zooming into the idle hole: the Python row shows a long data_to_device range, and inside it the runtime row shows cudaMemcpyAsync followed by a wait. The table, sorted by self_cpu_time_total, has aten::copy_ and aten::to near the top. Two causes combine. The DataLoader was created without pin_memory=True, so non_blocking=True silently became a synchronous copy from pageable memory; and the collate function tokenized raw text in the main process with num_workers=0. A second, smaller hole appears at the end of each step: a cudaStreamSynchronize under aten::item, caused by print(loss.item()) on every iteration.

The fixes are mundane: pre-tokenize the dataset, set num_workers=4 and pin_memory=True on the loader, and log the loss every 50 steps from a detached accumulator. The second profile shows the GPU stream nearly solid, with the CPU now running one step ahead, and the step time drops toward the ~600 ms of actual compute. Only now is it worth looking at kernels: grouping by input shape shows the LoRA adapters as many small matmuls whose launch cost is comparable to their run time, which is the point where torch.compile fusion starts to pay.

Multi-GPU jobs

In data- or tensor-parallel jobs, profile every rank, not just rank 0: the slow rank sets the pace for collectives, and its trace is the one that differs. The trace handler names files by worker, so pass a worker_name that includes the rank. In each trace, NCCL kernels appear on their own stream; time a rank spends inside an all-reduce kernel is often time spent waiting for a straggler, so a rank with long NCCL kernels may be the fast one. Compare the compute time before each collective across ranks to find the real laggard, and use HTA's overlap report to check that gradient all-reduce runs concurrently with backward compute rather than after it. Collective internals are covered in NCCL collectives.

Overhead and trade-offs

The profiler is not free, and its cost is uneven. Kernel tracing through CUPTI adds a small per-launch cost, which matters most exactly where launches are already the bottleneck, so a launch-bound decode loop looks somewhat worse under the profiler than it is. with_stack=True and record_shapes=True add host-side work per operator and inflate CPU self time. profile_memory=True adds allocator hooks. Practical rules follow from this: profile a few steps, never a whole epoch; keep one lean configuration (activities only) for timing comparisons and a rich configuration for attribution; never compare a profiled step time with an unprofiled one; and remember that torch.cuda.synchronize() calls you add to make timings neat will themselves remove the CPU-GPU overlap you are trying to see.

Failure modes

  • Profiling warmup. Without skip_first, the trace is dominated by compilation, autotuning and allocator growth that never recur.
  • Forgetting prof.step(). The schedule never advances, nothing is recorded, and the handler is never called.
  • Reading total time as cost. Sorting by total time puts wrapper ops on top; the work is in self time.
  • Trusting CPU timings of CUDA ops. A matmul's CPU time is its launch cost; its GPU time is in the CUDA columns.
  • Giant traces. A long active window produces files the viewer cannot open; three steps is usually enough.
  • Profiling rank 0 only. The slow rank is the interesting one.

What to do next

  1. Add record_function ranges around data transfer, forward, backward and the optimizer in your training loop, and leave them in: they cost little when profiling is off.
  2. Capture three steady-state steps with a schedule and skip_first, and keep the trace files next to the commit that produced them.
  3. Read the trace before the table: classify the step as compute-bound, launch-bound or host-stalled, then use the table to find the operator.
  4. Remove every per-step .item(), print and Python if on a GPU tensor in the hot path.
  5. Take one allocator snapshot of a full step and note peak activation memory.
  6. When the framework view runs out, move down to Nsight for LLMs and Nsight Systems and Nsight Compute.
Key takeaway: torch.profiler joins three streams, host operators, CUDA runtime calls and device kernels, through correlation ids, which is what lets you blame GPU time on a line of model code. Capture a few steady-state steps with a schedule, read the trace first to classify the step, then sort key_averages by self CUDA time and group by input shape to find the operator. Use allocator snapshots for memory, profile every rank in distributed jobs, and keep the profiler's own overhead in mind when timings look worse than expected.