NVIDIA Nsight Systems is a system-wide tracer. While your program runs it records, on one shared clock, the CUDA API calls each CPU thread makes, the kernels and memory copies each GPU executes, the NVTX ranges your code marks, operating-system calls such as thread waits and file reads, and optionally CPU samples and GPU hardware metrics. Its job is to answer one question: where does the time between the work go? It is the tool for a training step that takes 400 ms when the kernels add up to 250 ms.

Most people meet it as a GUI timeline, and the GPU profiling overview explains how to read those rows. This article treats it as a pipeline instead. A timeline is good for one step looked at once; a training job needs numbers you can compare across runs, ranks and commits. The command-line tool, nsys, produces a report you can summarise with built-in statistics, export to SQLite, query with your own scripts and process with multi-report recipes. By the end you will be able to capture a clean window from a PyTorch training loop, find idle GPU time with a script, attribute it to what the CPU was doing, and turn the result into a regression check.

Options change between releases. Every flag below appears in the current Nsight Systems User Guide, but check nsys --version and nsys profile --help on your machine before relying on newer ones such as the PyTorch and Python sampling options.

What Nsight Systems records

Nsight Systems records events with start and end timestamps and correlation IDs. A kernel launch on the CPU, such as cudaLaunchKernel, carries a correlation ID that also appears on the kernel's execution record on the GPU, which is how the tool draws the arrow from launch to execution and how your scripts can join the two. NVTX ranges are CPU-side intervals with names you choose; they are what make a trace readable in model terms such as forward, backward and optimizer step.

What it does not do is explain why a single kernel is slow. It shows that a GEMM took 3 ms, not whether it was limited by memory bandwidth or by tensor-core throughput. That is Nsight Compute's job, and the right order is always the same: Nsight Systems first to find where the time goes, Nsight Compute second, on the few kernels that matter.

Nsight Systems as a pipeline, not just a viewerAnnotated jobNVTX ranges + profiler start/stopnsys profiletrace cuda, nvtx, osrtreport.nsys-repbinary reportGUI timelineeyes on one stepnsys export -t sqlitereport.sqlitekernels, runtime, NVTX, StringIdsnsys statsSummary reportscuda_gpu_kern_sum, cuda_api_sumYour scriptsgaps, busy %CI regression gatefail if busy % or step time driftsnsys recipemulti-report analysis: gpu_gaps, nvtx_pace, nccl_sum
From capture to automated checks. The GUI is one consumer of the report; the stats reports, the SQLite export and recipes are the others.

Capturing a steady-state window

Profiling a whole training run produces enormous reports dominated by startup, data loading warm-up, compilation and the first iterations where allocators and autotuners are still settling. Capture a steady-state window instead. The simplest reliable way is to call the CUDA profiler start and stop functions from the training loop and tell nsys to record only between them.

import torch
import torch.cuda.nvtx as nvtx

WARMUP, CAPTURE = 20, 10

for step, batch in enumerate(loader):
    if step == WARMUP:
        torch.cuda.synchronize()
        torch.cuda.profiler.start()          # cudaProfilerStart
    nvtx.range_push("train_step")
    with nvtx.range("data_to_gpu"):
        x, y = batch[0].cuda(non_blocking=True), batch[1].cuda(non_blocking=True)
    with nvtx.range("forward"):
        loss = loss_fn(model(x), y)
    with nvtx.range("backward"):
        loss.backward()
    with nvtx.range("optimizer"):
        opt.step()
        opt.zero_grad(set_to_none=True)
    nvtx.range_pop()
    if step == WARMUP + CAPTURE - 1:
        torch.cuda.synchronize()
        torch.cuda.profiler.stop()           # cudaProfilerStop
        break
nsys profile \
  --trace=cuda,nvtx,osrt,cudnn,cublas \
  --capture-range=cudaProfilerApi --capture-range-end=stop \
  --sample=process-tree --cudabacktrace=all \
  --output=train_%q{RANK} --force-overwrite=true \
  python train.py

--capture-range=cudaProfilerApi makes collection wait for the start call, and --capture-range-end=stop ends collection at the stop call while letting the process continue. The alternative, --capture-range=nvtx with --nvtx-capture, starts on a named NVTX range without code changes to the profiler calls. --output expands %q{ENV_VAR} from the environment, so under a distributed launcher each rank writes its own report; %h and %p insert the hostname and process ID.

Add options deliberately, because each adds overhead or report size. --cudabacktrace records CPU call stacks on CUDA API calls and needs CPU sampling. The guide's PyTorch example also uses --pytorch=functions-trace-shapes,autograd-nvtx, --python-backtrace=cuda and --python-sampling=true, which attach Python stacks and PyTorch function ranges to the trace; these are recent options, so confirm your version supports them. --cuda-memory-usage=true tracks GPU memory per kernel and is documented as potentially causing significant runtime overhead; use it in a separate capture, not the one you time.

First numbers: nsys stats

Before opening the GUI, get the numbers. nsys stats runs summary reports over a report file and can write them as CSV for scripts.

nsys stats --report cuda_gpu_kern_sum --report cuda_api_sum --report nvtx_sum \
           --format csv --output . train_0.nsys-rep
ReportWhat it tells youFirst thing to look for
cuda_gpu_kern_sumTotal and average GPU time per kernel nameWhether the top kernels are the ones you expect (GEMMs, attention) or surprises such as copies and elementwise ops
cuda_api_sumCPU time inside each CUDA API functionLarge totals in synchronisation calls such as cudaStreamSynchronize or cudaDeviceSynchronize
nvtx_sumDuration statistics per NVTX rangeStep-time variance: a long tail in train_step means some steps stall
cuda_gpu_mem_time_sumTime in memory copies and memsets by kindHost-to-device copy time that should be overlapped with compute
nvtx_gpu_proj_sumNVTX ranges projected onto the GPU work they launchedGPU time of forward versus backward versus optimizer
osrt_sumTime in OS runtime callsLong waits on locks, condition variables or reads in the data path

These summaries already answer the first question for many jobs. If the kernel sum over the captured steps is much less than the step time multiplied by the step count, the GPU is idle part of every step, and the problem is not kernel speed. The next step finds where those idle periods are.

The SQLite export

Running nsys export -t sqlite report.nsys-rep (or passing --stats=true at capture time) produces a SQLite database alongside the report. The tables you need first:

  • CUPTI_ACTIVITY_KIND_KERNEL: one row per kernel execution, with start and end in nanoseconds, deviceId, streamId, correlationId, and shortName and demangledName as integer IDs into StringIds.
  • CUPTI_ACTIVITY_KIND_RUNTIME: CUDA runtime API calls on the CPU, with start, end, globalTid, correlationId and nameId.
  • NVTX_EVENTS: ranges and marks, with start, end, and a name either in text or as textId into StringIds.
  • StringIds: id and value; every repeated string is stored once here.
-- Top kernels by total GPU time
SELECT s.value AS kernel, COUNT(*) AS calls,
       SUM(k."end" - k.start) / 1e6 AS total_ms
FROM CUPTI_ACTIVITY_KIND_KERNEL k
JOIN StringIds s ON s.id = k.shortName
GROUP BY s.value ORDER BY total_ms DESC LIMIT 15;

Join names yourself through StringIds as shown. The documentation also describes convenience views with names already resolved, but the base tables are the stable contract across versions.

Scripted analysis: GPU busy time and idle gaps

The quantity that matters for training throughput is GPU busy time: the fraction of wall time during which at least one kernel is running on a device. The tempting shortcut, subtracting each kernel's end from the next kernel's start, is wrong as soon as kernels overlap on different streams, which they do whenever communication, copies or multiple streams are in play. Merge each device's kernel intervals first, then measure the holes between merged intervals.

import sqlite3
import sys
from collections import defaultdict


def merge(intervals):
    merged = []
    for s, e in sorted(intervals):
        if merged and s <= merged[-1][1]:
            merged[-1][1] = max(merged[-1][1], e)
        else:
            merged.append([s, e])
    return merged


def analyse(path, min_gap_us=100):
    db = sqlite3.connect(path)
    per_dev = defaultdict(list)
    for dev, s, e in db.execute(
            'SELECT deviceId, start, "end" FROM CUPTI_ACTIVITY_KIND_KERNEL'):
        per_dev[dev].append((s, e))
    # CPU-side runtime calls, to say what the host was doing during each gap
    calls = db.execute(
        'SELECT r.start, r."end", s.value FROM CUPTI_ACTIVITY_KIND_RUNTIME r '
        'JOIN StringIds s ON s.id = r.nameId').fetchall()
    for dev, ivs in sorted(per_dev.items()):
        m = merge(ivs)
        window = m[-1][1] - m[0][0]
        busy = sum(e - s for s, e in m)
        print(f"device {dev}: busy {100 * busy / window:.1f}% of {window / 1e6:.1f} ms")
        blame = defaultdict(float)
        for (_, e0), (s1, _) in zip(m, m[1:]):
            gap = s1 - e0
            if gap < min_gap_us * 1000:
                continue
            during = [n for cs, ce, n in calls if cs <= e0 + gap / 2 <= ce]
            blame[during[0] if during else "(no CUDA call: host busy elsewhere)"] += gap
        for name, ns in sorted(blame.items(), key=lambda kv: -kv[1])[:5]:
            print(f"  {ns / 1e6:8.1f} ms idle while host was in {name}")


if __name__ == "__main__":
    analyse(sys.argv[1])

The blame column is a heuristic: it reports which CUDA API call, if any, was in progress at the midpoint of each gap, across all threads. A gap during cudaStreamSynchronize means the host was waiting on the GPU earlier and only then launched more work; a gap with no CUDA call in progress means the host was doing something else entirely, such as Python, data loading or a lock. Add the NVTX range open at the same time to name the phase. For long captures the per-gap scan over calls is quadratic; sort calls by start and use bisection if it becomes slow. If copies occupy the device in a way you care about, add rows from the memcpy table to the intervals.

Worked example: an input-bound fine-tune

An illustrative case, with numbers chosen to show the method rather than measured on specific hardware. A single-GPU fine-tune reports 410 ms per step in nvtx_sum, but cuda_gpu_kern_sum sums to about 250 ms per step. The script prints a busy fraction near 61 percent, with most idle time in one gap at the start of each step, attributed to no CUDA call, and the NVTX range open during it is data_to_gpu. osrt_sum shows large totals in condition-variable waits on the main thread.

Two causes fit: the main thread is waiting for the next batch from too few loader workers, and the batch arrives in pageable memory, so the host-to-device copy cannot overlap with compute. The cuda_memcpy_async recipe, which identifies asynchronous copies that became synchronous because the memory is pageable, confirms the second. The fix is ordinary: pin_memory=True and more num_workers in the DataLoader, plus non_blocking=True on the copy, which the code above already had but which only helps once memory is pinned. Re-run the same capture and the same script; the claim is only accepted when the busy fraction rises and step time falls in the new report, not when the code looks right.

Recipes and multi-rank jobs

Recipes are Python analyses that ship with Nsight Systems and run over one or many reports, which is what multi-rank jobs need. They are invoked as nsys recipe <recipe-name> --input <reports> and write their output, often a Jupyter notebook plus data files, to a directory. Recipes depend on Python packages listed in the Analysis Guide; install them in the environment where you run the analysis, which need not be the training node. Useful ones for training:

RecipePurpose
gpu_gapsRegions where a GPU is idle longer than a threshold
cuda_api_syncSynchronisation APIs that block the host until GPU work completes
cuda_memcpy_asyncAsync copies that became synchronous because memory is pageable
nvtx_paceProgress and consistency of a named NVTX range across the run
cuda_gpu_kern_sumThe kernel summary computed across many reports
nccl_sumNCCL communication summary across communicators, sizes and ranks

For a multi-node job, nvtx_pace on train_step across all rank reports shows whether one rank is consistently slower, which in synchronous data-parallel training sets the pace for everyone. When communication time grows, the Nsight for LLMs article covers reading tensor-parallel collectives in the timeline.

Turning it into a regression check

Once the busy-fraction script exists, a regression check is a few more lines. Run a short, fixed capture on a known configuration in a nightly job, compute busy percent per device and median step time from NVTX_EVENTS, compare against a stored baseline, and fail when either moves beyond a tolerance you set from run-to-run noise. Store the report files as artefacts so a failing night can be opened in the GUI. This catches the regressions that end-to-end throughput dashboards blur: a new synchronising call, a data-path change, a kernel that stopped being fused.

Failure modes

  • Capturing warm-up. Autotuning, compilation and allocator growth dominate early steps. Always start the capture after warm-up.
  • Profiling a different program. Setting CUDA_LAUNCH_BLOCKING=1 or adding synchronisations for debugging changes overlap; the trace then explains a job that does not exist in production.
  • Huge reports. Long captures with backtraces and sampling produce reports that are slow to export and open. Capture a handful of steps.
  • Naive gap arithmetic. Subtracting consecutive kernel timestamps across streams counts overlapped work as idle or hides real gaps. Merge intervals.
  • Reading NVTX as GPU time. NVTX ranges are CPU intervals; with asynchronous execution the GPU work they launch can finish much later. Use the GPU projection report when you need GPU time per phase.
  • One rank only. Profiling rank 0 of a distributed job misses the straggler. Capture all ranks or at least a sample across nodes.

Trade-offs

Nsight Systems sees everything at once with low per-event cost, but it gives no per-kernel hardware detail and produces large files. The PyTorch profiler, described in the PyTorch profiler guide, knows operator names, shapes and allocator state natively and needs no external tool, but sees less of the operating system and other processes. Nsight Compute explains one kernel in depth at the cost of replaying it many times. Use Nsight Systems to decide where to look, and the other two to look closely.

What to do next

  1. Add NVTX ranges for data, forward, backward and optimizer to your training loop.
  2. Wrap a window of steps after warm-up in profiler start and stop calls.
  3. Capture with --capture-range=cudaProfilerApi and a per-rank output name.
  4. Run nsys stats and compare the kernel sum to step time.
  5. Export to SQLite and run the merged-interval busy script; note the top idle causes.
  6. Fix the largest cause, re-capture the same window and keep the before and after numbers.
  7. Turn the script into a nightly regression check with stored report artefacts.
Key takeaway: Treat Nsight Systems as a pipeline: capture a steady-state window between profiler start and stop calls, read the stats reports, export to SQLite, measure GPU busy time by merging kernel intervals per device, attribute the idle gaps to what the host was doing, and keep the script as a regression check.