A flame graph turns thousands of stack samples into a single picture. Each box is a function. Its width is the share of samples in which that function was on the stack, and its children sit on top of it. For an ordinary service the samples come from the CPU, so width means CPU time. An LLM training or serving process is different. Most of the useful work runs on a GPU, out of sight of any CPU sampler. The Python thread spends its time launching kernels, waiting on synchronisations, tokenising and scheduling. A naive CPU flame graph of a GPU-bound job therefore shows a wide tower of waiting and tells you almost nothing about where GPU time went.

This article covers three weights for the same stack (on-CPU, wall and GPU kernel time), how to produce each with py-spy, the PyTorch profiler and a short trace-folding script, a worked decode example, and the failure modes that make flame graphs lie. API names were checked against the PyTorch and py-spy sources in October 2026. General flame-graph reading and diffing is covered in continuous profiling in depth, so this page concentrates on what is specific to GPU-backed model code.

What a flame graph encodes

The input to every flame graph renderer is the folded or collapsed stack format, one line per unique stack, frames separated by semicolons, followed by a weight:

serve;engine_step;model_forward;LlamaDecoderLayer;attention 812
serve;engine_step;model_forward;LlamaDecoderLayer;mlp 1290
serve;engine_step;sample;Tensor.item 455

The renderer, usually Brendan Gregg's flamegraph.pl or the speedscope viewer, merges common prefixes into boxes. Three properties follow from that format, and each one is a common misreading.

  • The x-axis is not time. Siblings are sorted alphabetically so that identical stacks merge. For ordering you want a timeline or flame chart, such as Perfetto.
  • Width is a sum, not a duration. A box 30% wide could be one 30 ms call or 30,000 one-microsecond calls. Keep a call-count table beside the graph, because launch problems are about counts.
  • The weight decides the question. The same stacks weighted by CPU samples, by wall time or by GPU microseconds produce three different graphs. A frame that dominates one can be invisible in another.

Three weights for one stack

For an LLM process, pick the weight to match the question before you collect anything.

WeightCollected byAnswersBlind to
On-CPU samplespy-spy (default), perfIs the host the bottleneck: Python overhead, tokenizer, scheduler, data loading?Anything running on the GPU
Wall time, idle includedpy-spy --idleWhere does the thread wait: synchronisation, locks, queues, I/O?Why the wait is long
GPU kernel timePyTorch profiler, Kineto traceWhich model code launches the expensive kernels?Gaps between kernels, CPU work

The interesting LLM bugs sit at the boundary between these weights. A decode loop that calls .item() on a GPU tensor every token shows up as a wide synchronisation frame in the wall-time graph, as almost nothing in the CPU graph, and not at all in the GPU graph. The GPU sits idle while the host waits, and an idle GPU produces no kernel time to attribute. Compare the sum of GPU weight with wall time: if an eight-second capture holds three seconds of kernel time, the other five are gaps.

Three ways to weight the same Python stack, one folded format, one rendererpy-spy recordCPU samples, optional --idletorch.profilerwith_stack, with_modulesKineto trace JSONkernels + launch correlation--format rawfolded: a;b;c 123export_stacks()self_cuda_time_totalfold_trace.pykernel us per launch stackflamegraph.pl / speedscopeSVG or interactive viewdifffolded.plbefore vs after, per stepregressionsWeight = what the width meansCPU: time on a core. Wall (--idle): time a thread exists, waiting included. GPU: kernel time charged to the stack that launched it.Never mix weights in one graph. Normalise every graph to one step or one token before you compare two of them.
Figure 1. Three collectors, three weights, one folded format. The weight decides the question the graph can answer.

Host and wait time with py-spy

py-spy is a sampling profiler that reads the memory of a running CPython process from outside, so you can point it at a live serving process without restarting it. It needs ptrace permission, which in containers usually means SYS_PTRACE or running as the same user in the same PID namespace.

# 30 s of on-CPU samples from a running server, as an SVG flame graph
py-spy record --pid 4121 --duration 30 --rate 100 -o cpu.svg

# Same, but keep idle threads, so waits on locks and syncs are visible
py-spy record --pid 4121 --duration 30 --idle -o wall.svg

# Only threads holding the GIL: is Python itself the bottleneck?
py-spy record --pid 4121 --duration 30 --gil -o gil.svg

# Include C/C++ frames below Python (torch, CUDA runtime) where supported
py-spy record --pid 4121 --duration 30 --native --format raw -o cpu.folded

The raw format writes folded stacks, which you can diff and re-render. In an LLM server, four shapes recur. A wide --gil graph dominated by scheduler and request bookkeeping means the host cannot launch kernels fast enough. This is common in small-batch decode, where each step's kernels are short. A wide --native tower ending in cudaStreamSynchronize or cudaMemcpy under a Python line that reads a tensor value means a hidden sync. Wide detokenisation or JSON building on the engine thread means work that belongs on another thread. Wide tokenizer frames inside data-loader workers during training mean the input pipeline is starving the GPU.

The fix for a launch-bound decode is usually structural rather than a micro-optimisation: CUDA graphs, larger effective batches or a compiled step. CUDA streams and graphs covers the mechanism. The flame graph is how you prove the host was the limit before you pay for that change.

GPU-weighted stacks from the PyTorch profiler

To weight stacks by GPU time you need something that knows which CPU call launched which kernel. The PyTorch profiler records this. Each kernel carries a correlation identifier that links it to the CUDA runtime call that launched it. With with_stack=True the profiler also records the Python source location of each operator, and its export_stacks method writes folded stacks weighted by either "self_cpu_time_total" or "self_cuda_time_total".

import torch
from torch.profiler import profile, schedule, ProfilerActivity, record_function

sched = schedule(wait=2, warmup=3, active=5, repeat=1)   # skip compile/warm-up steps

with profile(
    activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA],
    schedule=sched,
    record_shapes=False,
    with_stack=True,       # Python file:line for each op
    with_modules=True,     # nn.Module hierarchy, e.g. model.layers.3.mlp
) as prof:
    for step in range(10):
        with record_function("decode_step"):
            logits = model(input_ids, past_key_values=cache).logits
            next_id = sample(logits)
        prof.step()

prof.export_stacks("gpu_stacks.folded", "self_cuda_time_total")
prof.export_chrome_trace("trace.json")   # keep the trace for fold_trace.py

Render with flamegraph.pl --title "GPU time" --countname us gpu_stacks.folded > gpu.svg. Older PyTorch tutorials also passed experimental_config=torch._C._profiler._ExperimentalConfig(verbose=True) to get fuller stacks. That is a private API whose behaviour can change between releases, so treat it as optional and check your version. Profile a handful of steady-state steps, because with_stack adds host overhead. In a torch.compile model many operators fuse into generated Triton kernels, whose names are not meaningful on their own, and the module hierarchy is what keeps the graph readable.

Folding a trace into GPU-weighted stacks

When you need control over the stack, for example engine phase, then module path, then kernel family, fold the exported trace yourself. In the Chrome-format trace PyTorch writes, GPU kernels have category kernel, launches have cuda_runtime, and CPU work appears as cpu_op, user_annotation and, with stacks enabled, python_function. Kernel and launch events share a correlation value in their args. Check those names against one of your own traces before relying on the script, because trace schemas move between versions.

import json, re, sys
from collections import defaultdict
from bisect import bisect_right

trace = json.load(open(sys.argv[1]))["traceEvents"]
X = [e for e in trace if e.get("ph") == "X"]
launch = {e["args"]["correlation"]: e for e in X
          if e.get("cat") == "cuda_runtime" and "correlation" in e.get("args", {})}
kernels = [e for e in X if e.get("cat") == "kernel"]

# CPU ranges per thread, sorted by start, to find what encloses a launch
ranges = defaultdict(list)
for e in X:
    if e.get("cat") in ("user_annotation", "cpu_op", "python_function"):
        ranges[(e["pid"], e["tid"])].append(e)
for r in ranges.values():
    r.sort(key=lambda e: e["ts"])

def family(name):
    n = name.lower()
    for key in ("gemm", "flash", "attn", "nccl", "triton", "elementwise", "reduce", "copy"):
        if key in n:
            return key
    return re.sub(r"<.*", "", name)[:40]

folded = defaultdict(float)
for k in kernels:
    l = launch.get(k["args"].get("correlation"))
    if l is None:
        folded["[unattributed];" + family(k["name"])] += k["dur"]
        continue
    rs = ranges[(l["pid"], l["tid"])]
    i = bisect_right([e["ts"] for e in rs], l["ts"])   # fine for a sketch; cache in real use
    stack = [e["name"] for e in rs[:i] if e["ts"] + e["dur"] >= l["ts"]]
    stack = [s.replace(";", ":") for s in stack][-12:]   # keep the innermost frames
    folded[";".join(stack + [family(k["name"])])] += k["dur"]

for stack, us in sorted(folded.items()):
    print(f"{stack} {int(us)}")

The output is GPU microseconds charged to the CPU stack that launched each kernel. Two details matter. First, when a kernel comes from a CUDA graph replay, the launch is the graph launch and not the original operator, so the stack stops at the replay call. Annotate captured regions with record_function before capture if you need them separated. Second, NCCL kernels in tensor-parallel runs overlap compute on another stream. Their width is time spent on the GPU, not time added to the step, so read them together with an overlap measurement such as the one in LLM GPU trace analysis.

Worked example: a decode step three times too slow

Here is a worked example with illustrative numbers. They show the method; your own model will produce different ones. A team serves an 8B model on one GPU. At batch size 4, decode runs at 38 ms per token, but a memory-bandwidth estimate (weights plus KV cache bytes divided by HBM bandwidth) suggests about 11 ms. They capture five steady-state steps three ways.

GraphWhat dominatedShare
GPU weight (fold_trace.py)attention + mlp GEMMs under model.layers.*~10.5 ms of kernel time per step
Wall time (py-spy --idle --native)sample() then Tensor.item then cudaStreamSynchronize~45% of engine-thread samples
On-CPU (py-spy --gil)engine_step then per-request stop-string check and detokenise~35% of samples

The GPU graph looks healthy: the expensive kernels are the ones that should be expensive, and they add up to roughly the estimate. That leaves about 27 ms per step outside the kernels. The wall-time graph shows where it goes. After sampling, the code calls .item() on every request's token to check for end-of-sequence, which forces a sync, and then runs Python stop-string matching and incremental detokenisation for each request before launching the next step. The GPU idles through all of it.

The fixes follow directly. Keep the sampled ids on the device and copy them to host asynchronously into pinned memory, one step behind. Move stop-string and detokenisation work to a separate thread that consumes those ids. Then capture the decode step in a CUDA graph to remove most of the per-kernel launch cost. Re-captured and diffed, the sync frame has gone from the wall graph and the GPU graph is nearly unchanged, which is correct: the kernels were never the problem.

Differential flame graphs for regressions

For regressions, fold two captures, normalise them, and render the difference:

# Normalise each capture to per-step weights first: divide by the number of steps profiled
python normalise.py before.folded --steps 5 > before.n
python normalise.py after.folded  --steps 5 > after.n
difffolded.pl before.n after.n | flamegraph.pl --title "after - before (GPU us/step)" > diff.svg

difffolded.pl emits each stack with both weights. flamegraph.pl then draws the second profile's shape and colours boxes red where weight grew and blue where it shrank. Its -n flag normalises the first profile to the second's total, which hides absolute changes, so per-step normalisation is usually what you want. Diff only like with like: same batch size, sequence lengths, parallel layout and weight type. For production regressions, collect folded stacks continuously at a low rate, as described in profiling in production, and diff a release window against the previous one.

Failure modes

Common ways a flame graph misleads you:

  • Asynchrony mistaken for cost. CUDA launches return immediately, so the Python line that launched a slow kernel looks cheap in a CPU graph, while the next synchronising line looks expensive. Charge GPU time with the correlation fold, never by CPU stack alone.
  • Warm-up in the window. Compilation, autotuning and allocator growth in the first steps can dominate. Use a profiler schedule and discard the early steps.
  • Sampling bias. At 100 Hz a 30-second capture holds about 3,000 samples per thread. Frames under 1% are noise. Raise the duration, not the rate, because a higher rate increases overhead and perturbs the very launch timing you are measuring.
  • Multi-rank confusion. Profile one rank per job unless you are comparing ranks, label the file with the rank, and never merge ranks into one graph. A straggler rank's waits look like collective time on every other rank.
  • Profiler overhead. with_stack, record_shapes and memory profiling all slow the host. A launch-bound decode can look worse under the profiler than it is. Confirm every conclusion with an unprofiled timing.

Trade-offs

ToolStrengthCost
py-spyAttaches to a live process, low overhead, no code changeCPU and wall only; GPU work is invisible
torch.profiler export_stacksGPU time per Python stack in a few linesHigher overhead, short windows, version-dependent stack detail
Custom trace foldYour own stack definition: phase, module, kernel familyYou maintain a parser against a moving schema
Nsight Systems timelineOrdering, overlap, gaps, multiple streamsNot aggregated; hard to compare runs

Use flame graphs to aggregate and compare, and timelines to explain ordering and gaps, as covered in Nsight for LLM workloads.

What to do next

  1. Write down the question first: host overhead, waiting, or GPU kernel cost. Pick the matching weight.
  2. Capture one steady-state window three ways: py-spy --gil, py-spy --idle --native, and export_stacks(..., "self_cuda_time_total") with with_modules=True.
  3. Compare total GPU weight per step with measured step time. If the gap is large, the problem lies outside the kernels, so start from the wall-time graph.
  4. Search the wall graph for synchronising frames under .item(), .cpu(), .tolist() and boolean checks on GPU tensors in the decode loop.
  5. Adapt fold_trace.py to your stack definition, validate its category and argument names on one trace, and commit it next to the model code.
  6. Store folded stacks, normalised per step, from every benchmark run, and diff each release against the last.
  7. Confirm every fix with an unprofiled step time before claiming the win.
Key takeaway: An LLM flame graph is only as useful as its weight. CPU samples show host overhead, wall time with idle threads shows waits and hidden syncs, and GPU time folded through launch correlation shows which model code owns the expensive kernels. Capture steady-state windows, normalise per step, diff releases, and confirm every conclusion with an unprofiled timing.