Metrics tell you that a service got slower or more expensive. Traces tell you which request and which hop. Neither tells you which line of code is burning the CPU. A profiler does, and a continuous profiler does it all the time, in production, on every instance, cheaply enough that nobody has to remember to turn it on before the incident.

The idea is old: interrupt the program at regular intervals, record the call stack, and count. What makes it continuous is the engineering around that idea: overhead near one percent, function names without shipping debug symbols to every host, storage that answers a week-long query in seconds, and comparisons that make a regression jump out. This article walks the pipeline stage by stage, with the statistics, data model and code you need to run one and trust it.

Advertisement

Sampling, and why the numbers are trustworthy

A sampling profiler does not measure each function call. It asks, many times per second, "what is this CPU executing right now?" and keeps a tally. If a function is on the stack in 3 percent of samples, it accounts for roughly 3 percent of CPU time. The accuracy is a counting problem, so it improves with the square root of the number of samples.

Do the arithmetic for a realistic setting. Parca Agent defaults to 19 samples per second per logical CPU. A service using 8 cores for one minute yields about 19 x 60 x 8 = 9,120 samples. A function at 2 percent of CPU gets about 182 of them; the standard deviation of that count is about the square root, 13.5, so the estimate is 2 percent plus or minus about 0.15 points. Over an hour, or across 50 replicas, the error becomes negligible. This is why continuous profiling can use low rates: it trades per-instance resolution for aggregation across time and fleet.

Odd rates such as 19 or 99 Hz keep sampling out of lockstep with periodic work at round frequencies, which would be systematically over- or under-counted. In-process profilers often sample faster: Go's CPU profiler uses 100 Hz. Higher rates help short-lived processes and short incidents, at proportional cost.

The pipeline

A continuous profiler is a pipeline: sample, unwind, symbolize, aggregate, store by labels, query and diffTimer interruptN Hz per CPUStack unwinderFP or DWARF tablesLocal aggregationstack -> countUploadevery 10-60 son the host: addresses only, cheap and boundedSymbolizerbuild ID -> debuginfoIngestdedupe stacks, attach labelsColumnar storagetime x labels x stack idpprof / OTLPQueryselector + time rangeMergesum samples per stackFlame graphDiff: A vs B
Host-side work is kept minimal: capture addresses, aggregate identical stacks, upload. Symbolization, storage and analysis happen centrally.

On the host, a timer fires, a handler captures the stack as instruction addresses, and identical stacks are aggregated into counts, so a minute of sampling becomes a few thousand unique stacks. Every 10 to 60 seconds the agent uploads that aggregate, tagged with labels such as service, version, pod and region.

Centrally, a symbolizer maps addresses to function, file and line using debug information looked up by the binary's build ID. Ingest deduplicates stacks across uploads, because the same few thousand stacks recur all day. Storage is columnar and indexed by time and labels. A query selects a label set and time range, sums sample counts per stack, and renders the merged profile as a flame graph or a diff against another range.

Collection comes in two families. In-runtime profilers, such as Go's pprof, the JVM's async-profiler or JDK Flight Recorder, and py-spy for Python, understand their runtime and produce clean stacks including interpreted and JIT-compiled frames. Whole-system eBPF profilers attach a program to a perf event, sample every process on the host without code changes, and must reconstruct stacks themselves; how such programs load and run is covered in eBPF observability architecture. Many fleets run both: eBPF for coverage, in-runtime profilers where deep language detail matters.

Advertisement

Unwinding: turning a register snapshot into a call stack

At the moment of the sample, the profiler has the instruction pointer and the stack pointer. Recovering the callers is called unwinding, and it is the hardest part of whole-system profiling.

Frame pointers. If code is compiled to keep a frame pointer register, each frame stores the caller's frame pointer and return address at a known place, forming a linked list. Walking it is a few memory reads per frame, cheap enough to do inside the kernel at interrupt time. The catch is that compilers commonly omit frame pointers to free a register, which breaks the chain at the first function built that way. Distributions have been reversing that: Fedora 38 and Ubuntu 24.04 enabled frame pointers by default in their packages, and you can do the same with -fno-omit-frame-pointer in your own builds. The measured cost is usually small, but measure it for your workload.

Unwind tables. Without frame pointers, the information lives in the binary's .eh_frame section, the same tables C++ exceptions use, which describe how to find the caller's frame at every instruction address. Interpreting them per sample in the kernel is too slow, so eBPF profilers preprocess the tables in user space into compact lookup structures and load them into maps that the sampling program consults. It works without rebuilding anything, at the price of agent memory and startup work per binary.

Runtimes. JIT-compiled code has no tables on disk, and an interpreter's native stack shows only the interpreter loop. Profilers handle these with runtime-specific unwinders that read the interpreter's own frame structures, or with symbol maps written by the runtime, such as the /tmp/perf-<pid>.map convention that JITs can emit for perf-compatible tools. If your flame graph shows a wide PyEval_EvalFrame or a block of hex addresses, this layer is missing.

Symbolization and the build ID

Production binaries are usually stripped, and shipping full debug information to every host is wasteful. So agents send addresses plus a mapping for each executable: its load address, its file offset and its build ID, a hash the linker embeds that uniquely identifies the build. The server looks up debug information by build ID, from your own upload step in CI or from a debuginfod server, and resolves addresses to function, file and line once per unique address rather than once per sample.

Build IDs make this robust: two hosts running the same release share symbols, and a hotfix gets a new ID, so names never come from the wrong build. Make uploading debug information a CI step for every release artifact. A profile that is 40 percent [unknown] is usually a missing upload, not a profiler bug.

The data model

Most tools speak, or can convert to, pprof's protobuf format. The design is normalized to keep repeated data small: strings live once in a string table, functions and locations are referenced by ID, and a sample is just a list of location IDs plus one value per sample type. The OpenTelemetry Profiles signal, which entered public alpha in 2026, defines an OTLP representation designed to round-trip with pprof and to carry trace and span links. A simplified view, followed by a converter to the folded-stack text format that flame graph tools consume:

# Simplified pprof structure (field names as in profile.proto)
# sample_type: [{type: "samples", unit: "count"}, {type: "cpu", unit: "nanoseconds"}]
# sample:      [{location_id: [12, 7, 3], value: [42, 420000000], label: [...]}]
#              location_id is leaf first: 12 is executing, 3 is the root caller
# location:    {id: 12, address: 0x4a3f10, line: [{function_id: 5, line: 88}]}
# function:    {id: 5, name: <string index>, filename: <string index>}
# string_table: ["", "samples", "count", "main.encode", ...]

def folded(profile):
    '''Turn a decoded pprof profile into 'root;...;leaf count' lines.'''
    names = {}
    for loc in profile.location:
        fn = profile.function[loc.line[0].function_id - 1] if loc.line else None
        names[loc.id] = profile.string_table[fn.name] if fn else hex(loc.address)
    out = {}
    for s in profile.sample:
        stack = ";".join(names[i] for i in reversed(s.location_id))
        out[stack] = out.get(stack, 0) + s.value[0]
    return out

The lookup above assumes dense function IDs starting at 1; real code should map ID to function explicitly. The folded format is the lingua franca: you can grep it, diff it and render it.

Storage, labels and cardinality

Stacks repeat enormously, so storage separates identity from values: unique stacks are stored once with an ID, and each upload becomes rows of time, label set, stack ID and count. The set of distinct stacks grows slowly, so a week compresses well.

Labels are where the cost hides. Service, version, region, pod and CPU architecture are useful, bounded dimensions. Request IDs, user IDs or URLs with embedded identifiers are not: each distinct value multiplies the rows that must be stored and merged. The reasoning is identical to the cardinality discipline in metrics architecture. If you need per-request attribution, link profiles to traces instead, as described below, rather than turning every request into a label value.

Reading flame graphs, and diffing them

In a flame graph each box is a function on a stack, its children sit above it, and its width is proportional to the samples in which it appeared. The x-axis is sorted alphabetically, not by time, so adjacent boxes say nothing about order. Read it top-down for self time, the plateaus where nothing sits above a function, and bottom-up for inclusive time, the wide towers.

Single profiles answer "where does time go?" Most production questions are "what changed?", and for those you diff. Normalize each profile to fractions of its own total, then compare per function, because absolute counts differ whenever traffic or replica counts differ.

def self_share(folded_profile):
    '''Fraction of all samples where each function is the leaf (self time).'''
    total = sum(folded_profile.values())
    share = {}
    for stack, n in folded_profile.items():
        leaf = stack.rsplit(";", 1)[-1]
        share[leaf] = share.get(leaf, 0) + n / total
    return share

def diff(before, after, min_delta=0.005):
    a, b = self_share(before), self_share(after)
    rows = [(f, a.get(f, 0), b.get(f, 0)) for f in set(a) | set(b)]
    rows = [r for r in rows if abs(r[2] - r[1]) >= min_delta]
    return sorted(rows, key=lambda r: r[2] - r[1], reverse=True)

Worked example. After release 2.4 of a checkout service, fleet CPU rose 18 percent at unchanged request rate. Diffing an hour of version 2.3 against an hour of 2.4, selected by the version label, showed JSON encoding moving from 4 to 15 percent of self time, almost all under a new reflection-based path in an audit-logging middleware that serialized the full request for every call. Logging a precomputed summary returned CPU to baseline, and the next diff confirmed it. Because both profiles already existed, nobody had to reproduce the load.

Linking profiles to traces

A trace says a span took 900 ms; a profile says what code ran. Joining them requires tagging samples with the span that was active when they were taken. In Go, pprof labels attach key-value pairs to a goroutine's CPU samples for the duration of a function:

import (
    "context"
    "runtime/pprof"
)

func handle(ctx context.Context, route string) {
    // Bounded label: safe to aggregate on.
    pprof.Do(ctx, pprof.Labels("route", route), func(ctx context.Context) {
        process(ctx)
    })
}

Tagging every sample with a span ID is exactly the unbounded label warned about above, so systems that offer span profiles store the span link alongside the sample without making it a queryable dimension, and fetch by span on demand. The OpenTelemetry Profiles design builds this link into the format, which matters if you already run distributed tracing: a slow span can open directly on the code that ran during it. Because the signal is still alpha, check your collector and backend versions before depending on it.

Overhead and operational guidance

Budget about one percent of CPU for profiling, and measure it: run the agent on half of a canary fleet and compare CPU per request. Sampling cost scales with rate times stack depth, so very deep stacks, common in reactive frameworks and recursive code, cost more than the headline rate suggests. Memory profiles work differently: Go's heap profiler records roughly one allocation per 512 KiB allocated by default, so it reports where memory was allocated, not a precise live-object census.

  • Profile types. CPU profiles show on-CPU time only. Latency spent waiting on locks, I/O or the scheduler needs off-CPU or wall-clock profiles; a service that is slow but idle looks innocent in a CPU flame graph.
  • CPU limits. A container throttled by its CFS quota shows the code it ran, not the time it spent throttled. Pair profiles with throttling metrics.
  • Retention. Keep full resolution for days and merged, downsampled profiles for weeks; regressions are usually found by comparing against last week's release.

Failure modes

SymptomLikely causeFix
Many stacks only one or two frames deepFrame pointers omitted, no unwind tables loadedEnable frame pointers or DWARF unwinding in the agent
Large [unknown] or hex blocksDebug info not uploaded for this build IDUpload symbols in CI for every artifact
Wide interpreter loop, no app functionsNo runtime-specific unwinderEnable language support or use an in-runtime profiler
Profiles for some pods missingAgent lacks permissions or cannot see the container's mount namespaceCheck agent privileges and process discovery
Diff shows everything movedComparing absolute counts across different trafficNormalize to shares before diffing
Storage cost grows linearly with trafficUnbounded label such as request IDRemove the label; use span links

Trade-offs

ChoiceGainsCosts
eBPF whole-system agentEvery process, no code changeUnwinding complexity, privileges, weaker runtime detail
In-runtime profilerAccurate language frames, allocation and lock profilesPer-language integration, one runtime at a time
Higher sampling rateResolution for short incidents and processesProportional CPU and storage
More labelsFiner slicingStorage and query cost multiply

What to do next

  1. Pick one service and run a profiler on a canary; measure the CPU overhead per request.
  2. Check stack quality: look for one-frame stacks, [unknown] blocks and bare interpreter loops, and fix unwinding or symbol upload first.
  3. Add a CI step that uploads debug information keyed by build ID for every release.
  4. Define a small, bounded label set: service, version, region, pod.
  5. Make a version-to-version diff part of every release review.
  6. Add off-CPU or wall-clock profiles for latency-bound services.
  7. If you trace, evaluate span-to-profile links, noting that the OpenTelemetry Profiles signal is still alpha.
Key takeaway: Continuous profiling is a pipeline: sample at a low, odd rate, unwind cheaply, ship addresses, symbolize centrally by build ID, store unique stacks once, and merge by labels at query time. The statistics are sound because samples aggregate over time and fleet, so low rates are fine. Most bad profiles are bad unwinding or missing symbols, and most regressions are found by diffing normalized profiles between versions. Keep labels bounded and link to traces for per-request detail.