A metric tells you a service is using 40 percent more CPU than yesterday. A trace tells you a request spent 300 ms inside one handler. Neither tells you which line of code is responsible. A CPU profile does: it is a statistical answer to where the time went, down to the function. Continuous profiling keeps that answer available all the time, for every instance, so you can ask the question after the fact instead of trying to reproduce the problem with a profiler attached.

How profilers sample, unwind stacks, symbolize and store data is covered in continuous profiling architecture. This page is the operational companion: how to choose a collection model, set it up per runtime, prove the overhead is acceptable, get the permissions right in containers, and turn profiles into a routine part of incident response.

Advertisement

What continuous profiling adds

Ad hoc profiling has three problems in production. The problem is often gone by the time someone attaches a profiler. The profile covers one instance, which may not be representative. And there is no baseline, so you cannot tell whether a hot function was always hot. Continuous profiling fixes all three by sampling every instance at a low rate all the time and storing the results with labels such as service, version and region.

That enables three questions you cannot answer otherwise: what changed between yesterday and today (a diff), what the whole fleet spends its CPU on (an aggregate, which drives cost work), and what one slow request was doing (when profiles are linked to traces). The first is the one that pays for the system during incidents; the second often pays for it in compute savings.

Choosing a collection model

There are two ways to get profiles out of a process, and many teams end up running both.

ModelHow it worksStrengthsWeaknesses
Language SDK (push)A library in the process runs the runtime's own profiler and uploads periodicallyRuntime-aware: heap, allocations, locks, goroutines; works without host privilegesCode change per service; one integration per language
Pull from endpointsA scraper fetches profiles from an HTTP endpoint such as Go's pprof handlersNo upload code; same model as metrics scrapingEndpoints must be reachable and protected
eBPF host agentOne privileged agent per node samples all processes from the kernelZero code change; covers native code and the kernelCPU-focused; needs privileges; interpreted runtimes need special unwinding support

A pragmatic default: deploy an eBPF agent for fleet-wide CPU coverage, and add a language SDK to the services where memory and contention matter. On standards: OpenTelemetry's profiles signal entered public alpha in March 2026, with pipeline support in Collector v0.148.0 and later and the donated eBPF profiler running as a collector receiver. The SIG advises against relying on it for critical production workloads yet, so treat it as something to evaluate rather than the foundation of a rollout today.

Advertisement

Per-runtime setup

Go has profiling built in. The CPU profiler samples at 100 Hz by default and the heap profiler records an allocation sample on average every 512 KiB. Mutex and block profiles are off until you enable them. A service that exposes pprof for pull, and optionally pushes with the Pyroscope SDK, looks like this:

import (
    "net/http"
    _ "net/http/pprof"           // registers /debug/pprof/* on DefaultServeMux
    "runtime"

    "github.com/grafana/pyroscope-go"
)

func startProfiling() {
    runtime.SetMutexProfileFraction(5) // record about 1 in 5 contention events
    runtime.SetBlockProfileRate(10000) // about one sample per 10us spent blocked

    // Pull model: serve pprof on a private port, never on the public listener.
    go http.ListenAndServe("127.0.0.1:6060", nil)

    // Push model: the SDK profiles in-process and uploads periodically.
    pyroscope.Start(pyroscope.Config{
        ApplicationName: "shop.checkout",
        ServerAddress:   "http://pyroscope:4040",
        ProfileTypes: []pyroscope.ProfileType{
            pyroscope.ProfileCPU,
            pyroscope.ProfileAllocSpace,
            pyroscope.ProfileInuseSpace,
        },
    })
}

Check the SDK's README for the current tag and profile-type options rather than copying field lists from older posts; they have changed over time.

The JVM has two production-grade options. Java Flight Recorder ships with the JDK and can run continuously with its default settings profile. async-profiler gives CPU, wall-clock, allocation and lock profiles without safepoint bias, for example asprof -e cpu -d 30 -f cpu.html <pid> for a one-off capture or as a Java agent for continuous use. Its wall mode is the one to reach for when threads are slow but not busy.

Python processes are usually profiled from outside with a sampling profiler such as py-spy, which reads interpreter state from another process and so needs ptrace permission. Native code (C, C++, Rust) is where eBPF agents shine, provided binaries keep frame pointers or unwind tables and symbols are available somewhere the backend can find them.

Overhead: set a budget and measure it

CPU sampling at around 100 Hz is designed to be cheap, and it usually is. Do not take that on trust: allocation and lock profiling at aggressive rates can cost far more, uploads compete for network and CPU on busy hosts, and the cost depends on your workload, runtime version and how deep your stacks are. A service with very deep stacks or millions of small allocations per second is exactly where a default setting surprises you. Set a budget, for example at most 2 percent CPU and no measurable change in p99 latency, and measure it with a canary:

  1. Pick two groups of instances of the same service on the same hardware and traffic share.
  2. Enable profiling on one group only, with the exact settings you intend to ship.
  3. Compare CPU per request, p50 and p99 latency, and memory over at least one daily traffic cycle.
  4. If the difference exceeds the budget, lower sampling rates or drop profile types, then repeat.

Ship a kill switch with the rollout: a config flag or environment variable that disables the profiler without a redeploy. You will want it the first time a runtime upgrade interacts badly with an agent.

Permissions in containers

Kernel-level sampling uses perf_event_open, which Docker's default seccomp profile restricts. The async-profiler documentation lists the options: relax or disable the seccomp profile (and possibly add SYS_ADMIN), use its --fdtransfer mode, or fall back to -e ctimer, which uses timers instead of perf events. On hosts, the kernel.perf_event_paranoid sysctl controls what unprivileged processes may sample.

eBPF agents run as a privileged DaemonSet. Treat that as a security decision, not a deployment detail: the agent can read memory and stacks of every process on the node. Run it in its own namespace, pin its image by digest, restrict who can change its configuration, and review it like any other privileged component. For language SDKs, the main risk is the pull endpoint: pprof handlers expose stack traces, command lines and, through goroutine dumps, sometimes values. Bind them to localhost or a private port and never to the public listener.

The rollout

Rolling continuous profiling into production1. Inventoryruntimes, kernels2. Stagingmeasure overhead3. CanaryA/B vs control4. Fleetalways onCollectionSDK push or eBPF agentProfile storelabels, retentionQuery and UIflame graph, diffGuardrailsoverhead budget, kill switchAccesssymbols, secretsIncident usedeploy diff, trace linkThe pipeline is simple; the rollout, guardrails and habits around it decide whether it is used.
Inventory and measure before going fleet-wide, and keep guardrails and access controls in place from the first canary.

Start with an inventory: which runtimes and versions you run, which kernels, which binaries are stripped. Stripped binaries and missing debug info are the most common reason a first flame graph is a wall of hex addresses. Upload symbols or debug info from CI for every build, keyed by build id, so the backend can symbolize later without shipping symbols to production hosts.

Then label consistently. Profiles are only useful in comparison, and comparison needs the same labels as your metrics and traces: service.name, service.version, environment, region, and instance. Avoid high-cardinality labels such as user ids for the same reason you avoid them in metrics; they fragment storage and make aggregation slow. Choose retention to match the questions: full resolution for a week or two covers incident diffs, and aggregated fleet views for a quarter cover cost work.

Worked example: a CPU regression after deploy

At 10:20 a dashboard shows checkout CPU per request up 35 percent since version 2026.10.02-b. Latency has crept up but no alert has fired. With continuous profiling the investigation is a query, not a reproduction:

  1. Select service shop.checkout, profile type CPU, and compare version 2026.10.02-a (baseline) with 2026.10.02-b over the same hour of traffic.
  2. Open the diff flame graph. Most frames are unchanged. One stack has grown from about 3 percent to about 24 percent of samples: renderReceipt -> formatCurrency -> regexp.MustCompile.
  3. Read the diff for the release: a refactor moved a regular expression compile from package initialisation into the function body, so it now compiles on every call.
  4. Move the compile back to a package-level variable, ship, and confirm in the next diff that the stack is back near its baseline share and CPU per request has recovered.

Two details made this fast. Comparing by version rather than by time window removed noise from traffic mix changes. And because the comparison was normalised to share of samples, a busier hour did not look like a regression.

Without profiles, the same incident usually goes differently: someone bisects the release on a staging host, tries to reproduce production traffic with a load generator, and attaches a profiler to a process that may not exhibit the problem at the load they managed to generate. That can take a day. The profile data that answers the question in minutes was produced by every production instance while the regression was happening; continuous profiling simply keeps it.

Memory and contention incidents

Memory growth. Use the in-use heap profile, not the allocation profile. Allocation profiles show churn, which drives garbage collection cost; in-use profiles show what is still referenced, which is what a leak looks like. Compare in-use heap at hour 1 and hour 12 after a restart: a stack whose retained bytes grow steadily is your suspect, frequently a cache without eviction or a map keyed by request id.

Slow but idle. When latency rises but CPU does not, a CPU profile will look normal because nothing is running. Switch to wall-clock (async-profiler wall), off-CPU, mutex or block profiles. A pool of 20 database connections shared by 200 goroutines shows up clearly in a block profile as time waiting in the pool's acquire path.

One slow request. If your profiler and tracer support span-linked profiles, a slow span can open the samples taken during it. Without that, filter profiles by instance and a narrow time window around the trace.

Failure modes

  • Unsymbolized frames. Stripped binaries or missing debug info produce addresses instead of names. Upload symbols from CI by build id.
  • Broken native stacks. Code compiled without frame pointers and without unwind data yields truncated stacks in eBPF profiles. Check a known binary end to end before trusting fleet views.
  • Wrong profile type. CPU profiles of an I/O-bound service show little; allocation profiles mistaken for in-use profiles send leak hunts in the wrong direction.
  • Exposed endpoints. A pprof handler on the public port leaks internals and lets anyone trigger expensive profiles.
  • Label sprawl. Per-request labels make queries slow and storage expensive.
  • Nobody looks. Profiles collected but never used in an incident or a cost review are pure overhead. Put a diff step in your incident runbook.

Trade-offs

DecisionOption AOption B
CoverageeBPF agent: every process, no code changeSDK: richer types, per-service work
Sampling rateHigher: finer detail, more overheadLower: cheaper, needs longer windows
RetentionLong: quarter-scale cost analysisShort: cheaper, incident diffs only
BackendSelf-hosted: data stays in houseHosted: no operations, per-volume pricing
StandardOTel profiles: future-proof formatAlpha today; vendor formats are mature

What to do next

  1. Inventory runtimes, kernel versions and stripped binaries across your services.
  2. Pick a collection model per runtime: eBPF agent for CPU coverage, SDKs where memory and locks matter.
  3. Set an overhead budget and measure it with a profiled and an unprofiled canary group over a full day.
  4. Ship a kill switch that disables profiling without a redeploy.
  5. Upload symbols from CI keyed by build id.
  6. Use the same service, version and environment labels as your metrics and traces.
  7. Lock down pprof and similar endpoints to private ports and treat eBPF agents as privileged components.
  8. Add a version-to-version diff flame graph step to your incident runbook.
  9. Review fleet-wide CPU profiles quarterly to find the top cost hotspots.

Keep learning: profiling architecture, eBPF for observability, distributed traces and async-profiler.

Key takeaway: Continuous profiling answers which code is responsible, for every instance, after the fact. Use an eBPF agent for broad CPU coverage and language SDKs where memory and contention matter, prove overhead against a budget with a canary, keep a kill switch, upload symbols from CI, and label profiles like your metrics. Treat agents and pprof endpoints as privileged, and make a version-to-version diff a standard incident step.