Most performance work starts with a question that dashboards cannot answer: where is the CPU going, or why is this request slow when nothing is saturated? Traces show which service and span took the time; metrics show that something changed. A profiler shows which functions in your code consumed the CPU, memory or waiting time. Google Cloud Profiler is a continuous statistical profiler: a small agent in each process periodically captures short profiles in production and uploads them, so you can look at last Tuesday's behaviour without having reproduced it.
This article explains how the agent and backend cooperate, which profile types each language supports, the documented overhead and retention, setup for Python, Go, Java and Node.js, how to read flame graphs and compare versions, a worked regression hunt, exporting profiles through the API, and the failure modes that leave teams staring at an empty UI.
How the agent and backend cooperate
Each process runs an agent library. The agent registers with the Cloud Profiler API, identifying its deployment: project, service name, service version and zone. Google's documentation describes the schedule as follows: each minute, on average, for each deployment and each profile type, the backend selects an agent to capture a profile, and a single profile usually represents data collected for 10 seconds. Heap and threads profiles are instantaneous snapshots rather than 10-second windows.
Two consequences follow. First, a fleet of 50 replicas does not run 50 profilers at once; roughly one replica per profile type works at a time, which is how overhead stays low. Second, you need volume: a single minute's profile is a sample of one replica for 10 seconds, so the UI aggregates many profiles over the selected time range, and conclusions from a handful of profiles are noisy.
Profile types, overhead and retention
Profile types differ by language, taken from Google's per-language setup pages:
| Language | Profile types | Notes |
|---|---|---|
| Go | CPU time, heap, allocated heap, contention (mutex), threads (goroutines) | Mutex profiling is opt-in via MutexProfiling: true |
| Java | CPU time, wall time, heap | Heap requires Java 11 or newer and is opt-in via -cprof_enable_heap_sampling=true |
| Node.js | Heap, wall time | Native module built with node-gyp; Node.js 14 or newer |
| Python | CPU time, wall time | Wall time covers the main thread only |
CPU time is time the CPU spent executing code. Wall time is elapsed time including waiting on locks, I/O and the network, so a function that blocks on a database call is large in wall time and small in CPU time. Heap is memory live at the moment of the snapshot; allocated heap is everything allocated during the window, including garbage, and is the right view for garbage-collection pressure. Contention is time spent waiting for locks; threads counts goroutines by stack, which exposes leaks.
Overhead, per the documentation: CPU and heap-allocation profiling costs less than 5 percent while data is being collected, and amortised over execution time and across replicas it is commonly less than 0.5 percent. Profile data is retained for 30 days.
Setup in four languages
Enable the API and grant the runtime's service account the Cloud Profiler Agent role:
gcloud services enable cloudprofiler.googleapis.com
gcloud projects add-iam-policy-binding $PROJECT \
--member="serviceAccount:api-runtime@$PROJECT.iam.gserviceaccount.com" \
--role="roles/cloudprofiler.agent"On GKE with Workload Identity, bind the role to the Google service account your Kubernetes service account maps to. Outside Google Cloud, supply a project ID and credentials explicitly. Then start the agent as early as possible in the process.
# Python: pip install google-cloud-profiler
import logging, os
import googlecloudprofiler
log = logging.getLogger(__name__)
def start_profiler():
try:
googlecloudprofiler.start(
service="orders-api",
service_version=os.environ.get("APP_VERSION", "dev"),
verbose=1, # 0-3; use 3 while debugging setup
)
except (ValueError, NotImplementedError) as exc:
log.warning("profiler not started: %s", exc) # never crash the app over profiling// Go: import "cloud.google.com/go/profiler"
cfg := profiler.Config{
Service: "orders-api",
ServiceVersion: version,
MutexProfiling: true,
}
if err := profiler.Start(cfg); err != nil {
log.Printf("profiler not started: %v", err)
}# Java: unpack the agent from
# https://storage.googleapis.com/cloud-profiler/java/latest/profiler_java_agent.tar.gz
java -agentpath:/opt/cprof/profiler_java_agent.so=-cprof_service=orders-api,-cprof_service_version=1.5.0 \
-jar orders-api.jar
// Node.js: npm install @google-cloud/profiler (first line of the entry point)
require('@google-cloud/profiler').start({serviceContext: {service: 'orders-api', version: '1.5.0'}});Set the service version from your build, never a constant. Version is the axis you will compare on when a release regresses, and a profile labelled "latest" for every build cannot answer the question "what changed".
Forking servers and short-lived processes
Pre-forking servers need care. The Python documentation notes that for uWSGI you should enable lazy-apps so the profiler initialises in each worker, and for wall profiling also py-call-osafterfork; it also requires that start be called from the main thread for wall profiling, and advises against wall profiling for applications with a large number of threads. The general rule for any forking server, including gunicorn, is to start the agent after the fork in each worker, not in the master, because a profiler thread started before fork() does not survive into the child. Verify by checking that profiles appear and that they contain your request handlers, not just the master's idle loop.
Serverless platforms that throttle CPU between requests can starve the agent's background work, and short-lived processes may exit before they are ever selected. If a service runs for less than a few minutes at a time, continuous profiling is the wrong tool; profile it with a local profiler under a realistic load instead. The Cloud Run deep dive explains the CPU allocation models that determine whether background threads get CPU.
Reading flame graphs and comparing versions
The UI renders aggregated profiles as a flame graph: each bar is a function, its width is the share of the metric attributed to that function and everything it called, and children sit below their callers. Wide bars near the top are where to look; a wide bar whose children are narrow has substantial self time, meaning the function's own code is expensive.
The tools that matter most are filters and comparison. Focus on a function to merge all its call paths into one view, which is essential when a helper is called from dozens of places. Hide framework frames to see your code. Compare two sets of profiles, for example version 1.4.2 against 1.5.0, and the graph colours frames by how much they grew or shrank. Read CPU and wall together: when wall time is far larger than CPU time for a handler, the problem is waiting, and you should look at contention or I/O rather than optimising loops.
Worked example: a CPU regression after a release
Scenario: a Go service, orders-api, runs 20 replicas. After release 1.5.0 the CPU request per pod is breached and the autoscaler adds 6 replicas; latency is fine, cost is not. Cloud Monitoring confirms CPU per request rose about 30%, and Cloud Trace shows no single slow span. Time for the profiler.
Step 1: CPU time, compare 1.5.0 against 1.4.2 over the last 24 hours. The comparison shows regexp.Compile grown from almost nothing to about 18% of CPU, called from a new validateSKU function. Someone wrote regexp.MustCompile(pattern) inside the function body, so the pattern is recompiled on every request.
Step 2: allocated heap, same comparison. The regexp compilation sites now dominate allocations. Back in the CPU view, the runtime's garbage-collection frames have grown too; that explains the rest of the increase, since compiling creates short-lived garbage the collector must clean up.
Step 3: the fix moves compilation to a package-level variable. After 1.5.1 ships, the same comparison shows regexp.Compile gone from the hot path and CPU per request back to the 1.4.2 baseline; the autoscaler returns to 20 replicas. Total investigation: under an hour, with no reproduction environment. That is the case for continuous profiling in one example.
Second example: a slow memory leak
CPU regressions are the easy case. Memory problems show up as pods being OOM-killed every few hours, and by the time anyone looks, the evidence has been restarted away. Continuous heap and threads profiles keep it.
Scenario: the same Go service's memory climbs steadily after each restart until the container limit kills it about every six hours. Open the heap profile type and narrow the time range to the last hour before a kill, then to the first hour after a restart. Because heap profiles are snapshots of live memory, comparing those two windows shows which allocation sites hold memory that never gets released. Here the growth sits under a cache keyed by customer ID that has no eviction.
Now open threads for the same windows. If the goroutine count also grows and the widest stacks sit in a channel receive or a network read, the leak is goroutines stuck waiting, each pinning its own buffers; that is a different fix (timeouts and cancellation through context.Context) from bounding a cache. Allocated heap, by contrast, would mostly show churn rather than retention, so it is the wrong view for a leak and the right one for garbage-collection CPU.
Two habits make this work. Keep memory limits high enough that a pod lives long enough to be sampled several times before it dies, and record the restart time so you know which window to compare. A leak that kills a pod in two minutes leaves only a profile or two behind, which is too little to trust.
Exporting profiles through the API
For offline analysis or long-term archiving beyond the 30-day retention, the v2 API exposes projects.profiles.list (GET https://cloudprofiler.googleapis.com/v2/projects/PROJECT/profiles), which requires the cloudprofiler.profiles.list permission and pages up to 1,000 profiles at a time. The returned profiles carry their bytes in pprof format, which go tool pprof and other pprof-compatible tools can read.
import base64, pathlib
from googleapiclient.discovery import build # pip install google-api-python-client
svc = build("cloudprofiler", "v2")
req = svc.projects().profiles().list(parent="projects/my-project", pageSize=1000)
out = pathlib.Path("profiles"); out.mkdir(exist_ok=True)
while req is not None:
resp = req.execute()
for prof in resp.get("profiles", []):
dep = prof["deployment"]
name = f'{dep["target"]}_{prof["profileType"]}_{prof["name"].rsplit("/", 1)[-1]}.pb.gz'
(out / name).write_bytes(base64.b64decode(prof["profileBytes"]))
req = svc.projects().profiles().list_next(req, resp)
# then: go tool pprof -top profiles/<file>.pb.gzCheck the field names against the API reference for your client version before relying on this in a pipeline, and filter by deployment labels so you export only the services you need.
Failure modes
- No profiles at all. API not enabled, missing
roles/cloudprofiler.agent, or no network path to the API. Run withverbose=3(Python) orDebugLogging: true(Go) and read the agent's log lines. Also check your runtime version against the supported list on the language's setup page. - Profiles of the wrong process. Forking servers profiling the master only; start the agent post-fork.
- Versions you cannot compare. Service version hard-coded or set to a mutable tag.
- Too few profiles. Low-traffic or short-lived services produce sparse data; widen the time range or accept that conclusions are rough.
- Blaming the profiler. Overhead concentrates on one replica at a time, so a latency blip on one pod every minute or so may be the agent; check it against the documented overhead before disabling profiling for everyone.
Trade-offs and alternatives
Compared with on-demand profiling (pprof endpoints, py-spy, async-profiler run by hand), Cloud Profiler's advantage is that the data already exists when the question arrives, and that versions can be compared without reproduction. Its limits are the fixed sampling cadence, 30-day retention, a language-specific set of profile types, and the need to instrument each process. Open-source continuous profilers and eBPF-based profilers can cover more languages without code changes and keep data as long as you pay to store it, at the cost of running the backend yourself. Profiling also complements rather than replaces Cloud Monitoring and Cloud Logging: metrics tell you when, traces tell you which service, profiles tell you which line.
What to do next
- Enable the Cloud Profiler API and grant
roles/cloudprofiler.agentto one service's runtime account. - Add the agent to that service with the build version as
service_version, wrapped so failures only log. - For forking servers, start the agent post-fork and confirm handler frames appear in profiles.
- After a day, open CPU and wall views side by side and note the three widest frames you own.
- On the next release, compare the new version against the previous one before closing the deploy ticket.
- If you need history beyond 30 days, schedule a
profiles.listexport to Cloud Storage.