Cloud Trace is Google Cloud's managed distributed-tracing backend. It stores spans, the timed records of individual operations, groups them into traces by a shared trace ID, and lets you see where the time in a request went across services, databases and queues. It is the natural choice when your services run on Cloud Run, GKE or Compute Engine and you already use Cloud Logging and Cloud Monitoring, because traces, logs and metrics can be tied together by the same identifiers.

This article covers Cloud Trace as an engineer wiring it into a real system meets it: the ways spans get in, what Cloud Run does automatically, how trace context crosses service boundaries, how to make every log line clickable through to its trace, the limits that shape instrumentation, and how to control cost with sampling. General tracing concepts are covered in distributed tracing architecture; the specifics here were checked against Google Cloud documentation in September 2026.

Advertisement

The model: traces, spans and where they are stored

A trace is identified by a 128-bit trace ID, written as 32 hexadecimal characters. Each span has a 64-bit span ID, a parent span ID, a name, start and end timestamps, a kind (server, client, producer, consumer or internal), a status and attributes. Cloud Trace accepts OpenTelemetry's data model directly, so the attribute names from the OpenTelemetry semantic conventions, such as http.route and db.system, are what you should use.

Spans are stored in a _Trace observability bucket and retained for 30 days. The ingestion window is bounded too: the Trace API rejects spans with timestamps more than 14 days in the past or more than 3 days in the future, which matters if you buffer spans on devices or replay them from files. In the console, Trace Explorer shows a latency distribution for the spans you filter to and opens individual traces as a waterfall, with linked log entries alongside.

Cloud Trace on Google Cloud: ingestion paths, storage and correlationCloud Run serviceautomatic request spansGKE / VM serviceOpenTelemetry SDKBatch / workerOpenTelemetry SDKOpenTelemetry Collectorbatch, sample, enrichOTLPOTLPtelemetry.googleapis.comOTLP, traces writer roleOTLP directexportCloud Trace storage_Trace bucket, 30 daysTrace Explorerlatency distribution, waterfallCloud Loggingtrace + spanId fieldslinkedThe same trace ID must reach every hop (traceparent) and every log line (logging.googleapis.com/trace).
Spans reach Cloud Trace directly over OTLP or through an OpenTelemetry Collector. Cloud Run adds request spans automatically. Logs carry the trace ID so Trace Explorer and Logs Explorer can link to each other.

Three ways to get spans in

OTLP to the Telemetry API. Google Cloud exposes an OpenTelemetry Protocol endpoint at telemetry.googleapis.com, accepting OTLP over gRPC or HTTP with protobuf. The caller needs the Cloud Telemetry Traces Writer role, roles/telemetry.tracesWriter. Because the data arrives in OpenTelemetry's own format, no translation step is needed, which Google notes can otherwise lose data. Applications set OTEL_EXPORTER_OTLP_ENDPOINT to that host and supply Google credentials; Java has an OpenTelemetry GCP authentication extension for this, and Python and Go use the Google auth libraries.

Through an OpenTelemetry Collector. Applications send OTLP to a local or cluster Collector, which batches, samples, enriches and exports. This is the better choice on GKE and for fleets, because authentication, sampling policy and retry behaviour live in one place rather than in every service. The contrib Collector has a googlecloud exporter, and Google also publishes a Google-built OpenTelemetry Collector distribution. Collector topologies are covered in OpenTelemetry Collector architecture.

The Cloud Trace API. The older API, with its own span format, still works and some legacy exporters use it. New instrumentation should use OTLP. The limits differ slightly between the two paths, so check which one your exporter uses when a span is rejected.

# otel-collector config: receive OTLP from pods, sample, export to Cloud Trace
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318
processors:
  memory_limiter:
    check_interval: 1s
    limit_percentage: 80
  resourcedetection:
    detectors: [gcp]          # adds project, region, cluster, instance attributes
  batch: {}
exporters:
  googlecloud: {}             # uses the pod's workload identity credentials
service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, resourcedetection, batch]
      exporters: [googlecloud]
Advertisement

What Cloud Run does for you, and what it does not

Cloud Run generates request traces automatically for incoming requests, with no code. Google documents the sampling rate as at most 0.1 requests per second per instance, that is roughly one request every ten seconds on each container instance, and the rate is not configurable. These automatic traces are not billed; spans from your own instrumentation are billed at normal Cloud Trace rates. Cloud Run also populates the standard W3C traceparent header on incoming requests, so an instrumented service can continue the trace Cloud Run started.

The automatic span tells you how long the request took at the platform edge, not why. To see database calls, outbound HTTP requests and queue publishes, instrument the service with OpenTelemetry. Two Cloud Run details bite here. If CPU is only allocated during request processing, a background exporter may not get CPU after the response is sent, so spans can be delayed or lost; flush the tracer provider before the response completes for short-lived work, or keep CPU always allocated (instance-based billing). And on shutdown, call the provider's shutdown in the SIGTERM handler so buffered spans are exported. See Cloud Run architecture for how instances and CPU allocation work.

Instrumenting a Python service

The example below configures the OpenTelemetry SDK with a parent-based sampler, so a service respects the sampling decision made by its caller and samples 5 percent of traces it starts itself, and exports OTLP to a Collector sidecar or agent. Framework instrumentation creates server spans for incoming requests; the manual span adds business context.

from opentelemetry import trace
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.sdk.trace.sampling import ParentBased, TraceIdRatioBased
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.instrumentation.flask import FlaskInstrumentor
from flask import Flask, request

app = Flask(__name__)

provider = TracerProvider(
    resource=Resource.create({"service.name": "checkout", "service.version": "2.14.0"}),
    sampler=ParentBased(root=TraceIdRatioBased(0.05)),
)
provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter(endpoint="localhost:4317", insecure=True)))
trace.set_tracer_provider(provider)
FlaskInstrumentor().instrument_app(app)
tracer = trace.get_tracer("checkout")

@app.post("/orders")
def create_order():
    with tracer.start_as_current_span("price_cart") as span:
        span.set_attribute("cart.items", len(request.json["items"]))
        total = price(request.json["items"])     # db and HTTP calls become child spans
    return {"total": total}

Propagation: traceparent and X-Cloud-Trace-Context

A trace only spans services if every hop forwards trace context. The standard is the W3C traceparent header, which carries the trace ID, the parent span ID in hexadecimal and a sampled flag. OpenTelemetry uses it by default. Google Cloud also has an older header, X-Cloud-Trace-Context, with the form TRACE_ID/SPAN_ID;o=OPTIONS, where the trace ID is hexadecimal but the span ID is a decimal integer and o=1 means sampled. Some older Google Cloud libraries and services still read or write it.

If a service in the chain only understands one of the two headers, traces break at that hop. Configure propagators to read and write both during a migration; Google publishes an OpenTelemetry propagator for the Cloud Trace header format for the major languages. Asynchronous hops need explicit handling: when publishing to Pub/Sub, inject context into message attributes, and on the consumer create a new trace with a span link to the producer's context rather than making a batch of messages children of one parent. Propagation pitfalls in general are covered in trace context propagation.

Correlating logs with traces

The biggest practical payoff from Cloud Trace is jumping from a slow span to the log lines written during it, and from an error log to the trace that produced it. Cloud Logging links a log entry to a trace through the LogEntry fields trace, spanId and traceSampled. On Cloud Run, GKE and other environments where a logging agent parses JSON written to stdout, you set them with special keys: logging.googleapis.com/trace, whose value should be projects/PROJECT_ID/traces/TRACE_ID, logging.googleapis.com/spanId, and logging.googleapis.com/trace_sampled, a boolean.

import json, sys
from opentelemetry import trace

PROJECT = "my-project"

def log(severity: str, message: str, **fields):
    ctx = trace.get_current_span().get_span_context()
    entry = {"severity": severity, "message": message, **fields}
    if ctx.is_valid:
        entry["logging.googleapis.com/trace"] = f"projects/{PROJECT}/traces/{ctx.trace_id:032x}"
        entry["logging.googleapis.com/spanId"] = f"{ctx.span_id:016x}"
        entry["logging.googleapis.com/trace_sampled"] = ctx.trace_flags.sampled
    sys.stdout.write(json.dumps(entry) + chr(10))

log("ERROR", "payment declined", order_id="o-81237", provider="acme")

Log correlation works even for unsampled requests: every request still has a trace ID, so all log lines for one request can be grouped by it, even when no spans were stored. That makes it cheap to keep tracing sampled low while still debugging individual requests through logs. See Cloud Logging for sinks and retention.

Limits that shape your instrumentation

The OTLP ingestion path documents the following per-span limits. They are generous for sensible instrumentation and easy to hit with careless instrumentation, such as recording each item of a large batch as an event or attaching a full request body as an attribute. There is also a project-level ingestion quota measured in bytes per minute; check the current value for your region on the quotas page rather than assuming.

Limit (OTLP / Telemetry API)ValueWhat it means for instrumentation
Attributes per span1,024put request context on spans, not every variable
Attribute key / value size512 bytes / 64 KiBnever attach whole payloads
Span name length1,024 bytesuse route templates, not raw URLs
Events per span256loops that add an event per item will be truncated
Links per span128fan-in batch spans must summarise, not link every message
Retention in the _Trace bucket30 daysexport or summarise anything you need longer

Sampling and cost

Cloud Trace is billed on spans ingested, beyond a monthly free allotment; check the current pricing page, because the unit prices change. The number of spans per request is the multiplier people underestimate. A request that touches five services, each emitting a server span, a few client spans and a dozen database spans, can produce 60 spans. At 2,000 requests per second, full sampling would be 120,000 spans per second, over 300 billion a month. At 5 percent head sampling it is about 6,000 spans per second.

Head sampling, decided when the root span starts, is cheap but blind: it keeps 5 percent of slow and failed requests along with 5 percent of the healthy ones. Tail sampling in a Collector gateway waits until a trace is complete and keeps all errors, all requests above a latency threshold and a small fraction of the rest, at the cost of routing all spans of a trace to the same Collector instance and holding them in memory for a decision window. A common split is low head sampling in services plus tail sampling for errors and slow requests; see tail-based sampling.

A worked investigation

The checkout service's p99 latency rose from 400 milliseconds to 1.8 seconds after a release. In Trace Explorer, filtering to the POST /orders server spans and selecting the slow tail of the latency distribution shows traces around 1.7 seconds. The waterfall for one of them shows price_cart taking 1.5 seconds, and inside it 48 sequential database spans of about 30 milliseconds each, one per cart item. The span count per trace had also tripled, which the billing report would have shown a day later. Fixing the N+1 query brought p99 back and cut span volume. The linked logs on the same trace showed no errors, which is why error-rate alerts did not fire; the regression was only visible as latency.

Failure modes

  • Broken traces. A proxy, queue or legacy library that drops traceparent splits one request into several unrelated traces. Test propagation across every hop in staging.
  • Missing permission. Without roles/telemetry.tracesWriter, or the equivalent for the exporter you use, exports fail with permission errors that appear only in the exporter's or Collector's own logs. Alert on exporter failure counts.
  • Lost spans at shutdown or after the response. Flush and shut down the tracer provider, especially on Cloud Run.
  • High-cardinality span names. Naming spans with raw URLs or IDs makes the latency distribution useless. Use route templates and put IDs in attributes.
  • Sensitive data in attributes. Spans are readable by anyone with trace viewer access. Scrub tokens, personal data and payloads in the SDK or in a Collector processor.
  • Clock skew. Child spans that appear to start before their parents come from host clocks disagreeing; trust durations within one host more than offsets between hosts.

What to do next

  1. Pick one ingestion path: OTLP to telemetry.googleapis.com for a few services, or an OpenTelemetry Collector for GKE and larger fleets.
  2. Grant roles/telemetry.tracesWriter to the service accounts that export, and alert on exporter errors.
  3. Instrument one critical request path with the OpenTelemetry SDK and a ParentBased sampler, and verify that a single trace spans every hop.
  4. Add the three logging.googleapis.com trace fields to your structured logger and confirm log lines link from Trace Explorer.
  5. Count spans per request and multiply by traffic before choosing a sampling rate; add tail sampling for errors and slow requests.
  6. Flush and shut down tracer providers in Cloud Run and batch workloads.
  7. Review span names and attributes for cardinality and sensitive data, and remember the 30-day retention.
Key takeaway: Cloud Trace stores OpenTelemetry spans for 30 days and links them to Cloud Logging through the trace ID. Send spans over OTLP to telemetry.googleapis.com or through a Collector, propagate traceparent across every hop, and write the logging.googleapis.com trace fields into every log line. Cloud Run's automatic traces are a sparse, free starting point; real diagnosis needs your own instrumentation, a sampling design sized from spans per request, and attention to the documented limits.