OpenTelemetry is a CNCF project that standardises how software produces telemetry: traces, metrics and logs, with profiles in development. It is not a backend. It defines an API that code calls, SDKs that implement it in each language, a wire protocol called OTLP, a vocabulary of attribute names called semantic conventions, and a Collector that receives, processes and forwards the data. Its value is that instrumentation is written once and the destination can change without touching the code.

Most teams meet OpenTelemetry through an auto-instrumentation agent, and things work until they do not: spans vanish under load, metrics carry an unexpected service name, or a library's spans never appear. Each of those is explained by the internal model. This page describes that model from first principles, with Python as the example language. Rolling OpenTelemetry out across a whole stack is covered in OpenTelemetry full stack, and the Collector pipeline in OpenTelemetry pipeline design.

Advertisement

The core idea: API separate from SDK

OpenTelemetry splits every language implementation into two packages. The API contains the interfaces your code and your libraries call: get a tracer, start a span, record a measurement, read the current context. The SDK is the implementation that decides what happens to those calls: sampling, batching, aggregation and export. If no SDK is installed and registered, API calls go to no-op implementations that do almost nothing.

This split is the reason library authors can instrument their code without forcing a telemetry backend, or even a telemetry dependency of any weight, on their users. An HTTP client library depends only on the API. The application owner installs and configures the SDK once at startup, and every library's calls start producing data. It also explains the most common surprise: a library is instrumented, but nothing appears, because the application never registered an SDK provider, or registered it after the library had already obtained its tracer from a no-op global in a language where that matters.

Inside one process: API, SDK and exportersYour code and librariescall the API onlyOpenTelemetry APITracer, Meter, Logger, ContextSDK: TracerProvider / MeterProvider / LoggerProvider + ResourceSamplerkeep or dropSpanProcessorbatch, queueMetricReaderperiodic collectViewsrename, bucketsExporters: OTLP over gRPC (4317) or HTTP (4318)Collectorreceive, process, exportBackendtraces, metrics, logs storeNo SDK installed:API calls are no-opswith negligible cost
Instrumented code talks only to the API. The SDK's providers apply sampling, processing and aggregation, and exporters send OTLP to a Collector or a backend.

Signals and their providers

Each signal has a provider, the SDK object that owns configuration and hands out instruments. A TracerProvider creates tracers, a MeterProvider creates meters, and a LoggerProvider creates loggers. Tracers, meters and loggers are obtained with a name and version identifying the instrumentation scope, normally the library that is instrumented, so the backend can tell spans from your payments module apart from spans emitted by the HTTP framework.

Traces are trees of spans, each with a trace id, span id, parent, name, kind, start and end time, attributes, events, links and a status. Metrics are measurements recorded on instruments: counters, up-down counters, histograms and gauges, synchronous or observed through callbacks; the SDK aggregates them in memory and exports aggregates, not individual measurements. Logs in OpenTelemetry are mostly bridged: the logs API is meant for logging library authors to connect existing frameworks, so application code keeps using its normal logger and the bridge adds trace and span ids. The metrics architecture article covers temporality and instrument choice in depth.

Maturity differs per language and per signal, so check the project's status page for your language rather than assuming. At the time of writing it lists traces and metrics as stable in Python, Java and Go, logs as stable in Java, release candidate in Go and still in development in Python.

Advertisement

The SDK pipeline: samplers, processors, readers, views

Inside the tracing SDK a span passes three stages. The sampler runs when the span starts and decides whether it is recorded and sampled. The default sampler, parentbased_always_on, follows the parent's decision and samples every root, and parentbased_traceidratio samples a fraction of roots while children follow their parent, which keeps traces whole across services. Span processors receive spans on start and end; the BatchSpanProcessor queues ended spans and exports them in batches on a background thread. The specification defaults are a 5,000 ms schedule delay, a 30,000 ms export timeout, a queue of 2,048 spans and batches of up to 512. When the queue is full, new spans are dropped. Finally an exporter serialises the batch and sends it.

In the metrics SDK, instruments record measurements into aggregations, a metric reader collects them, and an exporter sends them. The PeriodicExportingMetricReader collects every 60,000 ms by default. Views let the application owner change what an instrument produces without touching the code that records it: rename a metric, drop attributes to control cardinality, or change histogram bucket boundaries. Views are the right tool when a library's default buckets do not fit your latency range or when one attribute explodes the series count, a problem covered in metric cardinality management.

Limits protect the process. Each span by default keeps at most 128 attributes, 128 events and 128 links, and attribute values have no length limit unless you set one. Long SQL statements or request bodies in attributes are a common way to blow up exporter payloads, so set OTEL_ATTRIBUTE_VALUE_LENGTH_LIMIT if your instrumentation records them.

Resources and semantic conventions

A resource describes the entity producing telemetry: service name, version, environment, host, container, cloud region. It is attached once to the provider and applies to every span, metric and log the process emits. service.name is the one attribute every backend relies on; if it is not set, the SDK reports a name beginning unknown_service, which is how unattributed telemetry appears in shared backends. Resource detectors can fill in host, process, container and cloud attributes automatically.

Semantic conventions are the agreed attribute and metric names for common operations: HTTP, databases, messaging, RPC, cloud and more. They are what lets a backend build a service map or a latency dashboard without knowing which framework produced the data. For example, http.server.request.duration is a stable histogram measured in seconds, with http.request.method and url.scheme required and http.route and http.response.status_code added when available. Instrumentations written against conventions older than v1.20.0 can be switched with OTEL_SEMCONV_STABILITY_OPT_IN to emit the stable names, or both sets during a migration. Use the conventions for your own spans too whenever one exists, and choose a clear prefix for your own domain attributes.

OTLP: the wire protocol

OTLP is the protocol between SDKs, Collectors and backends. It carries all signals in protobuf over gRPC or HTTP, or as JSON over HTTP. The conventional ports are 4317 for gRPC and 4318 for HTTP, and over HTTP each signal has its own path: /v1/traces, /v1/metrics and /v1/logs. The specification recommends http/protobuf as the default protocol but allows SDKs to default to gRPC for backward compatibility, so always set OTEL_EXPORTER_OTLP_PROTOCOL explicitly rather than relying on a language's default. Exporters retry transient failures with exponential back-off and jitter, and the default export timeout is ten seconds.

Because every hop speaks OTLP, the usual deployment sends from the SDK to a local Collector agent, which adds resource attributes, batches and forwards to a gateway or a vendor. That keeps credentials and vendor choice out of application configuration. Context propagation between services, the traceparent header and baggage, is a separate mechanism configured through propagators; the default is tracecontext,baggage.

A minimal Python setup

The code below configures traces and metrics explicitly, which is the clearest way to see each component. It installs a resource, a parent-based ratio sampler, a batch span processor with an OTLP HTTP exporter, and a metric reader with a view that sets bucket boundaries for one histogram.

# pip install opentelemetry-sdk opentelemetry-exporter-otlp-proto-http
from opentelemetry import trace, metrics
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.sdk.trace.sampling import ParentBased, TraceIdRatioBased
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.export import PeriodicExportingMetricReader
from opentelemetry.sdk.metrics.view import View, ExplicitBucketHistogramAggregation
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.exporter.otlp.proto.http.metric_exporter import OTLPMetricExporter

resource = Resource.create({
    "service.name": "checkout",
    "service.version": "2026.10.1",
    "deployment.environment.name": "prod",
})

tp = TracerProvider(resource=resource, sampler=ParentBased(TraceIdRatioBased(0.10)))
tp.add_span_processor(BatchSpanProcessor(OTLPSpanExporter()))   # endpoint from env
trace.set_tracer_provider(tp)

latency_view = View(
    instrument_name="checkout.payment.duration",
    aggregation=ExplicitBucketHistogramAggregation(
        [0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10]),
)
mp = MeterProvider(
    resource=resource,
    metric_readers=[PeriodicExportingMetricReader(OTLPMetricExporter())],
    views=[latency_view],
)
metrics.set_meter_provider(mp)

Instrumented code then uses only the API. It obtains a tracer and a meter with a scope name and version, wraps the operation in a span, records low-cardinality attributes and records measurements. Errors set the span status and record the exception as an event.

import time
from opentelemetry import trace, metrics
from opentelemetry.trace import Status, StatusCode

tracer = trace.get_tracer("checkout.payments", "1.4.0")   # instrumentation scope
meter = metrics.get_meter("checkout.payments", "1.4.0")
pay_duration = meter.create_histogram("checkout.payment.duration", unit="s")
pay_failures = meter.create_counter("checkout.payment.failures")

def charge(order, gateway):
    start = time.monotonic()
    with tracer.start_as_current_span("charge card", record_exception=False) as span:
        span.set_attribute("payment.provider", gateway.name)    # low cardinality
        span.set_attribute("order.item_count", len(order.items))
        try:
            return gateway.charge(order.total, order.card_token)
        except gateway.Declined as exc:
            span.set_status(Status(StatusCode.ERROR, "declined"))
            span.record_exception(exc)
            pay_failures.add(1, {"payment.provider": gateway.name, "reason": "declined"})
            raise
        finally:
            pay_duration.record(time.monotonic() - start, {"payment.provider": gateway.name})

Note that the order id is not an attribute on the metrics. It would be acceptable on the span, which is one record, but on a metric it creates a new time series per order. The same rule decides most attribute choices: identifiers go on spans and logs, categories go on metrics.

Configuration through the environment

Every SDK reads a common set of environment variables, which lets platform teams configure services without code changes and lets the zero-code agents work at all. In Python, the opentelemetry-instrument wrapper installs the SDK from environment settings and applies the instrumentation packages that opentelemetry-bootstrap found for the installed libraries.

export OTEL_SERVICE_NAME=checkout
export OTEL_RESOURCE_ATTRIBUTES=service.version=2026.10.1,deployment.environment.name=prod
export OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf        # set it; SDK defaults differ
export OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-agent:4318
export OTEL_TRACES_SAMPLER=parentbased_traceidratio
export OTEL_TRACES_SAMPLER_ARG=0.10
export OTEL_PROPAGATORS=tracecontext,baggage            # the default, stated anyway
export OTEL_BSP_MAX_QUEUE_SIZE=8192                      # default 2048
export OTEL_METRIC_EXPORT_INTERVAL=30000                 # default 60000 ms

# Python zero-code instrumentation
pip install opentelemetry-distro opentelemetry-exporter-otlp
opentelemetry-bootstrap -a install      # adds instrumentations for installed libraries
opentelemetry-instrument python app.py

Explicit code configuration and environment configuration can conflict. Decide which one owns each setting and stick to it; a service that sets a sampler in code and another one in the environment will behave according to its language's precedence rules, which differ.

Worked example: the missing payment spans

A checkout service traces fine in staging, but in production during sales events the backend holds about a third fewer payment spans than the payment duration histogram counts, and the histogram is the ground truth because metrics are recorded for every request. The team checks the SDK's own diagnostics and finds dropped-span warnings from the batch processor. Sales peaks reach about 3,000 spans per second per process. The processor exports batches of up to 512 spans from one background thread, and each call to the remote vendor endpoint was taking about 400 ms, so export throughput topped out near 1,300 spans per second. The backlog grew by about 1,700 spans a second and the 2,048-span queue was full within about a second of the peak starting.

The fix has three parts. The exporter is pointed at a Collector agent on the same node, so each export is a local call. The queue is raised to 8,192 and the batch interval is left alone, giving headroom for bursts. And the sampler, which had been the default parentbased_always_on because nobody set it, is changed to a 10 percent parent-based ratio, while the payment histogram, which is not sampled, continues to see every request. The dropped-span warnings stop, and span counts match the sampling rate applied to the histogram's count. The lesson generalises: metrics are aggregated in process and are unaffected by trace sampling, so they are the right source for rates and latency, and traces are the right source for examples.

Failure modes

  • No SDK registered. Libraries call the API, the API is a no-op, nothing is exported, and no error is raised.
  • unknown_service. service.name is not set, so telemetry from many services merges under one name.
  • Silent drops. The batch processor queue fills during bursts or slow exports and drops spans; watch the SDK's own logs and the Collector's receiver metrics.
  • Broken traces. A service uses a different propagator, or an async hop does not carry context, and traces split at that boundary. See trace context propagation.
  • Cardinality explosion. User ids, URLs with ids, or raw error messages as metric attributes. Use views to drop them and http.route instead of the raw path.
  • Protocol mismatch. An exporter sends gRPC to the HTTP port or the reverse; set the protocol and endpoint together.
  • Lost data on shutdown. Short-lived processes exit before the batch is exported; call the providers' shutdown or force-flush before exit.

Trade-offs

ChoiceGainCost
Zero-code agentFast coverage of frameworks and clientsLess control; startup cost; harder to debug
Explicit SDK setup in codeEvery component visible and testableBoilerplate per service; must match platform defaults
Export to a local CollectorFast exports, central control of vendors and credentialsOne more component to run and monitor
Head sampling in the SDKCheap, simple, whole tracesRare errors may be dropped; tail sampling needs a Collector tier
Semantic conventions everywherePortable dashboards and service mapsMigration work when conventions change

What to do next

  1. Check that every service sets service.name, service.version and an environment attribute in its resource.
  2. Confirm an SDK provider is registered at startup before libraries are loaded, and that shutdown flushes it.
  3. Set OTEL_EXPORTER_OTLP_PROTOCOL and the endpoint explicitly, and send to a local Collector rather than straight to a vendor.
  4. Choose a sampler deliberately, normally parent-based ratio, and document the rate.
  5. Watch for dropped spans and size the batch processor queue for your burst rate.
  6. Use semantic convention names for HTTP, database and messaging telemetry, and plan the migration if your instrumentations predate the stable HTTP conventions.
  7. Add views to drop high-cardinality metric attributes and set bucket boundaries that fit your latency range.
  8. Check your language's signal status on the OpenTelemetry status page before relying on logs or profiles in production.
Key takeaway: OpenTelemetry separates an API that code and libraries call from an SDK that the application owner configures, so instrumentation is written once and the destination can change. Providers per signal own a resource, samplers, span processors, metric readers, views and exporters, and OTLP carries the results over gRPC on 4317 or HTTP on 4318. Set service.name and the other resource attributes, choose a sampler on purpose, size the batch queue for bursts, keep identifiers off metrics, follow semantic conventions, set the OTLP protocol explicitly, and check each signal's maturity in your language before depending on it.