Most teams start observability by installing a tool. A few weeks later they have a dashboard nobody opens, logs nobody can search, and an incident where the first forty minutes are spent finding out which service is slow. The tool was not the problem. Nobody decided what questions the system had to answer.
This guide builds observability from nothing in the order that pays off fastest: decide the questions, fix the naming, instrument one service end to end, put a collector in the middle, stand up storage for metrics, logs and traces, then build exactly one dashboard and two alerts. The stack used here is OpenTelemetry, Prometheus, Loki, Tempo and Grafana because they are open source and speak the same protocol, but every step maps onto a hosted vendor if you would rather pay than operate.
Start with questions, not tools
Write down the questions an on-call engineer must answer in the first ten minutes of an incident. For a typical web product they are short and concrete:
- Are users affected right now, and how many? (error rate and latency at the edge)
- Which service or dependency is the source? (per-service RED metrics, trace waterfalls)
- What changed? (deploy markers, version labels)
- What exactly happened to one bad request? (a trace and the logs that share its trace id)
- Is it getting worse, and how fast is the error budget burning? (SLO burn rate)
Each question names a signal. Questions 1, 2 and 5 are metrics: cheap, aggregated, good for alerting. Question 4 needs traces and logs: expensive per event, but they explain a single request. Question 3 needs labels on all of them. Anything you build that does not help answer one of these questions can wait.
Also write down what you will not do in the first two weeks: no custom business dashboards, no log analytics, no profiling, no anomaly detection. Those are worth having later, on top of a base that works.
The architecture
Services emit all three signals with the OpenTelemetry SDK over OTLP to a collector. The collector is the one place you change routing, sampling, redaction and batching without redeploying applications. It forwards metrics to Prometheus, logs to Loki and traces to Tempo. Grafana queries all three, and Prometheus rules feed Alertmanager.
Two design rules matter more than the choice of stores. First, applications never talk to a backend directly, so a backend migration never touches application code. Second, every signal carries the same resource attributes and every log line inside a request carries the trace id, so you can pivot from a latency spike to an example trace to its logs in three clicks.
Step 1: resource identity and naming
Before writing instrumentation, fix the attributes that identify where telemetry came from. OpenTelemetry calls them resource attributes. The minimum set:
| Attribute | Example | Why |
|---|---|---|
service.name | checkout | Becomes the Prometheus job label; the primary filter everywhere |
service.namespace | shop | Groups services; prefixed onto job as shop/checkout |
service.version | 2026.10.03-4f2a1c | Answers what changed |
service.instance.id | pod or host id | Becomes the instance label |
deployment.environment.name | prod | Keeps staging out of production alerts |
Set them with environment variables so every language uses the same mechanism, and generate those variables from your deploy tooling rather than typing them per service. A misspelt service.name is a missing service on every dashboard.
OTEL_SERVICE_NAME=checkout
OTEL_RESOURCE_ATTRIBUTES=service.namespace=shop,service.version=${GIT_SHA},deployment.environment.name=prod
OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
OTEL_TRACES_SAMPLER=parentbased_traceidratio
OTEL_TRACES_SAMPLER_ARG=0.2
Step 2: instrument one service end to end
Pick the service closest to users, usually the API gateway or the main web backend. Start with the language's automatic instrumentation for the HTTP framework, database client and outbound HTTP client; it produces server spans, client spans and the standard HTTP duration histogram without code changes. Then add the small amount of manual instrumentation that auto-instrumentation cannot know about. In Python the manual part looks like this:
import logging
from opentelemetry import trace, metrics
tracer = trace.get_tracer("checkout")
meter = metrics.get_meter("checkout")
orders = meter.create_counter("orders_placed", unit="1",
description="Orders accepted by checkout")
log = logging.getLogger("checkout")
def place_order(cart, payment):
with tracer.start_as_current_span("place_order") as span:
span.set_attribute("cart.items", len(cart.items)) # low-cardinality only
try:
charge = payment.charge(cart.total) # auto-instrumented client span
except PaymentDeclined as e:
span.set_attribute("payment.declined", True)
log.warning("payment declined", extra={"reason": e.code})
raise
orders.add(1, {"payment.method": payment.method})
log.info("order placed", extra={"order_id": charge.order_id})
return chargeThree habits to set now. Put identifiers such as order ids in logs and span attributes, never in metric labels, because every distinct label value creates a new time series. Log through the OpenTelemetry logging bridge (or include the current trace id in your structured log formatter) so every line inside a request carries trace_id. And record errors on the span, so a trace search for errors finds the request you are looking at.
Use parent-based ratio sampling from the start, as in the variables above. It keeps whole traces consistent across services because downstream services follow the decision made at the edge. Twenty percent is a starting point; tail sampling in the collector can come later.
Step 3: the collector
Run the OpenTelemetry Collector (the contrib distribution has the most components) as a service your applications can reach. This config receives OTLP, protects itself from memory exhaustion, batches, and fans out to the three stores:
receivers:
otlp:
protocols:
grpc: { endpoint: 0.0.0.0:4317 }
http: { endpoint: 0.0.0.0:4318 }
processors:
memory_limiter:
check_interval: 1s
limit_percentage: 80
spike_limit_percentage: 20
batch: {}
exporters:
otlphttp/metrics:
endpoint: http://prometheus:9090/api/v1/otlp # exporter appends /v1/metrics
otlphttp/logs:
endpoint: http://loki:3100/otlp # exporter appends /v1/logs
otlp/traces:
endpoint: tempo:4317
tls: { insecure: true }
service:
pipelines:
metrics: { receivers: [otlp], processors: [memory_limiter, batch], exporters: [otlphttp/metrics] }
logs: { receivers: [otlp], processors: [memory_limiter, batch], exporters: [otlphttp/logs] }
traces: { receivers: [otlp], processors: [memory_limiter, batch], exporters: [otlp/traces] }The otlphttp exporter takes a base URL and appends the per-signal path itself, which is why the Prometheus endpoint stops at /api/v1/otlp. Put memory_limiter first in every pipeline so a traffic spike makes the collector refuse data rather than crash. Component names and defaults do change between Collector releases, so pin the image version and read the changelog before upgrading.
Step 4: the stores
Prometheus accepts OTLP natively once started with --web.enable-otlp-receiver; it then serves /api/v1/otlp/v1/metrics. The receiver has no authentication, so keep it on a private network that only the collector can reach. The Prometheus OpenTelemetry guide also recommends an out-of-order window, because pushed samples can arrive late, and promoting the resource attributes you want as labels:
# prometheus.yml
storage:
tsdb:
out_of_order_time_window: 30m
otlp:
promote_resource_attributes:
- service.namespace
- service.version
- deployment.environment.namePrometheus maps service.name (prefixed by service.namespace when present) to the job label and service.instance.id to instance; the remaining resource attributes land on a target_info series you can join against. Promote only the few attributes you filter on constantly; every promoted attribute is copied onto every series.
Loki ingests OTLP logs at /otlp and indexes only a small set of labels, keeping the rest as structured metadata. That is the right trade for logs: you filter by service and environment, then search text. Tempo stores traces in object storage and is queried by trace id or TraceQL. Start all three with local disk, then move Loki and Tempo to object storage before you depend on them; it is cheaper and survives a node loss.
Step 5: one dashboard that answers the questions
Build a single service overview dashboard with a service selector and four rows: request rate, error ratio, latency percentiles, and saturation (CPU, memory, connection pool use). With the standard HTTP server histogram exported through OTLP, the queries look like this (exact metric and label names depend on your Prometheus translation settings, so confirm them in the metrics browser first):
# Rate
sum by (job) (rate(http_server_request_duration_seconds_count{job="shop/checkout"}[5m]))
# Error ratio (5xx)
sum(rate(http_server_request_duration_seconds_count{job="shop/checkout", http_response_status_code=~"5.."}[5m]))
/
sum(rate(http_server_request_duration_seconds_count{job="shop/checkout"}[5m]))
# p95 latency by route
histogram_quantile(0.95,
sum by (le, http_route) (rate(http_server_request_duration_seconds_bucket{job="shop/checkout"}[5m])))Turn on exemplars for the latency panel and configure the Tempo data source to link trace ids to Loki. That gives the three-click pivot: spike, exemplar trace, logs for that trace. Add deploy annotations from your CI system so question 3 answers itself.
Step 6: two alerts, both on symptoms
Define one availability SLO and one latency SLO for the edge service, for example 99.9 percent of requests succeed and 99 percent complete within 500 ms over 30 days. Alert on burn rate, how fast the error budget is being consumed, not on raw thresholds. The multiwindow pattern from the Google SRE workbook pages when the one-hour burn rate exceeds 14.4 (2 percent of a 30-day budget in an hour) and a five-minute window confirms it is still happening:
- alert: CheckoutAvailabilityFastBurn
expr: |
(job:http_errors:ratio_rate1h{job="shop/checkout"} > (14.4 * 0.001))
and
(job:http_errors:ratio_rate5m{job="shop/checkout"} > (14.4 * 0.001))
labels: { severity: page }
annotations:
runbook: "https://runbooks.internal/checkout-availability"The ratio_rate series are recording rules built from the error ratio query above. Add a slower ticket-level alert (six-hour window, burn rate 6) later. Do not page on CPU, memory or queue depth in week one; those are causes, and they belong on the dashboard, where they explain a symptom alert that already fired.
Worked example: the slow checkout
Two weeks in, the fast-burn latency alert fires at 14:05. The overview dashboard shows p95 on POST /checkout rising from 300 ms to 2.1 s while error rate is flat and request rate is normal. The deploy annotation shows no checkout release, but service.version on the payments service changed at 13:58.
Clicking an exemplar on the latency panel opens a trace: the place_order span takes 2.0 s, of which 1.9 s is a client span to payments, and inside payments a database span runs the same query forty times. The logs for that trace id show a new fraud-check code path. The cause, a query inside a loop introduced at 13:58, is identified in about six minutes and the payments team rolls back. None of it required anything beyond the six steps above.
What it costs
Estimate before you build, because observability bills grow quietly. A worked estimate for 20 services with 3 instances each, 500 requests per second in total:
| Signal | Volume driver | Rough size |
|---|---|---|
| Metrics | 60 instances x about 1,500 series each | about 90,000 active series |
| Traces | 500 rps x 20% sampled x 8 spans x about 1 KB | about 0.8 MB/s, 70 GB/day |
| Logs | 500 rps x 4 lines x about 400 bytes | about 0.8 MB/s, 70 GB/day |
Ninety thousand series is comfortable for one Prometheus server. Trace and log volume dominate storage, which is why object storage, compression and short default retention (for example 14 days for traces and debug logs, 30 days for info logs, 13 months for downsampled metrics) matter more than which store you chose. Re-run this estimate monthly against real ingestion numbers.
Failure modes
- Cardinality explosion. A user id or full URL in a metric label multiplies series until Prometheus runs out of memory. Use route templates, keep ids in logs and traces, and alert on series count.
- Broken trace context. Traces stop at a queue or a hand-rolled HTTP client because context was not propagated. Check that a trace crosses every hop of your main flow.
- Collector as a single point of failure. Run at least two replicas behind a load balancer and watch the collector's own metrics for refused and dropped data.
- Unauthenticated receivers. OTLP endpoints on Prometheus and Loki accept writes from anyone who can reach them. Keep them private.
- Alerting on causes. CPU alerts page at night and are usually harmless. Page on SLO burn; investigate causes from the dashboard.
- Dashboards nobody owns. Give each dashboard an owner and delete the ones not viewed in 90 days.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Self-hosted stack | Low licence cost, full control | You operate four stateful systems |
| Hosted vendor | No operations, fast start | Per-GB or per-series pricing that grows with traffic |
| Collector in the middle | Change routing without redeploys | One more service to run and scale |
| Head sampling at 20% | Predictable cost | Rare slow requests can be missed |
| Tail sampling later | Keeps every error trace | Collector must hold whole traces in memory |
Because everything goes through OTLP and a collector, the hosted-versus-self-hosted decision is reversible: you change exporter configuration, not applications.
What to do next
- Write the five on-call questions for your system and map each one to a signal.
- Define the resource attributes and generate the OTEL_ environment variables from your deploy tooling.
- Auto-instrument the edge service, add trace ids to its logs, and confirm one request appears as a trace with matching logs.
- Run two collector replicas with
memory_limiterandbatch, pinned to a specific version. - Enable the Prometheus OTLP receiver on a private network, with an out-of-order window and promoted attributes.
- Build the four-row service dashboard with exemplars and deploy annotations.
- Add one availability and one latency burn-rate alert, each linked to a runbook.
- Instrument the next service, then check a trace crosses the boundary between them.
- Run the cost estimate against real ingestion after two weeks and set retention to match.
To go further: the three pillars, in depth, OpenTelemetry across a full stack, collector architecture, burn-rate alerting, managing metric cardinality and writing runbooks.