"Observability" gets used as a synonym for a monitoring product, or for the three telemetry types: metrics, logs and traces. The three pillars article covers those signals, how to join them and what they cost. This article covers the idea underneath them. Observability is a property of a system: how well you can explain what it is doing, including behaviour nobody predicted, using only what it emits and without shipping new code.
That definition is practical. It gives you a test you can run against your own system and a design target for instrumentation. It also explains why some teams with a large telemetry bill still cannot answer "why are these checkouts slow?" We will work from the definition to the unit of data that best supports it (the wide event), the cardinality that makes it useful, the investigation loop that uses it, and the cost controls that keep it affordable. A worked example follows one real-shaped regression from alert to cause.
Observability is a property, not a product
The word comes from control theory. Rudolf Kalman defined a system as observable if its internal state can be determined from its outputs over time. The software version keeps the spirit. Your service has internal state (which code path ran, which cache missed, which tenant, which build, which feature flags), and the question is whether its outputs let you reconstruct that state for any given request after the fact.
Compare that with monitoring. Monitoring answers questions you thought of in advance: is the error rate above 1 percent, is the disk 90 percent full. You encode each question as a metric and an alert. That works for known failure modes and fails for novel ones, which in a distributed system are most of the interesting ones. When an incident is something nobody has seen before, the predefined dashboards show that something is wrong, but not what. Observability is how quickly you get from the symptom to the cause when the question is new. Monitoring is still needed; it decides when to start looking.
A test for whether your system is observable
You can score a system's observability by trying to answer questions about a real recent request or incident using only existing telemetry. Write down the time each one takes, and whether you needed to deploy code or add logging to answer it.
| Question | Needs | Typical gap |
|---|---|---|
| Which customers or tenants are affected? | Tenant ID on every request record | Only in logs of some services, or dropped from metrics for cardinality |
| Did it start with a deploy or flag change? | Build ID, version and flag state per request | Deploy markers on a dashboard but not on the data |
| Which dependency is slow for these requests? | Per-call timings tied to the request | Aggregated dependency latency with no link back to requests |
| What do the bad requests have in common? | Many fields on the same record, queryable together | Fields spread across three systems with different IDs |
| Is this new or has it happened before? | Retention long enough for a baseline | Seven days of raw data, aggregates only after that |
A system where an on-call engineer can answer all five in under ten minutes, without a deploy, is observable for practical purposes. If most answers start with "we would need to add logging", the gap is in instrumentation, not tooling.
Wide events: the unit of explanation
The data shape that answers those questions best is the wide event: one structured record per unit of work per service (an HTTP request, a queue message, a job run), emitted when the work finishes and carrying every field that might matter. That can be dozens to a few hundred fields. Stripe described the same idea years ago as canonical log lines: one summary line per request alongside the usual noisy logs. A trace span is a wide event with parent and child links. In OpenTelemetry the natural home is the server span for the request, enriched with attributes as the request runs.
The point is co-location. When tenant, build, endpoint, cache result, retry count, payload size, database time and status all sit on one record, any of them can be filtered or grouped against any other. Spread them across a metric, two log lines and a trace and you have to join across systems before asking anything. The structured logging guide covers schemas and redaction; here is the shape in code:
import time
from opentelemetry import trace
def wide_event_middleware(app):
def handler(request):
span = trace.get_current_span() # server span created by auto-instrumentation
ev = {"app.tenant_id": request.tenant_id,
"app.user_tier": request.user.tier,
"app.build_id": BUILD_ID,
"app.region": REGION,
"app.flags": ",".join(sorted(request.flags.enabled())),
"app.request_bytes": request.content_length or 0}
start = time.monotonic()
try:
response = app(request) # handlers add fields as they learn them:
ev["app.status"] = response.status # cache hit, rows scanned, retries, plan id ...
return response
except Exception as exc:
ev["app.error_type"] = type(exc).__name__
raise
finally:
ev["app.duration_ms"] = round((time.monotonic() - start) * 1000, 1)
ev.update(request.ctx.fields) # accumulated by inner code
span.set_attributes(ev)
return handlerInner code adds fields through a request-scoped dictionary (request.ctx.fields["app.cache"] = "miss") instead of writing log lines. A useful rule: if you would write a log line saying something happened during a request, make it a field on the request's event instead. Use the OpenTelemetry semantic convention names where they exist and a single prefix for your own; the OpenTelemetry guide explains the SDK pieces.
Cardinality and dimensionality
Two properties of the data decide what you can ask. Cardinality is the number of distinct values a field has: status code has a few dozen, user ID has millions. Dimensionality is how many fields each record carries. The most useful debugging fields are almost always high-cardinality: user, tenant, request, build, container, query fingerprint. Novel failures tend to hide in a specific combination, such as one tenant on one build in one region.
Time-series metrics cannot hold those fields. Every distinct label combination becomes its own series, so adding user ID to a request counter multiplies series count by the number of users and breaks the metrics backend. The cardinality management guide covers the budgets. Events do not have this problem, because each record is stored once whatever its values. Columnar stores (ClickHouse, Apache Pinot, purpose-built tracing backends) compress repeated values well and scan only the columns a query touches. So the design splits cleanly: low-cardinality metrics for alerting and SLOs, high-cardinality events for explanation, as the diagram shows.
The investigation loop
With wide events in place, investigation becomes a repeatable loop rather than a hunt through dashboards:
- Start from the symptom the alert describes: the SLO burning for p99 latency on
/checkout. - Isolate the bad population. Filter to the events that violate it (duration above the threshold, or errors).
- Compare it with the baseline. For every field, ask which values are over-represented among bad events compared with good ones in the same window.
- Form a hypothesis from the top differences and verify it by filtering on it: does the problem disappear when you exclude it, and does it appear only where it is present?
- Repeat inside the narrowed population until you reach something actionable.
Step 3 is the heart of the loop and it is easy to automate. Some vendors offer it as a product feature; the core is a few lines:
from collections import Counter
def over_represented(events, is_bad, fields, top=10):
# Rank (field, value) pairs by how much more common they are in bad events.
# Each event carries sample_rate, so counts are weighted back to real traffic.
bad, base = Counter(), Counter()
n_bad = n_base = 0
for e in events:
w = e.get("sample_rate", 1)
target = bad if is_bad(e) else base
if is_bad(e): n_bad += w
else: n_base += w
for f in fields:
for v in str(e.get(f)).split(","): # multi-valued, e.g. app.flags
target[(f, v)] += w
scores = []
for key, cnt in bad.items():
p_bad, p_base = cnt / n_bad, base[key] / max(n_base, 1)
scores.append((p_bad - p_base, key, round(p_bad, 3), round(p_base, 3)))
return sorted(scores, reverse=True)[:top]
Worked example: a checkout latency regression
Here is the loop on a realistic regression. At 14:10 the checkout latency SLO starts burning: p99 went from about 900 ms to 4 s, but the median barely moved. The dashboards show database latency up slightly and nothing else. Running the comparison over the last 30 minutes of /checkout events, with "bad" meaning slower than 2 s, gives:
score field value p_bad p_base
0.71 app.flags new_tax_engine 0.83 0.12
0.44 app.region eu-west-1 0.79 0.35
0.38 app.cart_items_gt20 true 0.47 0.09
0.05 app.build_id b-4471 0.52 0.47The build is barely over-represented, so this is not a bad deploy. The tax-engine flag is. Filtering to flag-on events and re-running inside that population shows app.tax.lookups climbing with cart size. A query over the event table, with one column per field, confirms the shape:
SELECT intDiv(cart_items, 10) * 10 AS cart_bucket,
sum(sample_rate) AS requests,
quantileExactWeighted(0.99)(duration_ms, sample_rate) AS p99_ms,
avg(tax_lookups) AS lookups
FROM events
WHERE route = '/checkout' AND has(splitByChar(',', flags), 'new_tax_engine')
AND timestamp > now() - INTERVAL 30 MINUTE
GROUP BY cart_bucket ORDER BY cart_bucket;Large carts make one tax lookup per line item, a classic N+1 pattern, and the flag's rollout to the EU that afternoon exposed it. The mitigation is to turn the flag off for carts over 20 items; the fix is a batched lookup. Total time from page to cause is about fifteen minutes, with no new code deployed to diagnose it. With only per-service metrics, the visible symptom (slightly slower database) would have pointed the team at the wrong system.
Keeping it affordable
Wide events are not free. At 20,000 requests per second and 2 KB per event, raw volume is 40 MB/s, about 3.5 TB a day before compression. Columnar compression often achieves large ratios on repetitive telemetry, but plan from your own measurements. Three controls keep it affordable without losing the ability to explain:
- Sample with weights. Keep every error and every slow request, and keep a fraction of healthy ones, such as 1 in 20. Record
sample_rateon each kept event and weight every count and quantile by it, as the code above does. The trace sampling guide covers head, tail and consistent sampling. - Derive metrics before sampling. Compute request counts and latency histograms from all traffic in the collector, so SLOs and alerts never depend on sampled data.
- Tier retention. Keep full events for days to a few weeks, and keep weighted aggregates longer for baselines.
Failure modes and trade-offs
- Telemetry without context. Lots of data, but tenant and build are missing from the records, so nothing can be sliced. Add identity and version fields before adding more volume.
- Unweighted sampled counts. Charts from sampled events under-count by the sample rate and over-weight errors. Always multiply by
sample_rate. - Cardinality pushed into metrics. Teams add user ID as a metric label to get high-cardinality answers and take the metrics backend down. Put it on events.
- Broken context propagation. A queue hop or thread pool drops the trace context and events cannot be joined across services.
- Sensitive data in fields. Wide events attract emails and tokens. Allowlist field names in the collector and hash identifiers where the raw value is not needed.
- Trade-off: richness against cost and privacy. Every field is a potential answer and a potential liability. Add fields that have answered a real question or plausibly will, and review the schema quarterly.
What to do next
- Score your system against the five-question test using last month's incident, and record what each answer needed.
- Add one wide event per request at the edge of each service, starting with tenant, build, region, flags, status and duration.
- Move "something happened" log lines inside requests into fields on that event.
- Store events in a columnar backend that can group by any field, and verify you can filter on user or tenant in seconds.
- Derive RED metrics and SLOs from unsampled traffic; sample events with recorded weights.
- Write the over-representation comparison as a saved query or script and use it in the next incident review.
- Add a collector allowlist for field names and redact or hash personal data.
- Re-run the five-question test after a quarter and compare times.