A chaos experiment injects a real fault into a real system and checks a prediction about what happens. Everything in that sentence depends on observability. You cannot state the prediction without a measurable steady state, you cannot stop a bad experiment without a signal that tells you it is going bad, and you cannot learn anything afterwards without telemetry that separates the injected fault from ordinary noise. Teams that skip this part are not doing chaos engineering. They are breaking things and watching Slack.

This article covers the observability side of chaos work. The general method (hypotheses, blast radius, game days, organisational buy-in) is covered in the chaos engineering guide. Here the focus is the plumbing: how to write steady state as queries, how to mark an experiment so every metric, span and dashboard knows it is running, how to abort automatically from error-budget burn, and how to use each experiment as a test of your alerts and dashboards as well as of your service. A worked example injects 200 ms of latency between a checkout service and its payments dependency.

Advertisement

Two systems under test, not one

Every chaos experiment tests two things at once. The first is the production system: does checkout keep its latency objective when payments slows down? The second is the observability system: when checkout degrades, does an alert fire, does it page the right team, does the dashboard show where the problem is, and can an engineer get from the page to the cause from the traces?

The second test is often the more valuable one. A real incident is the worst time to find out that your latency alert fires after the outage is over. An experiment shows the same gap on a quiet afternoon, with the fault start time known to the second, which turns "alerting seemed slow" into a number: time to detect, measured from injection to firing.

So write down two hypotheses for each experiment. A system hypothesis, such as "checkout p99 stays under 800 ms and the error ratio stays under 0.5% with 200 ms added to payments calls". And an observability hypothesis, such as "if the system hypothesis fails, the checkout latency burn alert fires within 5 minutes and the trace view attributes the time to the payments client span".

Steady state, written as queries

A steady state is a set of service level indicators with healthy ranges, not words ("checkout is fine") or resource metrics; users do not experience CPU. Use the same SLIs that back your service level objectives, so that "healthy" in an experiment means exactly what it means for your error budget. If you have no SLOs yet, start with SLIs and SLOs before running experiments.

Each SLI becomes a query the experiment controller can run every few seconds. For an HTTP service instrumented with a request counter and a latency histogram, three queries cover most cases:

# error ratio, last 2 minutes (short window: the controller needs fast feedback)
sum(rate(http_server_requests_total{service="checkout", code=~"5.."}[2m]))
  / sum(rate(http_server_requests_total{service="checkout"}[2m]))

# p99 latency in seconds from a histogram
histogram_quantile(0.99,
  sum by (le) (rate(http_server_request_duration_seconds_bucket{service="checkout"}[2m])))

# business signal: completed orders per second, compared with the same time last week
sum(rate(orders_completed_total[5m]))
  / sum(rate(orders_completed_total[5m] offset 1w))

The business signal matters because technical SLIs can stay green while the product breaks: a payments timeout rendered as a 200 "order pending" page keeps the error ratio at zero and stops revenue.

Before injecting anything, the controller runs every query and refuses to start if the system is already outside its steady state. An experiment that begins on a sick system tells you nothing and makes the sickness worse.

Advertisement

The architecture: a control loop through your telemetry

The experiment controller applies the fault, polls the steady-state queries, and either lets the experiment run to its planned end or reverts early. The fault injector is any tool that can apply and remove a fault with a time limit: Chaos Mesh or LitmusChaos on Kubernetes, AWS Fault Injection Service, a service-mesh fault filter, or a proxy such as Toxiproxy in staging. The controller must not be the only thing that can end a fault, so the fault itself carries a TTL that the injector enforces even if the controller process dies.

The chaos experiment is a control loop closed through your telemetryExperiment controllerhypothesis, scope, TTLFault injectorChaos Mesh, FIS, proxyapply / revertTarget servicecheckout to paymentsfaultTelemetrymetrics, traces, logsemitsBackendPrometheus, tracing storeSteady-state checkerSLI queries vs thresholdsqueryabort / continueDashboardsannotation: exp idannotateAlerting and on-calldid it fire? when?rulesExperiment recordTTD, SLI deltas, gapsfired at T+nresult
The controller never trusts its own view of the target. It reads the same backend your on-call engineers read, annotates dashboards, and records whether the alerting path noticed, and how quickly.

Two design rules follow. First, the checker must query the same backend and recording rules your alerts use; SLIs computed from the controller's own probes teach you about the probes. Second, its queries need short windows and a matching scrape interval. With 60-second scrapes and a 5-minute rate window, a 10-minute experiment can spend half its life hurting users before the guard trips. Use 15-second scrapes on the services in scope, or read the guard signal from the mesh proxy.

Making the experiment visible in every signal

During and after an experiment, someone will look at a graph and ask whether a spike was the experiment or a real problem. Answer that question in the data, not in a chat thread. Tag the experiment into each telemetry type:

  • Dashboards: write an annotation at fault start and end, tagged with the experiment id, target and fault type. Grafana's HTTP API accepts a POST to /api/annotations with time, timeEnd, tags and text; most backends have an equivalent.
  • Metrics: export a gauge such as chaos_experiment_active{experiment="exp-0412", target="payments"} set to 1 while the fault is applied. Alert rules and recording rules can then join on it, and post-incident queries can filter it out. Do not add an experiment label to every request metric; that multiplies series cardinality for a signal that is almost always zero.
  • Traces: when the injector works at the request level (a mesh fault filter or an in-process fault library), set a span attribute on affected spans, and propagate the experiment id in W3C baggage so downstream services can tag their spans too. There is no OpenTelemetry semantic convention for this, so pick one team-wide name, for example chaos.experiment.id, and document it. See trace context propagation for how baggage travels.
  • Logs: the injector and controller log structured start, abort and stop events with the same id.

Network-level faults, such as delay applied with tc netem in a pod's network namespace, happen below the application and cannot mark spans. Rely on the gauge and annotations there.

Abort guards driven by error-budget burn

A fixed threshold such as "abort if the error ratio exceeds 1%" is easy to write and hard to justify. Tie the guard to the error budget instead, because the budget is the agreed price of unreliability. Start from the arithmetic. With a 99.9% availability objective over 30 days, the budget is 0.1% of requests, the equivalent of 43.2 minutes of total outage. An experiment that touches 5% of traffic for 10 minutes, and fails every touched request in the worst case, spends 0.05 x 10 = 0.5 outage-minutes, about 1.2% of the monthly budget. That worst case is a reasonable ceiling for one experiment.

Express the guard as a burn rate, the observed error ratio over the allowed ratio; burn-rate alerting explains the multi-window form used for paging. For an experiment, use a short window and a cap derived from your per-experiment budget, and stop the moment it is exceeded:

def run_experiment(exp, prom, injector, annotations):
    # 1. refuse to start on a system that is already unhealthy
    for name, query in exp.steady_state.items():
        if not exp.healthy(name, prom.query(query)):
            return record(exp, "not_started", reason=name)

    ann = annotations.start(exp.id, tags=["chaos", exp.target, exp.fault.kind])
    handle = injector.apply(exp.fault, scope=exp.scope, ttl=exp.max_duration)  # TTL is the backstop
    outcome, reason = "completed", None
    try:
        deadline = now() + exp.duration
        while now() < deadline:
            sleep(exp.poll_interval)                  # e.g. 10 s, at least the scrape interval
            burn = prom.query(exp.burn_rate_query) or 0.0
            if burn > exp.max_burn_rate:              # budget guard
                outcome, reason = "aborted", f"burn {burn:.1f}"
                break
            for name, query in exp.steady_state.items():
                if not exp.healthy(name, prom.query(query)):
                    outcome, reason = "aborted", name
                    break
            if outcome == "aborted":
                break
    finally:
        injector.revert(handle)                       # idempotent; safe to call twice
        annotations.end(ann)
    return record(exp, outcome, reason=reason, alerts=alerts_fired_since(ann.start))

Two details matter. A query that returns no data must count as a failure; the missing burn rate is treated as zero here only because the steady-state checks treat a missing value as unhealthy. And the revert sits in finally because an exception while polling is exactly when you need the fault removed.

Worked example: 200 ms between checkout and payments

The service: checkout calls payments synchronously during order placement, with a 1-second client timeout and one retry. The SLO is 99.9% of checkout requests succeeding and p99 under 800 ms. Current p99 is about 450 ms, of which payments contributes roughly 120 ms. The question: what happens when payments gets 200 ms slower, which its own team says happens during their deploys?

The fault, on half of the checkout pods:

apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
  name: exp-0412-payments-latency
  namespace: shop
spec:
  action: delay
  mode: fixed-percent
  value: "50"                 # half the checkout pods; the rest are the control group
  selector:
    namespaces: [shop]
    labelSelectors:
      app: checkout
  direction: to
  target:
    mode: all
    selector:
      namespaces: [shop]
      labelSelectors:
        app: payments
  delay:
    latency: "200ms"
    jitter: "20ms"
  duration: "10m"             # the injector removes the fault after this even if nobody else does

Compare experiment pods against control pods over the same minutes, grouping the latency query by pod. Diurnal traffic and concurrent deploys affect both groups equally and cancel out.

An illustrative outcome, of the kind these experiments commonly produce:

SignalControl podsExperiment podsHypothesis
p99 latency460 ms1,140 msFailed (limit 800 ms)
Error ratio0.02%0.04%Held
Payments client retries per second0.19.5Not predicted
Checkout latency burn alertn/aFired at T+7m40sFailed (expected under 5 min)
Trace attributionn/aRetry spans missingFailed

The interesting finding is not the latency itself. The added 200 ms pushed a slice of payments calls past the 1-second client timeout, the client retried, and those requests paid the timeout plus a second call: tail latency rose by far more than 200 ms. The traces could not show this, because the HTTP client library created one span for the logical call and none for each attempt, so an engineer would have seen one mysteriously slow span. The alert fired late because it used a 30-minute window alone. Three fixes follow: a per-attempt span with a lower retry budget, a short-window condition on the burn alert, and a client timeout set from measured latency.

Chaos as a test suite for observability

Once the plumbing exists, run experiments whose main purpose is to test detection. For each alert that matters, ask which fault should trigger it, inject that fault at a small scope, and measure. Record four numbers per run: time to detect (injection to alert firing), time to notify (firing to the page reaching a person), whether the alert routed to the owning team, and whether the linked dashboard or runbook pointed at the cause. Keep them in a table, re-run quarterly, and treat a regression like any other failing test.

Good observability faults: killing the metrics exporter, filling a log shipper's disk, dropping spans at the collector. The first finds a common silent failure: alert expressions like error_ratio > 0.01 return nothing when the series disappears, and nothing fires. Pair each critical alert with an absent() rule or a heartbeat check.

Link every result to the runbook it exercised. If the page arrived but the runbook sent the engineer to the wrong dashboard, the experiment failed even though the alert worked.

Failure modes

FailureWhat goes wrongMitigation
Guard reads stale dataLong scrape or rate windows trip the abort minutes lateShort windows and 10 to 15 s scrapes on in-scope services
No data treated as healthyThe fault breaks telemetry and the guard sees silence as successMissing value counts as unhealthy; absent() rules
Controller dies mid-runFault stays applied indefinitelyInjector-enforced TTL; a dead-man check on the controller
Experiment hidden in aggregates5% scope moves the global p99 too little to seeCompare experiment and control groups, not global series
Spikes later mistaken for incidentsNo marker in metrics or dashboardsAnnotations plus an active-experiment gauge
Cardinality blow-upExperiment label added to every request metricOne gauge per experiment, not a label on all series
Alerts suppressed during runsSilencing everything hides the observability testRoute experiment pages to a test receiver instead of silencing

Operational guidance and trade-offs

Prove the control loop in staging, then move to production at tiny scope: staging traffic rarely exercises the retry and timeout paths production finds, and the guard and TTL are what make production runs acceptable. Run in working hours with the owning team aware, never during an incident or freeze. Rather than silencing alerts, route in-scope pages to a test receiver that records firing time.

The trade-off is fidelity against risk. A control group buys back clarity at small scope, which usually beats a larger blast radius.

What to do next

  1. Write the steady state for one critical service as two or three SLI queries plus one business signal, using the same recording rules as your alerts.
  2. Check the scrape interval and rate windows on the services in scope, and shorten them where the guard would react in minutes.
  3. Add experiment markers: dashboard annotations, an active-experiment gauge, and a documented span attribute and baggage key.
  4. Build or adopt a controller with a pre-flight health check, a burn-rate guard, missing-data-as-failure and a revert in a finally block, and set a TTL on every fault.
  5. Run one small latency experiment with a control group, and write down both the system and the observability hypothesis first.
  6. Record time to detect, time to notify, routing and runbook accuracy for each critical alert, and re-run them on a schedule.
  7. Add an absent() or heartbeat rule for every alert whose expression can silently return no data.
Key takeaway: Chaos engineering is only as good as the observability around it. Write the steady state as the same SLI queries your SLOs use, mark every experiment in dashboards, metrics, traces and logs, and let an error-budget burn guard plus an injector-enforced TTL end bad runs automatically. Treat each experiment as two tests, of the service and of the alerting path, and measure time to detect from the known injection time so detection gaps show up on a quiet afternoon instead of during an incident.