A chaos experiment injects a real fault into a real system and checks a prediction about what happens. Everything in that sentence depends on observability. You cannot state the prediction without a measurable steady state, you cannot stop a bad experiment without a signal that tells you it is going bad, and you cannot learn anything afterwards without telemetry that separates the injected fault from ordinary noise. Teams that skip this part are not doing chaos engineering. They are breaking things and watching Slack.
This article covers the observability side of chaos work. The general method (hypotheses, blast radius, game days, organisational buy-in) is covered in the chaos engineering guide. Here the focus is the plumbing: how to write steady state as queries, how to mark an experiment so every metric, span and dashboard knows it is running, how to abort automatically from error-budget burn, and how to use each experiment as a test of your alerts and dashboards as well as of your service. A worked example injects 200 ms of latency between a checkout service and its payments dependency.
Two systems under test, not one
Every chaos experiment tests two things at once. The first is the production system: does checkout keep its latency objective when payments slows down? The second is the observability system: when checkout degrades, does an alert fire, does it page the right team, does the dashboard show where the problem is, and can an engineer get from the page to the cause from the traces?
The second test is often the more valuable one. A real incident is the worst time to find out that your latency alert fires after the outage is over. An experiment shows the same gap on a quiet afternoon, with the fault start time known to the second, which turns "alerting seemed slow" into a number: time to detect, measured from injection to firing.
So write down two hypotheses for each experiment. A system hypothesis, such as "checkout p99 stays under 800 ms and the error ratio stays under 0.5% with 200 ms added to payments calls". And an observability hypothesis, such as "if the system hypothesis fails, the checkout latency burn alert fires within 5 minutes and the trace view attributes the time to the payments client span".
Steady state, written as queries
A steady state is a set of service level indicators with healthy ranges, not words ("checkout is fine") or resource metrics; users do not experience CPU. Use the same SLIs that back your service level objectives, so that "healthy" in an experiment means exactly what it means for your error budget. If you have no SLOs yet, start with SLIs and SLOs before running experiments.
Each SLI becomes a query the experiment controller can run every few seconds. For an HTTP service instrumented with a request counter and a latency histogram, three queries cover most cases:
# error ratio, last 2 minutes (short window: the controller needs fast feedback)
sum(rate(http_server_requests_total{service="checkout", code=~"5.."}[2m]))
/ sum(rate(http_server_requests_total{service="checkout"}[2m]))
# p99 latency in seconds from a histogram
histogram_quantile(0.99,
sum by (le) (rate(http_server_request_duration_seconds_bucket{service="checkout"}[2m])))
# business signal: completed orders per second, compared with the same time last week
sum(rate(orders_completed_total[5m]))
/ sum(rate(orders_completed_total[5m] offset 1w))The business signal matters because technical SLIs can stay green while the product breaks: a payments timeout rendered as a 200 "order pending" page keeps the error ratio at zero and stops revenue.
Before injecting anything, the controller runs every query and refuses to start if the system is already outside its steady state. An experiment that begins on a sick system tells you nothing and makes the sickness worse.
The architecture: a control loop through your telemetry
The experiment controller applies the fault, polls the steady-state queries, and either lets the experiment run to its planned end or reverts early. The fault injector is any tool that can apply and remove a fault with a time limit: Chaos Mesh or LitmusChaos on Kubernetes, AWS Fault Injection Service, a service-mesh fault filter, or a proxy such as Toxiproxy in staging. The controller must not be the only thing that can end a fault, so the fault itself carries a TTL that the injector enforces even if the controller process dies.
Two design rules follow. First, the checker must query the same backend and recording rules your alerts use; SLIs computed from the controller's own probes teach you about the probes. Second, its queries need short windows and a matching scrape interval. With 60-second scrapes and a 5-minute rate window, a 10-minute experiment can spend half its life hurting users before the guard trips. Use 15-second scrapes on the services in scope, or read the guard signal from the mesh proxy.
Making the experiment visible in every signal
During and after an experiment, someone will look at a graph and ask whether a spike was the experiment or a real problem. Answer that question in the data, not in a chat thread. Tag the experiment into each telemetry type:
- Dashboards: write an annotation at fault start and end, tagged with the experiment id, target and fault type. Grafana's HTTP API accepts a POST to
/api/annotationswithtime,timeEnd,tagsandtext; most backends have an equivalent. - Metrics: export a gauge such as
chaos_experiment_active{experiment="exp-0412", target="payments"}set to 1 while the fault is applied. Alert rules and recording rules can then join on it, and post-incident queries can filter it out. Do not add an experiment label to every request metric; that multiplies series cardinality for a signal that is almost always zero. - Traces: when the injector works at the request level (a mesh fault filter or an in-process fault library), set a span attribute on affected spans, and propagate the experiment id in W3C baggage so downstream services can tag their spans too. There is no OpenTelemetry semantic convention for this, so pick one team-wide name, for example
chaos.experiment.id, and document it. See trace context propagation for how baggage travels. - Logs: the injector and controller log structured start, abort and stop events with the same id.
Network-level faults, such as delay applied with tc netem in a pod's network namespace, happen below the application and cannot mark spans. Rely on the gauge and annotations there.
Abort guards driven by error-budget burn
A fixed threshold such as "abort if the error ratio exceeds 1%" is easy to write and hard to justify. Tie the guard to the error budget instead, because the budget is the agreed price of unreliability. Start from the arithmetic. With a 99.9% availability objective over 30 days, the budget is 0.1% of requests, the equivalent of 43.2 minutes of total outage. An experiment that touches 5% of traffic for 10 minutes, and fails every touched request in the worst case, spends 0.05 x 10 = 0.5 outage-minutes, about 1.2% of the monthly budget. That worst case is a reasonable ceiling for one experiment.
Express the guard as a burn rate, the observed error ratio over the allowed ratio; burn-rate alerting explains the multi-window form used for paging. For an experiment, use a short window and a cap derived from your per-experiment budget, and stop the moment it is exceeded:
def run_experiment(exp, prom, injector, annotations):
# 1. refuse to start on a system that is already unhealthy
for name, query in exp.steady_state.items():
if not exp.healthy(name, prom.query(query)):
return record(exp, "not_started", reason=name)
ann = annotations.start(exp.id, tags=["chaos", exp.target, exp.fault.kind])
handle = injector.apply(exp.fault, scope=exp.scope, ttl=exp.max_duration) # TTL is the backstop
outcome, reason = "completed", None
try:
deadline = now() + exp.duration
while now() < deadline:
sleep(exp.poll_interval) # e.g. 10 s, at least the scrape interval
burn = prom.query(exp.burn_rate_query) or 0.0
if burn > exp.max_burn_rate: # budget guard
outcome, reason = "aborted", f"burn {burn:.1f}"
break
for name, query in exp.steady_state.items():
if not exp.healthy(name, prom.query(query)):
outcome, reason = "aborted", name
break
if outcome == "aborted":
break
finally:
injector.revert(handle) # idempotent; safe to call twice
annotations.end(ann)
return record(exp, outcome, reason=reason, alerts=alerts_fired_since(ann.start))Two details matter. A query that returns no data must count as a failure; the missing burn rate is treated as zero here only because the steady-state checks treat a missing value as unhealthy. And the revert sits in finally because an exception while polling is exactly when you need the fault removed.
Worked example: 200 ms between checkout and payments
The service: checkout calls payments synchronously during order placement, with a 1-second client timeout and one retry. The SLO is 99.9% of checkout requests succeeding and p99 under 800 ms. Current p99 is about 450 ms, of which payments contributes roughly 120 ms. The question: what happens when payments gets 200 ms slower, which its own team says happens during their deploys?
The fault, on half of the checkout pods:
apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
name: exp-0412-payments-latency
namespace: shop
spec:
action: delay
mode: fixed-percent
value: "50" # half the checkout pods; the rest are the control group
selector:
namespaces: [shop]
labelSelectors:
app: checkout
direction: to
target:
mode: all
selector:
namespaces: [shop]
labelSelectors:
app: payments
delay:
latency: "200ms"
jitter: "20ms"
duration: "10m" # the injector removes the fault after this even if nobody else doesCompare experiment pods against control pods over the same minutes, grouping the latency query by pod. Diurnal traffic and concurrent deploys affect both groups equally and cancel out.
An illustrative outcome, of the kind these experiments commonly produce:
| Signal | Control pods | Experiment pods | Hypothesis |
|---|---|---|---|
| p99 latency | 460 ms | 1,140 ms | Failed (limit 800 ms) |
| Error ratio | 0.02% | 0.04% | Held |
| Payments client retries per second | 0.1 | 9.5 | Not predicted |
| Checkout latency burn alert | n/a | Fired at T+7m40s | Failed (expected under 5 min) |
| Trace attribution | n/a | Retry spans missing | Failed |
The interesting finding is not the latency itself. The added 200 ms pushed a slice of payments calls past the 1-second client timeout, the client retried, and those requests paid the timeout plus a second call: tail latency rose by far more than 200 ms. The traces could not show this, because the HTTP client library created one span for the logical call and none for each attempt, so an engineer would have seen one mysteriously slow span. The alert fired late because it used a 30-minute window alone. Three fixes follow: a per-attempt span with a lower retry budget, a short-window condition on the burn alert, and a client timeout set from measured latency.
Chaos as a test suite for observability
Once the plumbing exists, run experiments whose main purpose is to test detection. For each alert that matters, ask which fault should trigger it, inject that fault at a small scope, and measure. Record four numbers per run: time to detect (injection to alert firing), time to notify (firing to the page reaching a person), whether the alert routed to the owning team, and whether the linked dashboard or runbook pointed at the cause. Keep them in a table, re-run quarterly, and treat a regression like any other failing test.
Good observability faults: killing the metrics exporter, filling a log shipper's disk, dropping spans at the collector. The first finds a common silent failure: alert expressions like error_ratio > 0.01 return nothing when the series disappears, and nothing fires. Pair each critical alert with an absent() rule or a heartbeat check.
Link every result to the runbook it exercised. If the page arrived but the runbook sent the engineer to the wrong dashboard, the experiment failed even though the alert worked.
Failure modes
| Failure | What goes wrong | Mitigation |
|---|---|---|
| Guard reads stale data | Long scrape or rate windows trip the abort minutes late | Short windows and 10 to 15 s scrapes on in-scope services |
| No data treated as healthy | The fault breaks telemetry and the guard sees silence as success | Missing value counts as unhealthy; absent() rules |
| Controller dies mid-run | Fault stays applied indefinitely | Injector-enforced TTL; a dead-man check on the controller |
| Experiment hidden in aggregates | 5% scope moves the global p99 too little to see | Compare experiment and control groups, not global series |
| Spikes later mistaken for incidents | No marker in metrics or dashboards | Annotations plus an active-experiment gauge |
| Cardinality blow-up | Experiment label added to every request metric | One gauge per experiment, not a label on all series |
| Alerts suppressed during runs | Silencing everything hides the observability test | Route experiment pages to a test receiver instead of silencing |
Operational guidance and trade-offs
Prove the control loop in staging, then move to production at tiny scope: staging traffic rarely exercises the retry and timeout paths production finds, and the guard and TTL are what make production runs acceptable. Run in working hours with the owning team aware, never during an incident or freeze. Rather than silencing alerts, route in-scope pages to a test receiver that records firing time.
The trade-off is fidelity against risk. A control group buys back clarity at small scope, which usually beats a larger blast radius.
What to do next
- Write the steady state for one critical service as two or three SLI queries plus one business signal, using the same recording rules as your alerts.
- Check the scrape interval and rate windows on the services in scope, and shorten them where the guard would react in minutes.
- Add experiment markers: dashboard annotations, an active-experiment gauge, and a documented span attribute and baggage key.
- Build or adopt a controller with a pre-flight health check, a burn-rate guard, missing-data-as-failure and a revert in a finally block, and set a TTL on every fault.
- Run one small latency experiment with a control group, and write down both the system and the observability hypothesis first.
- Record time to detect, time to notify, routing and runbook accuracy for each critical alert, and re-run them on a schedule.
- Add an absent() or heartbeat rule for every alert whose expression can silently return no data.