Alert fatigue is what happens when responders receive so many alerts that do not need them that they stop treating alerts as information. They acknowledge without looking, mute channels, and develop a private sense of which alerts are real. The danger is not the annoyance; it is that a genuine incident arrives looking exactly like the hundred false ones before it, and gets the same reflexive acknowledgement.
Most advice on the subject is about policy: page only on symptoms, set a page budget, review alerts weekly. The related articles linked below cover that. This article is about finding out precisely which alerts are noisy and why, using data you already have. Every alerting system records when each alert started and stopped. From those timestamps alone you can compute six noise signatures, and each signature points at a specific, mechanical fix. You can run the analysis this afternoon, without waiting for anyone to label pages as useful or useless.
Why fatigue is a detection failure
Think of every page as a test with a false-positive rate. If a responder receives forty pages a week and two correspond to real problems, the rational prior for any new page is that it is probably noise. People adapt to that prior; they cannot help it. The measurable symptoms are rising time-to-acknowledge, acknowledgements that arrive within seconds (too fast for anyone to have looked), alerts left firing for days, and private silences that outlive the problem they were created for.
So the goal is raising the probability that a page means something (precision) while keeping the probability that a real incident pages someone (recall). Every fix below is judged by both.
Build the episode table
The unit of analysis is an episode: one continuous firing interval of one alert instance, where an instance is a rule plus its full label set. You can assemble episodes from any of three places. A Prometheus server exposes the synthetic ALERTS series with an alertstate label, so a range query over it yields start and end times. An Alertmanager webhook receiver can log every firing and resolved notification with its fingerprint. Your paging tool records when each incident was triggered, acknowledged and resolved, which adds the human dimension.
Join them into one row per episode: rule name, a fingerprint of the labels, start, end and, if it paged, the acknowledgement time. Thirty days is usually enough. Keep the raw table; the before and after comparison is the evidence that a fix worked.
from collections import defaultdict
from itertools import combinations
# episodes: one row per firing interval of one alert instance (rule + label set)
# {"rule": str, "fp": str, "start": float, "end": float, "acked_at": float | None}
def signatures(episodes, flap_gap_s=600, fast_s=300, bucket_s=300):
by_rule = defaultdict(list)
for e in episodes:
by_rule[e["rule"]].append(e)
report = {}
for rule, eps in by_rule.items():
eps.sort(key=lambda e: (e["fp"], e["start"]))
refires = sum(1 for a, b in zip(eps, eps[1:])
if a["fp"] == b["fp"] and b["start"] - a["end"] < flap_gap_s)
fast = sum(1 for e in eps
if e["end"] - e["start"] < fast_s and e["acked_at"] is None)
report[rule] = {"episodes": len(eps),
"flap_rate": refires / len(eps),
"resolve_before_ack": fast / len(eps)}
# co-firing: Jaccard similarity of the 5-minute buckets each rule was firing in
buckets = defaultdict(set)
for e in episodes:
for b in range(int(e["start"] // bucket_s), int(e["end"] // bucket_s) + 1):
buckets[e["rule"]].add(b)
pairs = []
for r1, r2 in combinations(buckets, 2):
inter = len(buckets[r1] & buckets[r2])
if inter:
pairs.append((inter / len(buckets[r1] | buckets[r2]), r1, r2))
return report, sorted(pairs, reverse=True)[:20]The function computes, per rule, how often an instance re-fires shortly after resolving and how often an episode ends on its own before anyone acknowledges it, plus a list of rule pairs that tend to fire in the same time buckets. The six patterns below come from these measurements plus two simple joins on the same table: hour-of-day concentration and overlap with deploy windows. None of this needs anyone's opinion. Sort rules by episode count, read the signatures for the top twenty, and match each to one of the patterns below.
Signature 1: flapping
A rule flaps when its condition hovers around the threshold, so the alert fires, resolves and fires again for the same instance within minutes. A flap rate above roughly 0.3, meaning three in ten episodes are re-fires of something that just resolved, is a strong sign. Each re-fire can notify again, so one wobbling metric produces a stream of pages for one underlying condition.
The fixes all work at the rule level. Smooth the input by using a longer rate window or avg_over_time, so momentary spikes do not cross the line. Add a for duration so the condition must hold on every evaluation for a period before the alert fires. And add keep_firing_for, which keeps the alert firing for a set time after the condition was last true, so a brief dip below the threshold does not resolve it and start a new episode.
groups:
- name: checkout
rules:
- alert: CheckoutLatencyHigh
# smooth the signal, then require it to persist before firing
expr: |
histogram_quantile(0.99,
sum by (le) (rate(http_request_duration_seconds_bucket{service="checkout"}[10m]))
) > 0.8
for: 10m # must be true on every evaluation for 10 minutes
keep_firing_for: 15m # stay firing for 15 minutes after the last true evaluation
labels:
severity: pageThe cost is detection delay: this rule now needs ten minutes of sustained breach to fire. Choose the duration from the impact model; if users can tolerate ten minutes of slow checkout before you must act, a ten-minute for loses nothing. Prometheus has no built-in hysteresis with separate fire and clear thresholds, so keep_firing_for is the simplest substitute.
Signature 2: resolves before anyone acts
Some alerts fire, page someone and resolve by themselves within a few minutes, before an acknowledgement. If more than half of a rule's episodes look like that, the rule is describing transient conditions the system already recovers from: a retry storm that drains, a pod restart that succeeds, a brief replication lag. The page woke someone for nothing they could have done.
There are two fixes, and the choice depends on whether the transient matters in aggregate. If it does not, lengthen the for duration past the typical self-recovery time, which the episode table tells you directly. If it does, because a pod restarting forty times a day signals a real defect even though each restart recovers, change the alert's severity so it opens a ticket or posts to a channel instead of paging, and alert on the rate of occurrences over a day rather than on each one. Routing by severity label is covered in the routing article linked below; the point here is that the signature tells you which rules need it.
Signature 3: the same incident from several sources
Large organisations accumulate overlapping monitors. A database outage can page through a Prometheus rule, a cloud provider's managed alarm, a synthetic checker and an application error tracker, all for one event. Co-firing analysis exposes these: two rules from different sources with a Jaccard similarity above about 0.7 over their firing buckets almost always describe the same thing.
The fix is at the source. For each user-facing symptom, choose one owner signal that pages, usually the one closest to the user, such as an SLO burn on the service's own request metrics. Demote the others to context: they still fire and appear on the incident, but as information, not as separate pages. Where several tools must feed the pager, give them a shared deduplication key derived from the affected service, so the pager merges them into one incident. Many paging tools accept such a key on incoming events; check how yours does it.
Signature 4: storms and cascades
A storm is many different rules firing within minutes of each other because they share a cause. When a database primary fails, every service that depends on it raises error-rate, latency and queue-depth alerts. Co-firing pairs that cluster around one shared dependency, rather than one service, are the signature.
Two fixes apply. First, group notifications so one message lists the affected alerts instead of sending one each; Alertmanager does this with grouping labels and timers. Second, suppress the downstream alerts while the upstream cause is known to be firing, which Alertmanager calls inhibition:
inhibit_rules:
- source_matchers:
- alertname = "DatabasePrimaryDown"
target_matchers:
- severity = "page"
- depends_on = "orders-db"
equal: [cluster]This requires every dependent alert to carry a label naming the dependency, which is worth adding systematically from a service catalogue rather than by hand. Be conservative: inhibit only alerts whose impact is fully explained by the source alert, and keep the inhibited alerts visible in the incident view so the responder can see the blast radius.
Signature 5: ratios at low traffic
Error-ratio alerts behave badly when traffic is low. At three in the morning a service might handle ten requests in a five-minute window; one failure is a ten percent error ratio and fires an alert tuned for daytime volume. In the episode table these show up as episodes concentrated in low-traffic hours with short durations.
The fix is a minimum-evidence guard: fire only if the ratio is high and the absolute number of failures is large enough to matter. The and operator in PromQL keeps results from the left side only where the right side also has a result:
- alert: CheckoutErrorRatioHigh
expr: |
(
sum(rate(http_requests_total{service="checkout", code=~"5.."}[15m]))
/
sum(rate(http_requests_total{service="checkout"}[15m]))
) > 0.05
and
sum(increase(http_requests_total{service="checkout", code=~"5.."}[15m])) >= 20
for: 5mPick the floor from the impact: twenty failed checkouts in fifteen minutes is worth waking someone; two is not. For services where low traffic is normal, consider alerting on SLO burn rate over longer windows, which naturally averages over sparse periods, or add synthetic traffic so the denominator never collapses.
Signature 6: absent and stale data
The last signature cuts both ways. Rules that compare a metric against a threshold silently stop evaluating when the metric disappears, because there is nothing to compare, so a dead exporter looks like a healthy system. Teams then add absence alerts, and those often become the noisiest rules of all, firing on every deploy, scrape timeout and target relabel.
Treat data presence as its own concern. Use absent_over_time or an up == 0 check with a for long enough to cover routine restarts, aggregate it per job rather than per instance, and route it to the team that owns the monitoring pipeline rather than to the service on-call. In the episode table, absence alerts that coincide with deploy windows are a clear sign the duration is too short.
Worked example: one month of a checkout platform
The following numbers are illustrative but typical of what the analysis turns up. A checkout platform pages 212 times in thirty days across 38 rules. The episode table shows that five rules account for 71 percent of pages:
| Rule | Pages | Signature | Fix | Pages after |
|---|---|---|---|---|
| CheckoutLatencyHigh | 48 | flap rate 0.46 | 10m window, for 10m, keep_firing_for 15m | 6 |
| PodRestarting | 39 | 81% resolve before ack | ticket severity, daily rate alert | 0 (4 tickets) |
| OrdersDbConnErrors | 26 | co-fires with RDS alarm, Jaccard 0.83 | demote cloud alarm to context | 9 |
| PaymentErrorRatio | 22 | 18 of 22 between 01:00 and 05:00 | absolute floor of 20 failures | 3 |
| ExporterDown | 15 | 13 of 15 in deploy windows | for 15m, route to platform team | 2 (to platform) |
The five changes take two days of work. Pages from those rules fall from 150 to 20, 18 of them to the checkout rotation, and total pages from 212 to 82. The check that matters comes next: the team replays the month's two real incidents against the new rules and confirms both would still have paged, one of them four minutes later than before. That delay is accepted explicitly and written into the change description.
Trade-offs and failure modes
- Detection delay. Every
forduration and floor delays or hides some true positives. Replay past incidents against new rules, and record the delay you accepted. - Over-inhibition. A broad inhibit rule can silence the alert that would have revealed a second, unrelated failure. Scope inhibitions with
equallabels and keep suppressed alerts visible. - Demoted alerts nobody reads. Moving noise to a channel or ticket queue only helps if someone triages it. Give that queue an owner and a review cadence, or delete the rule.
- Anomaly detection as a cure. Replacing thresholds with learned baselines often adds noise, because seasonal patterns, deploys and traffic shifts look anomalous. Use it for investigation dashboards before you let it page.
- Silences as a fix. Long silences hide both noise and real problems. Treat every silence older than a day as a defect in a rule.
What to do next
- Export thirty days of alert episodes from your rule engine, router or pager into one table.
- Run the signature function and list the twenty rules with the most episodes alongside their flap rate, resolve-before-ack ratio and top co-firing partner.
- Fix the top five by signature: smoothing and durations for flapping, severity changes for self-resolving alerts, owner signals for duplicates, inhibition for cascades, floors for low-traffic ratios.
- Replay recent real incidents against each changed rule and record any added detection delay.
- Add dependency labels from your service catalogue so inhibition can be scoped precisely.
- Re-run the analysis weekly and track total pages, top-five share and the five signature rates over time.
Related reading: alerting that does not burn out on-call for page budgets, outcome records and the weekly review, alerting architecture for routing, grouping, inhibition and timers in detail, burn-rate SLO alerting for symptom alerts that tolerate low traffic, SLIs, SLOs and error budgets for choosing the signal that should own a page, and on-call runbooks for making the remaining pages actionable.