On-call burnout rarely comes from one terrible incident. It comes from the steady drip: the disk alert that fires every Tuesday during backups, the CPU warning nobody acts on, the third duplicate page for an outage already being handled, the 3 a.m. alert that resolves itself before the laptop opens. Each one is small. Together they teach engineers that pages are noise, slow down the response to the page that matters, and make good people leave the rotation.
The fix is not a better paging tool; it is treating the alert set as a product with users (the people on call), a quality metric (how often a page needed a human) and a release process (every rule change goes through review). This article covers that process: a page budget, an outcome record for every page, per-rule precision, symptom-based paging rules, a CI lint, and the weekly review that deletes rules. The mechanics of routing, grouping and inhibition are in alerting architecture, in depth, and rotation design and handoffs are in on-call architecture; this article is about keeping the volume and quality of pages humane.
Why noisy alerting is a reliability problem, not a comfort problem
A page is an interrupt with a cost: the context switch during the day, the lost sleep at night, and the slower response to the next page because the responder has learned that most pages do not matter. That last cost is the dangerous one. If 80 percent of pages are noise, a rational responder triages before acting, and the true incident waits behind that habit. Alert fatigue is measured in minutes added to real incidents, not in grumbling.
There is also a retention cost that shows up months later. Experienced engineers opt out of the rotation, the rotation shrinks, each remaining person carries more shifts, and the noise per person rises. Teams that never measure page load usually discover it through an exit interview.
Set a page budget
A budget turns "too many pages" into a number you can fall short of or exceed. Pick it the way you pick an SLO: from what the humans can sustain, not from what the system currently produces. A common starting point is no more than two pages per 12-hour shift on average and no more than two pages per week during sleeping hours, with every page expected to need action. Write it down beside the rotation.
The budget does two things. It gives the weekly review a target, so a week with 30 pages is an incident of the alerting system itself and gets the same attention as a missed SLO. And it changes the default when someone proposes a new paging rule: the question becomes "what does this displace" rather than "could this ever be useful". A rule that would be nice to know about goes to a ticket queue, which has no budget problem because nobody gets woken for it.
Record the outcome of every page
You cannot improve precision you do not measure, and paging tools record when a page fired and who acknowledged it, not whether it mattered. Add one step to the end of every page: the responder records the outcome, in a form, a chat command, or a required field when resolving. Keep the vocabulary small enough that nobody argues about it:
| Outcome | Meaning | Counts as actionable |
|---|---|---|
| fixed | A human changed something and the symptom went away | Yes |
| mitigated | A human reduced impact; follow-up work remains | Yes |
| escalated | Needed another team, who acted | Yes |
| no_action | Real condition, but nothing a human needed to do | No |
| duplicate | Another page already covered this incident | No |
| false_positive | The condition was not real (bad data, broken exporter) | No |
Store the outcome with the alert name, fire time, acknowledgement and duration, from the paging tool's export or an Alertmanager webhook receiver that writes one row per notification. A month of this data is worth more than any amount of debate about which alerts are good, because it replaces opinion with a per-rule precision: actionable pages divided by pages.
The feedback loop
With outcomes recorded, the alert set becomes a closed loop: rules fire, pages carry outcomes into a history, the history produces a weekly report, the review turns the report into pull requests against the rules, and the budget tells you whether you are done. The loop is cheap to run, about thirty minutes a week, and it is the single habit that separates teams whose paging improves from teams whose paging only grows.
"""Weekly on-call report from a page log: one row per page, outcome filled in by the responder.
Columns: alert, fired_at (ISO, local time), acked, resolved_after_min, outcome
outcome is one of: fixed, mitigated, escalated, no_action, duplicate, false_positive
"""
import csv, collections
from datetime import datetime
ACTIONABLE = {"fixed", "mitigated", "escalated"}
rows = list(csv.DictReader(open("pages_last_7d.csv")))
per_rule = collections.defaultdict(lambda: collections.Counter())
for r in rows:
hour = datetime.fromisoformat(r["fired_at"]).hour
s = per_rule[r["alert"]]
s["pages"] += 1
s["actionable"] += r["outcome"] in ACTIONABLE
s["night"] += hour >= 22 or hour < 7
s["self_resolved"] += float(r["resolved_after_min"]) < 5 and r["outcome"] == "no_action"
total = sum(s["pages"] for s in per_rule.values())
night = sum(s["night"] for s in per_rule.values())
act = sum(s["actionable"] for s in per_rule.values())
print(f"pages={total} night={night} actionable={act / max(total, 1):.0%}")
print(f"{'rule':38} {'pages':>5} {'night':>5} {'precision':>9} {'self-res':>8}")
for name, s in sorted(per_rule.items(), key=lambda kv: -kv[1]["pages"]):
precision = s["actionable"] / s["pages"]
flag = " <- review" if precision < 0.5 or s["self_resolved"] > s["pages"] / 2 else ""
print(f"{name:38} {s['pages']:>5} {s['night']:>5} {precision:>9.0%} {s['self_resolved']:>8}{flag}")The report ranks rules by pages and flags two patterns: precision below 50 percent, and rules that mostly resolve on their own within five minutes with no action, which is the signature of a threshold without a sufficient for duration or a metric that flaps.
Page on symptoms, ticket on causes
Most noise comes from paging on causes: CPU above 80 percent, a pod restarting, a queue above some depth. Causes are real but often harmless, because the system absorbs them, and a single real incident produces many of them at once. Symptoms are what users experience: errors, latency, failed jobs, stale data. A symptom alert fires only when someone is affected, and one incident produces one symptom.
The strongest symptom alerts are tied to an error budget. A multi-window burn-rate rule pages when the service is consuming its 30-day budget fast enough to matter, and stays quiet for a brief spike that the budget can absorb; burn-rate alerting derives the windows and thresholds, and SLIs, SLOs and error budgets shows how to build the ratio. The rule file below pages on such a symptom and sends a real but non-urgent cause, a disk that will fill within a day, to a ticket.
groups:
- name: checkout-symptoms
rules:
- alert: CheckoutErrorBudgetFastBurn
expr: |
(sum(rate(http_requests_total{service="checkout",code=~"5.."}[1h]))
/ sum(rate(http_requests_total{service="checkout"}[1h]))) > (14.4 * 0.001)
and
(sum(rate(http_requests_total{service="checkout",code=~"5.."}[5m]))
/ sum(rate(http_requests_total{service="checkout"}[5m]))) > (14.4 * 0.001)
for: 2m
labels:
severity: page
team: payments
annotations:
summary: "Checkout is burning its 30-day error budget 14x too fast"
runbook_url: "https://runbooks.example.internal/checkout/error-budget-burn"
dashboard: "https://grafana.example.internal/d/checkout"
- alert: CheckoutDiskWillFillIn24h
expr: predict_linear(node_filesystem_avail_bytes{job="checkout-db"}[6h], 24 * 3600) < 0
for: 30m
labels:
severity: ticket # real, but nobody needs to wake up for it
team: payments
annotations:
summary: "Checkout DB disk projected to fill within 24h"
runbook_url: "https://runbooks.example.internal/checkout/disk"Causes still belong on dashboards and in tickets, where they help diagnosis. A few causes do deserve a page because they predict an outage that symptoms will only show when it is too late, such as a certificate expiring tomorrow or a disk filling within hours; give those generous lead time so they can be tickets during the day rather than pages at night.
Lint paging rules in CI
Rules live in git, so the cheapest protection is a check that runs on every change. Prometheus's promtool check rules validates syntax; the policy checks are yours. The script below enforces that every alert has an owning team and a known severity, that every paging rule has a runbook and a for duration, and that paging rules do not key on raw resource metrics. It will be wrong sometimes; when it is, the exception is a visible comment in review rather than a silent new page.
"""CI gate: every paging rule must be owned, documented and symptom-shaped."""
import sys, yaml, pathlib
REQUIRED_LABELS = {"severity", "team"}
SEVERITIES = {"page", "ticket", "info"}
errors = []
for path in pathlib.Path("alerts").rglob("*.yml"):
for group in yaml.safe_load(path.read_text())["groups"]:
for rule in group.get("rules", []):
if "alert" not in rule:
continue # recording rule
name, labels = rule["alert"], rule.get("labels", {})
notes = rule.get("annotations", {})
where = f"{path}:{name}"
if missing := REQUIRED_LABELS - labels.keys():
errors.append(f"{where}: missing labels {sorted(missing)}")
if labels.get("severity") not in SEVERITIES:
errors.append(f"{where}: severity must be one of {sorted(SEVERITIES)}")
if labels.get("severity") == "page":
if "runbook_url" not in notes:
errors.append(f"{where}: paging rule without runbook_url")
if "for" not in rule:
errors.append(f"{where}: paging rule without 'for' (pages on one bad sample)")
if any(w in rule["expr"] for w in ("node_cpu", "node_load", "container_memory")):
errors.append(f"{where}: paging on a resource cause; page on a user symptom")
print("\n".join(errors) or "alert lint ok")
sys.exit(1 if errors else 0)The runbook requirement matters more than it looks. A page whose responder has to work out from scratch what the alert means and what to do is slow and stressful; on-call runbooks covers what to put in them. If nobody can write a runbook for an alert, that is usually evidence it should not page.
The weekly review
Thirty minutes, the outgoing and incoming on-call plus whoever owns the noisiest rules. Go down the report from the top. For each flagged rule, choose one of four actions and open the pull request in the meeting:
- Delete. No outcome in the last month was actionable and nobody can name the incident it would catch. Deletion is the most underused action in alerting.
- Tune. The condition is real but the threshold or duration is wrong: lengthen
for, widen the window, or move to a burn-rate form. - Demote. Real and worth knowing, but not urgent: change severity to ticket.
- Fix the system. The alert is right and keeps firing because the underlying problem is real; file the engineering work, and give it priority, because it is costing sleep every week.
Track two numbers on a chart the team sees: pages per week against the budget, and overall precision. Also look at duplicates: if one incident routinely produces several pages, grouping or inhibition needs work, which is a routing change rather than a rule change.
Worked example: from 61 pages a week to 9
A payments team with a six-person rotation logs outcomes for two weeks and gets an average of 61 pages a week, 19 of them between 22:00 and 07:00, with 21 percent precision. The report puts five rules at the top, accounting for 48 pages:
| Rule | Pages/week | Precision | Decision | Pages after |
|---|---|---|---|---|
| HostCPUHigh (above 80% for 1m) | 17 | 0% | Delete; latency SLO covers user impact | 0 |
| DBDiskAbove85Percent | 11 | 9% | Replace with fill-in-24h prediction, as a ticket | 0 |
| CheckoutErrorRateAbove1Percent (no for) | 9 | 33% | Replace with multi-window burn-rate rule | 2 |
| PodRestarting | 6 | 0% | Demote to ticket; alert on availability instead | 0 |
| PaymentProviderTimeouts | 5 | 80% | Keep; add runbook link | 4 |
Two of the five were symptoms in the wrong shape, two were causes, and one was good. After those changes and inhibition of duplicate pages during a known provider outage, the next four weeks average 9 pages, 2 at night, at 78 percent precision, which is inside a budget of 2 per shift across 14 shifts. The team did not lose coverage: the burn-rate rule caught a real checkout regression in week three, 6 minutes after deploy, and that page was not buried among 60 others.
Trade-offs and failure modes
- Under-alerting. Cutting pages can remove the only signal for a real failure. Before deleting a rule, check that a symptom alert would have fired for every actionable page it produced; if not, write that symptom alert first.
- Outcome fatigue. If recording an outcome takes more than ten seconds, people stop. Make it one click or one command, and default nothing.
- Gaming precision. Labelling everything "fixed" makes the numbers lie. Review a sample of outcomes each week, and treat a too-clean report with suspicion.
- Slow symptoms. Burn-rate rules over long windows detect slow burns late by design. Keep a fast window for sharp outages and accept that a slow degradation surfaces as a ticket first.
- Ownerless rules. Alerts written by someone who has left page someone who cannot fix them. The team label plus a quarterly ownership check keeps every rule tied to a team that can act.
What to do next
- Write down a page budget for your rotation, with a separate limit for sleeping hours.
- Add a required outcome field to every page and keep a month of history.
- Run the weekly report; list the five rules with the most pages and their precision.
- Take each to a decision this week: delete, tune, demote or fix the system.
- Add the CI lint so every paging rule has an owner, runbook and for duration.
- Replace cause-based paging rules on your main services with a burn-rate symptom rule.
- Put pages per week and precision on a chart beside your SLOs, and review it weekly.