An alert is a promise: when this condition is true, the right person will hear about it, once, with enough context to act. Keeping that promise takes a pipeline with several moving parts, and most alerting pain (pages at 3 a.m. for nothing, a storm of 400 notifications for one outage, or worse, silence during a real incident) comes from misunderstanding one of them. This article walks the pipeline end to end using Prometheus and Alertmanager as the concrete example, because their design is widely copied and their configuration makes every stage explicit.
What should page and how to express service-level objectives as burn-rate conditions are separate questions, covered in on-call architecture and burn-rate alerting. Here the focus is the machinery: how rules are evaluated, how alerts are routed, grouped, muted and deduplicated, how the system stays up when a component fails, and how you test all of it before production does.
The pipeline in one picture
There are two separated responsibilities. The rule evaluator (Prometheus itself, or the ruler component in horizontally scaled systems) decides whether a condition holds. Alertmanager decides who hears about it, when and how often. Keeping them apart is what lets you change on-call routing without touching detection logic, silence a noisy alert during maintenance without editing rules, and run several evaluators feeding one notification layer.
The contract between them is simple: the evaluator periodically sends the full set of currently active alerts, each a set of labels plus annotations and timestamps. Alertmanager identifies an alert by its labels, so labels are the API. A label you add on a rule (team, severity) is something routing can match on; a value that changes on every evaluation, such as a number formatted into a label, creates a new alert identity each time and breaks deduplication. Put changing values in annotations.
Rule evaluation: pending, firing and resolved
Each rule group is evaluated on an interval (the global evaluation_interval or a per-group interval). For every series the expression returns, the rule creates an alert instance. A new instance starts in pending; if it remains in the result for the for duration it becomes firing and is sent to Alertmanager. When it drops out of the result, it is resolved, and the evaluator reports the resolution so receivers can be told.
groups:
- name: checkout-slo
interval: 30s
rules:
- alert: CheckoutHighErrorBudgetBurn
expr: |
(sum(rate(http_requests_total{service="checkout",code=~"5.."}[5m]))
/ sum(rate(http_requests_total{service="checkout"}[5m]))) > (14.4 * 0.001)
and
(sum(rate(http_requests_total{service="checkout",code=~"5.."}[1h]))
/ sum(rate(http_requests_total{service="checkout"}[1h]))) > (14.4 * 0.001)
for: 2m # must stay true across evaluations before firing
keep_firing_for: 10m # do not resolve on one good scrape; stops flapping
labels:
severity: page
team: payments
annotations:
summary: "Checkout is burning its 30-day error budget fast"
runbook_url: "https://runbooks.internal/checkout/high-burn"
dashboard: "https://grafana.internal/d/checkout"
- alert: Watchdog
expr: vector(1) # always firing; its absence is the alert
labels:
severity: noneTwo fields shape behaviour over time. for suppresses blips: a condition has to hold across evaluations before anyone is disturbed, at the cost of detection delay. keep_firing_for is the mirror image: according to the Prometheus documentation it keeps the alert firing for the given duration after the condition was last met, which prevents flapping and false resolutions caused by missing data. Without it, a single missed scrape during an outage can resolve and re-fire the alert, sending two pages and confusing the incident timeline.
Prometheus also writes a synthetic ALERTS series with alertname and alertstate labels for every pending or firing alert. Recording it lets you build dashboards of alert history and, more usefully, measure your own alerting: how often each alert fires, how long it stays firing, and which alerts fire without anyone acting. Expensive expressions belong in recording rules so the alerting rule stays cheap and evaluates on time.
Routing: a tree, first match wins
Alertmanager's configuration is a routing tree. Each incoming alert enters at the root and descends into the first child route whose matchers it satisfies; it keeps descending as deep as matches go, and the deepest matching node's receiver handles it. Setting continue: true on a route lets matching proceed to later siblings too, which is how one alert can go to both a pager and a team channel. Child routes inherit grouping and timing settings from their parent unless they override them.
route:
receiver: default-ticket
group_by: [alertname, cluster, service]
group_wait: 30s # collect the first burst before the first notification
group_interval: 5m # wait before notifying about NEW alerts in an existing group
repeat_interval: 4h # re-send an unchanged, still-firing group
routes:
- receiver: watchdog-heartbeat
matchers: [alertname="Watchdog"]
repeat_interval: 1m
- receiver: payments-pager
matchers: [team="payments", severity="page"]
continue: true # also let the next sibling route see it
- receiver: payments-chat
matchers: [team="payments"]
- receiver: infra-pager
matchers: [severity="page"]
mute_time_intervals: [planned-maintenance] # defined under time_intervals:
inhibit_rules:
- source_matchers: [alertname="ClusterUnreachable"]
target_matchers: [severity=~"page|ticket"]
equal: [cluster] # only mute alerts from the SAME clusterOrder is significant and a frequent source of silent bugs. In the example, a paging payments alert reaches the payments pager and, because of continue, the payments chat route. Without continue it would stop at the pager. A paging alert from any other team falls through to infra-pager, and anything that matches nothing lands on the root's ticket receiver, so there is always a destination. Routes can also carry mute_time_intervals and active_time_intervals for maintenance windows and business-hours-only receivers.
Grouping and the three timers
Grouping is what turns an outage into one notification instead of hundreds. Alerts that share the values of the route's group_by labels form one aggregation group, and the group is notified as a unit. The special value group_by: ['...'] groups by all labels, which effectively disables aggregation.
Three timers then control when notifications go out; the documented defaults are 30 seconds, 5 minutes and 4 hours.
| Timer | Default | Meaning | Tune it when |
|---|---|---|---|
group_wait | 30s | Delay before the first notification for a new group, to collect the initial burst | You need faster first pages (lower) or bursts arrive slowly (higher) |
group_interval | 5m | Delay before notifying about new alerts added to a group that already notified | Follow-up notifications feel too chatty or too slow |
repeat_interval | 4h | Re-send an unchanged group that is still firing | Long-running known issues re-page too often |
Worked example. A node failure takes out 40 pods across 6 services in one cluster. Each service has a high-error alert. With group_by: [alertname, cluster, service] you get six groups and, after about 30 seconds, six notifications, one per affected service. Grouping by [alertname, cluster] instead yields one notification listing six services, which is better for the person paged and worse if each service has a different owner who must be routed separately. Pick group labels that match how responsibility is divided. When a seventh service starts failing two minutes later, it joins its group and is announced at the next group_interval tick, not immediately.
Muting: inhibition and silences
Alertmanager has two mechanisms for suppressing notifications for alerts that are genuinely firing, and they solve different problems.
Inhibition is automatic and rule-based: while an alert matching source_matchers fires, alerts matching target_matchers are muted, provided the labels listed in equal have the same values on both. In the example, a ClusterUnreachable alert mutes paging and ticket alerts from the same cluster, because 40 symptoms of one root cause help nobody. The equal list is essential; without it, one unreachable cluster would mute alerts from every cluster in the company.
Silences are manual and time-boxed: a person creates a set of matchers with an expiry, a creator and a comment, typically for planned maintenance or a known issue being worked on. Silences are shared across the Alertmanager cluster. Require a comment that links a ticket, keep durations short, and review silences that are about to expire, because an expired silence on an unresolved issue re-pages someone who has no context.
Muted alerts are still visible in Alertmanager's interface and API, which matters during an incident: responders should see what was suppressed, not just what notified.
High availability without load balancing
Alerting must survive the failures it reports, so both halves run redundantly. For the evaluator, run two identical Prometheus replicas with the same rules; both will fire the same alerts. For Alertmanager, run a cluster of typically three peers that gossip their silences and their notification log, the record of which group was notified to which receiver and when.
The non-obvious rule is that each evaluator must send every alert to every Alertmanager peer directly, not through a load balancer; the Prometheus documentation is explicit on this. Deduplication works because every peer holds the same alerts and they agree, via the notification log and a staggered wait based on each peer's position in the cluster, that only one of them sends a given notification. Behind a load balancer, each peer sees a different subset of alerts, groups differ between peers, and you get duplicates or gaps. The design deliberately prefers an occasional duplicate notification during a network partition to a missed one.
The alert identity also has to match across evaluator replicas. If each replica adds a distinct external label such as replica, drop it before sending alerts (alert relabelling) or the two copies become two different alerts and everyone gets paged twice.
Who watches the watchers
A broken alerting pipeline looks exactly like a healthy, quiet system. The standard defence is the watchdog, or dead man's switch: an alert whose expression is vector(1), so it always fires. It is routed to an external heartbeat service with a short repeat_interval, and that service pages through an independent path when the heartbeat stops. One alert thereby tests rule evaluation, delivery to Alertmanager, routing and outbound notification together.
Add meta-monitoring of the components themselves: rule evaluation failures and evaluations that take longer than their interval, alerts the evaluator failed to send, Alertmanager notification failures by integration, and cluster membership size. Alert on these to a channel the main pipeline does not depend on.
Testing rules and routes in CI
Alert rules are code that only runs during emergencies, which makes them exactly the code that rots. Test them like code. promtool test rules replays synthetic series through your rule files and asserts which alerts are firing at given times; amtool validates the Alertmanager configuration and shows which receivers a given label set would reach.
# checkout_test.yml, run with: promtool test rules checkout_test.yml
rule_files: [checkout.rules.yml]
evaluation_interval: 30s
tests:
- interval: 30s
input_series:
- series: 'http_requests_total{service="checkout",code="500"}'
values: '0+30x200' # 1 error/s
- series: 'http_requests_total{service="checkout",code="200"}'
values: '0+30x200' # 1 ok/s -> 50% errors
alert_rule_test:
- eval_time: 1m
alertname: CheckoutHighErrorBudgetBurn
exp_alerts: [] # still pending: "for: 2m" not yet satisfied
- eval_time: 70m
alertname: CheckoutHighErrorBudgetBurn
exp_alerts:
- exp_labels: {severity: page, team: payments}
exp_annotations: # annotations are compared too
summary: "Checkout is burning its 30-day error budget fast"
runbook_url: "https://runbooks.internal/checkout/high-burn"
dashboard: "https://grafana.internal/d/checkout"
# Routing is tested too:
# amtool check-config alertmanager.yml
# amtool config routes test --config.file=alertmanager.yml team=payments severity=page
# -> payments-pager,payments-chatRun both in CI on every change to rules or routing, alongside promtool check rules for syntax. Add a lint that every rule with a paging severity carries a runbook_url annotation and a team label, because an alert nobody can route or act on is noise by construction. The runbook architecture article covers what the linked page should contain.
Failure modes and trade-offs
| Failure | Cause | Fix |
|---|---|---|
| Duplicate pages | Load balancer in front of Alertmanager, or a replica label on alerts | Send to all peers; drop replica labels in alert relabelling |
| Notification storm | Over-specific group_by; no inhibition for root causes | Group by ownership labels; inhibit symptoms with equal |
| Flapping resolve and re-fire | Gaps in data; no keep_firing_for | Add keep_firing_for; alert on absent data separately |
| Alert never fires | Expression returns nothing when the target is down | Pair with an absent() or up == 0 rule |
| Wrong team paged | Route order or a missing continue | Test routes with amtool in CI |
| Silent pipeline | Evaluator or notifier broken | Watchdog to an external heartbeat |
The main trade-off is speed against noise. Shorter for and group_wait detect sooner and wake people for blips; longer ones are calmer and slower. Resolve it per severity: pages need high precision and can tolerate a minute or two of delay; tickets can be both slow and sensitive.
What to do next
- Audit every alerting rule for a stable label set, a
severityandteamlabel, and a runbook link in annotations. - Add
keep_firing_forto paging alerts that flap, and an absence rule for each critical target. - Draw your routing tree, check route order and
continue, and make sure the root route has a catch-all receiver. - Choose
group_bylabels that match ownership, and add inhibition rules withequalfor known root causes. - Point every evaluator at every Alertmanager peer, remove any load balancer between them, and drop replica labels from alerts.
- Deploy a watchdog alert to an external heartbeat, and run promtool and amtool tests in CI on every rule or routing change.