On-call is the agreement that someone will respond, within minutes and at any hour, when a production service stops serving its users. It is a system, made of alert rules, routing, a paging service, rotations, runbooks and the people in them. Like any system it can be designed well or badly. A badly designed one pages people for things that do not matter, misses the things that do, and burns out the team that carries the pager.

This article treats on-call as an architecture. It follows the path from a signal to a human, sets out rules for what is allowed to page, shows SLO burn-rate alerts and Alertmanager routing as real configuration, sizes a rotation with published load limits, and gives handoff and review practices that keep the system improving. Running an incident once it has been declared is covered in incident response architecture; this page is about everything that gets a prepared person to the keyboard.

Advertisement

What on-call is for

On-call has one job: to shorten the time between a user-visible problem starting and a competent person starting to fix it. Every design decision should be judged against that goal and against its cost, which is human attention, sleep and goodwill.

Two consequences follow. First, a page is expensive and must be reserved for problems that need a human now. Anything that can wait until morning goes to a ticket queue. Second, on-call is not a substitute for engineering. If the same alert fires every week, the fix is a code or capacity change, not a better runbook for the person who keeps getting woken up. Google's SRE book makes this explicit: it caps on-call at no more than 25% of an SRE's time and aims to spend at least 50% on engineering work.

The pipeline from signal to human

A page travels through several systems, and each one is a place it can be delayed, duplicated or lost. The diagram shows the common shape: metrics and probes feed alert rules; an alert router such as Prometheus Alertmanager groups, deduplicates and routes alerts; a paging service applies schedules and escalation policies and notifies the person on call; and if nobody acknowledges in time, the page escalates.

From signal to human: what may page, where it routes, and what happens when nobody answersMetrics, probesSLIs, synthetic checksLogs, tracesfor diagnosis, not pagingAlert rulesburn rate, symptomsAlertmanagergroup, dedupe, inhibitPaging serviceschedule, escalationTicket queuenext business dayPrimaryack within minutesSecondaryif no ackManagerfinal escalationRunbook + dashboardlinked from the pageHandoff + reviewload metrics, fixesWatchdogexternal dead-man checkpagetickettimeouttimeoutopenshumans query thesealways-firing alertpages if heartbeat stopsEvery arrow is a place a page can be lost: test the whole path, not each box.
Signals feed alert rules. The router sends urgent alerts to a paging service and the rest to a ticket queue. Unacknowledged pages escalate from primary to secondary to manager. An always-firing watchdog alert proves the path itself is alive.

The last box matters most and is the one most often missing. If the alerting system itself fails, nothing pages. The standard defence is a watchdog, sometimes called a dead-man's switch: an alert that is always firing, routed to an external service that pages if the heartbeat stops arriving. It must run outside the failure domain it watches, on different infrastructure and ideally with a different provider.

Advertisement

What is allowed to page

Write the rules down and apply them to every new alert. A good set has four tests:

  1. User-visible. The alert describes a symptom users experience, such as errors, latency or missing data, not a cause such as high CPU. Causes belong on dashboards, where the person paged will look for them.
  2. Urgent. If nobody acts in the next hour, real harm results. Otherwise it is a ticket.
  3. Actionable. The person paged can do something about it, and the alert links to a runbook that says what.
  4. Owned. It routes to the team that can fix it, not to a central team that will forward it.

Cause-based alerts are tempting because they fire earlier, but they are noisy: CPU can be high while users are fine, and users can be failing while CPU is normal. The exception is a resource that will cause an outage on a predictable schedule, such as a disk filling or a certificate expiring. Alert on those with enough lead time to make them tickets rather than pages. How to write the runbook each page links to is covered in runbook architecture.

SLO burn-rate alerts

The best single source of pages is the error budget. If a service has a 99.9% availability SLO over 30 days, it can fail 0.1% of requests in that period. The burn rate is how fast you are spending that budget: a burn rate of 1 uses exactly the whole budget in 30 days, and a burn rate of 14.4 uses 2% of it in one hour.

Alerting on a single short window is noisy, and a single long window is slow to fire and slow to clear. The SRE workbook's chapter on alerting on SLOs recommends multiwindow, multi-burn-rate alerts, and gives these starting parameters:

Budget consumedLong windowShort windowBurn rateAction
2%1 hour5 minutes14.4Page
5%6 hours30 minutes6Page
10%3 days6 hours1Ticket

Each alert fires only when both windows exceed the threshold. The long window shows the problem is significant; the short window shows it is still happening, so the alert resolves quickly after a fix. As Prometheus rules for a 99.9% SLO:

groups:
- name: checkout-slo
  rules:
  - record: job:slo_errors:ratio_rate5m
    expr: |
      sum(rate(http_requests_total{job="checkout",code=~"5.."}[5m]))
      / sum(rate(http_requests_total{job="checkout"}[5m]))
  # ...the same recording rule for 30m, 1h and 6h windows...
  - alert: CheckoutErrorBudgetFastBurn
    expr: |
      job:slo_errors:ratio_rate1h > (14.4 * 0.001)
      and job:slo_errors:ratio_rate5m > (14.4 * 0.001)
    labels:
      severity: page
      team: payments
    annotations:
      summary: "Checkout is burning error budget 14x too fast"
      runbook_url: "https://runbooks.internal/checkout/error-burn"

Low-traffic services need care: with a few requests a minute, one failure is a large fraction and the 5-minute window becomes noisy. Synthetic traffic, longer windows, or a minimum request count in the expression all help. Error-budget policy, the decisions that follow from a spent budget, is covered in error budgets.

Routing, grouping and escalation

The router turns a stream of alerts into a small number of notifications to the right people. Three features do most of the work: grouping, so 200 failing pods produce one page instead of 200; routing on labels, so a team's alerts reach that team's schedule; and inhibition, so a broad outage suppresses the narrower alerts it causes.

route:
  receiver: team-tickets
  group_by: [alertname, service]
  group_wait: 30s        # wait to batch alerts that start together
  group_interval: 5m     # minimum gap between updates for a group
  repeat_interval: 4h    # re-notify if still firing
  routes:
  - matchers: ['severity="page"', 'team="payments"']
    receiver: payments-pager
inhibit_rules:
- source_matchers: ['alertname="RegionDown"']
  target_matchers: ['severity="page"']
  equal: [region]
receivers:
- name: payments-pager
  pagerduty_configs:
  - routing_key: <from secret store>
- name: team-tickets
  webhook_configs:
  - url: https://tickets.internal/alertmanager

Escalation lives in the paging service. A typical policy notifies the primary immediately, escalates to the secondary if the page is not acknowledged within 5 to 15 minutes, and then to a manager or the whole team. Acknowledging stops escalation but does not resolve the alert. Keep the policy as configuration in version control alongside the alert rules, so changes are reviewed like code.

Rotation design and sizing

A rotation has a shape, a shift length and a size. Common shapes are a single-site rotation, where the same people cover nights, and follow-the-sun, where two or more sites in different time zones each cover their daytime. Most teams run a primary and a secondary. Shifts of one week are common because they limit handoffs; 12-hour shifts are used with follow-the-sun.

The SRE book gives two limits that make sizing a calculation. On load: dealing with an incident, including root-cause analysis, remediation and follow-up, takes about six hours, so a 12-hour shift should average no more than two incidents. On time: with the 25% cap, a single-site team with week-long primary and secondary shifts needs at least eight engineers, and a dual-site team needs at least six per site.

Worked example. A payments team of five runs a single-site, week-long primary and secondary rotation. Each engineer is on call two weeks in five, or 40% of their time, well over the cap. Last month the rotation received 46 pages, 30 of them outside working hours, and 19 were the same flapping disk-latency alert. The plan:

  1. Delete the disk-latency page and replace it with a capacity ticket; that removes 41% of the load at once.
  2. Replace the remaining cause-based pages with the two burn-rate alerts above.
  3. Join the rotation with a sister team of four that owns related services, giving nine engineers and on-call time of 2 in 9, about 22%.
  4. Re-measure after a month: target fewer than two incidents per 12-hour shift and fewer than one night page per week.

Merging rotations only works if everyone can act on every page, so it comes with shared runbooks and a training rotation where new members shadow a primary before carrying the pager alone.

Shift mechanics: handoff and review

A handoff turns one person's context into the next person's starting point. Make it a short written note plus a 15-minute conversation, using the same template each time:

On-call handoff: payments, week 40
Open incidents:   none
Ongoing issues:   checkout p99 elevated since Tue deploy; ticket PAY-812
Pages this week:  5 (2 at night); 1 false positive -> alert PR #431
Risky changes:    ledger schema migration Thu 14:00, owner: Priya
Silences active:  disk alerts on db-7 until Fri (hardware swap)
Asks:             review runbook for card-network timeouts

Hold a weekly on-call review where the outgoing primary walks through every page. For each one, decide: was it actionable, was the runbook right, and what change stops it recurring? Track the answers. The review is where on-call improves; without it, the same pages repeat indefinitely.

Measure load explicitly: pages per shift, night pages per shift, time to acknowledge, the fraction of pages that were actionable, and the share of pages from the top three alerts. Publish the numbers. They are the evidence a team needs to get engineering time for reliability work, and they show early when a rotation is becoming unsustainable.

Failure modes and trade-offs

  • Alert fatigue. Too many non-actionable pages train people to acknowledge without looking, so the real page gets the same treatment.
  • Silent pipeline. The alerting system, paging integration or a phone's do-not-disturb setting fails, and nothing pages. Use a watchdog and send a test page at every handoff.
  • Stale silences. A silence set during maintenance outlives it and hides a real outage. Give every silence an expiry and an owner.
  • Hero culture. One expert answers everything, others never learn, and the rotation collapses when that person leaves. Escalate to them, but let the primary lead.
  • Orphaned alerts. Alerts routed to a central team that cannot fix the service lose time to forwarding.
  • No recovery time. People who were up at 3 a.m. are expected at a 9 a.m. meeting. Give time off after night pages, and compensate on-call with pay or time in lieu.

Follow-the-sun removes most night pages but adds handoffs and needs enough people in each site. Primary-plus-secondary doubles the people on call at any time but makes missed pages rare. Week-long shifts reduce handoffs but make bad weeks longer. Choose deliberately and revisit with the load numbers. Rehearsing the system before you need it is covered in game days.

What to do next

  1. Export the last 90 days of pages and compute pages per shift, night pages, time to acknowledge and the share from the top three alerts.
  2. Apply the four paging tests to every alert; demote failures to tickets or dashboards, and delete the rest.
  3. Define an SLO for each critical user journey and add the two paging burn-rate alerts and the ticket alert.
  4. Make sure every page links to a runbook and routes to the owning team; put alert rules, routing and escalation policy in version control.
  5. Add an external watchdog and send a test page through the full path at each handoff.
  6. Size the rotation against the 25% cap and two incidents per 12-hour shift; merge rotations or move to follow-the-sun if the numbers do not work.
  7. Start a written handoff and a weekly review, and agree on compensation and recovery time with management.
  8. Read running an SLO program to set the objectives your pages depend on.
Key takeaway: On-call is a system: alert rules, routing, a paging service, rotations, runbooks and people. Page only for urgent, user-visible, actionable problems that a named team owns, preferably through multiwindow burn-rate alerts on SLOs. Group and inhibit alerts, escalate unacknowledged pages, and watch the paging path itself with an external watchdog. Size rotations against published limits, about two incidents per 12-hour shift and no more than 25% of time on call, and use handoffs, weekly reviews and load metrics to make every bad week the last of its kind.