Most writing about on-call is about the system around the responder: which alerts may page, how burn-rate alerts work, how rotations are sized. That design matters, and the on-call architecture guide covers it. This page is about the other side: what one person does from the moment their shift starts to the moment they hand it over, and what a team does so that person can do it well at three in the morning.
The core idea is simple and easy to forget under pressure. On call, your job is to restore service, not to understand the failure. Understanding comes later, with daylight, colleagues and a postmortem. Everything below follows from that: prepare so the first minutes are mechanical, mitigate with reversible actions before diagnosing, escalate early, communicate on a schedule and leave a clean trail for whoever comes next.
What on-call asks of one person
During a shift you are the person who answers for a set of services when automation cannot. That means being reachable within the response time your team agreed, having the access and tools to act, knowing where the runbooks are and having the authority to take the mitigating actions they describe without asking permission. If any of those is missing, the rotation is asking you to carry risk you cannot discharge, and the fix belongs to the team rather than to your willpower.
It also means an explicit split of duties. While a page is open you respond; between pages you usually have lighter duties such as triaging tickets, tidying alerts that fired without need and improving runbooks. Interrupt work is real work: plan the sprint knowing the on-call person will not deliver feature work at full speed.
Before the shift: a readiness check
Most slow responses start before the page: an expired VPN certificate, a laptop that needs an update, a production credential that was never granted. Catch them at the start of the shift with a script rather than a memory. The script below is deliberately boring; every FAIL is something you would otherwise discover during an incident.
#!/usr/bin/env bash
# oncall-preflight.sh: run at the start of every shift; fix every FAIL now, not during a page
set -u
fail=0
check() {
if eval "$2" >/dev/null 2>&1; then echo "ok $1"; else echo "FAIL $1"; fail=1; fi
}
check "VPN reaches internal" "curl -sf --max-time 5 https://internal.example.com/healthz"
check "prod cluster access" "kubectl --context prod get ns kube-system"
check "cloud credentials" "aws sts get-caller-identity"
check "dashboards reachable" "curl -sf --max-time 5 https://grafana.example.com/api/health"
check "runbooks up to date" "git -C ~/runbooks pull --ff-only"
check "feature flag CLI login" "flags whoami"
exit $failPair the script with a few human checks: send yourself a test page and confirm the phone makes noise with do-not-disturb on, read the previous shift's log, check the deploy calendar for risky changes and confirm who your secondary and escalation contacts are this week. If a change freeze, a launch or a migration overlaps your shift, find out now who owns it.
The loop for one page
Every page runs through the same loop. Acknowledge first, because an unacknowledged page escalates and pulls in someone else. Assess impact in user terms: which users, which functions, how badly, since when. Decide whether this is an incident, which in most teams means opening a channel and assigning roles; when unsure, declare, because closing an unneeded incident costs little and running a real one informally costs a lot. Then mitigate, verify on the user-facing indicator that the symptom is gone, communicate, and leave a record. The guide to running an incident covers the roles once more than one person is involved.
The first fifteen minutes
- Minute 0 to 1. Acknowledge. Open the alert and its linked runbook and dashboard. Note the time in a scratch file or the channel.
- Minute 1 to 5. Confirm the symptom is real on the user-facing indicator, not only on the alerting metric. Check scope: one region or all, one customer or many, one endpoint or the whole service. Look at the change log for the last few hours: deploys, config pushes, feature flags, infrastructure changes and upstream provider status.
- Minute 5. Decide whether to declare. If users are affected and the cause is not obviously already resolving, declare, open the channel and post the first update, even if it only says what you are seeing and when the next update will come.
- Minute 5 to 15. Apply the first plausible reversible mitigation from the ladder below. If nothing on the ladder applies, or you have tried one step and the symptom persists, escalate. Do not spend these minutes reading code.
The fifteen-minute frame is a time-box, not a target from any standard. Its purpose is to stop the most common failure of a lone responder: twenty minutes of quiet investigation while users stay broken and nobody else knows.
Mitigate before you diagnose
Mitigations differ in how much you need to know before using them and how easy they are to undo. Prefer the ones you can apply with little knowledge and undo quickly. Each rung of the ladder below is a question to ask in order.
| Rung | Question | Typical action | Undo |
|---|---|---|---|
| 1 | Did something change just before it started? | Roll back the deploy, revert the config, turn off the flag. | Roll forward again. |
| 2 | Is the damage limited to one zone, region or cell? | Drain or fail over away from it. | Restore traffic. |
| 3 | Is it load-related? | Scale out, shed low-priority traffic, enable rate limits. | Scale back, lift limits. |
| 4 | Is one dependency failing? | Switch it to a degraded mode, serve cache, disable the feature that calls it. | Re-enable. |
| 5 | Is a single bad input or tenant driving it? | Block or isolate that input or tenant. | Unblock. |
| 6 | None of the above? | Escalate and keep gathering evidence. | Nothing to undo. |
Rollback deserves its first place. A large share of incidents begin shortly after a change, and reverting is usually safe even when you are not sure the change caused the problem, provided the team keeps rollbacks cheap and routine. If your service cannot be rolled back quickly, that is a reliability gap to fix after the shift, not a reason to skip the rung. Write down every action with its time, including ones that did not help; the postmortem needs them.
Escalating without hesitation
Escalation is part of the loop, not a confession. Escalate when the impact is large, when you have tried the obvious mitigation without success, when the problem is in a system you do not own, or when you notice you are tired or stuck. A good escalation message lets the other person act without a call to understand it:
Paging you for: checkout latency, prod eu-west, started 02:31 UTC.
Impact: p99 checkout latency far above target; some payments timing out.
Tried: rolled back 02:29 payments-client config (02:44) -> no change.
Seen: errors concentrated on calls to the card-auth provider; provider status page green.
Need: someone who knows the card-auth integration. Channel: #inc-2026-10-03-checkoutThe same structure works for status updates to stakeholders: what is affected, what is being done, and when the next update will come. Post updates on the promised schedule even when nothing has changed; silence reads as nobody working on it.
Worked example: a page at 02:40
A checkout latency alert pages at 02:40. The responder acknowledges and opens the linked dashboard: the user-facing p99 latency for checkout rose sharply at 02:31, only in one region, and the error rate on payment calls is climbing. The change log shows a config push to the payments client at 02:29. That is rung one, so at 02:44 they revert the config and post a declaration in a new incident channel.
By 02:50 the symptom is unchanged. The rollback did not help, so the config push was probably a coincidence. Rung two applies, since the problem is regional: the runbook allows shifting checkout traffic to the neighbouring region, and the responder does so at 02:53 while sending the escalation message above to the payments secondary. Checkout latency returns to normal as traffic moves, and the neighbouring region absorbs the load. At 03:05 the payments engineer finds the card-auth provider is degraded in that region only; their status page updates twenty minutes later.
Service was restored at 02:55 through a mitigation that needed no diagnosis. The cause was found later by someone with the right context, and the responder spent no time reading payment code. The shift log records the timeline, the reverted config (which was reapplied in the morning) and two follow-ups: alert on provider error rates per region, and make regional failover for checkout automatic.
Handoffs and the shift log
Keep a running shift log and use it for the handoff, ideally in a short synchronous conversation as well as in writing. The next person should be able to read it in two minutes and know what is still open.
Shift log: payments on-call, Wed 09:00 -> Thu 09:00
Open:
- INC checkout-latency eu-west: mitigated 02:55 (traffic shifted). Traffic still on eu-central.
Owner for shift-back decision: payments team, morning stand-up.
- Disk-usage warning on ledger-db-3, rising slowly; ticket opened, not urgent.
Changes I made: reverted payments-client config 02:44, reapplied 08:30 after review.
Noisy alerts: queue-depth-warn fired 4 times, never actionable -> ticket to tune.
Watch out for: card-auth provider maintenance window Thursday 22:00 UTC.Noisy and unhelpful alerts belong in the log too. They feed the weekly alert review described in alerting that does not burn out on-call, which is how a rotation gets quieter over time instead of louder.
Onboarding new responders
Nobody should take their first page alone. A common progression is shadowing, where the newcomer receives the same pages as the primary and watches how they are handled; then reverse shadowing, where the newcomer leads and the experienced responder watches and steps in if needed; then a first solo shift with a named backup. Between stages, run game days: inject a realistic failure in a staging or controlled production setting and let the newcomer work it through the real tools and runbooks. Each game day usually exposes a broken link, a missing permission or a runbook step that no longer matches reality, which is as valuable as the practice. Runbooks themselves should follow the runbook guide so a tired newcomer can execute them.
Keeping it sustainable
Sustainable on-call is measured, not felt. Track pages per shift, pages outside working hours, time spent on interrupts and how many pages led to action. Watch the trend rather than any single shift, and compare it against a page budget the team agreed on. When the trend rises, the response is engineering work on alerts and reliability, not asking people to cope.
Protect recovery. After a night with significant pages, the responder should start late or hand off the next day; teams that write this down get it, and teams that leave it to individuals do not. Where the team spans time zones, a follow-the-sun rotation removes most night pages entirely, at the cost of more handoffs, which makes the shift log even more important. Compensation and time-off policies vary by organisation and jurisdiction, so agree them explicitly rather than assuming.
Failure modes
- Hero mode. One person investigates silently for an hour. Fix with the fifteen-minute time-box and a norm that escalation is expected.
- Diagnosis before mitigation. The responder reads logs while a rollback would have fixed it. Put the ladder in every runbook's first lines.
- Access discovered missing during a page. Run the readiness check every shift.
- Silent stakeholders. No updates, so others open side channels and interrupt the responder. Post on a schedule.
- Lost context at handoff. A mitigation is left in place and nobody owns undoing it. Every open item in the log gets an owner.
- Rising load nobody acts on. Pages creep up for months. Review the numbers every week.
What to do next
- Write a readiness script for your rotation and run it at the start of your next shift.
- Add the mitigation ladder to the top of your most-used runbooks, with the exact rollback and failover commands.
- Agree a time-box after which the responder must escalate, and say so in the on-call policy.
- Adopt a shift log template and use it at the next handoff.
- Schedule a game day for the newest member of the rotation.
- Start tracking pages per shift and out-of-hours pages, and review them weekly.