An on-call rotation is a schedule that decides which human is woken when an alert fires, who helps them, and how long they can sustain it. For LLM serving on GPUs, the rotation carries more than the usual web-service load: hardware faults on expensive accelerators, hung collective communication in multi-GPU model replicas, memory pressure from the key-value cache, model rollouts that are fast but wrong, and cost spikes that are incidents in their own right.
This article covers the human and schedule side of that work: what should page whom, how to size a rotation with arithmetic rather than hope, which rotation shapes suit which teams, how to set escalation timeouts, what a GPU-fleet handoff must contain, and how to measure whether the rotation is burning people out. Alert rule design, runbooks and incident communications have their own articles, and general on-call practice that applies to any service is covered in the on-call architecture guide, all linked at the end. This article stays with what GPU serving changes. You will finish with a small rotation generator, a burden report and a checklist.
What makes GPU serving on-call different
Four properties of GPU LLM serving change how a rotation should be built.
- Hardware faults are routine. At fleet scale some GPU will report an Xid error, an uncorrectable memory error or a degraded interconnect link most weeks. NVIDIA's Xid 79, a GPU that has fallen off the bus, takes a whole model replica down if the replica spans several GPUs. Most of these should trigger automation that cordons and drains the node, not a human page.
- Failures are often hangs, not crashes. A tensor-parallel replica waiting on a collective that one rank will never join looks healthy to a process check and dead to a user. Detection must be outside-in, and the person paged must recognise the pattern.
- Quality incidents are real incidents. A new model or quantisation can keep latency perfect while answers degrade. Somebody who understands evaluations must be reachable, which usually means a model-owner escalation tier.
- Capacity is slow to add. GPU capacity cannot be scaled in seconds like stateless CPU services, so on-call decisions include shedding load, shrinking context limits or failing over to a smaller model, each a product decision that must be pre-approved.
The practical result is a split between three expertise areas: the platform (nodes, drivers, schedulers, networking), the inference engine (batching, cache, parallelism, kernels) and the model (quality, safety, rollouts). Few people are strong in all three, so the rotation design must route each page to the right one.
Sizing a rotation with arithmetic
Start with numbers, because rotation problems are usually arithmetic problems. Google's SRE book suggests a ceiling of about two incidents per 12-hour shift, so that each can be handled properly and followed up, and suggests at least eight engineers for a single-site 24/7 rotation or six per site for a two-site rotation, with no more than about a quarter of an engineer's time on call. Treat these as sanity bounds, not laws.
Measure three inputs from the last eight to twelve weeks of history: pages per week after deduplication (several alerts about one cause count once), the fraction of pages that land outside local working hours, and the median and tail time to resolve. Then compute pages per shift, night pages per person per month and hours of incident work per on-call week. If pages per shift exceed the ceiling, no rotation shape will fix it: the alerting or the system must change first, and that work should be owned by the rotation's team.
A useful habit is to classify every page after the fact as actionable and urgent, actionable but could wait, or not actionable. The last two categories should become tickets or be deleted, and their share is the fastest lever on burden you have.
Rotation shapes
| Shape | How it works | Good for | Costs |
|---|---|---|---|
| Single-site weekly | One primary 24/7 for a week, one secondary | Small teams, low night load | Night pages land on one person |
| Follow-the-sun | Each site covers its daytime; handoffs twice daily | Two or more sites near 12 hours apart | Handoff quality, more people |
| Split day/night | Different people for day and night within one site | Night load concentrated | Night shift is unpopular; pay it |
| Tiered by expertise | Platform primary; engine and model owners as escalation | GPU serving with distinct specialities | Escalation latency, more rosters |
| Business hours plus best effort | Paging only in hours; nights unstaffed | Internal or non-critical services | Night incidents wait |
For most LLM serving teams the workable default is a platform primary and secondary, follow-the-sun if you have two sites, with engine and model owners as named escalation rosters that are paged only by the primary or by specific high-confidence alerts. Weekly rotations with a handoff on a weekday morning are the most common choice; shorter shifts reduce fatigue but multiply handoffs.
A rotation generator and burden report
A rotation generator is short enough to own yourself, and owning it means fairness is measured rather than assumed. The sketch below produces a weekly primary and secondary schedule, avoids putting anyone on twice in a row or in their blackout weeks, and then reports burden per person from historical page timestamps.
from collections import Counter
from datetime import date, timedelta
def build_rotation(people, start, weeks, blackout):
"""people: ordered list; blackout: {name: {week_index, ...}}. Returns [(week_start, primary, secondary)]."""
sched, load, last = [], Counter(), None
for w in range(weeks):
free = [p for p in people if w not in blackout.get(p, set()) and p != last]
free.sort(key=lambda p: (load[p], people.index(p))) # least-loaded first, stable order
primary, secondary = free[0], free[1]
load[primary] += 2 # primary weighs double
load[secondary] += 1
sched.append((start + timedelta(weeks=w), primary, secondary))
last = primary
return sched
def burden(sched, pages, tz_offset_h, night=(22, 7)):
"""pages: list of UTC datetimes (deduplicated). Counts pages and night pages per primary."""
total, nights = Counter(), Counter()
for ts in pages:
for wk, primary, _ in sched:
if wk <= ts.date() < wk + timedelta(weeks=1):
local_h = (ts.hour + tz_offset_h[primary]) % 24
total[primary] += 1
if local_h >= night[0] or local_h < night[1]:
nights[primary] += 1
break
return total, nightsPublish the burden report every month next to the schedule. If one person carries far more night pages than others, the cause is usually a recurring weekday batch job or a regional traffic peak, and it is better fixed than rotated around.
Escalation policy
Escalation policy decides what happens when the primary does not respond or needs help. Set the acknowledgement timeout from data: take the 95th percentile of historical time-to-acknowledge for pages that were acknowledged, and set the timeout a little above it, typically five minutes. A shorter value wakes secondaries for nothing; a much longer one adds directly to customer impact.
escalation:
service: llm-serving-prod
steps:
- notify: primary # platform rotation
ack_timeout_min: 5
- notify: [primary, secondary]
ack_timeout_min: 10
- notify: inference-engine-owner # batching, cache, parallelism
ack_timeout_min: 15
- notify: incident-commander
direct_routes: # high-confidence alerts that skip the platform tier
- alert: eval_regression_canary
notify: model-owner
- alert: spend_burn_critical
notify: [primary, finops-owner]
auto_remediate_first: # automation acts before any human page
- alert: gpu_xid_fatal
action: cordon_and_drain_node
page_if: replicas_healthy_fraction < 0.8The final block is the most valuable for GPU fleets. A single failed GPU should cost a replica, which the scheduler replaces; only when healthy capacity falls below what the traffic needs should a human be woken. Route such policies through review like code.
Handoffs for a GPU fleet
Handoffs fail silently: the outgoing engineer knows a node is half-broken, the incoming one does not. For GPU fleets the handoff note should be a short, structured document generated partly from systems and partly by hand, and reviewed in a fifteen-minute overlap call at each shift change.
HANDOFF 2026-10-06 14:00 UTC from: site A primary to: site B primary
Open incidents ......... 1 (latency spike, eu-west, mitigated by load shed, root cause unknown)
Cordoned nodes ......... 3 (2 awaiting vendor RMA, 1 under NVLink investigation)
Rollouts in flight ..... model v14 canary at 10% in us-east; eval gate due 18:00 UTC
Capacity headroom ...... us-east 22%, eu-west 9% (below 15% target), ap-south 31%
Temporary changes ...... max context capped at 32k in eu-west until 22:00 UTC; revert step in ticket
Noisy alerts ........... kv_cache_evictions_high flapping; silenced until 20:00 UTC, ticket open
Watch for .............. a batch-inference customer starts a large job at 16:00 UTCEvery temporary mitigation needs an expiry and a revert step, because the most common self-inflicted incident is a forgotten load-shed or context cap that quietly hurts users for days.
Worked example: moving to follow-the-sun
A team of twelve, six in each of two sites about eleven hours apart, runs a GPU serving platform. Over ten weeks, history shows 21 deduplicated pages a week, 40 percent of them outside the receiving engineer's local daytime under the current single-site 24/7 rotation. Under that rotation a primary week means about 8 night pages, more than one a night, and two of the twelve have asked to leave the rotation.
Moving to follow-the-sun puts each site on its own daytime. Pages split roughly in half, so each site's primary handles about 10.5 a week, 1.5 per 12-hour shift, within the two-per-shift ceiling, and night pages fall to near zero except during the handoff edges. Each engineer is primary one week in six. Triage of the ten weeks also shows that 7 of the 21 weekly pages were single-GPU Xid events that the scheduler would have handled anyway. Converting them to auto-cordon with a capacity-based page cuts the load to 14 a week, about one per shift, which leaves time for the follow-up work that reduces it further. These numbers are illustrative, but the order of work is the lesson: measure, cut non-actionable pages, then reshape the rotation.
Onboarding, pay and burnout
Bring new people in through shadowing: one rotation as a shadow who receives every page but has no responsibility, then one as primary with an experienced reverse-shadow. Pair this with game days on a staging cluster where you inject the GPU-specific failures: a drained node, a hung collective, a cache-exhaustion storm, a canary with a quality regression. Pay for on-call time in money or time off according to your local rules and make it predictable. Track burden openly: pages per shift, night pages per person, time to acknowledge, and the share of pages that were not actionable. A rising non-actionable share is the earliest warning that a rotation is heading for burnout.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
| Hero rotation | One expert answers everything | Tiered rosters, runbooks, shadowing |
| Page flood | Many alerts per cause | Deduplicate by cause, group alerts |
| Hardware pages humans | Every Xid wakes someone | Auto-cordon; page on capacity |
| Lost context | Incoming primary unaware of cordons or caps | Structured handoff with expiries |
| Unowned quality | Latency fine, answers wrong | Model-owner route on eval canaries |
| Silent no-ack | Phone on mute, page lost | Re-page plus secondary at measured timeout |
| Burnout | People leave the rotation | Burden report, night pay, cut non-actionable pages |
What to do next
- Export ten weeks of pages, deduplicate by cause and compute pages per shift and night pages per person.
- Classify every page as urgent, deferrable or non-actionable, and delete or ticket the last two.
- Automate cordon-and-drain for single-GPU faults and page on remaining healthy capacity instead.
- Pick a rotation shape from your sites and load; add engine and model owner escalation rosters.
- Set acknowledgement timeouts from the measured 95th percentile, and review them quarterly.
- Adopt a structured handoff note with expiries for every temporary mitigation.
- Run shadowing and a GPU-failure game day before anyone takes a first primary shift.
- Publish a monthly burden report and act on the non-actionable share.
Related reading on this site: SLO burn-rate alerts for GPU serving, a symptom-keyed runbook index, incident communication channels, outside-in synthetic monitoring and on-call architecture for any service.