An on-call rotation is a schedule that decides which human is woken when an alert fires, who helps them, and how long they can sustain it. For LLM serving on GPUs, the rotation carries more than the usual web-service load: hardware faults on expensive accelerators, hung collective communication in multi-GPU model replicas, memory pressure from the key-value cache, model rollouts that are fast but wrong, and cost spikes that are incidents in their own right.

This article covers the human and schedule side of that work: what should page whom, how to size a rotation with arithmetic rather than hope, which rotation shapes suit which teams, how to set escalation timeouts, what a GPU-fleet handoff must contain, and how to measure whether the rotation is burning people out. Alert rule design, runbooks and incident communications have their own articles, and general on-call practice that applies to any service is covered in the on-call architecture guide, all linked at the end. This article stays with what GPU serving changes. You will finish with a small rotation generator, a burden report and a checklist.

What makes GPU serving on-call different

Four properties of GPU LLM serving change how a rotation should be built.

  • Hardware faults are routine. At fleet scale some GPU will report an Xid error, an uncorrectable memory error or a degraded interconnect link most weeks. NVIDIA's Xid 79, a GPU that has fallen off the bus, takes a whole model replica down if the replica spans several GPUs. Most of these should trigger automation that cordons and drains the node, not a human page.
  • Failures are often hangs, not crashes. A tensor-parallel replica waiting on a collective that one rank will never join looks healthy to a process check and dead to a user. Detection must be outside-in, and the person paged must recognise the pattern.
  • Quality incidents are real incidents. A new model or quantisation can keep latency perfect while answers degrade. Somebody who understands evaluations must be reachable, which usually means a model-owner escalation tier.
  • Capacity is slow to add. GPU capacity cannot be scaled in seconds like stateless CPU services, so on-call decisions include shedding load, shrinking context limits or failing over to a smaller model, each a product decision that must be pre-approved.

The practical result is a split between three expertise areas: the platform (nodes, drivers, schedulers, networking), the inference engine (batching, cache, parallelism, kernels) and the model (quality, safety, rollouts). Few people are strong in all three, so the rotation design must route each page to the right one.

Sizing a rotation with arithmetic

Start with numbers, because rotation problems are usually arithmetic problems. Google's SRE book suggests a ceiling of about two incidents per 12-hour shift, so that each can be handled properly and followed up, and suggests at least eight engineers for a single-site 24/7 rotation or six per site for a two-site rotation, with no more than about a quarter of an engineer's time on call. Treat these as sanity bounds, not laws.

Measure three inputs from the last eight to twelve weeks of history: pages per week after deduplication (several alerts about one cause count once), the fraction of pages that land outside local working hours, and the median and tail time to resolve. Then compute pages per shift, night pages per person per month and hours of incident work per on-call week. If pages per shift exceed the ceiling, no rotation shape will fix it: the alerting or the system must change first, and that work should be owned by the rotation's team.

A useful habit is to classify every page after the fact as actionable and urgent, actionable but could wait, or not actionable. The last two categories should become tickets or be deleted, and their share is the fastest lever on burden you have.

Rotation shapes

ShapeHow it worksGood forCosts
Single-site weeklyOne primary 24/7 for a week, one secondarySmall teams, low night loadNight pages land on one person
Follow-the-sunEach site covers its daytime; handoffs twice dailyTwo or more sites near 12 hours apartHandoff quality, more people
Split day/nightDifferent people for day and night within one siteNight load concentratedNight shift is unpopular; pay it
Tiered by expertisePlatform primary; engine and model owners as escalationGPU serving with distinct specialitiesEscalation latency, more rosters
Business hours plus best effortPaging only in hours; nights unstaffedInternal or non-critical servicesNight incidents wait
Follow-the-sun: two sites, two 12-hour daytime shifts, one handoff each wayUTC000306091215182124Site AA primary 02:00-14:00 UTC (08:00-20:00 local)Site BB primary 14:00-02:00 UTChandoff A to BB to ASecondary at each site covers the same hours. Nobody is primary at night if the offset is near 12 hours.
A two-site follow-the-sun layout. Local times assume Site A is at UTC+6; choose shift boundaries from your actual offsets.

For most LLM serving teams the workable default is a platform primary and secondary, follow-the-sun if you have two sites, with engine and model owners as named escalation rosters that are paged only by the primary or by specific high-confidence alerts. Weekly rotations with a handoff on a weekday morning are the most common choice; shorter shifts reduce fatigue but multiply handoffs.

A rotation generator and burden report

A rotation generator is short enough to own yourself, and owning it means fairness is measured rather than assumed. The sketch below produces a weekly primary and secondary schedule, avoids putting anyone on twice in a row or in their blackout weeks, and then reports burden per person from historical page timestamps.

from collections import Counter
from datetime import date, timedelta

def build_rotation(people, start, weeks, blackout):
    """people: ordered list; blackout: {name: {week_index, ...}}. Returns [(week_start, primary, secondary)]."""
    sched, load, last = [], Counter(), None
    for w in range(weeks):
        free = [p for p in people if w not in blackout.get(p, set()) and p != last]
        free.sort(key=lambda p: (load[p], people.index(p)))   # least-loaded first, stable order
        primary, secondary = free[0], free[1]
        load[primary] += 2                                    # primary weighs double
        load[secondary] += 1
        sched.append((start + timedelta(weeks=w), primary, secondary))
        last = primary
    return sched

def burden(sched, pages, tz_offset_h, night=(22, 7)):
    """pages: list of UTC datetimes (deduplicated). Counts pages and night pages per primary."""
    total, nights = Counter(), Counter()
    for ts in pages:
        for wk, primary, _ in sched:
            if wk <= ts.date() < wk + timedelta(weeks=1):
                local_h = (ts.hour + tz_offset_h[primary]) % 24
                total[primary] += 1
                if local_h >= night[0] or local_h < night[1]:
                    nights[primary] += 1
                break
    return total, nights

Publish the burden report every month next to the schedule. If one person carries far more night pages than others, the cause is usually a recurring weekday batch job or a regional traffic peak, and it is better fixed than rotated around.

Escalation policy

Escalation policy decides what happens when the primary does not respond or needs help. Set the acknowledgement timeout from data: take the 95th percentile of historical time-to-acknowledge for pages that were acknowledged, and set the timeout a little above it, typically five minutes. A shorter value wakes secondaries for nothing; a much longer one adds directly to customer impact.

Escalation ladder for a GPU serving pagePage primaryt = 0Re-page + secondaryno ack by 5 minEngine or model ownerno ack 15 min / needs expertiseIncident commandercustomer impact > 30 minTimeouts come from measured time-to-acknowledge, not from habit. Expertise escalation is allowed at any step.
Time-based steps plus an expertise path: the primary may page an engine or model owner at any time.
escalation:
  service: llm-serving-prod
  steps:
    - notify: primary            # platform rotation
      ack_timeout_min: 5
    - notify: [primary, secondary]
      ack_timeout_min: 10
    - notify: inference-engine-owner   # batching, cache, parallelism
      ack_timeout_min: 15
    - notify: incident-commander
  direct_routes:                 # high-confidence alerts that skip the platform tier
    - alert: eval_regression_canary
      notify: model-owner
    - alert: spend_burn_critical
      notify: [primary, finops-owner]
  auto_remediate_first:          # automation acts before any human page
    - alert: gpu_xid_fatal
      action: cordon_and_drain_node
      page_if: replicas_healthy_fraction < 0.8

The final block is the most valuable for GPU fleets. A single failed GPU should cost a replica, which the scheduler replaces; only when healthy capacity falls below what the traffic needs should a human be woken. Route such policies through review like code.

Handoffs for a GPU fleet

Handoffs fail silently: the outgoing engineer knows a node is half-broken, the incoming one does not. For GPU fleets the handoff note should be a short, structured document generated partly from systems and partly by hand, and reviewed in a fifteen-minute overlap call at each shift change.

HANDOFF  2026-10-06 14:00 UTC   from: site A primary   to: site B primary
Open incidents ......... 1 (latency spike, eu-west, mitigated by load shed, root cause unknown)
Cordoned nodes ......... 3 (2 awaiting vendor RMA, 1 under NVLink investigation)
Rollouts in flight ..... model v14 canary at 10% in us-east; eval gate due 18:00 UTC
Capacity headroom ...... us-east 22%, eu-west 9% (below 15% target), ap-south 31%
Temporary changes ...... max context capped at 32k in eu-west until 22:00 UTC; revert step in ticket
Noisy alerts ........... kv_cache_evictions_high flapping; silenced until 20:00 UTC, ticket open
Watch for .............. a batch-inference customer starts a large job at 16:00 UTC

Every temporary mitigation needs an expiry and a revert step, because the most common self-inflicted incident is a forgotten load-shed or context cap that quietly hurts users for days.

Worked example: moving to follow-the-sun

A team of twelve, six in each of two sites about eleven hours apart, runs a GPU serving platform. Over ten weeks, history shows 21 deduplicated pages a week, 40 percent of them outside the receiving engineer's local daytime under the current single-site 24/7 rotation. Under that rotation a primary week means about 8 night pages, more than one a night, and two of the twelve have asked to leave the rotation.

Moving to follow-the-sun puts each site on its own daytime. Pages split roughly in half, so each site's primary handles about 10.5 a week, 1.5 per 12-hour shift, within the two-per-shift ceiling, and night pages fall to near zero except during the handoff edges. Each engineer is primary one week in six. Triage of the ten weeks also shows that 7 of the 21 weekly pages were single-GPU Xid events that the scheduler would have handled anyway. Converting them to auto-cordon with a capacity-based page cuts the load to 14 a week, about one per shift, which leaves time for the follow-up work that reduces it further. These numbers are illustrative, but the order of work is the lesson: measure, cut non-actionable pages, then reshape the rotation.

Onboarding, pay and burnout

Bring new people in through shadowing: one rotation as a shadow who receives every page but has no responsibility, then one as primary with an experienced reverse-shadow. Pair this with game days on a staging cluster where you inject the GPU-specific failures: a drained node, a hung collective, a cache-exhaustion storm, a canary with a quality regression. Pay for on-call time in money or time off according to your local rules and make it predictable. Track burden openly: pages per shift, night pages per person, time to acknowledge, and the share of pages that were not actionable. A rising non-actionable share is the earliest warning that a rotation is heading for burnout.

Failure modes

FailureSymptomFix
Hero rotationOne expert answers everythingTiered rosters, runbooks, shadowing
Page floodMany alerts per causeDeduplicate by cause, group alerts
Hardware pages humansEvery Xid wakes someoneAuto-cordon; page on capacity
Lost contextIncoming primary unaware of cordons or capsStructured handoff with expiries
Unowned qualityLatency fine, answers wrongModel-owner route on eval canaries
Silent no-ackPhone on mute, page lostRe-page plus secondary at measured timeout
BurnoutPeople leave the rotationBurden report, night pay, cut non-actionable pages

What to do next

  1. Export ten weeks of pages, deduplicate by cause and compute pages per shift and night pages per person.
  2. Classify every page as urgent, deferrable or non-actionable, and delete or ticket the last two.
  3. Automate cordon-and-drain for single-GPU faults and page on remaining healthy capacity instead.
  4. Pick a rotation shape from your sites and load; add engine and model owner escalation rosters.
  5. Set acknowledgement timeouts from the measured 95th percentile, and review them quarterly.
  6. Adopt a structured handoff note with expiries for every temporary mitigation.
  7. Run shadowing and a GPU-failure game day before anyone takes a first primary shift.
  8. Publish a monthly burden report and act on the non-actionable share.

Related reading on this site: SLO burn-rate alerts for GPU serving, a symptom-keyed runbook index, incident communication channels, outside-in synthetic monitoring and on-call architecture for any service.

Key takeaway: A sustainable LLM on-call rotation starts with measured page load, not with a schedule. Cut non-actionable pages, let automation absorb single-GPU faults, choose a rotation shape that keeps nights quiet, route engine and model problems to named owners, set timeouts from data, and hand off with explicit state so the next person inherits the fleet rather than a mystery.