Most migrations that go badly do not fail on technology. The facade worked, the backfill copied every row, the change capture kept up. They fail because nobody wrote down what 'ready' meant, the cutover happened at the end of a long day on someone's judgement, the rollback path had never been run, and the old system was switched off a week before anyone noticed the nightly report had depended on it. A migration playbook is the operating procedure that prevents those failures: a fixed sequence of phases, each guarded by measurable gates, with named owners, a rehearsed rollback, and one explicit point of no return placed as late as possible.

This article is about running the migration. The technical machinery underneath it, a routing facade, backfill plus change data capture, shadow reads and comparison, is covered in the companion migration architecture article; here we assume those mechanisms exist and focus on how to decide, coordinate and recover. The worked example is moving a payments ledger from a self-managed database to a managed one, but the playbook applies equally to a service rewrite, a queue replacement or a vendor switch.

Advertisement

Why a playbook rather than a plan

A project plan answers 'when will it be done'. A playbook answers 'what do we do next, given what we now observe'. That difference matters because migrations are full of surprises that a schedule cannot absorb: a mismatch rate that will not go to zero, a consumer nobody knew about, a latency regression at the 99th percentile. A playbook turns each surprise into a predefined branch: stay in the phase, roll back one phase, or abort.

The playbook has four parts. Phases are stable intermediate states the system can sit in indefinitely. Gates separate phases and consist of entry criteria (may we start?) and exit criteria (are we done?). Runbooks describe the exact steps for each transition, forward and back. Roles and communications say who decides, who executes, and who must be told.

The property you are buying is that every forward step is small, measured and reversible until a deliberate final step. If you cannot describe how to undo a step, it is not yet a playbook step; it is a leap.

The phases and their gates

A migration playbook: phases separated by gates, with one marked point of no return0. Prepareinventory, owners1. Build + syncnew system, backfill2. Shadowcompare, no impact3. Ramp reads1% to 100%4. Switch writesold becomes follower5. Stop reverse syncpoint of no return6. Decommissiondelete, close outgate G5: soak period cleanG1G2G3Every gate = entry criteria + exit criteriameasured by queries, signed off by a named ownerBefore G5 every step is reversible by a routing change. After G5, rollback means a new migration.The playbook's job is to make that boundary explicit, late, and deliberate.
Six phases separated by gates. Phases 1 to 4 can each be undone by a routing or flag change because the old system stays current. Stopping reverse synchronisation is the single point of no return.
PhaseSystem stateExit gate (examples)Undo
0. PrepareOld only; inventory of every reader, writer and jobAll consumers listed with owners; success metrics agreedNothing to undo
1. Build and syncNew system populated by backfill, kept current by change captureBackfill complete; replication lag p99 under targetDrop the new copy
2. ShadowProduction reads also sent to new; results compared, discardedMismatch rate below threshold for 24h; latency within budgetTurn shadowing off
3. Ramp readsA growing share of reads served from new100% reads on new for a soak period with error and latency parityFlag reads back to old
4. Switch writesNew is primary; reverse sync keeps old currentSoak clean; reconciliation clean; no consumers on oldFlag writes back; reverse sync means old is current
5. Stop reverse syncOld frozenBusiness sign-off; final reconciliation archivedNone: a new migration
6. DecommissionOld deletedData retention obligations met; costs closedRestore from archive only

Phase 0 is the one teams skip and regret. The inventory must come from evidence, not memory: database connection logs, query logs grouped by client, network flow logs and the service catalog. Anything that touches the old system and is not on the list will discover the migration for you, usually at the worst moment.

Advertisement

Go/no-go as code

A gate is only useful if it is evaluated the same way every time. Write the criteria as data, back each one with a query, and have a script produce the verdict. The meeting then reviews the script's output instead of debating impressions.

# gates.yaml: go/no-go criteria as code, evaluated by a script, not by a meeting's mood
G3_ramp_reads:
  owner: payments-platform-oncall
  entry:
    - shadow_mismatch_rate_24h < 0.0001        # from the comparison job
    - backfill_complete == true
    - replication_lag_p99_seconds < 5
    - rollback_drill_passed_within_days <= 7
  exit:                                         # must hold at 100% reads for the soak
    - error_rate_new <= error_rate_old * 1.1
    - latency_p99_new_ms <= latency_p99_old_ms * 1.2
    - soak_hours >= 72
  abort_if:                                     # any one triggers the rollback runbook
    - error_rate_new > error_rate_old * 2 for 10m
    - data_mismatch_detected == true
import operator, re, yaml

OPS = {"<": operator.lt, "<=": operator.le, ">": operator.gt, ">=": operator.ge, "==": operator.eq}

def evaluate(gate_name, metrics, path="gates.yaml"):
    """Return (go, failures). metrics is a dict filled from Prometheus/SQL queries."""
    gate = yaml.safe_load(open(path))[gate_name]
    failures = []
    for rule in gate["entry"]:
        m = re.match(r"(\w+)\s*(<=|>=|==|<|>)\s*(\S+)", rule)
        name, op, raw = m.groups()
        want = {"true": True, "false": False}.get(raw, None)
        want = float(raw) if want is None else want
        have = metrics.get(name)
        if have is None:
            failures.append(f"{name}: no data (a missing metric is a NO-GO, never a pass)")
        elif not OPS[op](have, want):
            failures.append(f"{name}: have {have}, need {op} {want}")
    return (not failures, failures)

go, why = evaluate("G3_ramp_reads", collect_metrics())
print("GO" if go else "NO-GO\n" + "\n".join(why))

Three rules make this robust. First, a missing metric is a NO-GO: if the comparison job silently stopped, a gate that treats 'no mismatches reported' as success will wave a broken migration through. Second, criteria compare the new system against the old one measured over the same window, not against an absolute number written months ago; traffic changes. Third, the evaluator's output is attached to the change ticket, so afterwards anyone can see exactly what was true when the decision was made.

Humans still decide. The gate owner can refuse a GO for reasons outside the metrics, such as a holiday freeze or an unrelated incident, but cannot override a NO-GO without writing down why and who agreed.

Roles and communications

RoleResponsibility
Migration leadOwns the playbook, schedules phases, runs gate reviews
Gate owner (per gate)Signs off GO; usually the on-call lead of the owning team
ExecutorPerforms runbook steps during transitions; never the same person as the observer
ObserverWatches dashboards and abort conditions during transitions; can call abort alone
Consumer liaisonsOne contact per downstream team from the Phase 0 inventory
Rollback ownerOn call through every transition and soak; has rehearsed the rollback

Separating executor and observer is cheap and effective: the person typing commands is the worst-placed person to notice a dashboard turning red. Give the observer unilateral authority to abort; an abort that turns out to be unnecessary costs an evening, while a hesitation can cost data.

Communicate on a fixed rhythm. Announce each transition at least a day ahead to every consumer liaison, post start, completion and any abort in one channel, and keep a single status document that states the current phase, the next gate, and its date. Most migration confusion comes from two teams holding different beliefs about which system is primary today.

The cutover runbook

Each transition gets a runbook with times relative to T, exact commands or flag names, verification steps with their expected values, and the abort branch. Here is the most dangerous transition, switching writes, for the ledger example.

CUTOVER RUNBOOK: switch writes (phase 4)          T = scheduled start, low-traffic window
T-24h  Announce window in #migrations and status page draft; confirm rollback owner on call
T-60m  Run gate evaluator G4-entry; attach output to the change ticket. NO-GO -> stop here
T-15m  Freeze deploys to both systems; snapshot old store; record replication position
T+0    Flip write flag to NEW (single config change, audited); reverse sync NEW -> OLD running
T+5m   Verify: write error rate, reverse-sync lag, row counts on both sides for last 5 minutes
T+15m  Verify: business canary (one real order end to end, reconciled in both systems)
T+60m  Declare "writes on NEW, reversible"; unfreeze deploys to NEW only
Abort  Any abort_if condition: flip write flag back to OLD, confirm reverse sync drained,
       open incident, keep both systems frozen until the comparison job is clean

Notice what the runbook does not contain: judgement calls without criteria, or steps that change two things at once. The write switch is one audited flag change. Freezing deploys removes a whole class of confounding changes. The business canary, a real transaction followed through both systems, catches mistakes that technical metrics miss, such as a currency rounding difference that leaves every row present but wrong. For flag mechanics and safe ramps, see feature flag delivery.

The rollback decision tree

Rollback must be decided quickly, so decide the logic in advance. At any moment during phases 2 to 4, an observed problem walks this tree:

  1. Is data being corrupted or lost right now? Abort immediately to the previous phase. Do not diagnose first.
  2. Is an abort_if condition met? Abort. These were agreed when everyone was calm.
  3. Is a user-facing SLO burning faster than its budget allows? Step back one ramp level (for example from 50% to 10% reads) and diagnose.
  4. Is it a non-user-facing anomaly, such as higher cost or an internal job slower? Hold the current phase, open a ticket, and add a criterion to the next gate.

Rolling back must be as rehearsed as rolling forward. Run the rollback in staging, then in production during a low-traffic window before Phase 4, and record the time it took; that time is an input to your abort thresholds. A rollback path that has never run is a hypothesis. The general principles are covered in rollback strategy.

Rollback also has a data cost that the playbook should name. After writes switch, rolling back is only clean if reverse synchronisation has kept the old system current. Monitor reverse-sync lag as a first-class abort condition: if it grows unbounded, you are silently losing your safety net.

The point of no return

Every migration has a step after which going back means running a new migration in the other direction. Usually it is stopping reverse synchronisation, deleting old data, or allowing the new system to accept writes the old schema cannot represent. A good playbook names this step, places it as late as possible, and puts its own gate in front of it.

Pushing the point of no return later has a cost. Reverse sync must be built and operated, both systems are paid for, and new features that need the new schema must wait or be written to be backwards compatible. Teams often discover mid-migration that product work is blocked on a schema change that reverse sync cannot carry. Decide upfront how long the soak after the write switch will be, typically one to four weeks covering at least one month-end or other business cycle, and refuse schema changes that break reverse sync until the gate passes.

Decommissioning

Decommissioning is part of the migration, not a cleanup task for later. Until the old system is gone, it costs money, carries security patches, and tempts someone to read from it. The steps: confirm from access logs that nothing has connected for the agreed period; revoke credentials before deleting anything, so a forgotten consumer fails loudly while recovery is still possible; archive data according to retention obligations; delete infrastructure; and remove the facade's old route and flags so the code no longer branches. Close the migration with a short review covering what the gates caught, what they missed, and which criteria to reuse; the format in change management works well.

Failure modes

FailureHow it happensPrevention
Unknown consumer breaksInventory built from memoryBuild Phase 0 from connection and query logs
Gate passed on missing dataComparison job died; zero mismatches reportedMissing metric is a NO-GO
Rollback fails when neededNever rehearsed; reverse sync laggingProduction rollback drill; reverse-sync lag as an abort condition
Two changes at onceSchema change shipped with the write switchDeploy freeze during transitions
Endless dual runningNo decommission gate or ownerDecommission is a phase with a date and an owner
Silent semantic driftRows match but meaning differs (rounding, time zones)Business canary and domain-level reconciliation

What to do next

  1. Write the phase table for your migration, with the undo action for each phase, and mark the point of no return.
  2. Build the Phase 0 inventory from logs, assign an owner to every consumer, and appoint consumer liaisons.
  3. Express each gate's entry, exit and abort criteria in a file and implement an evaluator that treats missing data as NO-GO.
  4. Write runbooks for every transition, forward and back, and rehearse the rollback in production before switching writes.
  5. Assign executor, observer and rollback owner roles for each transition, and give the observer abort authority.
  6. Schedule decommissioning as a phase with a date, and close with a review that improves the next playbook.
Key takeaway: A migration playbook turns a risky project into a sequence of small, reversible steps. Define stable phases, guard each with entry, exit and abort criteria evaluated as code (missing data is a NO-GO), and write and rehearse a runbook for every transition in both directions. Separate executor and observer, communicate on a fixed rhythm, push the single point of no return as late as possible behind its own gate, and finish by decommissioning the old system on a schedule.