Most migrations that go badly do not fail on technology. The facade worked, the backfill copied every row, the change capture kept up. They fail because nobody wrote down what 'ready' meant, the cutover happened at the end of a long day on someone's judgement, the rollback path had never been run, and the old system was switched off a week before anyone noticed the nightly report had depended on it. A migration playbook is the operating procedure that prevents those failures: a fixed sequence of phases, each guarded by measurable gates, with named owners, a rehearsed rollback, and one explicit point of no return placed as late as possible.
This article is about running the migration. The technical machinery underneath it, a routing facade, backfill plus change data capture, shadow reads and comparison, is covered in the companion migration architecture article; here we assume those mechanisms exist and focus on how to decide, coordinate and recover. The worked example is moving a payments ledger from a self-managed database to a managed one, but the playbook applies equally to a service rewrite, a queue replacement or a vendor switch.
Why a playbook rather than a plan
A project plan answers 'when will it be done'. A playbook answers 'what do we do next, given what we now observe'. That difference matters because migrations are full of surprises that a schedule cannot absorb: a mismatch rate that will not go to zero, a consumer nobody knew about, a latency regression at the 99th percentile. A playbook turns each surprise into a predefined branch: stay in the phase, roll back one phase, or abort.
The playbook has four parts. Phases are stable intermediate states the system can sit in indefinitely. Gates separate phases and consist of entry criteria (may we start?) and exit criteria (are we done?). Runbooks describe the exact steps for each transition, forward and back. Roles and communications say who decides, who executes, and who must be told.
The property you are buying is that every forward step is small, measured and reversible until a deliberate final step. If you cannot describe how to undo a step, it is not yet a playbook step; it is a leap.
The phases and their gates
| Phase | System state | Exit gate (examples) | Undo |
|---|---|---|---|
| 0. Prepare | Old only; inventory of every reader, writer and job | All consumers listed with owners; success metrics agreed | Nothing to undo |
| 1. Build and sync | New system populated by backfill, kept current by change capture | Backfill complete; replication lag p99 under target | Drop the new copy |
| 2. Shadow | Production reads also sent to new; results compared, discarded | Mismatch rate below threshold for 24h; latency within budget | Turn shadowing off |
| 3. Ramp reads | A growing share of reads served from new | 100% reads on new for a soak period with error and latency parity | Flag reads back to old |
| 4. Switch writes | New is primary; reverse sync keeps old current | Soak clean; reconciliation clean; no consumers on old | Flag writes back; reverse sync means old is current |
| 5. Stop reverse sync | Old frozen | Business sign-off; final reconciliation archived | None: a new migration |
| 6. Decommission | Old deleted | Data retention obligations met; costs closed | Restore from archive only |
Phase 0 is the one teams skip and regret. The inventory must come from evidence, not memory: database connection logs, query logs grouped by client, network flow logs and the service catalog. Anything that touches the old system and is not on the list will discover the migration for you, usually at the worst moment.
Go/no-go as code
A gate is only useful if it is evaluated the same way every time. Write the criteria as data, back each one with a query, and have a script produce the verdict. The meeting then reviews the script's output instead of debating impressions.
# gates.yaml: go/no-go criteria as code, evaluated by a script, not by a meeting's mood
G3_ramp_reads:
owner: payments-platform-oncall
entry:
- shadow_mismatch_rate_24h < 0.0001 # from the comparison job
- backfill_complete == true
- replication_lag_p99_seconds < 5
- rollback_drill_passed_within_days <= 7
exit: # must hold at 100% reads for the soak
- error_rate_new <= error_rate_old * 1.1
- latency_p99_new_ms <= latency_p99_old_ms * 1.2
- soak_hours >= 72
abort_if: # any one triggers the rollback runbook
- error_rate_new > error_rate_old * 2 for 10m
- data_mismatch_detected == trueimport operator, re, yaml
OPS = {"<": operator.lt, "<=": operator.le, ">": operator.gt, ">=": operator.ge, "==": operator.eq}
def evaluate(gate_name, metrics, path="gates.yaml"):
"""Return (go, failures). metrics is a dict filled from Prometheus/SQL queries."""
gate = yaml.safe_load(open(path))[gate_name]
failures = []
for rule in gate["entry"]:
m = re.match(r"(\w+)\s*(<=|>=|==|<|>)\s*(\S+)", rule)
name, op, raw = m.groups()
want = {"true": True, "false": False}.get(raw, None)
want = float(raw) if want is None else want
have = metrics.get(name)
if have is None:
failures.append(f"{name}: no data (a missing metric is a NO-GO, never a pass)")
elif not OPS[op](have, want):
failures.append(f"{name}: have {have}, need {op} {want}")
return (not failures, failures)
go, why = evaluate("G3_ramp_reads", collect_metrics())
print("GO" if go else "NO-GO\n" + "\n".join(why))Three rules make this robust. First, a missing metric is a NO-GO: if the comparison job silently stopped, a gate that treats 'no mismatches reported' as success will wave a broken migration through. Second, criteria compare the new system against the old one measured over the same window, not against an absolute number written months ago; traffic changes. Third, the evaluator's output is attached to the change ticket, so afterwards anyone can see exactly what was true when the decision was made.
Humans still decide. The gate owner can refuse a GO for reasons outside the metrics, such as a holiday freeze or an unrelated incident, but cannot override a NO-GO without writing down why and who agreed.
Roles and communications
| Role | Responsibility |
|---|---|
| Migration lead | Owns the playbook, schedules phases, runs gate reviews |
| Gate owner (per gate) | Signs off GO; usually the on-call lead of the owning team |
| Executor | Performs runbook steps during transitions; never the same person as the observer |
| Observer | Watches dashboards and abort conditions during transitions; can call abort alone |
| Consumer liaisons | One contact per downstream team from the Phase 0 inventory |
| Rollback owner | On call through every transition and soak; has rehearsed the rollback |
Separating executor and observer is cheap and effective: the person typing commands is the worst-placed person to notice a dashboard turning red. Give the observer unilateral authority to abort; an abort that turns out to be unnecessary costs an evening, while a hesitation can cost data.
Communicate on a fixed rhythm. Announce each transition at least a day ahead to every consumer liaison, post start, completion and any abort in one channel, and keep a single status document that states the current phase, the next gate, and its date. Most migration confusion comes from two teams holding different beliefs about which system is primary today.
The cutover runbook
Each transition gets a runbook with times relative to T, exact commands or flag names, verification steps with their expected values, and the abort branch. Here is the most dangerous transition, switching writes, for the ledger example.
CUTOVER RUNBOOK: switch writes (phase 4) T = scheduled start, low-traffic window
T-24h Announce window in #migrations and status page draft; confirm rollback owner on call
T-60m Run gate evaluator G4-entry; attach output to the change ticket. NO-GO -> stop here
T-15m Freeze deploys to both systems; snapshot old store; record replication position
T+0 Flip write flag to NEW (single config change, audited); reverse sync NEW -> OLD running
T+5m Verify: write error rate, reverse-sync lag, row counts on both sides for last 5 minutes
T+15m Verify: business canary (one real order end to end, reconciled in both systems)
T+60m Declare "writes on NEW, reversible"; unfreeze deploys to NEW only
Abort Any abort_if condition: flip write flag back to OLD, confirm reverse sync drained,
open incident, keep both systems frozen until the comparison job is cleanNotice what the runbook does not contain: judgement calls without criteria, or steps that change two things at once. The write switch is one audited flag change. Freezing deploys removes a whole class of confounding changes. The business canary, a real transaction followed through both systems, catches mistakes that technical metrics miss, such as a currency rounding difference that leaves every row present but wrong. For flag mechanics and safe ramps, see feature flag delivery.
The rollback decision tree
Rollback must be decided quickly, so decide the logic in advance. At any moment during phases 2 to 4, an observed problem walks this tree:
- Is data being corrupted or lost right now? Abort immediately to the previous phase. Do not diagnose first.
- Is an abort_if condition met? Abort. These were agreed when everyone was calm.
- Is a user-facing SLO burning faster than its budget allows? Step back one ramp level (for example from 50% to 10% reads) and diagnose.
- Is it a non-user-facing anomaly, such as higher cost or an internal job slower? Hold the current phase, open a ticket, and add a criterion to the next gate.
Rolling back must be as rehearsed as rolling forward. Run the rollback in staging, then in production during a low-traffic window before Phase 4, and record the time it took; that time is an input to your abort thresholds. A rollback path that has never run is a hypothesis. The general principles are covered in rollback strategy.
Rollback also has a data cost that the playbook should name. After writes switch, rolling back is only clean if reverse synchronisation has kept the old system current. Monitor reverse-sync lag as a first-class abort condition: if it grows unbounded, you are silently losing your safety net.
The point of no return
Every migration has a step after which going back means running a new migration in the other direction. Usually it is stopping reverse synchronisation, deleting old data, or allowing the new system to accept writes the old schema cannot represent. A good playbook names this step, places it as late as possible, and puts its own gate in front of it.
Pushing the point of no return later has a cost. Reverse sync must be built and operated, both systems are paid for, and new features that need the new schema must wait or be written to be backwards compatible. Teams often discover mid-migration that product work is blocked on a schema change that reverse sync cannot carry. Decide upfront how long the soak after the write switch will be, typically one to four weeks covering at least one month-end or other business cycle, and refuse schema changes that break reverse sync until the gate passes.
Decommissioning
Decommissioning is part of the migration, not a cleanup task for later. Until the old system is gone, it costs money, carries security patches, and tempts someone to read from it. The steps: confirm from access logs that nothing has connected for the agreed period; revoke credentials before deleting anything, so a forgotten consumer fails loudly while recovery is still possible; archive data according to retention obligations; delete infrastructure; and remove the facade's old route and flags so the code no longer branches. Close the migration with a short review covering what the gates caught, what they missed, and which criteria to reuse; the format in change management works well.
Failure modes
| Failure | How it happens | Prevention |
|---|---|---|
| Unknown consumer breaks | Inventory built from memory | Build Phase 0 from connection and query logs |
| Gate passed on missing data | Comparison job died; zero mismatches reported | Missing metric is a NO-GO |
| Rollback fails when needed | Never rehearsed; reverse sync lagging | Production rollback drill; reverse-sync lag as an abort condition |
| Two changes at once | Schema change shipped with the write switch | Deploy freeze during transitions |
| Endless dual running | No decommission gate or owner | Decommission is a phase with a date and an owner |
| Silent semantic drift | Rows match but meaning differs (rounding, time zones) | Business canary and domain-level reconciliation |
What to do next
- Write the phase table for your migration, with the undo action for each phase, and mark the point of no return.
- Build the Phase 0 inventory from logs, assign an owner to every consumer, and appoint consumer liaisons.
- Express each gate's entry, exit and abort criteria in a file and implement an evaluator that treats missing data as NO-GO.
- Write runbooks for every transition, forward and back, and rehearse the rollback in production before switching writes.
- Assign executor, observer and rollback owner roles for each transition, and give the observer abort authority.
- Schedule decommissioning as a phase with a date, and close with a review that improves the next playbook.