Most migrations that fail do not fail on the new system. They fail in the space between systems: rows that were never copied, writes that landed on one side only, a cutover that could not be undone, or a comparison nobody ran. Whether you are moving a monolith's orders module into a service, a database onto a new engine, or a service onto a new data model, the hard part is the same. Two systems must hold the same truth for a while, and traffic must move between them in steps you can reverse.

This article describes the architecture that makes that safe: a facade that owns routing, a data path built on backfill plus change data capture, a comparator that proves equivalence on real traffic, and a cutover sequence with a defined rollback at every step. It does not cover choosing a migration strategy for a whole portfolio, which is cloud migration patterns, or switching a deploy between two identical environments, which is blue-green deployment.

Advertisement

The moving parts

Five components appear in almost every successful live migration. Naming them early makes the plan concrete and shows what must be built before any traffic moves.

  • Facade. A single point that receives every request for the migrating capability and decides which system serves it. It can be an API gateway rule, a reverse proxy or a module inside the old codebase. Clients never change their endpoints.
  • Data path. A copy of existing data (backfill) plus a continuous stream of changes (change data capture, CDC) from the current source of truth to the other system.
  • Comparator. Something that sends the same read to both systems and records differences, plus offline checksums over the stored data.
  • Control plane. Feature flags that set, per route, which system is primary and what share of traffic moves, changeable in seconds without a deploy.
  • Reverse path. After writes move, a change stream from new back to old, so rollback does not lose data.
Live migration: one facade, two systems, a change stream and a comparatorClientsunchanged URLsFacade / routerreads flag per routeLegacy systemsource of truth firstNew systemsource of truth laterFlag control planepercent per routeComparatorshadow diffs, metricsprimaryshadowbackfill + CDCreverse CDCboth resultsClients see one endpoint throughout. Data flows old to new until writes cut over, then new to oldfor as long as rollback must stay possible. Every step is reversible except the last.
The migration architecture. The facade and flags decide routing; CDC keeps the two stores converged in whichever direction is current.

Phases, and what makes each one reversible

PhaseWhat changesRollbackExit criterion
0. InventoryNothing; list endpoints, tables, consumers, jobsn/aEvery reader and writer of the data is known
1. FacadeAll traffic passes the facade, still to legacyRemove the ruleLatency and errors unchanged
2. SyncBackfill, then CDC old to newStop CDC, drop new dataLag stays within seconds; checksums match
3. ShadowSampled reads also go to new; results comparedSet sample rate to 0Mismatch rate below target for a week
4. Read cutoverReads served by new, 1% then 10, 50, 100Flag back to legacyError budget intact at 100%
5. Write cutoverNew becomes source of truth; reverse CDC onFlip back; reverse CDC kept old currentStable through a full business cycle
6. DecommissionReverse CDC off, legacy removedNone: point of no returnNobody has read legacy for N weeks

The order matters. Reads move before writes because a wrong read is visible and fixable while a lost write is not. And the point of no return is pushed to the very end, after the new system has carried full load long enough to trust it.

Advertisement

The facade and routing

The facade is the strangler fig pattern made concrete: new functionality grows around the old system until the old one can be cut away. Put it where routing decisions are cheap and observable. A gateway works for HTTP APIs split by path. For logic buried in a monolith, an in-process interface with two implementations behind a flag is often simpler than extracting network calls first.

Route on a stable key rather than on each request at random. Hashing a customer or tenant ID into buckets keeps a given customer on one system, which avoids confusing read-after-write gaps where a customer writes to one system and reads from the other. It also lets you start with internal accounts. The flag mechanics, including bucketing, are covered in feature flag delivery architecture.

Moving data: backfill plus change capture, not dual writes

The tempting shortcut is dual writes: change the application to write to both stores. It fails in ways that are hard to see. If the second write fails after the first succeeded, the stores diverge and nothing records it. If two requests update the same row concurrently, the two stores can apply them in different orders. Retrying makes it worse. The problem and its standard fix are covered in the outbox pattern.

The more reliable design keeps one source of truth and derives the other store from its change log. For a relational source that means logical decoding of the write-ahead log, described in CDC via logical decoding. The sequence is: record the current log position, start a consistent snapshot, backfill it into the new store, then stream changes from the recorded position. Backfill and stream overlap, so the same row can arrive twice or out of order. The apply step must therefore be idempotent and version-aware:

# Apply one change event to the new store. Idempotent and order-safe per key.
def apply(event):
    # event: {"key": ..., "op": "upsert"|"delete", "row": {...}, "version": source_lsn}
    current = new_store.get_version(event["key"])          # None if absent
    if current is not None and current >= event["version"]:
        metrics.inc("cdc.skipped_stale")                   # replay or backfill overlap
        return
    if event["op"] == "delete":
        new_store.tombstone(event["key"], version=event["version"])
    else:
        new_store.upsert(event["key"], transform(event["row"]), version=event["version"])

The version is the source's log position or a row version column. Any write older than what the target holds is skipped, so replays and overlap are harmless, and a crash anywhere can be recovered by restarting from the last committed position. Deletes must be kept as tombstones with a version, or a late backfill copy of the deleted row will resurrect it.

Worked example: 400 million order rows, backfilled at 20,000 rows per second, take 400,000,000 / 20,000 = 20,000 seconds, about 5.6 hours. The source keeps writing at 2,000 changes per second, so about 40 million changes accumulate and the log must be retained for the whole backfill plus margin. After backfill, the stream catches up at the apply rate minus the write rate, here 18,000 per second, so the backlog clears in about 37 minutes. Measure all three numbers before the migration, because a log retention that is too short turns a clean plan into a restart.

Shadow reads and comparison

Data that matches row for row can still produce different answers, because the new code computes things differently. Shadow reads test the full path on real traffic: the facade serves every request from legacy, and for a sample it also calls the new system off the request path and compares the results.

import random

def get_order(order_id, ctx):
    route = flags.route("orders.read", ctx)                # "legacy", "shadow" or "new"
    if route == "new":
        return new_api.get_order(order_id)
    result = legacy_api.get_order(order_id)
    if route == "shadow" and random.random() < flags.sample_rate("orders.read"):
        executor.submit(compare, order_id, result)         # off the request path
    return result                                          # callers always get legacy

def compare(order_id, legacy_result):
    try:
        candidate = new_api.get_order(order_id, timeout=0.5)
    except Exception as e:
        metrics.inc("shadow.error", tags={"type": type(e).__name__})
        return
    diff = diff_normalised(legacy_result, candidate,
                           ignore={"etag", "served_by"},       # known noise
                           round_money=2, utc_times=True)
    metrics.inc("shadow.match" if not diff else "shadow.mismatch")
    if diff:
        mismatch_log.write(order_id=order_id, fields=sorted(diff))

Three rules keep shadowing useful. Never let the shadow call slow down or fail the real request. Normalise known noise such as timestamps, ordering of unordered lists and money rounding before comparing, or real differences drown in false ones. And classify mismatches by field and cause, so the mismatch rate falls as bugs are fixed rather than being explained away. Shadowing writes is much riskier, since side effects such as emails or payments would happen twice, so shadow writes only against a stub or with side effects disabled.

Cutting over reads, then writes

Read cutover is a flag change per route, stepped through percentages with the monitoring of a canary release; see canary release architecture for automated analysis. At each step compare error rates, latency and business metrics such as checkout completion between the two cohorts.

Write cutover is the step that changes the source of truth, and it needs a short fence so that no write lands on the old side after the new side takes over. The usual sequence is: stop accepting writes for the affected keys for a few seconds (return a retryable error or queue them), wait until CDC lag reaches zero, flip the write route, start reverse CDC from new to old, then reopen writes. Routing by tenant bucket lets you do this one bucket at a time, which shrinks each pause and each blast radius.

Reverse CDC is the price of a real rollback. Without it, flipping back after a day of writes on the new system means losing that day, so in practice nobody flips back. With it, legacy stays current and rollback remains a flag change.

Validating that the data really matches

Shadow reads cover what traffic touches. Offline validation covers everything else. Row counts are the minimum but miss changed values, so use bucketed checksums computed identically on both sides:

-- Bucketed checksum, run on both databases and compared bucket by bucket.
SELECT mod(order_id, 1024)                               AS bucket,
       count(*)                                          AS row_count,
       sum(hashtext(concat_ws('|', order_id, coalesce(status, '~'), coalesce(total_cents::text, '~')))) AS checksum
FROM   orders
WHERE  updated_at < :cutoff          -- same cutoff on both sides
GROUP  BY 1
ORDER  BY 1;

Buckets whose checksums differ narrow the search from 400 million rows to a few hundred thousand, which you then compare row by row. Use a cutoff on the update time so in-flight changes do not show up as false mismatches, and make NULLs explicit: a plain concatenation with a NULL column yields NULL, and sum() silently skips it. If the schemas differ, compute the checksum over the transformed values, and use the same hash function on both sides; the hashtext function here is PostgreSQL-specific. Finally, check business invariants, such as every order having a customer or balances summing to the ledger, which catch transformation bugs that equality checks cannot.

Trade-offs: when the full architecture is worth it

ApproachBest whenCost
Big-bang cutover in a maintenance windowSmall data, tolerant users, simple rollback by restoreDowntime; everything is tested at once
Incremental with facade and CDCLarge data, no downtime allowed, long-lived systemWeeks of running two systems; CDC and comparator to build
Per-tenant movesMulti-tenant data that partitions cleanlyRouting table to maintain; cross-tenant queries get harder

The incremental architecture is not free. Running two systems doubles on-call surface, and the change stream and comparator are real software that needs tests. If the data fits in a restore that takes an hour and the business accepts an hour of downtime, a rehearsed big-bang cutover with a backup is often the cheaper and safer choice.

Failure modes

  • An unlisted writer. A batch job or another team's service writes to the old database directly, bypassing the facade. CDC catches its writes before cutover, but after cutover they land on a store nobody reads. Find every writer in phase 0, or revoke write grants on legacy after cutover.
  • CDC lag spikes. A bulk update produces millions of changes; the new side falls behind and shadow mismatches soar. Alert on lag and pause cutover steps while it is high.
  • Log retention loss. The replication slot or log is removed before the stream catches up, and the only fix is a new backfill.
  • Non-deterministic transforms. A transform using the current time or a random ID produces different values on replay; derive everything from the source row.
  • Rollback that was never tested. Reverse CDC exists on paper but has never carried traffic. Rehearse the flip back in staging and once in production at low percentage.

What to do next

  1. Inventory every reader and writer of the data you are moving.
  2. Put a facade in front of the capability with routing by a stable key.
  3. Measure backfill rate, write rate and log retention and do the arithmetic.
  4. Build the idempotent apply step and a bucketed checksum job.
  5. Run shadow reads until the mismatch rate meets a written target.
  6. Rehearse write cutover and the rollback flip, then execute bucket by bucket.
Key takeaway: A safe live migration has a facade that owns routing, one source of truth at any moment, a change stream that keeps the other store converged, a comparator that proves equivalence on real traffic, and a rollback path at every step until the very last. Prefer backfill plus change capture to dual writes, make the apply step idempotent and version-aware, move reads before writes, and keep reverse sync running until you are sure.