Most migrations that fail do not fail on the new system. They fail in the space between systems: rows that were never copied, writes that landed on one side only, a cutover that could not be undone, or a comparison nobody ran. Whether you are moving a monolith's orders module into a service, a database onto a new engine, or a service onto a new data model, the hard part is the same. Two systems must hold the same truth for a while, and traffic must move between them in steps you can reverse.
This article describes the architecture that makes that safe: a facade that owns routing, a data path built on backfill plus change data capture, a comparator that proves equivalence on real traffic, and a cutover sequence with a defined rollback at every step. It does not cover choosing a migration strategy for a whole portfolio, which is cloud migration patterns, or switching a deploy between two identical environments, which is blue-green deployment.
The moving parts
Five components appear in almost every successful live migration. Naming them early makes the plan concrete and shows what must be built before any traffic moves.
- Facade. A single point that receives every request for the migrating capability and decides which system serves it. It can be an API gateway rule, a reverse proxy or a module inside the old codebase. Clients never change their endpoints.
- Data path. A copy of existing data (backfill) plus a continuous stream of changes (change data capture, CDC) from the current source of truth to the other system.
- Comparator. Something that sends the same read to both systems and records differences, plus offline checksums over the stored data.
- Control plane. Feature flags that set, per route, which system is primary and what share of traffic moves, changeable in seconds without a deploy.
- Reverse path. After writes move, a change stream from new back to old, so rollback does not lose data.
Phases, and what makes each one reversible
| Phase | What changes | Rollback | Exit criterion |
|---|---|---|---|
| 0. Inventory | Nothing; list endpoints, tables, consumers, jobs | n/a | Every reader and writer of the data is known |
| 1. Facade | All traffic passes the facade, still to legacy | Remove the rule | Latency and errors unchanged |
| 2. Sync | Backfill, then CDC old to new | Stop CDC, drop new data | Lag stays within seconds; checksums match |
| 3. Shadow | Sampled reads also go to new; results compared | Set sample rate to 0 | Mismatch rate below target for a week |
| 4. Read cutover | Reads served by new, 1% then 10, 50, 100 | Flag back to legacy | Error budget intact at 100% |
| 5. Write cutover | New becomes source of truth; reverse CDC on | Flip back; reverse CDC kept old current | Stable through a full business cycle |
| 6. Decommission | Reverse CDC off, legacy removed | None: point of no return | Nobody has read legacy for N weeks |
The order matters. Reads move before writes because a wrong read is visible and fixable while a lost write is not. And the point of no return is pushed to the very end, after the new system has carried full load long enough to trust it.
The facade and routing
The facade is the strangler fig pattern made concrete: new functionality grows around the old system until the old one can be cut away. Put it where routing decisions are cheap and observable. A gateway works for HTTP APIs split by path. For logic buried in a monolith, an in-process interface with two implementations behind a flag is often simpler than extracting network calls first.
Route on a stable key rather than on each request at random. Hashing a customer or tenant ID into buckets keeps a given customer on one system, which avoids confusing read-after-write gaps where a customer writes to one system and reads from the other. It also lets you start with internal accounts. The flag mechanics, including bucketing, are covered in feature flag delivery architecture.
Moving data: backfill plus change capture, not dual writes
The tempting shortcut is dual writes: change the application to write to both stores. It fails in ways that are hard to see. If the second write fails after the first succeeded, the stores diverge and nothing records it. If two requests update the same row concurrently, the two stores can apply them in different orders. Retrying makes it worse. The problem and its standard fix are covered in the outbox pattern.
The more reliable design keeps one source of truth and derives the other store from its change log. For a relational source that means logical decoding of the write-ahead log, described in CDC via logical decoding. The sequence is: record the current log position, start a consistent snapshot, backfill it into the new store, then stream changes from the recorded position. Backfill and stream overlap, so the same row can arrive twice or out of order. The apply step must therefore be idempotent and version-aware:
# Apply one change event to the new store. Idempotent and order-safe per key.
def apply(event):
# event: {"key": ..., "op": "upsert"|"delete", "row": {...}, "version": source_lsn}
current = new_store.get_version(event["key"]) # None if absent
if current is not None and current >= event["version"]:
metrics.inc("cdc.skipped_stale") # replay or backfill overlap
return
if event["op"] == "delete":
new_store.tombstone(event["key"], version=event["version"])
else:
new_store.upsert(event["key"], transform(event["row"]), version=event["version"])The version is the source's log position or a row version column. Any write older than what the target holds is skipped, so replays and overlap are harmless, and a crash anywhere can be recovered by restarting from the last committed position. Deletes must be kept as tombstones with a version, or a late backfill copy of the deleted row will resurrect it.
Worked example: 400 million order rows, backfilled at 20,000 rows per second, take 400,000,000 / 20,000 = 20,000 seconds, about 5.6 hours. The source keeps writing at 2,000 changes per second, so about 40 million changes accumulate and the log must be retained for the whole backfill plus margin. After backfill, the stream catches up at the apply rate minus the write rate, here 18,000 per second, so the backlog clears in about 37 minutes. Measure all three numbers before the migration, because a log retention that is too short turns a clean plan into a restart.
Shadow reads and comparison
Data that matches row for row can still produce different answers, because the new code computes things differently. Shadow reads test the full path on real traffic: the facade serves every request from legacy, and for a sample it also calls the new system off the request path and compares the results.
import random
def get_order(order_id, ctx):
route = flags.route("orders.read", ctx) # "legacy", "shadow" or "new"
if route == "new":
return new_api.get_order(order_id)
result = legacy_api.get_order(order_id)
if route == "shadow" and random.random() < flags.sample_rate("orders.read"):
executor.submit(compare, order_id, result) # off the request path
return result # callers always get legacy
def compare(order_id, legacy_result):
try:
candidate = new_api.get_order(order_id, timeout=0.5)
except Exception as e:
metrics.inc("shadow.error", tags={"type": type(e).__name__})
return
diff = diff_normalised(legacy_result, candidate,
ignore={"etag", "served_by"}, # known noise
round_money=2, utc_times=True)
metrics.inc("shadow.match" if not diff else "shadow.mismatch")
if diff:
mismatch_log.write(order_id=order_id, fields=sorted(diff))Three rules keep shadowing useful. Never let the shadow call slow down or fail the real request. Normalise known noise such as timestamps, ordering of unordered lists and money rounding before comparing, or real differences drown in false ones. And classify mismatches by field and cause, so the mismatch rate falls as bugs are fixed rather than being explained away. Shadowing writes is much riskier, since side effects such as emails or payments would happen twice, so shadow writes only against a stub or with side effects disabled.
Cutting over reads, then writes
Read cutover is a flag change per route, stepped through percentages with the monitoring of a canary release; see canary release architecture for automated analysis. At each step compare error rates, latency and business metrics such as checkout completion between the two cohorts.
Write cutover is the step that changes the source of truth, and it needs a short fence so that no write lands on the old side after the new side takes over. The usual sequence is: stop accepting writes for the affected keys for a few seconds (return a retryable error or queue them), wait until CDC lag reaches zero, flip the write route, start reverse CDC from new to old, then reopen writes. Routing by tenant bucket lets you do this one bucket at a time, which shrinks each pause and each blast radius.
Reverse CDC is the price of a real rollback. Without it, flipping back after a day of writes on the new system means losing that day, so in practice nobody flips back. With it, legacy stays current and rollback remains a flag change.
Validating that the data really matches
Shadow reads cover what traffic touches. Offline validation covers everything else. Row counts are the minimum but miss changed values, so use bucketed checksums computed identically on both sides:
-- Bucketed checksum, run on both databases and compared bucket by bucket.
SELECT mod(order_id, 1024) AS bucket,
count(*) AS row_count,
sum(hashtext(concat_ws('|', order_id, coalesce(status, '~'), coalesce(total_cents::text, '~')))) AS checksum
FROM orders
WHERE updated_at < :cutoff -- same cutoff on both sides
GROUP BY 1
ORDER BY 1;Buckets whose checksums differ narrow the search from 400 million rows to a few hundred thousand, which you then compare row by row. Use a cutoff on the update time so in-flight changes do not show up as false mismatches, and make NULLs explicit: a plain concatenation with a NULL column yields NULL, and sum() silently skips it. If the schemas differ, compute the checksum over the transformed values, and use the same hash function on both sides; the hashtext function here is PostgreSQL-specific. Finally, check business invariants, such as every order having a customer or balances summing to the ledger, which catch transformation bugs that equality checks cannot.
Trade-offs: when the full architecture is worth it
| Approach | Best when | Cost |
|---|---|---|
| Big-bang cutover in a maintenance window | Small data, tolerant users, simple rollback by restore | Downtime; everything is tested at once |
| Incremental with facade and CDC | Large data, no downtime allowed, long-lived system | Weeks of running two systems; CDC and comparator to build |
| Per-tenant moves | Multi-tenant data that partitions cleanly | Routing table to maintain; cross-tenant queries get harder |
The incremental architecture is not free. Running two systems doubles on-call surface, and the change stream and comparator are real software that needs tests. If the data fits in a restore that takes an hour and the business accepts an hour of downtime, a rehearsed big-bang cutover with a backup is often the cheaper and safer choice.
Failure modes
- An unlisted writer. A batch job or another team's service writes to the old database directly, bypassing the facade. CDC catches its writes before cutover, but after cutover they land on a store nobody reads. Find every writer in phase 0, or revoke write grants on legacy after cutover.
- CDC lag spikes. A bulk update produces millions of changes; the new side falls behind and shadow mismatches soar. Alert on lag and pause cutover steps while it is high.
- Log retention loss. The replication slot or log is removed before the stream catches up, and the only fix is a new backfill.
- Non-deterministic transforms. A transform using the current time or a random ID produces different values on replay; derive everything from the source row.
- Rollback that was never tested. Reverse CDC exists on paper but has never carried traffic. Rehearse the flip back in staging and once in production at low percentage.
What to do next
- Inventory every reader and writer of the data you are moving.
- Put a facade in front of the capability with routing by a stable key.
- Measure backfill rate, write rate and log retention and do the arithmetic.
- Build the idempotent apply step and a bucketed checksum job.
- Run shadow reads until the mismatch rate meets a written target.
- Rehearse write cutover and the rollback flip, then execute bucket by bucket.