Most Hadoop-to-cloud migrations do not fail because DistCp is slow or because a Hive query will not run on Spark. They fail as programs. Scope keeps growing, nobody owns the 400 jobs whose author left, the old cluster can never be switched off because one team still needs it, and the dual-run bill keeps rising until finance asks why the cloud costs more than the data centre did.
This playbook is about running the migration, not executing it. The mechanics of moving data, metadata, jobs and security are covered in the Hadoop to cloud migration guide; read it alongside this one. Here you get the phase structure and the evidence each gate needs, a scoring model that decides what happens to every workload, a cost model that makes the dual-run window visible, how to run waves as a factory, a cutover runbook with rollback triggers, and how to actually finish by switching the cluster off.
Why migrations stall
Look at stalled programs and the same patterns recur. The inventory was a spreadsheet of tables, not of workloads and their consumers, so dependencies appear during cutover. Everything was treated as must-migrate, so dead jobs consumed as much effort as revenue-critical ones. Success was measured by data moved, not by workloads switched off on-premises, so the program reported 90 percent done for a year. And there was no date or decision rule for decommissioning, so the cluster kept running for the last few stragglers.
The fix is to treat the migration as a sequence of decisions with evidence. Each phase ends with a gate: a short meeting where named owners show specific artefacts, and the program either moves on or does not. The article on Hadoop's decline covers why organisations leave; this one assumes the decision is made and focuses on finishing.
Six phases and their gates
| Phase | Goal | Exit evidence | Typical length |
|---|---|---|---|
| 1 Assess | Know every workload, its owner, consumers and value | Workload register with disposition and owner for 100 percent of jobs | 4-8 weeks |
| 2 Mobilize | Make the target safe and fundable | Landing zone security sign-off, network path tested, budget and team in place | 4-8 weeks |
| 3 Pilot wave | Prove the method on one real domain | Pilot cut over, rollback rehearsed, runbook and estimates corrected | 6-10 weeks |
| 4 Waves | Move the rest at a steady cadence | Each wave meets the cutover exit criteria below | Months, by size |
| 5 Decommission | Stop paying for the old estate | No jobs on cluster for the agreed quiet period, archive verified, contracts ended | 4-12 weeks |
| 6 Optimize | Make the new estate cheaper than the old | Cost per workload tracked and trending down | Ongoing |
The durations are planning ranges from typical programs, not benchmarks; your own pilot is what calibrates them. Two rules make gates work. First, a gate is pass or fail on its evidence, not on effort spent. Second, the program sponsor attends, because the gates are where scope and budget decisions are made.
Assess: give every workload a disposition
The unit of migration is a workload: a scheduled job or pipeline plus the tables it writes and the consumers that read them. Pull workloads from the scheduler (Oozie coordinators, Airflow DAGs, cron), join them to YARN application history for resource use and to the metastore and audit logs for the tables they touch. Then give each one of four dispositions.
- Retire: no consumer read its outputs in the last 90 days, or a newer pipeline duplicates it. This is the cheapest migration there is.
- Rehost: runs nearly unchanged on managed Spark or Hive in the cloud, with path and configuration changes only.
- Replatform: moves to a different engine or table format, such as Hive to Spark SQL on Iceberg, with moderate rewriting.
- Refactor: needs redesign, for example MapReduce code, HBase-coupled jobs or anything relying on HDFS-specific behaviour such as atomic directory renames.
from dataclasses import dataclass
@dataclass
class Workload:
name: str
days_since_output_read: int
consumers: int
engine: str # "hive", "spark", "mapreduce", "pig", "hbase"
uses_hdfs_rename_semantics: bool
loc: int # lines of job code
business_tier: int # 1 = revenue critical ... 3 = internal
def disposition(w: Workload) -> str:
if w.days_since_output_read > 90 and w.business_tier > 1:
return "retire" # confirm with the owner before deleting
if w.engine in ("mapreduce", "pig", "hbase") or w.uses_hdfs_rename_semantics:
return "refactor"
if w.engine == "hive" and w.loc > 500:
return "replatform"
return "rehost"
def wave_priority(w: Workload) -> int:
# Lower score = earlier wave. Low risk and low effort first; tier-1 workloads after the pilot has proven the method
effort = {"rehost": 1, "replatform": 2, "refactor": 3}.get(disposition(w), 0)
return effort * 10 + (4 - w.business_tier) * 3 + min(w.consumers, 10)The thresholds are starting points to tune with owners, not rules. In practice the retire bucket is where programs win or lose time: every workload retired is one that never needs testing, dual running or a cutover window. Treat retirement as a formal decision with an owner's sign-off and a period during which the data is kept read-only, so that if a consumer was missed the output can be restored quickly.
The business case: model the dual-run window
A migration business case compares the steady-state cost of the old estate with the new one. The number that sinks programs is neither of those. It is the overlap: months in which you pay for the cluster, its support contracts and data centre space, plus growing cloud spend. A small model makes that visible and gives the sponsor a reason to fund retirement and decommissioning work.
def program_cost(months, onprem_monthly, cloud_steady_monthly, migrated_by_month,
decommission_month, one_off_migration_cost):
# migrated_by_month[m] = fraction of workload running in cloud at month m (0..1)
total = one_off_migration_cost
for m in range(months):
cloud = cloud_steady_monthly * migrated_by_month[m]
onprem = onprem_monthly if m < decommission_month else 0
total += cloud + onprem
return total
# Illustrative inputs only -- replace with your own contracts and estimates
months = 24
ramp = [min(1.0, max(0.0, (m - 3) / 12)) for m in range(months)] # migrate months 3-15
base = dict(months=months, onprem_monthly=180_000, cloud_steady_monthly=140_000,
migrated_by_month=ramp, one_off_migration_cost=900_000)
on_time = program_cost(decommission_month=16, **base)
late = program_cost(decommission_month=22, **base)
print(f"decommission month 16: {on_time:,.0f}")
print(f"decommission month 22: {late:,.0f} (+{late - on_time:,.0f})")With those illustrative inputs, slipping decommissioning by six months adds six months of on-premises cost, 1.08 million, while the cloud bill is unchanged. That is the lever the sponsor controls. The model also shows why the cloud steady state must be estimated honestly: if the 140,000 figure assumes autoscaling, spot capacity and tiered storage that nobody has built, the business case is fiction. Re-run the model at every gate with actual numbers.
Mobilize: landing zone and team before data
- Landing zone. Accounts or projects per environment, network connectivity sized from measured throughput, encryption keys, a catalog, IAM roles per workload class, logging and budgets with alerts. Security sign-off is gate evidence, not a later task.
- Identity mapping. Decide how Kerberos principals and Ranger policies map to cloud roles and catalog grants before any data lands; the Ranger article covers what those policies express today.
- Orchestration target. Choose the scheduler that replaces Oozie early, because every rehosted job needs a new definition; the Airflow with Hadoop guide shows the common pattern.
- Team topology. A small platform team owns the landing zone and tooling; a migration factory team moves waves; each domain supplies a workload owner who signs off on validation. Write the RACI down: the factory moves the workload, but only the owner accepts it.
- Freeze policy. Agree that new pipelines are built only in the cloud from the start of Mobilize, so the inventory stops growing.
Waves: run the migration as a factory
A wave is a set of workloads that share tables or consumers and can cut over together, typically one domain or one table family. Size waves so a team can finish one in two to four weeks; bigger waves hide problems until cutover. Order them by the priority score, with the pilot chosen to be representative but not tier 1.
Each wave follows the same pipeline: bulk and incremental data copy, catalog registration, job port, dual run, comparison, cutover, then a quiet period. The mechanics of each step are in the advanced DistCp article and the execution guide. What the playbook adds is a weekly dashboard that the program reviews: workloads by disposition and state, workloads switched off on-premises (the only progress metric that counts), dual-run comparisons failing, and actual versus planned cloud spend.
Cutover exit criteria for a workload should be specific: outputs match on row counts and agreed column checksums for a set number of consecutive runs, downstream consumers confirm they read from the new location, SLAs are met for the same period, and the on-premises job is disabled, not deleted, with its output directories made read-only.
The cutover runbook, with rollback triggers
T-7d Go/no-go: dual-run comparisons green for N runs; owner sign-off; consumers listed
T-2d Freeze schema changes on affected tables; announce window to consumers
T-1d Final incremental sync; verify counts and checksums; snapshot on-prem tables
T-0 Pause on-prem schedules for the wave
Final delta sync; register final partitions in the cloud catalog
Switch consumers (views, DNS, connection configs) to cloud tables
Enable cloud schedules; watch the first run end to end
T+1d Compare first cloud outputs with the T-1d snapshot baseline; consumer smoke tests
T+7d Quiet period ends: disable on-prem jobs permanently; make on-prem data read-only
T+30d Archive or delete on-prem copies per retention policy
ROLLBACK if any of:
- first cloud run fails and cannot be fixed inside the agreed window (e.g. 4h)
- output mismatch beyond tolerance on a tier-1 table
- a consumer cannot read the new location and has no workaround
Rollback = re-point consumers to on-prem, resume on-prem schedules, replay any cloud-only writesThe rollback steps only work if writes during the window are known. That is why the pause comes before the final sync, and why consumers switch through an indirection you control, such as a view, a catalog alias or a configuration value, rather than through hard-coded paths in fifty notebooks. Rehearse rollback in the pilot wave; a rollback that has never run is a hope, not a plan.
Decommission: actually switching it off
Decommissioning is a phase with its own work, not a celebration. Its steps: confirm from YARN and audit logs that no jobs or reads hit the cluster during the quiet period; take a final archive of anything retention or legal hold requires, verify it can be read from the archive tier, and record where it lives; revoke service accounts and Kerberos principals; end support and licence contracts on their notice dates, which often need months of lead time, so start that clock at G3; then wipe disks to your media-sanitisation standard and return or recycle hardware.
Keep a short list of who asked for an exception and why. Exceptions are where clusters go to live forever. The usual answer is a time-boxed small cluster or a read-only archive in object storage with a query engine on top, both cheaper than the old estate.
Optimize: make the new bill smaller than the old one
Cloud cost after migration is driven by habits carried over from a fixed-capacity cluster. Long-running clusters sized for peak, no lifecycle policies on raw data, every table queried by full scans because it was never partitioned or converted to a table format, and no owner per cost line. Set guardrails in phase 6: mandatory tags mapping spend to workloads and owners, budgets per domain, autoscaling or job-scoped clusters, spot or preemptible capacity for retry-tolerant batch, storage tiering for cold partitions, and compaction for tables. Converting the hottest tables to Iceberg, as described in the Iceberg and Trino stack article, tends to pay off in both query cost and operability.
Failure modes and trade-offs
| Failure | Early signal | Countermeasure |
|---|---|---|
| Endless 90 percent | Data moved rises, workloads switched off flat | Report only switched-off workloads as progress |
| Dual-run cost overrun | Cloud spend grows while on-prem stays fixed | Fund retirement; set the decommission date at G2 |
| Hidden consumers | Breakage reports after cutover | Audit-log driven consumer lists; read-only quiet period |
| Big-bang pressure | Waves growing to cover whole platforms | Cap wave size by team capacity |
| Lift-and-shift cost shock | Cloud bill above on-prem at steady state | Plan replatforming for heavy workloads; FinOps from day one |
| Owner fatigue | Sign-offs slipping | Sponsor-chaired gates; small, frequent asks |
The central trade-off is speed against modernisation. Rehosting everything is fastest and shortens dual run, but carries old inefficiencies into a metered environment. Refactoring everything delivers a better platform but extends the overlap. Most programs rehost the long tail quickly, replatform the expensive core, and retire aggressively.
What to do next
- Export the scheduler's job list and YARN history, and build the workload register with owners and last-read dates.
- Run the disposition scoring and review the retire list with owners first.
- Put your real contract and cloud estimates into the cost model and agree a decommission month with the sponsor.
- Write gate evidence for G1 to G5 and put the gates in the calendar.
- Pick a pilot domain, write its T-minus runbook from the template above, and rehearse rollback.
- Start support and licence notice periods early, and track switched-off workloads as the headline metric.