Most Hadoop-to-cloud migrations do not fail because DistCp is slow or because a Hive query will not run on Spark. They fail as programs. Scope keeps growing, nobody owns the 400 jobs whose author left, the old cluster can never be switched off because one team still needs it, and the dual-run bill keeps rising until finance asks why the cloud costs more than the data centre did.

This playbook is about running the migration, not executing it. The mechanics of moving data, metadata, jobs and security are covered in the Hadoop to cloud migration guide; read it alongside this one. Here you get the phase structure and the evidence each gate needs, a scoring model that decides what happens to every workload, a cost model that makes the dual-run window visible, how to run waves as a factory, a cutover runbook with rollback triggers, and how to actually finish by switching the cluster off.

Advertisement

Why migrations stall

Look at stalled programs and the same patterns recur. The inventory was a spreadsheet of tables, not of workloads and their consumers, so dependencies appear during cutover. Everything was treated as must-migrate, so dead jobs consumed as much effort as revenue-critical ones. Success was measured by data moved, not by workloads switched off on-premises, so the program reported 90 percent done for a year. And there was no date or decision rule for decommissioning, so the cluster kept running for the last few stragglers.

The fix is to treat the migration as a sequence of decisions with evidence. Each phase ends with a gate: a short meeting where named owners show specific artefacts, and the program either moves on or does not. The article on Hadoop's decline covers why organisations leave; this one assumes the decision is made and focuses on finishing.

Six phases and their gates

A Hadoop exit as a program: six phases, each closed by a gate1 Assessinventory, score2 Mobilizelanding zone, team3 Pilot waveone domain4 Wavesfactory cadence5 Decommissionfreeze, archive6 OptimizeFinOpsG1G2G3G4G5Gate evidence (examples)G1: every workload has a disposition and an ownerG2: landing zone passes security review; budget approvedG3: pilot cut over, rollback rehearsed, runbook updatedG4: zero jobs on cluster for N weeks; archive verifiedG5: hardware and licences returned; cost per job trackedDual-run cost windowstarts when the first cloud workload runsends only when the cluster is switched offevery month of delay pays for both estatesthe program's main financial riskshorten it with retirement, not heroicsExecution detail (data, metadata, jobs, security lanes) lives inside phases 3 and 4.This playbook is about deciding, sequencing, funding and finishing.
Figure 1. Six phases, each closed by a gate with concrete evidence. The dual-run window, when you pay for both estates, opens at the first cloud workload and closes only at decommission, so the plan is built to shorten it.
PhaseGoalExit evidenceTypical length
1 AssessKnow every workload, its owner, consumers and valueWorkload register with disposition and owner for 100 percent of jobs4-8 weeks
2 MobilizeMake the target safe and fundableLanding zone security sign-off, network path tested, budget and team in place4-8 weeks
3 Pilot waveProve the method on one real domainPilot cut over, rollback rehearsed, runbook and estimates corrected6-10 weeks
4 WavesMove the rest at a steady cadenceEach wave meets the cutover exit criteria belowMonths, by size
5 DecommissionStop paying for the old estateNo jobs on cluster for the agreed quiet period, archive verified, contracts ended4-12 weeks
6 OptimizeMake the new estate cheaper than the oldCost per workload tracked and trending downOngoing

The durations are planning ranges from typical programs, not benchmarks; your own pilot is what calibrates them. Two rules make gates work. First, a gate is pass or fail on its evidence, not on effort spent. Second, the program sponsor attends, because the gates are where scope and budget decisions are made.

Advertisement

Assess: give every workload a disposition

The unit of migration is a workload: a scheduled job or pipeline plus the tables it writes and the consumers that read them. Pull workloads from the scheduler (Oozie coordinators, Airflow DAGs, cron), join them to YARN application history for resource use and to the metastore and audit logs for the tables they touch. Then give each one of four dispositions.

  • Retire: no consumer read its outputs in the last 90 days, or a newer pipeline duplicates it. This is the cheapest migration there is.
  • Rehost: runs nearly unchanged on managed Spark or Hive in the cloud, with path and configuration changes only.
  • Replatform: moves to a different engine or table format, such as Hive to Spark SQL on Iceberg, with moderate rewriting.
  • Refactor: needs redesign, for example MapReduce code, HBase-coupled jobs or anything relying on HDFS-specific behaviour such as atomic directory renames.
from dataclasses import dataclass

@dataclass
class Workload:
    name: str
    days_since_output_read: int
    consumers: int
    engine: str              # "hive", "spark", "mapreduce", "pig", "hbase"
    uses_hdfs_rename_semantics: bool
    loc: int                 # lines of job code
    business_tier: int       # 1 = revenue critical ... 3 = internal

def disposition(w: Workload) -> str:
    if w.days_since_output_read > 90 and w.business_tier > 1:
        return "retire"                      # confirm with the owner before deleting
    if w.engine in ("mapreduce", "pig", "hbase") or w.uses_hdfs_rename_semantics:
        return "refactor"
    if w.engine == "hive" and w.loc > 500:
        return "replatform"
    return "rehost"

def wave_priority(w: Workload) -> int:
    # Lower score = earlier wave. Low risk and low effort first; tier-1 workloads after the pilot has proven the method
    effort = {"rehost": 1, "replatform": 2, "refactor": 3}.get(disposition(w), 0)
    return effort * 10 + (4 - w.business_tier) * 3 + min(w.consumers, 10)

The thresholds are starting points to tune with owners, not rules. In practice the retire bucket is where programs win or lose time: every workload retired is one that never needs testing, dual running or a cutover window. Treat retirement as a formal decision with an owner's sign-off and a period during which the data is kept read-only, so that if a consumer was missed the output can be restored quickly.

The business case: model the dual-run window

A migration business case compares the steady-state cost of the old estate with the new one. The number that sinks programs is neither of those. It is the overlap: months in which you pay for the cluster, its support contracts and data centre space, plus growing cloud spend. A small model makes that visible and gives the sponsor a reason to fund retirement and decommissioning work.

def program_cost(months, onprem_monthly, cloud_steady_monthly, migrated_by_month,
                 decommission_month, one_off_migration_cost):
    # migrated_by_month[m] = fraction of workload running in cloud at month m (0..1)
    total = one_off_migration_cost
    for m in range(months):
        cloud = cloud_steady_monthly * migrated_by_month[m]
        onprem = onprem_monthly if m < decommission_month else 0
        total += cloud + onprem
    return total

# Illustrative inputs only -- replace with your own contracts and estimates
months = 24
ramp = [min(1.0, max(0.0, (m - 3) / 12)) for m in range(months)]   # migrate months 3-15
base = dict(months=months, onprem_monthly=180_000, cloud_steady_monthly=140_000,
            migrated_by_month=ramp, one_off_migration_cost=900_000)

on_time = program_cost(decommission_month=16, **base)
late    = program_cost(decommission_month=22, **base)
print(f"decommission month 16: {on_time:,.0f}")
print(f"decommission month 22: {late:,.0f}  (+{late - on_time:,.0f})")

With those illustrative inputs, slipping decommissioning by six months adds six months of on-premises cost, 1.08 million, while the cloud bill is unchanged. That is the lever the sponsor controls. The model also shows why the cloud steady state must be estimated honestly: if the 140,000 figure assumes autoscaling, spot capacity and tiered storage that nobody has built, the business case is fiction. Re-run the model at every gate with actual numbers.

Mobilize: landing zone and team before data

  • Landing zone. Accounts or projects per environment, network connectivity sized from measured throughput, encryption keys, a catalog, IAM roles per workload class, logging and budgets with alerts. Security sign-off is gate evidence, not a later task.
  • Identity mapping. Decide how Kerberos principals and Ranger policies map to cloud roles and catalog grants before any data lands; the Ranger article covers what those policies express today.
  • Orchestration target. Choose the scheduler that replaces Oozie early, because every rehosted job needs a new definition; the Airflow with Hadoop guide shows the common pattern.
  • Team topology. A small platform team owns the landing zone and tooling; a migration factory team moves waves; each domain supplies a workload owner who signs off on validation. Write the RACI down: the factory moves the workload, but only the owner accepts it.
  • Freeze policy. Agree that new pipelines are built only in the cloud from the start of Mobilize, so the inventory stops growing.

Waves: run the migration as a factory

A wave is a set of workloads that share tables or consumers and can cut over together, typically one domain or one table family. Size waves so a team can finish one in two to four weeks; bigger waves hide problems until cutover. Order them by the priority score, with the pilot chosen to be representative but not tier 1.

Each wave follows the same pipeline: bulk and incremental data copy, catalog registration, job port, dual run, comparison, cutover, then a quiet period. The mechanics of each step are in the advanced DistCp article and the execution guide. What the playbook adds is a weekly dashboard that the program reviews: workloads by disposition and state, workloads switched off on-premises (the only progress metric that counts), dual-run comparisons failing, and actual versus planned cloud spend.

Cutover exit criteria for a workload should be specific: outputs match on row counts and agreed column checksums for a set number of consecutive runs, downstream consumers confirm they read from the new location, SLAs are met for the same period, and the on-premises job is disabled, not deleted, with its output directories made read-only.

The cutover runbook, with rollback triggers

T-7d   Go/no-go: dual-run comparisons green for N runs; owner sign-off; consumers listed
T-2d   Freeze schema changes on affected tables; announce window to consumers
T-1d   Final incremental sync; verify counts and checksums; snapshot on-prem tables
T-0    Pause on-prem schedules for the wave
       Final delta sync; register final partitions in the cloud catalog
       Switch consumers (views, DNS, connection configs) to cloud tables
       Enable cloud schedules; watch the first run end to end
T+1d   Compare first cloud outputs with the T-1d snapshot baseline; consumer smoke tests
T+7d   Quiet period ends: disable on-prem jobs permanently; make on-prem data read-only
T+30d  Archive or delete on-prem copies per retention policy

ROLLBACK if any of:
  - first cloud run fails and cannot be fixed inside the agreed window (e.g. 4h)
  - output mismatch beyond tolerance on a tier-1 table
  - a consumer cannot read the new location and has no workaround
Rollback = re-point consumers to on-prem, resume on-prem schedules, replay any cloud-only writes

The rollback steps only work if writes during the window are known. That is why the pause comes before the final sync, and why consumers switch through an indirection you control, such as a view, a catalog alias or a configuration value, rather than through hard-coded paths in fifty notebooks. Rehearse rollback in the pilot wave; a rollback that has never run is a hope, not a plan.

Decommission: actually switching it off

Decommissioning is a phase with its own work, not a celebration. Its steps: confirm from YARN and audit logs that no jobs or reads hit the cluster during the quiet period; take a final archive of anything retention or legal hold requires, verify it can be read from the archive tier, and record where it lives; revoke service accounts and Kerberos principals; end support and licence contracts on their notice dates, which often need months of lead time, so start that clock at G3; then wipe disks to your media-sanitisation standard and return or recycle hardware.

Keep a short list of who asked for an exception and why. Exceptions are where clusters go to live forever. The usual answer is a time-boxed small cluster or a read-only archive in object storage with a query engine on top, both cheaper than the old estate.

Optimize: make the new bill smaller than the old one

Cloud cost after migration is driven by habits carried over from a fixed-capacity cluster. Long-running clusters sized for peak, no lifecycle policies on raw data, every table queried by full scans because it was never partitioned or converted to a table format, and no owner per cost line. Set guardrails in phase 6: mandatory tags mapping spend to workloads and owners, budgets per domain, autoscaling or job-scoped clusters, spot or preemptible capacity for retry-tolerant batch, storage tiering for cold partitions, and compaction for tables. Converting the hottest tables to Iceberg, as described in the Iceberg and Trino stack article, tends to pay off in both query cost and operability.

Failure modes and trade-offs

FailureEarly signalCountermeasure
Endless 90 percentData moved rises, workloads switched off flatReport only switched-off workloads as progress
Dual-run cost overrunCloud spend grows while on-prem stays fixedFund retirement; set the decommission date at G2
Hidden consumersBreakage reports after cutoverAudit-log driven consumer lists; read-only quiet period
Big-bang pressureWaves growing to cover whole platformsCap wave size by team capacity
Lift-and-shift cost shockCloud bill above on-prem at steady statePlan replatforming for heavy workloads; FinOps from day one
Owner fatigueSign-offs slippingSponsor-chaired gates; small, frequent asks

The central trade-off is speed against modernisation. Rehosting everything is fastest and shortens dual run, but carries old inefficiencies into a metered environment. Refactoring everything delivers a better platform but extends the overlap. Most programs rehost the long tail quickly, replatform the expensive core, and retire aggressively.

What to do next

  1. Export the scheduler's job list and YARN history, and build the workload register with owners and last-read dates.
  2. Run the disposition scoring and review the retire list with owners first.
  3. Put your real contract and cloud estimates into the cost model and agree a decommission month with the sponsor.
  4. Write gate evidence for G1 to G5 and put the gates in the calendar.
  5. Pick a pilot domain, write its T-minus runbook from the template above, and rehearse rollback.
  6. Start support and licence notice periods early, and track switched-off workloads as the headline metric.
Key takeaway: A Hadoop exit succeeds as a program, not as a copy job. Give every workload an owner and a disposition, and retire aggressively. Model the dual-run window so the cost of a late decommission is visible from day one. Move in small waves with explicit cutover and rollback criteria, and rehearse rollback in the pilot. Measure progress only by workloads switched off on-premises. Plan decommissioning, including contract notice periods, as its own phase, then run FinOps so the new estate ends up cheaper than the one you left.