Apache Hadoop

Apache Hadoop

Deep technical articles on this topic.

124Articles
124Topics covered
Articles in this category

All 103 articles, sorted alphabetically

ARTICLE · 001

Apache Ambari, in depth: writing custom services, alerts as code, the metrics system, upgrading Ambari itself and rebuilding a lost server

Extending and operating Apache Ambari beyond the basics: how a command reaches an agent and a Python command script, writing a custom service with met…

Read article →
ARTICLE · 002

Airflow on Hadoop, in depth: edge-node workers, Kerberos, WebHDFS sensors, Spark on YARN and idempotent daily partitions

Run Apache Airflow as the orchestrator for a Kerberized Hadoop cluster: the Airflow 3 components and where Hadoop clients live, routing tasks to edge-…

Read article →
ARTICLE · 003

Apache Atlas, in depth: the type system, hooks and the Kafka notification path, lineage, classification propagation, search and running it in production

How Apache Atlas captures and serves metadata for a Hadoop data platform: the type system, the Hive hook and the ATLAS_HOOK and ATLAS_ENTITIES Kafka t…

Read article →
ARTICLE · 004

Apache NiFi, in depth: FlowFiles, repositories, back pressure and clustering

How Apache NiFi works and how to run it: FlowFiles and the content, FlowFile and provenance repositories, the ProcessSession contract with a custom pr…

Read article →
ARTICLE · 005

Apache Ozone, in depth: a day-2 operations playbook for snapshots, quotas, balancing, node lifecycle, upgrades and HDFS migration

Running Apache Ozone after the install: bucket layouts and quotas, snapshots and snapshot diff, trash, the container balancer, maintenance versus deco…

Read article →
ARTICLE · 006

Cloudera Manager, in depth: server, agents and supervisord, parcels, stale configuration, the Management Service, the REST API and a rolling restart worked end to end

How Cloudera Manager runs a Hadoop cluster: the server, database and agents, supervisord and process directories, parcels, role config groups and stal…

Read article →
ARTICLE · 007

Delta Lake, in depth: running it on a Hadoop estate, safe commits on HDFS and S3, migrating Hive tables, who can read it, maintenance, and Delta versus Iceberg

Delta Lake for Hadoop teams: what the transaction log changes compared with a Hive directory table, why commits are safe on HDFS and need a DynamoDB L…

Read article →
ARTICLE · 008

DistCp

How DistCp uses a MapReduce job to parallelize large HDFS-to-HDFS or HDFS-to-cloud copies. Covers throughput tuning, incremental sync, and gotchas for…

Read article →
ARTICLE · 009

Apache Flume, in depth: events, channel transactions, two-tier topologies, HDFS sink tuning and the failure modes of log ingestion

How Apache Flume moves log data into HDFS: events and agents, the source-channel-sink transaction contract and at-least-once delivery, memory versus f…

Read article →
ARTICLE · 010

Hadoop 3 Key Features, in depth: erasure coding, multiple NameNodes, new ports, shaded clients, YARN resource types and what changed after 3.0

What Hadoop 3 changed and what each change costs to adopt: erasure coding with real storage arithmetic, three or more NameNodes, the moved default por…

Read article →
ARTICLE · 011

Apache Ambari, in depth: server, agents, stacks, blueprints and the REST API for Hadoop operations

How Apache Ambari provisions and operates Hadoop clusters: the server, agents and database, the desired-state model, stack and service definitions, th…

Read article →
ARTICLE · 012

Hadoop Backup and Disaster Recovery: HDFS Snapshots, DistCp and Metadata Backups

Build a Hadoop backup and DR plan: a threat model showing why replication is not backup, RPO and RTO tiers, HDFS snapshots and trash, an incremental D…

Read article →
ARTICLE · 013

HDFS Balancer architecture

Deep-dive on the HDFS Balancer and Disk Balancer: threshold policy, source/proxy/target replica moves, placement invariants, bandwidth governors, iter…

Read article →
ARTICLE · 014

Hadoop Benchmarks, in depth: TestDFSIO, TeraSort and NameNode throughput, run and read correctly

How to benchmark a Hadoop cluster layer by layer: which tool measures which layer, running and reading TestDFSIO, TeraGen, TeraSort and TeraValidate, …

Read article →
ARTICLE · 015

Hadoop Capacity Planning, in depth: sizing storage, NameNode memory, YARN compute and recovery bandwidth into one node count

A working method for sizing a Hadoop cluster: turn ingest rate, retention and growth into physical HDFS bytes under replication and erasure coding, es…

Read article →
ARTICLE · 016

YARN Capacity Scheduler

Deep-dive on the YARN Capacity Scheduler: queue hierarchy and capacity guarantees, elastic borrowing and max capacities, preemption to reclaim, user l…

Read article →
ARTICLE · 017

HDFS NameNode Checkpointing: fsimage, Edit Log, Fast Restarts

How the HDFS NameNode checkpoints: the fsimage snapshot and edit log, JournalNodes in HA, the Secondary/Standby merge, and what breaks when checkpoint…

Read article →
ARTICLE · 018

Hadoop Cloud Migration Playbook, in depth: phase gates, workload disposition scoring, the dual-run cost model, wave governance, cutover runbooks and decommissioning

The program-level playbook for leaving an on-premises Hadoop cluster: six phases with gate evidence, a scoring model to retire, rehost, replatform or …

Read article →
ARTICLE · 019

Hadoop Cluster Planning, in depth: node roles, HA quorums, rack and network layout, disks and growth

A practical guide to Hadoop cluster layout once sizing is done: master, worker, edge and utility node classes, NameNode, JournalNode, ZooKeeper and Re…

Read article →
ARTICLE · 020

Hadoop Cost Management, in depth: unit costs, showback from YARN and HDFS metering, finding cold data, and the storage and compute levers that actually cut spend

How to manage the running cost of a Hadoop cluster: build a rate card from total cost, meter compute from YARN vcoreSeconds and memorySeconds, meter s…

Read article →
ARTICLE · 021

Why Hadoop Is Declining, in depth: the architectural causes, what still belongs on HDFS and YARN, and how to decide whether to leave

An engineering explanation of Hadoop's decline: coupled storage and compute, data locality versus fast networks, NameNode heap limits, MapReduce …

Read article →
ARTICLE · 022

Hadoop DataNode Decommission, in depth: admin states, throttles and stuck nodes

How the HDFS NameNode decommissions a DataNode: admin states, host files and the JSON provider, the admin monitor and its throttles, throughput arithm…

Read article →
ARTICLE · 023

Hadoop Node Decommissioning

Graceful DataNode removal. Block re-replication.

Read article →
ARTICLE · 024

HDFS Disk Balancer vs Balancer: Fixing DataNode Disk Skew

HDFS disk balancer vs the Hadoop balancer: why disks inside one DataNode drift apart, hdfs diskbalancer -plan, -execute and -query, throttling, volume…

Read article →
ARTICLE · 025

Hadoop DistCp Architecture: Cross-Cluster HDFS Replication

How Hadoop DistCp replicates HDFS data between clusters: copy listings and dynamic splits, snapshot-diff incremental sync, bandwidth throttling and CR…

Read article →
ARTICLE · 026

Apache Druid Architecture, in depth: segments, ingestion and handoff, the broker query path, and running it without surprises

How Apache Druid works inside: the Coordinator, Overlord, Broker, Router, Historical and ingestion services, deep storage, the metadata store and ZooK…

Read article →
ARTICLE · 027

Hadoop Ecosystem in 2026, in depth: a component status map, Java baselines, catalogs and how to choose a stack

A component-by-component map of the Hadoop ecosystem in 2026: which projects are active, dormant or retired (with release dates), why Java 17 now sets…

Read article →
ARTICLE · 028

HDFS centralized cache management architecture

Deep-dive on HDFS centralized cache management: cache pools and directives with quotas, the NameNode cache manager driving DataNodes to mmap/mlock blo…

Read article →
ARTICLE · 029

HDFS Federation architecture

Deep-dive on HDFS Federation: multiple independent NameNodes with per-namespace block pools over shared DataNodes, ViewFS client mount tables vs Route…

Read article →
ARTICLE · 030

HDFS quotas architecture

Deep-dive on HDFS quotas: namespace (object-count) quotas, space (replicated-byte) quotas, storage-type quotas, how QuotaCounts are cached on director…

Read article →
ARTICLE · 031

HDFS Router-Based Federation (RBF): Architecture Deep-Dive

How HDFS Router-Based Federation scales past the single-NameNode ceiling with stateless routers, a State Store mount table, and one global namespace o…

Read article →
ARTICLE · 032

HDFS Snapshots Explained: Copy-on-Write, .snapshot, snapshotDiff

How Hadoop HDFS snapshots work: allowSnapshot and createSnapshot, read-only copy-on-write views under .snapshot, NameNode diff lists, snapshotDiff for…

Read article →
ARTICLE · 033

The HDFS write pipeline, in depth: block allocation, packets, placement, hflush and hsync, and recovery

How HDFS writes data: chain replication, addBlock and placement, packets and checksums, the data and ack queues, hflush versus hsync, pipeline and lea…

Read article →
ARTICLE · 034

Hadoop HDFS and YARN Architecture in Depth

How HDFS and YARN work together: blocks, NameNode and DataNodes, write and read pipelines, QJM high availability, ResourceManager, NodeManagers and Ap…

Read article →
ARTICLE · 035

Hive architecture, in depth: following one query through HiveServer2, the metastore, Tez on YARN, the NameNode and the final commit

Hive's architecture traced through the Hadoop layers one query touches: HiveServer2 sessions and impersonation, metastore calls during compilatio…

Read article →
ARTICLE · 036

Impala architecture, in depth: how Impala lives on a Hadoop cluster, from HDFS locality and short-circuit reads to the shared metastore and memory outside YARN

How Impala's architecture fits a Hadoop cluster: impalads co-located with DataNodes, block locations from the NameNode, the shared Hive Metastore…

Read article →
ARTICLE · 037

JVM Tuning for Hadoop Services, in depth: per-daemon heaps, Hadoop 3 environment variables, reading JvmPauseMonitor, and how pauses become failovers

JVM tuning for long-running Hadoop daemons: what fills the NameNode, DataNode, JournalNode, ZKFC, ResourceManager and NodeManager heaps, where Hadoop …

Read article →
ARTICLE · 038

Hadoop Kerberos architecture

Deep-dive on Kerberos for Hadoop: principals and realms, the AS/TGS ticket exchange, keytabs for unattended services, delegation tokens for distribute…

Read article →
ARTICLE · 039

Hadoop Log Aggregation, in depth: how YARN collects container logs, the HDFS layout, file formats, rolling uploads, policies and retention

YARN log aggregation explained: local container log directories, the NodeManager upload pipeline, the bucketed remote layout and its 1777 permissions,…

Read article →
ARTICLE · 040

Hadoop Migration to Cloud, in depth: inventory from the fsimage, data and metadata lanes, Iceberg conversion, validation and a cutover that can be rolled back

An execution playbook for moving a Hadoop cluster to cloud object storage: inventory from the fsimage, logical versus raw sizing and transfer arithmet…

Read article →
ARTICLE · 041

Hadoop Monitoring, in depth: metrics2 and the /jmx endpoint, the NameNode, DataNode and YARN signals that matter, alert rules and a poller you can run

How to monitor a Hadoop cluster from first principles: where metrics come from (metrics2 sources and sinks, the /jmx servlet), the NameNode, DataNode,…

Read article →
ARTICLE · 042

ZKFC and HDFS NameNode HA: Failover, Quorum Journal, Fencing

How HDFS NameNode high availability works: ZKFC (ZKFailoverController) and ZooKeeper election, Quorum Journal Manager epochs, fencing, and Observer re…

Read article →
ARTICLE · 043

Apache Ozone architecture, in depth: Ozone Manager, SCM, containers, Ratis pipelines, bucket layouts and operations

How Apache Ozone is built: why it separates the namespace (Ozone Manager) from block space (Storage Container Manager), volumes, buckets and keys, FSO…

Read article →
ARTICLE · 044

HDFS rack awareness architecture

Deep-dive on HDFS rack awareness: the topology resolution script and NameNode network tree, the default one-local-plus-two-remote-rack block placement…

Read article →
ARTICLE · 045

Apache Ranger

How Ranger centralizes access policies across HDFS, Hive, HBase, Kafka, and Solr — with tag-based rules, row/column-level security, and comprehensive …

Read article →
ARTICLE · 046

Hadoop Rolling Upgrade, in depth: upgrading HDFS and YARN without downtime, the rollback image, downgrade versus rollback, and a runbook

How a Hadoop rolling upgrade works and how to run one safely: why HDFS needs a rollback image, the prepare, NameNode, DataNode and finalize steps with…

Read article →
ARTICLE · 047

Hadoop Security Best Practices: Kerberos, Ranger, Encryption and Auditing

A layered hardening programme for Hadoop: Kerberos principals, auth_to_local and delegation tokens, HDFS ACLs, Ranger and YARN LinuxContainerExecutor,…

Read article →
ARTICLE · 048

HDFS short-circuit reads -- bypassing the DataNode for local data

Deep-dive on HDFS short-circuit reads: the normal DataNode read path, local reads with co-located compute, short-circuit reading the block file direct…

Read article →
ARTICLE · 049

Spark External Shuffle Service: Dynamic Allocation, Push Shuffle

Why Spark needs an external shuffle service: shuffle data that outlives executors, dynamic allocation, push-based shuffle, spill and disaggregated shu…

Read article →
ARTICLE · 050

Speculative execution -- racing around stragglers

Deep-dive on speculative execution: the straggler problem (a job finishes only when all tasks finish), straggler detection, launching a speculative du…

Read article →
ARTICLE · 051

Hadoop Upgrade Patterns, in depth: express, rolling, side-by-side and canary upgrades, ordering rules, rehearsal and rollback points

How to choose and run a Hadoop upgrade: what an upgrade really changes, the five patterns (express, rolling, side-by-side, canary, client-staged), ord…

Read article →
ARTICLE · 052

When Hadoop Still Makes Sense, in depth: measuring your own cluster, deciding per workload, and staying on Hadoop without getting stuck

How to decide, with evidence from your own cluster, which workloads should stay on Hadoop: sampling YARN utilisation from the ResourceManager REST API…

Read article →
ARTICLE · 053

YARN Fair Scheduler Explained: Fair Share, DRF and Preemption

How the Hadoop YARN Fair Scheduler shares a cluster: weighted queues, fair share from live demand, Dominant Resource Fairness, minShare, maxShare, pre…

Read article →
ARTICLE · 054

YARN node labels architecture

Deep-dive on YARN node labels: mapping nodes into named partitions so the scheduler places GPU, high-memory, licensed, and spot workloads on the right…

Read article →
ARTICLE · 055

ZooKeeper in Hadoop in Depth: Who Uses It, What They Store and How to Run It Safely

ZooKeeper as shared Hadoop infrastructure: znodes, sessions, ephemeral nodes and watches from first principles; how HDFS ZKFC, YARN ResourceManager, r…

Read article →
ARTICLE · 056

HDFS Blocks and Replication, in depth: block anatomy, rack-aware placement, the redundancy monitor and failure recovery

How HDFS blocks and replicas work: on-disk layout, generation stamps, default placement, heartbeats and block reports, the redundancy monitor's p…

Read article →
ARTICLE · 057

HDFS Checkpoints and JournalNodes

The fsimage, edit-log segments, and JournalNode quorum at artifact level: transaction IDs, rolling, segment recovery, checkpoint triggers, image trans…

Read article →
ARTICLE · 058

HDFS DataNode, in depth: on-disk layout, replica states, heartbeats and block reports, scanners and disk failure handling

Inside the HDFS DataNode: block files and checksums on disk, replica states, heartbeat and block report timing, the streaming data path, block and dir…

Read article →
ARTICLE · 059

HDFS Encryption Zones: Transparent Encryption with Hadoop KMS

How HDFS encryption zones encrypt data at rest: zone keys in the Hadoop KMS, per-file data keys stored as xattrs, AES counter mode, key ACLs, zone set…

Read article →
ARTICLE · 060

HDFS Erasure Coding

How Reed-Solomon erasure coding replaces three-way replication for cold data. Covers the RS(6,3) scheme, striping, reconstruction IO cost, and when to…

Read article →
ARTICLE · 061

HDFS High Availability

How HDFS achieves NameNode HA with an active/standby pair, a JournalNode quorum for the shared edit log, ZooKeeper for leader election, and fencing to…

Read article →
ARTICLE · 062

HDFS NameNode Deep Dive, in depth: namespace, block map, edit log group commit, RPC and heap sizing

Inside the HDFS NameNode: what it persists and what it rebuilds, the namespace lock, the write RPCs, edit log group commit, block reports, heap sizing…

Read article →
ARTICLE · 063

HDFS Overview, in depth: architecture, blocks, write and read paths, fault tolerance and trade-offs

How HDFS works: NameNode and DataNodes, blocks and rack-aware replication, the write pipeline and read path, failure handling, a capacity worked examp…

Read article →
ARTICLE · 064

HDFS Permissions and ACLs: POSIX Bits, Mask Entry, Ranger

How HDFS checks access: POSIX-style permission bits, extended ACLs with access and default entries, the ACL mask entry, setfacl and getfacl, and Apach…

Read article →
ARTICLE · 065

HDFS Safe Mode: NameNode Read-Only Startup and How to Leave It

Why the HDFS NameNode starts in safe mode, the block-report threshold (0.999) that ends it, dfsadmin -safemode commands, and the risks of forcing a le…

Read article →
ARTICLE · 066

HDFS Small Files Problem

Why HDFS is fundamentally optimized for large files, how millions of small files exhaust NameNode heap, and the standard techniques to consolidate the…

Read article →
ARTICLE · 067

HDFS Trash, in depth: how delete becomes a rename, checkpoints and the emptier, retention arithmetic, what bypasses trash, and running it safely

A practical guide to HDFS trash: per-user and per-zone trash roots, Current and timestamped checkpoints, the NameNode emptier loop, how fs.trash.inter…

Read article →
ARTICLE · 068

HDFS Performance Tuning

HDFS tuning that moves numbers: block size versus NameNode heap and parallelism, RPC handler counts, short-circuit reads, hedged reads, transfer threa…

Read article →
ARTICLE · 069

Azure HDInsight, in depth: the 5.1 service, cluster anatomy, storage and metastores, autoscale, security and when to choose something else

A practical guide to Azure HDInsight 5.1: support status and retired versions, node roles, ADLS Gen2 and external metastores, a worked Spark deploymen…

Read article →
ARTICLE · 070

Iceberg + Trino, in depth: catalogs, query planning, row-level writes and table maintenance from the SQL engine

How Trino works with Apache Iceberg tables: what the table format and the engine each own, choosing a catalog, designing tables in Trino SQL, how the …

Read article →
ARTICLE · 071

MapReduce Combiner, in depth: the algebra that makes it safe, where Hadoop runs it, and how to prove it helped

A complete guide to the Hadoop MapReduce combiner: why it exists, the associative and commutative contract, the three places Hadoop may invoke it (spi…

Read article →
ARTICLE · 072

MapReduce Distributed Cache, in depth: shipping side data, localization, visibility and map-side joins

How the Hadoop MapReduce distributed cache works and how to use it well: addCacheFile, addCacheArchive and classpath methods, -files, -archives and -l…

Read article →
ARTICLE · 073

MapReduce Is Dead, in depth: what replaced the engine, which of its ideas still run everything, and how to retire the last jobs

Why the Hadoop MapReduce engine lost to DAG and SQL engines, which of its ideas live on inside Spark, Tez and every MPP database, where MapReduce jobs…

Read article →
ARTICLE · 074

MapReduce Map Phase, in depth: input splits, record boundaries, locality, the Mapper lifecycle and map-only jobs

How Hadoop's map phase works from first principles: how FileInputFormat computes splits and the 10% slop rule, how LineRecordReader handles recor…

Read article →
ARTICLE · 075

Hadoop Mapper, in depth: the run() contract, Context and counters, map-only jobs, map-side joins, chaining, Streaming and testing

How to write Hadoop MapReduce mappers well: the Mapper class contract and its run() loop, a worked log-parsing mapper, using Context for counters, con…

Read article →
ARTICLE · 076

MapReduce Output Committer, in depth: the commit protocol, attempt arbitration, the v1 and v2 algorithms, AM restart recovery and the manifest committer

How Hadoop MapReduce turns many task attempts into exactly one visible output: the OutputCommitter contract, how the AppMaster lets only one speculati…

Read article →
ARTICLE · 077

MapReduce Overview, in depth: how a Hadoop job flows from submission through map, shuffle and reduce to committed output

An end-to-end guide to Hadoop MapReduce on YARN: the programming model, the client, ResourceManager, MRAppMaster and ShuffleHandler, split sizing, the…

Read article →
ARTICLE · 078

MapReduce Partitioner, in depth: the getPartition contract, hashing hazards, total order, secondary sort and skew

How the Hadoop MapReduce partitioner works and how to write one that is correct and balanced: where getPartition runs, its contract, HashPartitioner a…

Read article →
ARTICLE · 079

MapReduce Reduce Phase, in depth: Fetchers, the Merge Manager, Secondary Sort and Output Commit

Inside a MapReduce reduce task: shuffle fetchers, in-memory and on-disk merges, the final merge, the Reducer API and Writable reuse, grouping comparat…

Read article →
ARTICLE · 080

MapReduce Reducer, in depth: the reduce contract, aggregation algebra, joins, top-N, skew, side outputs and determinism under retries

How to write correct Hadoop MapReduce reducers: the reduce() contract, which aggregations can double as combiners, a worked revenue and top-N reducer,…

Read article →
ARTICLE · 081

MapReduce Shuffle and Sort, in depth: the map-side buffer, spills and index files, the ShuffleHandler, and diagnosing a slow shuffle from counters

How Hadoop MapReduce moves and sorts intermediate data byte by byte: the collection buffer and its metadata, spill and merge, IFile index records, the…

Read article →
ARTICLE · 082

MapReduce Tuning, in depth: a counter-driven method for memory, parallelism and bytes moved

How to tune Hadoop MapReduce jobs from evidence: container memory versus heap and the 0.8 ratio, split sizing and small files, reducer count and skew,…

Read article →
ARTICLE · 083

Modern Lakehouse Architecture, in depth: the layers, the atomic commit on object storage, catalogs, ingestion, table maintenance and migrating from Hive tables

A format-neutral guide to lakehouse architecture: what each layer (object storage, Parquet, Iceberg, Delta Lake or Hudi, the catalog, engines) is resp…

Read article →
ARTICLE · 084

Apache Oozie, in depth: workflows, the launcher, coordinators and data triggers, bundles, retries and reruns, and running a retired scheduler safely

How Apache Oozie works and how to operate it now that it is retired: workflow, coordinator and bundle layers, the server, database and YARN launcher, …

Read article →
ARTICLE · 085

Ozone vs HDFS, in depth: where metadata lives, which semantics differ, probing capabilities from code, and choosing per workload

A decision guide for Apache Ozone versus HDFS: NameNode heap versus Ozone Manager and SCM, namespace sizing arithmetic, a semantics table for rename, …

Read article →
ARTICLE · 086

Ranger Policy Deep Dive: policy JSON anatomy, resource matching, the evaluation algorithm worked by hand, policies as code and debugging denials

Apache Ranger policies in depth: every field of a policy document, resource matching with wildcards, recursion, excludes and the USER macro, the evalu…

Read article →
ARTICLE · 087

Apache Sqoop, in depth: how imports become map tasks, incremental loads, non-atomic exports, and moving off a retired tool

How Apache Sqoop 1.4.7 really works: the client-side metadata and codegen step, the map-only MapReduce job, how split ranges come from a min/max bound…

Read article →
ARTICLE · 088

WebHDFS and HttpFS Explained: HDFS REST API, 307 Redirect, Knox

How the WebHDFS REST API works: /webhdfs/v1 URLs with op=OPEN and CREATE, the 307 redirect to a DataNode, HttpFS and Knox gateways, SPNEGO, delegation…

Read article →
ARTICLE · 089

YARN ApplicationMaster

The YARN ApplicationMaster end to end: submission and launch, register/allocate/finish on ApplicationMasterProtocol, the heartbeat allocation loop, lo…

Read article →
ARTICLE · 090

YARN cgroups Container Isolation

Linux cgroups for CPU + memory limits.

Read article →
ARTICLE · 091

YARN Containers, in depth: launch contexts, localization, container executors, the kill sequence, exit statuses and debugging on the NodeManager

What a YARN container actually is on a worker node: the allocation and its token, the ContainerLaunchContext, how resource localization and its public…

Read article →
ARTICLE · 092

YARN Docker Container Support, in depth: the runtime, container-executor policy, user identity, Spark on Docker and failure modes

How YARN runs containers inside Docker images: the DockerLinuxContainerRuntime launch path, yarn-site versus the root-owned container-executor.cfg pol…

Read article →
ARTICLE · 093

YARN Fair Scheduler, in depth: a worked allocation file, fair-share arithmetic, placement rules, preemption tuning, AM-share starvation and migrating to the Capacity Scheduler

A configuration and tuning guide to the Hadoop YARN Fair Scheduler: how instantaneous and steady fair shares are computed, a worked fair-scheduler.xml…

Read article →
ARTICLE · 094

YARN Node Labels and Placement, in depth: partitions, node attributes and placement constraints working together on one cluster

How to control where YARN containers run by combining node-label partitions, node attributes and placement constraints: what each tool answers, a work…

Read article →
ARTICLE · 095

YARN NodeManager, in depth: capacity advertising, heartbeats, health checks, work-preserving restart and graceful decommission

How the Hadoop YARN NodeManager works as a daemon: sizing memory-mb and vcores with a worked Spark example, the heartbeat protocol with the ResourceMa…

Read article →
ARTICLE · 096

YARN Opportunistic Containers, in depth: execution types, centralised and distributed allocation, NodeManager queuing, kills and promotion

How YARN opportunistic containers work: GUARANTEED versus OPPORTUNISTIC execution types, centralised and AMRMProxy-based distributed allocation, the d…

Read article →
ARTICLE · 097

YARN, in Depth: The Resource Model, How Requests Become Containers, Sizing Nodes, Scheduler Choice and Running It in Production

An operator's overview of Hadoop YARN: ResourceManager, NodeManagers and ApplicationMasters, the memory and vcore resource model, request normali…

Read article →
ARTICLE · 098

YARN Preemption

How YARN schedulers kill containers to reclaim resources and enforce queue guarantees: preemption mechanics, grace periods, configuration in Fair and …

Read article →
ARTICLE · 099

YARN Queue Hierarchies, in depth: designing the tree, capacity arithmetic down the path, ACL inheritance, placement and changing queues live

How to design and operate a YARN Capacity Scheduler queue hierarchy: what parents and leaves do, a worked capacity calculation on a 1,000 GB cluster, …

Read article →
ARTICLE · 100

YARN ResourceManager, in depth: services, the event dispatcher, the heartbeat path, HA with ZooKeeper and work-preserving restart

How the YARN ResourceManager works inside: its four RPC services and ports, the AsyncDispatcher and the application, attempt, container and node state…

Read article →
ARTICLE · 101

YARN Resource Types (GPU, FPGA), in depth: declaring, scheduling and isolating accelerators

How the YARN resource model schedules GPUs and FPGAs: resource-types.xml, units and allocation limits, why the DominantResourceCalculator is required,…

Read article →
ARTICLE · 102

YARN Timeline Server, in depth: the data model, versions 1, 1.5 and 2, publishing from an ApplicationMaster, and running it without it running you

How the YARN Timeline Server stores application history: the entity, event, primary-filter and domain model, the v1 LevelDB server, v1.5's HDFS w…

Read article →
ARTICLE · 103

ZooKeeper's Role in Hadoop, in depth: the failover time budget, the YARN state store, and a game day that proves failover works

What ZooKeeper actually decides in a Hadoop cluster and what it does not, how long HDFS and YARN failover really takes and why, the ResourceManager st…

Read article →