All 97 articles, sorted alphabetically
Spark Accumulators, in depth: the closure problem, how updates travel, exactly-once only in actions, custom AccumulatorV2, PySpark rules and DataFrame.observe
How Spark accumulators really work: why driver variables do not update from tasks, the driver and executor update path, the actions-only exactly-once …
Read article →Spark AQE architecture
Deep-dive on Adaptive Query Execution: query stages at shuffle boundaries, MapOutputStatistics-driven re-planning, partition coalescing and advisory s…
Read article →Spark Arrow, in depth: the columnar format, IPC batches, toPandas and createDataFrame, type mapping and memory
How Apache Arrow works inside Apache Spark: the columnar memory layout of validity, offset and data buffers, the IPC stream format, the paths that mov…
Read article →Spark + Avro, in depth: container files, schema resolution, type mapping and Confluent-framed Kafka messages
How Avro encodes records and why the writer schema is mandatory, the container file layout, reader and writer schema resolution, Spark 4.x options and…
Read article →Spark barrier execution mode architecture
Deep-dive on Spark barrier execution mode: gang scheduling that launches all N tasks of a stage at once so they can run MPI-style all-to-all communica…
Read article →Spark runtime bloom-filter join architecture
Deep-dive on Spark's runtime bloom-filter (dynamic filtering) join: why shuffle dominates cost, building a bloom filter from the small side'…
Read article →Spark Broadcast Variables, in depth: closures vs broadcast, TorrentBroadcast, memory, lifecycle and refresh
How Spark broadcast variables work and how to use them safely: why closures ship data per task, the Broadcast API, how TorrentBroadcast splits values …
Read article →Spark Broadcast Hints, in depth: what BROADCAST costs on the driver and executors, sizing the build side, hard limits, timeouts and AQE
What happens after Spark accepts a BROADCAST hint: the build-side job, collection to the driver, HashedRelation construction and torrent broadcast; wh…
Read article →Spark broadcast joins -- avoiding the shuffle for large-small joins
Deep-dive on Spark broadcast joins: the join/shuffle problem, broadcasting the small side to avoid shuffling the large side, the auto-broadcast thresh…
Read article →Spark cache() and persist(), in depth: blocks, the BlockManager, executor loss, dynamic allocation and sizing a cache from measurements
What happens underneath cache() and persist() in Spark: how a cached partition becomes a named block, the BlockManager's memory and disk stores, …
Read article →Spark Cache Strategy, in depth: counting reuse, a cost model, where to cut the cache point, lifecycle and staleness
How to decide what to cache in Spark: counting reuse from plans, a cost model with measured times, cutting the cache after narrowing and before fan-ou…
Read article →Spark coalesce vs repartition, in depth: physical plans, upstream parallelism, balance, determinism and output files
What Spark's coalesce and repartition really do: narrow merge versus full shuffle, the physical plans, how coalesce silently shrinks upstream par…
Read article →Spark Columnar Processing, in depth: ColumnarBatch, vectorized readers, row and column transitions, and columnar plugins
How columnar processing works in open-source Spark: ColumnVector and ColumnarBatch, the vectorized Parquet and ORC readers, ColumnarToRow and RowToCol…
Read article →Spark DataFrame Complex Types, in depth: structs, arrays, maps, higher-order functions and nested performance
How to work with Spark DataFrame complex types: struct, array and map representation and access, explode versus higher-order functions (transform, fil…
Read article →Spark Connect architecture
Deep-dive on Apache Spark Connect: the DataFrame API as an unresolved protobuf logical-plan builder, the gRPC channel, the Connect server's analy…
Read article →Spark Connected Components, in depth: min-label propagation, large-star / small-star, randomized contraction and the traps that sink production runs
How connected components works on Spark: GraphX min-label Pregel, GraphFrames two_phase large-star / small-star and randomized contraction, plus check…
Read article →Spark DAG (Directed Acyclic Graph), in depth: how the DAGScheduler builds, reuses, skips and recovers stages
How Apache Spark turns an action into a directed acyclic graph of stages: the DAGScheduler event loop, building result and shuffle map stages from dep…
Read article →Spark Data Skew Handling, in depth: finding hot keys, what AQE splits for you, salting, two-stage aggregation and write-side skew
A practical guide to data skew in Apache Spark: why one oversized partition sets the runtime of a whole stage, how to find hot keys in the Spark UI an…
Read article →Spark DataSource V2 architecture
Deep-dive on Spark DataSource V2: TableCatalog and DDL, capability sets, the ScanBuilder pushdown negotiation and the unhandled-filter contract, stati…
Read article →Spark on Databricks, in depth: the runtime, compute types and access modes, the life of a query, jobs as code, tuning and cost
What changes when Apache Spark runs on Databricks: the Databricks Runtime and Photon, all-purpose, jobs, serverless and SQL compute, standard versus d…
Read article →Spark DataFrame API, in depth: column expressions, null logic, joins, aggregations, windows and the plans they build
A practical guide to the Spark DataFrame API: lazy plans versus actions, column expressions, select and withColumn, three-valued null logic and ANSI m…
Read article →Spark DataFrame and Dataset
How DataFrames add schema and columnar optimization to Spark, how Datasets extend DataFrames with type safety, and when to use each.
Read article →Spark Dataset API, in depth: typed transformations, groupByKey and KeyValueGroupedDataset, joinWith, the Aggregator API, and what each does to your physical plan
A practical guide to the typed side of Apache Spark's Dataset API in Scala and Java: map, flatMap and mapPartitions, groupByKey with mapGroups, f…
Read article →Spark Debugging, in depth: reading distributed stack traces, isolating bad records, executor deaths and wrong results
A practical method for debugging Apache Spark jobs: why errors surface at actions, how to read Job aborted traces and find the real cause, a taxonomy …
Read article →Delta Lake Architecture in Depth
A 2500-word walkthrough of Delta Lake: transaction log, Parquet data files, checkpoints, OPTIMIZE + Z-ORDER, MERGE, time travel, streaming, VACUUM.
Read article →Spark DStreams (Legacy), in depth: micro-batches, receivers, Kafka direct streams, checkpoint recovery and migrating off
How the deprecated Spark Streaming DStream engine works and how to operate or retire it: batch-to-RDD model, receivers versus the Kafka direct stream,…
Read article →Spark dynamic allocation architecture
Deep-dive on Spark dynamic allocation: the ExecutorAllocationManager's backlog and idle timers, exponential ramp-up, external shuffle service vs …
Read article →Spark Encryption, in depth: TLS and AES for RPC, local disk I/O encryption, UI TLS and Parquet column keys
How to encrypt everything a Spark application touches: the separate switches for RPC traffic (TLS or the legacy AES protocol), temporary data on local…
Read article →Spark Execution Architecture in Depth: plans, stages, tasks, shuffle and memory
How a Spark action becomes jobs, stages and tasks: Catalyst and AQE, the DAG and task schedulers, executor memory, shuffle mechanics, a sized worked e…
Read article →Spark EXPLAIN Plans
A comprehensive guide to Apache Spark EXPLAIN plans: reading logical and physical execution plans, understanding Catalyst optimizer rules, interpretin…
Read article →Spark external shuffle service architecture - decoupling shuffle from executor lifetime
Deep-dive on Spark's external shuffle service: sort-based shuffle file layout and index, Netty zero-copy serving, surviving executor death for dy…
Read article →Spark GPU Acceleration, in depth: how the cuDF plugin runs SQL and DataFrames on GPUs, how to size memory, and when it pays
How GPU acceleration for Apache Spark works: plan rewriting by the NVIDIA cuDF plugin (formerly the RAPIDS Accelerator), columnar batches, PCIe and pi…
Read article →Spark GraphFrames, in depth: DataFrame graphs, motif queries, Pregel and running graph algorithms at scale
A practical guide to GraphFrames on Apache Spark: installation under the io.graphframes coordinates and graphframes-py, building graphs from tables, m…
Read article →Spark GraphX, in depth: the property graph, vertex-cut partitioning, aggregateMessages, Pregel and running it in production
How Spark GraphX works and when to use it: the property graph and triplet model, vertex-cut storage with routing tables and replicated vertex views, t…
Read article →Spark HBase Connector, in depth
The Apache hbase-spark connector in depth: HBaseContext and one connection per executor, bulkPut, bulkGet and hbaseRDD, the DataFrame source and hbase…
Read article →Spark History Server, in depth: event logs, replay, the disk and hybrid stores, rolling logs, compaction and running it in production
How the Spark History Server works and how to run it well: the scan loop and listing database, replaying event logs into the KV store, sizing memory a…
Read article →Spark + Hudi architecture
Spark on Apache Hudi: the timeline, copy-on-write versus merge-on-read, index-driven upserts, snapshot and incremental reads, and table services.
Read article →Spark + Iceberg architecture
Iceberg from the Spark side: catalog configuration, DataFrameWriterV2 and SQL writes, MERGE with copy-on-write vs merge-on-read, time travel, partitio…
Read article →Spark and JDBC, in depth: partitioned reads, pushdown, fetch size, type mapping and safe writes
How the Spark JDBC data source really works: driver-side schema discovery, the WHERE clauses generated from partitionColumn and bounds, predicate list…
Read article →Spark Join Strategy Hints, in depth: how the planner picks a join, when hints win, and when they are ignored
A practical guide to Spark SQL join strategy hints: BROADCAST, MERGE, SHUFFLE_HASH and SHUFFLE_REPLICATE_NL, how join selection works with and without…
Read article →Spark JSON Read/Write, in depth: JSON Lines, schema inference, corrupt records, nested data and VARIANT
How Spark reads and writes JSON: JSON Lines versus multiLine documents and splittability, the cost and pitfalls of schema inference, PERMISSIVE, DROPM…
Read article →Spark + Jupyter Integration, in depth: where the driver lives, and why it decides everything
Running Spark from Jupyter: the three driver placements (kernel as driver, Livy and sparkmagic, Spark Connect), session lifecycle and getOrCreate trap…
Read article →Koalas → Spark Pandas API Migration, in depth: the renames, the per-release breaks up to Spark 4.0, a codemod and a parity-test harness
A practical playbook for moving code from Koalas (databricks.koalas) to the pandas API on Spark (pyspark.pandas): what was renamed, which behaviours c…
Read article →Spark Kryo Serialization, in depth: where spark.serializer applies, registration, buffers and failure modes
Spark Kryo serialization explained: which bytes spark.serializer controls and which it does not, how Kryo encodes objects, registering classes and cus…
Read article →Spark on Kubernetes architecture
Spark on Kubernetes in depth: the driver pod and headless service, RBAC, pod templates, OOMKilled versus memory overhead, shuffle tracking, PVC reuse …
Read article →Apache Livy, in depth: running Spark over REST with interactive sessions, batches, impersonation and recovery
How Apache Livy runs Spark work over REST: interactive sessions with a remote driver and REPL versus batch submission, the sessions, statements and ba…
Read article →Spark unified memory management architecture
Deep-dive on Spark memory: the reserved/user/unified heap layout, unified execution and storage with borrow-and-evict, the eviction asymmetry, spill-t…
Read article →Spark Memory Tuning, in depth: container anatomy, telling heap OOM, container kills, driver OOM and spill apart, and a worked executor sizing
A symptom-first guide to Spark memory tuning: what makes up an executor container, how to tell heap OOM, container kills, driver OOM and spill apart, …
Read article →Spark ML Overview, in depth: what the library contains, how distributed training really executes, and when to use something else
A practical map of Apache Spark's machine learning library: spark.ml versus the RDD-based spark.mllib, the algorithm families, how training runs …
Read article →Spark ML Classification, in depth: the classifier catalog, distributed training, rawPrediction vs probability, thresholds and class imbalance
How classification works in Spark ML: the nine classifiers and which are binary only, how logistic regression and trees train across executors, rawPre…
Read article →Spark ML Cross-Validation, in depth: how CrossValidator splits, fits and selects, fold design for grouped and time-ordered data, and reading the metrics honestly
What Spark's CrossValidator actually does with your DataFrame: the seeded fold split, per-fold caching, the driver thread pool, argmax selection …
Read article →Spark ML Feature Engineering, in depth: point-in-time features, imputation, scaling, target encoding and selection
A practical guide to feature engineering with Spark ML: building leak-free point-in-time aggregates, time-based splits, imputing and scaling numeric f…
Read article →Spark ML Model Persistence, in depth: the saved-model layout, custom stages, cross-version loading and safe promotion
How Spark ML saves and loads models: the metadata, stages and data directories, a train-save-validate-load cycle, custom stage rules, tuning models, t…
Read article →Spark ML Pipelines, in depth: the fit/transform contract, leak-free tuning, custom stages, persistence and scoring at scale
A practical guide to pyspark.ml Pipelines: how Estimators, Transformers and PipelineModels work, a full worked churn example, handling nulls and unsee…
Read article →Spark ML Recommendation, in depth: ALS collaborative filtering from first principles to production top-K serving
How Spark ML's alternating least squares recommender works: the factorization objective, what one iteration solves, implicit-feedback confidence,…
Read article →Spark ML Regression, in depth: solvers, GLMs, boosted trees, a worked delivery-time model and the failures that matter
How Spark ML regressors train on a cluster: normal vs L-BFGS solvers, elastic net, GLM families for counts and skew, GBT with early stopping, a full P…
Read article →Spark ML Transformers and Estimators, in depth: the fit and transform contract, Params resolution, Pipeline.fit and a learned custom stage
How Spark ML Transformers and Estimators really work: the fit/transform contract, Params precedence and copy, how Pipeline.fit walks stages, a quantil…
Read article →Spark MLlib, in depth: the feature, statistics, similarity-search and pattern-mining toolkit, evaluation, and leaving the RDD API
A practical guide to the parts of Spark MLlib around the model: which feature transformers learn state and what they compute, hashing versus vocabular…
Read article →Apache Spark Overview
What Spark is, how RDDs and DataFrames give in-memory distributed computing, and where Spark fits versus MapReduce, Flink, and modern warehouses.
Read article →Spark PageRank
Classic algorithm on GraphFrames.
Read article →Spark Pandas API, in depth: how pyspark.pandas maps pandas semantics onto Spark plans, the default index, ordering, apply_batch, caching and when to drop to PySpark
A first-principles guide to the pandas API on Spark (pyspark.pandas): the internal frame and lazy plans, the three default index types and their costs…
Read article →Spark Pandas UDFs, in depth: type-hint dispatch, struct columns, the batch-boundary trap, iterator setup, grouped aggregates and cogroup
What a Spark pandas UDF actually receives and must return: the four type-hint shapes, StructType as pandas DataFrame, why batch-local statistics are a…
Read article →Spark Parquet Read + Write, in depth
How Spark reads and writes Parquet: row groups, column chunks, pages and the footer, split planning, filter pushdown and what defeats it, a worked lay…
Read article →Spark Partition Pruning vs Predicate Pushdown, Static and Dynamic
How Spark partition pruning skips data: static pruning on partition columns, dynamic pruning from join values, predicate pushdown, min/max file skippi…
Read article →Spark Partitioning in Depth: Read Splits, Shuffle Partitions, Write Layout and Skew, With the Arithmetic
How Spark decides partition counts at the three boundaries of a job: the file split formula with maxPartitionBytes and openCostInBytes worked through,…
Read article →Spark Persistence, in depth: what to cache, storage levels, materialization, eviction and when to let go
How persist() and cache() really work in Apache Spark: lazy marking, block storage on executors, storage levels and their defaults across RDD, Dataset…
Read article →Spark Photon architecture
Deep-dive on Photon: a C++ vectorized execution engine beneath Spark SQL. How the planner splits the physical plan into native and JVM operators, how …
Read article →PySpark Performance, in depth: the JVM-Python boundary, UDF cost, Arrow and vectorised UDFs, and a worked rewrite
Why PySpark jobs are slow and how to make them fast: where Python runs in Spark, what a row-at-a-time UDF costs, Arrow and vectorised UDF families wit…
Read article →Spark RDD lineage architecture
Deep-dive on Spark RDD lineage: lazy transformations building a DAG, narrow vs wide dependencies, stage cutting at shuffle boundaries, recomputation o…
Read article →Spark S3 optimization, in depth: S3A committers, read tuning, request limits and file layout
How to run Spark well on Amazon S3: why rename-based commits are slow and unsafe, how the S3A directory, partitioned and magic committers work, the ex…
Read article →Spark Salting for Skew Joins, in depth: choosing the salt count, salting only hot keys, which join types survive, and when AQE is enough
A hands-on guide to salting skewed joins in Spark: measuring hot keys, sizing the salt count from rows per task, selective salting with a broadcast ho…
Read article →Spark Secrets Management, in depth: where credentials leak in a Spark application and how to deliver them safely
A practical guide to secrets in Apache Spark: the paths a credential takes through spark-submit, driver, executors, event logs and the UI, what spark.…
Read article →Spark Security
Securing Apache Spark clusters: Kerberos authentication, Ranger authorization, TLS encryption, encrypted shuffle, network isolation, and audit logging…
Read article →Spark shuffle architecture
Deep-dive on Spark shuffle: map output + partitioner + external shuffle service + reduce fetch, plus push-based shuffle and AQE.
Read article →Spark Sort-Merge Join, in depth: the merge algorithm, buffered matches, plans, bucketing and skew
How Spark executes a sort-merge join, based on the Spark 3.5 source: shuffle and sort requirements, streamed and buffered sides by join type, the Exte…
Read article →Spark SQL
How Spark SQL provides ANSI-compatible SQL over any DataFrame source (Parquet, Delta, JDBC, Kafka), the query engine architecture, and Thrift server.
Read article →Spark SQL Config Tuning, in depth: where settings live, how to verify them, partition arithmetic, joins, skew and a worked tuning pass
How to tune Spark SQL configuration as an engineering surface: configuration layers and precedence, runtime versus static settings, reading effective …
Read article →Spark SQL Functions, in depth: how built-ins resolve and run, NULL and ANSI semantics, try_ functions, time zones, aggregates and UDF trade-offs
How Spark SQL built-in functions are resolved, optimized and code-generated; NULL rules, ANSI mode and the try_ family, time-zone traps, deterministic…
Read article →Spark SQL optimizer architecture
Deep-dive on Catalyst optimizer: analyzer, rule-based rewrites, cost model, physical planner, Tungsten codegen, AQE, and extensions.
Read article →Jobs, Stages and Tasks in Spark: How Spark Splits Work
Jobs, stages and tasks in Spark: where stage boundaries fall, what sets partition count, task vs stage retry, speculation, AQE, and how to find stragg…
Read article →Spark Statistics + CBO, in depth: collecting statistics, estimation formulas, join reordering and AQE
How Spark SQL statistics and the cost-based optimizer work: ANALYZE TABLE variants, where statistics are stored, filter and join estimation formulas, …
Read article →Spark Structured Streaming Deduplication, in depth: dropDuplicates, watermarks, dropDuplicatesWithinWatermark, state sizing and idempotent sinks
How to remove duplicate events in Spark Structured Streaming: where duplicates come from, how the streaming dedup operator uses state, dropDuplicates …
Read article →Spark Structured Streaming Architecture in Depth
How Spark Structured Streaming works inside: the micro-batch loop and checkpoint, watermarks, the state store, exactly-once sources and sinks, trigger…
Read article →Spark Streaming Deduplication, in depth: choosing the key, measuring the horizon, tiered dedup and custom state
Designing deduplication for Spark Structured Streaming: where duplicates come from, event ID vs offset vs content-hash keys, measuring the duplicate-d…
Read article →Spark Streaming foreachBatch, in depth: the micro-batch contract, batchId idempotency, MERGE upserts, multi-sink fan-out and failure modes
How Spark Structured Streaming's foreachBatch really works: where your function runs in the micro-batch lifecycle, why it is at-least-once by def…
Read article →Spark Structured Streaming + Kafka: Offsets and Exactly-Once
Spark Structured Streaming with Kafka: where offsets live, checkpoint offset management, end-to-end exactly-once, rate limits, and what you cannot cha…
Read article →Spark Streaming Kafka Source, in depth
The Structured Streaming Kafka source as a component: per-batch offset planning, consumer pools, starting-offset precedence, every option for Spark 4.…
Read article →Spark Streaming Sinks, in depth: the commit sequence, output modes, the file-sink metadata log, Kafka duplicates and choosing a sink
How Structured Streaming sinks really work: the offsets-to-commits sequence that makes replays possible, the Spark 4.2 sink and output-mode matrix, th…
Read article →Spark Streaming Sources, in depth: the replayable-offset contract, file source options, test sources and custom Python sources
How Structured Streaming sources work: offsets and the write-ahead offset log, the built-in file, Kafka, table, rate, rate-micro-batch and socket sour…
Read article →Spark Streaming Triggers, in depth: default, fixed-interval, available-now, continuous and Real-time, and how to choose an interval
How Structured Streaming triggers work: the micro-batch plan, offset log and commit loop; documented semantics of default, ProcessingTime, Once, Avail…
Read article →Spark Structured Streaming Watermarks and Late Data
How Spark Structured Streaming watermarks work: event time, withWatermark, the rule for dropping late rows, window state eviction, output modes, strea…
Read article →Spark Structured Streaming state
Deep-dive on Spark Structured Streaming state management: keyed state stores, checkpointing for exactly-once recovery, watermarks bounding state and h…
Read article →Spark Triangle Count, in depth: wedges, degree ordering, GraphX and GraphFrames, and a skew-aware DataFrame implementation
How to count triangles at scale in Spark: definitions, the degree-ordered O(m^1.5) algorithm with a worked example, GraphX and GraphFrames operators, …
Read article →Spark Tungsten -- pushing performance to the metal
Deep-dive on Spark's Project Tungsten: the JVM overhead problem (object memory, GC, virtual calls), off-heap managed memory, compact binary rows …
Read article →Spark UI in Depth: How It Collects Its Data, How to Read Every Tab, and How to Diagnose Slow Jobs
A practical guide to the Apache Spark web UI: the listener bus and status store behind it, retention limits and dropped events, reading the Jobs, Stag…
Read article →Spark whole-stage code generation architecture
Deep-dive on Spark SQL whole-stage code generation: replacing the Volcano iterator model with a fused Janino-compiled Java loop via the produce/consum…
Read article →Spark Wide vs Narrow Dependencies, in depth: the dependency classes, how to read them in a plan, how to turn a wide dependency into a narrow one and what each costs on failure
A first-principles guide to narrow and wide dependencies in Apache Spark: what the terms mean precisely, the Dependency classes behind them, which ope…
Read article →