Apache Spark

Apache Spark

Deep technical articles on this topic.

138Articles
138Topics covered
Articles in this category

All 97 articles, sorted alphabetically

ARTICLE · 01

Spark Accumulators, in depth: the closure problem, how updates travel, exactly-once only in actions, custom AccumulatorV2, PySpark rules and DataFrame.observe

How Spark accumulators really work: why driver variables do not update from tasks, the driver and executor update path, the actions-only exactly-once …

Read article →
ARTICLE · 02

Spark AQE architecture

Deep-dive on Adaptive Query Execution: query stages at shuffle boundaries, MapOutputStatistics-driven re-planning, partition coalescing and advisory s…

Read article →
ARTICLE · 03

Spark Arrow, in depth: the columnar format, IPC batches, toPandas and createDataFrame, type mapping and memory

How Apache Arrow works inside Apache Spark: the columnar memory layout of validity, offset and data buffers, the IPC stream format, the paths that mov…

Read article →
ARTICLE · 04

Spark + Avro, in depth: container files, schema resolution, type mapping and Confluent-framed Kafka messages

How Avro encodes records and why the writer schema is mandatory, the container file layout, reader and writer schema resolution, Spark 4.x options and…

Read article →
ARTICLE · 05

Spark barrier execution mode architecture

Deep-dive on Spark barrier execution mode: gang scheduling that launches all N tasks of a stage at once so they can run MPI-style all-to-all communica…

Read article →
ARTICLE · 06

Spark runtime bloom-filter join architecture

Deep-dive on Spark's runtime bloom-filter (dynamic filtering) join: why shuffle dominates cost, building a bloom filter from the small side'…

Read article →
ARTICLE · 07

Spark Broadcast Variables, in depth: closures vs broadcast, TorrentBroadcast, memory, lifecycle and refresh

How Spark broadcast variables work and how to use them safely: why closures ship data per task, the Broadcast API, how TorrentBroadcast splits values …

Read article →
ARTICLE · 08

Spark Broadcast Hints, in depth: what BROADCAST costs on the driver and executors, sizing the build side, hard limits, timeouts and AQE

What happens after Spark accepts a BROADCAST hint: the build-side job, collection to the driver, HashedRelation construction and torrent broadcast; wh…

Read article →
ARTICLE · 09

Spark broadcast joins -- avoiding the shuffle for large-small joins

Deep-dive on Spark broadcast joins: the join/shuffle problem, broadcasting the small side to avoid shuffling the large side, the auto-broadcast thresh…

Read article →
ARTICLE · 10

Spark cache() and persist(), in depth: blocks, the BlockManager, executor loss, dynamic allocation and sizing a cache from measurements

What happens underneath cache() and persist() in Spark: how a cached partition becomes a named block, the BlockManager's memory and disk stores, …

Read article →
ARTICLE · 11

Spark Cache Strategy, in depth: counting reuse, a cost model, where to cut the cache point, lifecycle and staleness

How to decide what to cache in Spark: counting reuse from plans, a cost model with measured times, cutting the cache after narrowing and before fan-ou…

Read article →
ARTICLE · 12

Spark coalesce vs repartition, in depth: physical plans, upstream parallelism, balance, determinism and output files

What Spark's coalesce and repartition really do: narrow merge versus full shuffle, the physical plans, how coalesce silently shrinks upstream par…

Read article →
ARTICLE · 13

Spark Columnar Processing, in depth: ColumnarBatch, vectorized readers, row and column transitions, and columnar plugins

How columnar processing works in open-source Spark: ColumnVector and ColumnarBatch, the vectorized Parquet and ORC readers, ColumnarToRow and RowToCol…

Read article →
ARTICLE · 14

Spark DataFrame Complex Types, in depth: structs, arrays, maps, higher-order functions and nested performance

How to work with Spark DataFrame complex types: struct, array and map representation and access, explode versus higher-order functions (transform, fil…

Read article →
ARTICLE · 15

Spark Connect architecture

Deep-dive on Apache Spark Connect: the DataFrame API as an unresolved protobuf logical-plan builder, the gRPC channel, the Connect server's analy…

Read article →
ARTICLE · 16

Spark Connected Components, in depth: min-label propagation, large-star / small-star, randomized contraction and the traps that sink production runs

How connected components works on Spark: GraphX min-label Pregel, GraphFrames two_phase large-star / small-star and randomized contraction, plus check…

Read article →
ARTICLE · 17

Spark DAG (Directed Acyclic Graph), in depth: how the DAGScheduler builds, reuses, skips and recovers stages

How Apache Spark turns an action into a directed acyclic graph of stages: the DAGScheduler event loop, building result and shuffle map stages from dep…

Read article →
ARTICLE · 18

Spark Data Skew Handling, in depth: finding hot keys, what AQE splits for you, salting, two-stage aggregation and write-side skew

A practical guide to data skew in Apache Spark: why one oversized partition sets the runtime of a whole stage, how to find hot keys in the Spark UI an…

Read article →
ARTICLE · 19

Spark DataSource V2 architecture

Deep-dive on Spark DataSource V2: TableCatalog and DDL, capability sets, the ScanBuilder pushdown negotiation and the unhandled-filter contract, stati…

Read article →
ARTICLE · 20

Spark on Databricks, in depth: the runtime, compute types and access modes, the life of a query, jobs as code, tuning and cost

What changes when Apache Spark runs on Databricks: the Databricks Runtime and Photon, all-purpose, jobs, serverless and SQL compute, standard versus d…

Read article →
ARTICLE · 21

Spark DataFrame API, in depth: column expressions, null logic, joins, aggregations, windows and the plans they build

A practical guide to the Spark DataFrame API: lazy plans versus actions, column expressions, select and withColumn, three-valued null logic and ANSI m…

Read article →
ARTICLE · 22

Spark DataFrame and Dataset

How DataFrames add schema and columnar optimization to Spark, how Datasets extend DataFrames with type safety, and when to use each.

Read article →
ARTICLE · 23

Spark Dataset API, in depth: typed transformations, groupByKey and KeyValueGroupedDataset, joinWith, the Aggregator API, and what each does to your physical plan

A practical guide to the typed side of Apache Spark's Dataset API in Scala and Java: map, flatMap and mapPartitions, groupByKey with mapGroups, f…

Read article →
ARTICLE · 24

Spark Debugging, in depth: reading distributed stack traces, isolating bad records, executor deaths and wrong results

A practical method for debugging Apache Spark jobs: why errors surface at actions, how to read Job aborted traces and find the real cause, a taxonomy …

Read article →
ARTICLE · 25

Delta Lake Architecture in Depth

A 2500-word walkthrough of Delta Lake: transaction log, Parquet data files, checkpoints, OPTIMIZE + Z-ORDER, MERGE, time travel, streaming, VACUUM.

Read article →
ARTICLE · 26

Spark DStreams (Legacy), in depth: micro-batches, receivers, Kafka direct streams, checkpoint recovery and migrating off

How the deprecated Spark Streaming DStream engine works and how to operate or retire it: batch-to-RDD model, receivers versus the Kafka direct stream,…

Read article →
ARTICLE · 27

Spark dynamic allocation architecture

Deep-dive on Spark dynamic allocation: the ExecutorAllocationManager's backlog and idle timers, exponential ramp-up, external shuffle service vs …

Read article →
ARTICLE · 28

Spark Encryption, in depth: TLS and AES for RPC, local disk I/O encryption, UI TLS and Parquet column keys

How to encrypt everything a Spark application touches: the separate switches for RPC traffic (TLS or the legacy AES protocol), temporary data on local…

Read article →
ARTICLE · 29

Spark Execution Architecture in Depth: plans, stages, tasks, shuffle and memory

How a Spark action becomes jobs, stages and tasks: Catalyst and AQE, the DAG and task schedulers, executor memory, shuffle mechanics, a sized worked e…

Read article →
ARTICLE · 30

Spark EXPLAIN Plans

A comprehensive guide to Apache Spark EXPLAIN plans: reading logical and physical execution plans, understanding Catalyst optimizer rules, interpretin…

Read article →
ARTICLE · 31

Spark external shuffle service architecture - decoupling shuffle from executor lifetime

Deep-dive on Spark's external shuffle service: sort-based shuffle file layout and index, Netty zero-copy serving, surviving executor death for dy…

Read article →
ARTICLE · 32

Spark GPU Acceleration, in depth: how the cuDF plugin runs SQL and DataFrames on GPUs, how to size memory, and when it pays

How GPU acceleration for Apache Spark works: plan rewriting by the NVIDIA cuDF plugin (formerly the RAPIDS Accelerator), columnar batches, PCIe and pi…

Read article →
ARTICLE · 33

Spark GraphFrames, in depth: DataFrame graphs, motif queries, Pregel and running graph algorithms at scale

A practical guide to GraphFrames on Apache Spark: installation under the io.graphframes coordinates and graphframes-py, building graphs from tables, m…

Read article →
ARTICLE · 34

Spark GraphX, in depth: the property graph, vertex-cut partitioning, aggregateMessages, Pregel and running it in production

How Spark GraphX works and when to use it: the property graph and triplet model, vertex-cut storage with routing tables and replicated vertex views, t…

Read article →
ARTICLE · 35

Spark HBase Connector, in depth

The Apache hbase-spark connector in depth: HBaseContext and one connection per executor, bulkPut, bulkGet and hbaseRDD, the DataFrame source and hbase…

Read article →
ARTICLE · 36

Spark History Server, in depth: event logs, replay, the disk and hybrid stores, rolling logs, compaction and running it in production

How the Spark History Server works and how to run it well: the scan loop and listing database, replaying event logs into the KV store, sizing memory a…

Read article →
ARTICLE · 37

Spark + Hudi architecture

Spark on Apache Hudi: the timeline, copy-on-write versus merge-on-read, index-driven upserts, snapshot and incremental reads, and table services.

Read article →
ARTICLE · 38

Spark + Iceberg architecture

Iceberg from the Spark side: catalog configuration, DataFrameWriterV2 and SQL writes, MERGE with copy-on-write vs merge-on-read, time travel, partitio…

Read article →
ARTICLE · 39

Spark and JDBC, in depth: partitioned reads, pushdown, fetch size, type mapping and safe writes

How the Spark JDBC data source really works: driver-side schema discovery, the WHERE clauses generated from partitionColumn and bounds, predicate list…

Read article →
ARTICLE · 40

Spark Join Strategy Hints, in depth: how the planner picks a join, when hints win, and when they are ignored

A practical guide to Spark SQL join strategy hints: BROADCAST, MERGE, SHUFFLE_HASH and SHUFFLE_REPLICATE_NL, how join selection works with and without…

Read article →
ARTICLE · 41

Spark JSON Read/Write, in depth: JSON Lines, schema inference, corrupt records, nested data and VARIANT

How Spark reads and writes JSON: JSON Lines versus multiLine documents and splittability, the cost and pitfalls of schema inference, PERMISSIVE, DROPM…

Read article →
ARTICLE · 42

Spark + Jupyter Integration, in depth: where the driver lives, and why it decides everything

Running Spark from Jupyter: the three driver placements (kernel as driver, Livy and sparkmagic, Spark Connect), session lifecycle and getOrCreate trap…

Read article →
ARTICLE · 43

Koalas → Spark Pandas API Migration, in depth: the renames, the per-release breaks up to Spark 4.0, a codemod and a parity-test harness

A practical playbook for moving code from Koalas (databricks.koalas) to the pandas API on Spark (pyspark.pandas): what was renamed, which behaviours c…

Read article →
ARTICLE · 44

Spark Kryo Serialization, in depth: where spark.serializer applies, registration, buffers and failure modes

Spark Kryo serialization explained: which bytes spark.serializer controls and which it does not, how Kryo encodes objects, registering classes and cus…

Read article →
ARTICLE · 45

Spark on Kubernetes architecture

Spark on Kubernetes in depth: the driver pod and headless service, RBAC, pod templates, OOMKilled versus memory overhead, shuffle tracking, PVC reuse …

Read article →
ARTICLE · 46

Apache Livy, in depth: running Spark over REST with interactive sessions, batches, impersonation and recovery

How Apache Livy runs Spark work over REST: interactive sessions with a remote driver and REPL versus batch submission, the sessions, statements and ba…

Read article →
ARTICLE · 47

Spark unified memory management architecture

Deep-dive on Spark memory: the reserved/user/unified heap layout, unified execution and storage with borrow-and-evict, the eviction asymmetry, spill-t…

Read article →
ARTICLE · 48

Spark Memory Tuning, in depth: container anatomy, telling heap OOM, container kills, driver OOM and spill apart, and a worked executor sizing

A symptom-first guide to Spark memory tuning: what makes up an executor container, how to tell heap OOM, container kills, driver OOM and spill apart, …

Read article →
ARTICLE · 49

Spark ML Overview, in depth: what the library contains, how distributed training really executes, and when to use something else

A practical map of Apache Spark's machine learning library: spark.ml versus the RDD-based spark.mllib, the algorithm families, how training runs …

Read article →
ARTICLE · 50

Spark ML Classification, in depth: the classifier catalog, distributed training, rawPrediction vs probability, thresholds and class imbalance

How classification works in Spark ML: the nine classifiers and which are binary only, how logistic regression and trees train across executors, rawPre…

Read article →
ARTICLE · 51

Spark ML Cross-Validation, in depth: how CrossValidator splits, fits and selects, fold design for grouped and time-ordered data, and reading the metrics honestly

What Spark's CrossValidator actually does with your DataFrame: the seeded fold split, per-fold caching, the driver thread pool, argmax selection …

Read article →
ARTICLE · 52

Spark ML Feature Engineering, in depth: point-in-time features, imputation, scaling, target encoding and selection

A practical guide to feature engineering with Spark ML: building leak-free point-in-time aggregates, time-based splits, imputing and scaling numeric f…

Read article →
ARTICLE · 53

Spark ML Model Persistence, in depth: the saved-model layout, custom stages, cross-version loading and safe promotion

How Spark ML saves and loads models: the metadata, stages and data directories, a train-save-validate-load cycle, custom stage rules, tuning models, t…

Read article →
ARTICLE · 54

Spark ML Pipelines, in depth: the fit/transform contract, leak-free tuning, custom stages, persistence and scoring at scale

A practical guide to pyspark.ml Pipelines: how Estimators, Transformers and PipelineModels work, a full worked churn example, handling nulls and unsee…

Read article →
ARTICLE · 55

Spark ML Recommendation, in depth: ALS collaborative filtering from first principles to production top-K serving

How Spark ML's alternating least squares recommender works: the factorization objective, what one iteration solves, implicit-feedback confidence,…

Read article →
ARTICLE · 56

Spark ML Regression, in depth: solvers, GLMs, boosted trees, a worked delivery-time model and the failures that matter

How Spark ML regressors train on a cluster: normal vs L-BFGS solvers, elastic net, GLM families for counts and skew, GBT with early stopping, a full P…

Read article →
ARTICLE · 57

Spark ML Transformers and Estimators, in depth: the fit and transform contract, Params resolution, Pipeline.fit and a learned custom stage

How Spark ML Transformers and Estimators really work: the fit/transform contract, Params precedence and copy, how Pipeline.fit walks stages, a quantil…

Read article →
ARTICLE · 58

Spark MLlib, in depth: the feature, statistics, similarity-search and pattern-mining toolkit, evaluation, and leaving the RDD API

A practical guide to the parts of Spark MLlib around the model: which feature transformers learn state and what they compute, hashing versus vocabular…

Read article →
ARTICLE · 59

Apache Spark Overview

What Spark is, how RDDs and DataFrames give in-memory distributed computing, and where Spark fits versus MapReduce, Flink, and modern warehouses.

Read article →
ARTICLE · 60

Spark PageRank

Classic algorithm on GraphFrames.

Read article →
ARTICLE · 61

Spark Pandas API, in depth: how pyspark.pandas maps pandas semantics onto Spark plans, the default index, ordering, apply_batch, caching and when to drop to PySpark

A first-principles guide to the pandas API on Spark (pyspark.pandas): the internal frame and lazy plans, the three default index types and their costs…

Read article →
ARTICLE · 62

Spark Pandas UDFs, in depth: type-hint dispatch, struct columns, the batch-boundary trap, iterator setup, grouped aggregates and cogroup

What a Spark pandas UDF actually receives and must return: the four type-hint shapes, StructType as pandas DataFrame, why batch-local statistics are a…

Read article →
ARTICLE · 63

Spark Parquet Read + Write, in depth

How Spark reads and writes Parquet: row groups, column chunks, pages and the footer, split planning, filter pushdown and what defeats it, a worked lay…

Read article →
ARTICLE · 64

Spark Partition Pruning vs Predicate Pushdown, Static and Dynamic

How Spark partition pruning skips data: static pruning on partition columns, dynamic pruning from join values, predicate pushdown, min/max file skippi…

Read article →
ARTICLE · 65

Spark Partitioning in Depth: Read Splits, Shuffle Partitions, Write Layout and Skew, With the Arithmetic

How Spark decides partition counts at the three boundaries of a job: the file split formula with maxPartitionBytes and openCostInBytes worked through,…

Read article →
ARTICLE · 66

Spark Persistence, in depth: what to cache, storage levels, materialization, eviction and when to let go

How persist() and cache() really work in Apache Spark: lazy marking, block storage on executors, storage levels and their defaults across RDD, Dataset…

Read article →
ARTICLE · 67

Spark Photon architecture

Deep-dive on Photon: a C++ vectorized execution engine beneath Spark SQL. How the planner splits the physical plan into native and JVM operators, how …

Read article →
ARTICLE · 68

PySpark Performance, in depth: the JVM-Python boundary, UDF cost, Arrow and vectorised UDFs, and a worked rewrite

Why PySpark jobs are slow and how to make them fast: where Python runs in Spark, what a row-at-a-time UDF costs, Arrow and vectorised UDF families wit…

Read article →
ARTICLE · 69

Spark RDD lineage architecture

Deep-dive on Spark RDD lineage: lazy transformations building a DAG, narrow vs wide dependencies, stage cutting at shuffle boundaries, recomputation o…

Read article →
ARTICLE · 70

Spark S3 optimization, in depth: S3A committers, read tuning, request limits and file layout

How to run Spark well on Amazon S3: why rename-based commits are slow and unsafe, how the S3A directory, partitioned and magic committers work, the ex…

Read article →
ARTICLE · 71

Spark Salting for Skew Joins, in depth: choosing the salt count, salting only hot keys, which join types survive, and when AQE is enough

A hands-on guide to salting skewed joins in Spark: measuring hot keys, sizing the salt count from rows per task, selective salting with a broadcast ho…

Read article →
ARTICLE · 72

Spark Secrets Management, in depth: where credentials leak in a Spark application and how to deliver them safely

A practical guide to secrets in Apache Spark: the paths a credential takes through spark-submit, driver, executors, event logs and the UI, what spark.…

Read article →
ARTICLE · 73

Spark Security

Securing Apache Spark clusters: Kerberos authentication, Ranger authorization, TLS encryption, encrypted shuffle, network isolation, and audit logging…

Read article →
ARTICLE · 74

Spark shuffle architecture

Deep-dive on Spark shuffle: map output + partitioner + external shuffle service + reduce fetch, plus push-based shuffle and AQE.

Read article →
ARTICLE · 75

Spark Sort-Merge Join, in depth: the merge algorithm, buffered matches, plans, bucketing and skew

How Spark executes a sort-merge join, based on the Spark 3.5 source: shuffle and sort requirements, streamed and buffered sides by join type, the Exte…

Read article →
ARTICLE · 76

Spark SQL

How Spark SQL provides ANSI-compatible SQL over any DataFrame source (Parquet, Delta, JDBC, Kafka), the query engine architecture, and Thrift server.

Read article →
ARTICLE · 77

Spark SQL Config Tuning, in depth: where settings live, how to verify them, partition arithmetic, joins, skew and a worked tuning pass

How to tune Spark SQL configuration as an engineering surface: configuration layers and precedence, runtime versus static settings, reading effective …

Read article →
ARTICLE · 78

Spark SQL Functions, in depth: how built-ins resolve and run, NULL and ANSI semantics, try_ functions, time zones, aggregates and UDF trade-offs

How Spark SQL built-in functions are resolved, optimized and code-generated; NULL rules, ANSI mode and the try_ family, time-zone traps, deterministic…

Read article →
ARTICLE · 79

Spark SQL optimizer architecture

Deep-dive on Catalyst optimizer: analyzer, rule-based rewrites, cost model, physical planner, Tungsten codegen, AQE, and extensions.

Read article →
ARTICLE · 80

Jobs, Stages and Tasks in Spark: How Spark Splits Work

Jobs, stages and tasks in Spark: where stage boundaries fall, what sets partition count, task vs stage retry, speculation, AQE, and how to find stragg…

Read article →
ARTICLE · 81

Spark Statistics + CBO, in depth: collecting statistics, estimation formulas, join reordering and AQE

How Spark SQL statistics and the cost-based optimizer work: ANALYZE TABLE variants, where statistics are stored, filter and join estimation formulas, …

Read article →
ARTICLE · 82

Spark Structured Streaming Deduplication, in depth: dropDuplicates, watermarks, dropDuplicatesWithinWatermark, state sizing and idempotent sinks

How to remove duplicate events in Spark Structured Streaming: where duplicates come from, how the streaming dedup operator uses state, dropDuplicates …

Read article →
ARTICLE · 83

Spark Structured Streaming Architecture in Depth

How Spark Structured Streaming works inside: the micro-batch loop and checkpoint, watermarks, the state store, exactly-once sources and sinks, trigger…

Read article →
ARTICLE · 84

Spark Streaming Deduplication, in depth: choosing the key, measuring the horizon, tiered dedup and custom state

Designing deduplication for Spark Structured Streaming: where duplicates come from, event ID vs offset vs content-hash keys, measuring the duplicate-d…

Read article →
ARTICLE · 85

Spark Streaming foreachBatch, in depth: the micro-batch contract, batchId idempotency, MERGE upserts, multi-sink fan-out and failure modes

How Spark Structured Streaming's foreachBatch really works: where your function runs in the micro-batch lifecycle, why it is at-least-once by def…

Read article →
ARTICLE · 86

Spark Structured Streaming + Kafka: Offsets and Exactly-Once

Spark Structured Streaming with Kafka: where offsets live, checkpoint offset management, end-to-end exactly-once, rate limits, and what you cannot cha…

Read article →
ARTICLE · 87

Spark Streaming Kafka Source, in depth

The Structured Streaming Kafka source as a component: per-batch offset planning, consumer pools, starting-offset precedence, every option for Spark 4.…

Read article →
ARTICLE · 88

Spark Streaming Sinks, in depth: the commit sequence, output modes, the file-sink metadata log, Kafka duplicates and choosing a sink

How Structured Streaming sinks really work: the offsets-to-commits sequence that makes replays possible, the Spark 4.2 sink and output-mode matrix, th…

Read article →
ARTICLE · 89

Spark Streaming Sources, in depth: the replayable-offset contract, file source options, test sources and custom Python sources

How Structured Streaming sources work: offsets and the write-ahead offset log, the built-in file, Kafka, table, rate, rate-micro-batch and socket sour…

Read article →
ARTICLE · 90

Spark Streaming Triggers, in depth: default, fixed-interval, available-now, continuous and Real-time, and how to choose an interval

How Structured Streaming triggers work: the micro-batch plan, offset log and commit loop; documented semantics of default, ProcessingTime, Once, Avail…

Read article →
ARTICLE · 91

Spark Structured Streaming Watermarks and Late Data

How Spark Structured Streaming watermarks work: event time, withWatermark, the rule for dropping late rows, window state eviction, output modes, strea…

Read article →
ARTICLE · 92

Spark Structured Streaming state

Deep-dive on Spark Structured Streaming state management: keyed state stores, checkpointing for exactly-once recovery, watermarks bounding state and h…

Read article →
ARTICLE · 93

Spark Triangle Count, in depth: wedges, degree ordering, GraphX and GraphFrames, and a skew-aware DataFrame implementation

How to count triangles at scale in Spark: definitions, the degree-ordered O(m^1.5) algorithm with a worked example, GraphX and GraphFrames operators, …

Read article →
ARTICLE · 94

Spark Tungsten -- pushing performance to the metal

Deep-dive on Spark's Project Tungsten: the JVM overhead problem (object memory, GC, virtual calls), off-heap managed memory, compact binary rows …

Read article →
ARTICLE · 95

Spark UI in Depth: How It Collects Its Data, How to Read Every Tab, and How to Diagnose Slow Jobs

A practical guide to the Apache Spark web UI: the listener bus and status store behind it, retention limits and dropped events, reading the Jobs, Stag…

Read article →
ARTICLE · 96

Spark whole-stage code generation architecture

Deep-dive on Spark SQL whole-stage code generation: replacing the Volcano iterator model with a fused Janino-compiled Java loop via the produce/consum…

Read article →
ARTICLE · 97

Spark Wide vs Narrow Dependencies, in depth: the dependency classes, how to read them in a plan, how to turn a wide dependency into a narrow one and what each costs on failure

A first-principles guide to narrow and wide dependencies in Apache Spark: what the terms mean precisely, the Dependency classes behind them, which ope…

Read article →