Encryption in Spark is not a single switch. A Spark application moves data across four different paths: RPC connections between the driver, executors and shuffle services; temporary files on executor disks; HTTP traffic to the UI and History Server; and the tables it finally reads and writes. Each path has its own configuration, its own defaults and its own failure modes, and enabling one leaves the others exactly as they were. Every one of them is off by default.

This article explains each layer from first principles, shows a hardened configuration and how to prove it works, and covers what the settings cost and how they break. Property names and defaults are taken from the Apache Spark security guide and the Parquet data source guide as of October 2026. Delivering the keys and passwords these settings need is a separate problem, covered in Spark secrets management.

Advertisement

The four paths and what protects each

Start with a map. The driver and executors talk over Netty-based RPC: task launches, status updates, broadcast pieces and, most importantly, shuffle blocks fetched between executors or from the external shuffle service. Executors write shuffle output, sort spills, and disk-persisted cache and broadcast blocks to local directories. Humans and tools reach the driver UI and the History Server over HTTP. Finally, jobs read and write tables in HDFS or an object store.

Where Spark data travels, and which setting protects each hopDriverscheduler, block managerExecutorstasks, shuffle, cacheBrowserUI and History ServerLocal diskshuffle, spills, cached blocksTable storageParquet on HDFS or S3 or GCSRPC: spark.ssl.rpc or AESspark.ssl.uispark.io.encryptionParquet column keysplus storage-side encryptionEach hop has its own switch. Turning one on does not turn on any other.The external shuffle service is another RPC endpoint and must match the executors.
Four paths, four independent switches. The table storage path is protected by the file format and the storage system, not by Spark core.
PathSettingDefaultProtects against
RPC, preferredspark.ssl.rpc.enabledfalseSniffing and tampering on the cluster network
RPC, legacyspark.network.crypto.enabledfalseSniffing; tampering only with the GCM cipher
Local diskspark.io.encryption.enabledfalseReading shuffle and spill files from disks or snapshots
UI and History Serverspark.ssl.ui.*, spark.ssl.historyServer.*falseSniffing SQL text, plans and environment values
Table dataParquet columnar encryptionoffAnyone with storage access but no key

Encryption on the wire also depends on authentication. spark.authenticate makes every internal connection prove knowledge of a shared application secret, and the legacy AES protocol requires it. Without authentication, anyone who can reach an executor port can talk to it, so turn it on first.

RPC encryption with TLS

The Spark security guide now describes two mutually exclusive ways to encrypt RPC, and prefers TLS. Spark's RPC layer uses Netty, and Netty's standard TLS support encrypts and authenticates every connection with certificates, as any other TLS service would. It is configured under the spark.ssl.rpc namespace, with the usual keystore and truststore settings, plus RPC-only options: openSSLEnabled with PEM privateKey and certChain files to use OpenSSL instead of the JDK provider, and trustStoreReloadingEnabled for long-running standalone clusters.

The detail that catches people out: setting spark.ssl.enabled turns on TLS for the UI namespaces but deliberately does not turn it on for RPC. You must set spark.ssl.rpc.enabled explicitly. The docs say this was done to give a safe upgrade path. If both TLS and the AES protocol are enabled, TLS wins, AES is not used, and Spark logs a warning.

# spark-defaults.conf: TLS for RPC, UI and History Server
spark.authenticate                         true
spark.authenticate.secret.file             /var/run/secrets/spark/auth-secret
# fixed ports, so firewalls and the packet-capture check below know where to look
spark.driver.port                          7078
spark.blockManager.port                    7079

# RPC TLS is NOT switched on by spark.ssl.enabled; it needs its own flag
# passwords: render from your secret store at submit time, never commit them
spark.ssl.rpc.enabled                      true
spark.ssl.rpc.protocol                     TLSv1.3
spark.ssl.rpc.keyStore                     /var/run/secrets/spark/rpc.jks
spark.ssl.rpc.keyStorePassword             FROM_SECRET_STORE
spark.ssl.rpc.trustStore                   /var/run/secrets/spark/truststore.jks
spark.ssl.rpc.trustStorePassword           FROM_SECRET_STORE

# UI over HTTPS
spark.ssl.ui.enabled                       true
spark.ssl.ui.protocol                      TLSv1.3
spark.ssl.ui.keyStore                      /var/run/secrets/spark/ui.jks
spark.ssl.ui.keyStorePassword              FROM_SECRET_STORE

# Temporary data on local disk
spark.io.encryption.enabled                true
spark.io.encryption.keySizeBits            256

Every endpoint must agree. Executors fetch shuffle blocks from each other and, with dynamic allocation, from the external shuffle service, which runs in the NodeManager on YARN or as its own daemon. It needs the same RPC settings and a certificate the executors trust, or fetches fail at the first shuffle.

Advertisement

The legacy AES protocol

Before TLS support, Spark shipped a bespoke scheme: spark.network.crypto.enabled derives session keys from the shared authentication secret and encrypts the stream with AES through the commons-crypto library. Many clusters still run it, and it is fine if configured correctly. Two defaults are not.

The cipher defaults to AES/CTR/NoPadding, which encrypts but does not authenticate. An attacker on the network cannot read the data but can flip bits in it undetected. That default is the subject of CVE-2025-55039. The fix is AES/GCM/NoPadding, available since 4.0.0, 3.5.2 and 3.4.4. Second, spark.network.crypto.authEngineVersion defaults to 1, which skips a key derivation step after the key exchange; version 2 applies it, and the docs recommend it. The two versions are mutually incompatible, so change every client and every shuffle service together.

# Legacy AES protocol, only where TLS is not yet possible
spark.authenticate                         true
spark.network.crypto.enabled               true
# the default, AES/CTR/NoPadding, is unauthenticated
spark.network.crypto.cipher                AES/GCM/NoPadding
# versions 1 and 2 cannot talk to each other: change every endpoint together
spark.network.crypto.authEngineVersion     2
# only once every shuffle service is upgraded
spark.network.crypto.saslFallback          false

An even older option, SASL encryption via spark.authenticate.enableSaslEncryption, survives for compatibility with old shuffle services. spark.network.crypto.saslFallback defaults to true so new clients can still authenticate to them. Once every service is upgraded, set it to false and set spark.network.sasl.serverAlwaysEncrypt so no unencrypted SASL connection is accepted.

Local disk I/O encryption

Shuffle is where plaintext lands on disk. With spark.io.encryption.enabled, Spark encrypts temporary data written to local disks: shuffle files, shuffle spills, and data blocks stored on disk for both caching and broadcast variables. The key is generated per application with the spark.io.encryption.keygen.algorithm (default HmacSHA1) key generator, at spark.io.encryption.keySizeBits of 128, 192 or 256 bits, default 128.

Two consequences follow. First, the key is sent to executors over RPC, which is why the guide strongly recommends enabling RPC encryption with it; encrypted disks with a plaintext key on the wire protect little. Second, the key lives only as long as the application, so leftover block manager directories from a crashed job are unreadable, which is exactly what you want from scratch data.

What it does not cover matters as much. The guide is explicit: application output written with APIs such as saveAsTable or saveAsHadoopFile is not included, and temporary files your own code creates may not be. Event logs, which contain SQL text and configuration, are written to their own directory and are not on the list either, so protect that directory with storage encryption and access control.

The UI and History Server

The UI is a data leak with a web page attached. The SQL tab shows query text with literal values, the Environment tab shows every configuration value that the redaction regex misses, and the History Server keeps both for as long as event logs are retained. The spark.ssl.ui and spark.ssl.historyServer namespaces serve them over HTTPS, and spark.ssl.standalone does the same for standalone master and worker UIs.

Namespaced settings inherit from the default spark.ssl.* values, except the port, which must be set per namespace. TLS stops eavesdropping but not access, so pair it with ACLs and an authenticating filter or proxy, as described in the Spark UI guide.

Data at rest: Parquet columnar encryption

Everything so far protects data while the application runs. The tables it produces need their own protection. Storage-level encryption, such as HDFS transparent encryption zones or object store server-side encryption with a customer-managed key, protects against stolen disks and is nearly free, but anyone with read access to the path reads plaintext. Parquet modular encryption, supported in Spark 3.2 and later with Parquet 1.12 and later, encrypts inside the file, so access to the bytes is no longer enough.

It uses envelope encryption. Parquet generates a random data encryption key (DEK) for each encrypted column and file, and encrypts those DEKs with master keys held in your KMS. By default it double-wraps: DEKs are encrypted with key encryption keys, which are themselves wrapped by the KMS and cached in executor memory, so a job with thousands of files does not call the KMS thousands of times. parquet.encryption.double.wrapping set to false switches to single wrapping.

from pyspark.sql import SparkSession

spark = (SparkSession.builder
    .config("spark.hadoop.parquet.crypto.factory.class",
            "org.apache.parquet.crypto.keytools.PropertiesDrivenCryptoFactory")
    .config("spark.hadoop.parquet.encryption.kms.client.class",
            "com.example.kms.VaultTransitKmsClient")      # your KmsClient implementation
    .getOrCreate())

orders = spark.read.table("raw.orders")

(orders.write
    .option("parquet.encryption.column.keys", "pii-key:email,phone;card-key:card_number")
    .option("parquet.encryption.footer.key", "footer-key")
    .mode("overwrite")
    .parquet("s3a://lake/curated/orders/"))

# Readers need no options: key metadata travels in the file, and the KMS decides
# whether this identity may unwrap each column key.
spark.read.parquet("s3a://lake/curated/orders/").select("order_id", "total").show()

The payoff is per-column access control enforced by the KMS. An analyst whose identity can unwrap footer-key but not pii-key can read order_id and total, and gets an error only when they select email. You must implement Parquet's KmsClient interface, with wrapKey, unwrapKey and initialize, for your KMS. The InMemoryKMS in the Parquet test jar is for demos only. Table formats built on Parquet, such as Iceberg, may have their own encryption features; check their documentation rather than assuming these options pass through.

Worked example: hardening a Spark on Kubernetes pipeline

A team runs nightly Spark 4 jobs on Kubernetes that join clickstream with a customer table containing email addresses and write a curated table to S3. An audit found three problems. The authentication secret was visible to anyone who could read pod specs, because on Kubernetes Spark passes it to executors in environment variables. Shuffle files containing email addresses sat unencrypted on node-local volumes. And the curated table was protected only by bucket encryption, so every data engineer with bucket access could read the emails.

The fix, in order. They mounted the authentication secret from a Kubernetes Secret and pointed spark.authenticate.secret.file at it. They issued a certificate per namespace from their internal CA, mounted the keystore and truststore, and enabled spark.ssl.rpc.enabled. They enabled I/O encryption with 256-bit keys. Finally, they wrote the curated table with email under a PII key that only the marketing service account can unwrap. The configuration above is essentially theirs. On Kubernetes there is usually no external shuffle service, so executors serve shuffle blocks directly and only the Spark pods needed the RPC settings.

They then proved each layer with the checks below, run once in staging and kept as a CI smoke test.

# Run each check once with the setting OFF and confirm it reports a leak first.
# 1. RPC: capture traffic on the pinned ports during a shuffle; count readable class names
n=$(sudo timeout 60 tcpdump -i any -A 'tcp port 7078 or tcp port 7079' 2>/dev/null \
      | grep -c 'org.apache.spark')
echo "plaintext hits: $n"            # expect many with TLS off, 0 with TLS on

# 2. Local disk: after a shuffle-heavy stage, search block files for a known value.
#    Block manager dirs live under spark.local.dir (or the executor's local volume).
dirs=$(find "${SPARK_LOCAL_DIRS:-/tmp}" -maxdepth 3 -type d -name 'blockmgr-*' 2>/dev/null)
[ -n "$dirs" ] || { echo "no blockmgr dirs found: check path, test inconclusive"; exit 2; }
grep -rl --binary-files=text 'alice@example.com' $dirs && echo LEAK || echo ok

# 3. Driver log: SSL and AES both on means AES is ignored, with a warning
grep -i 'ssl' driver.log | grep -i 'aes\|warn'

What it costs

TLS and AES-GCM both use AES instructions that modern CPUs accelerate, so the cipher itself is rarely the bottleneck. The measurable costs are elsewhere: CPU spent on shuffle-heavy stages, loss of some zero-copy transfer paths once bytes must pass through a cipher, and handshake time for many short connections. I/O encryption adds a cipher pass over every spill and shuffle write and read. Measure on your own workload by running the same shuffle-heavy job with and without each setting and comparing stage times in the UI; published figures from other hardware and versions will not tell you much.

Parquet columnar encryption costs little CPU but adds KMS calls, which double wrapping keeps rare, and it disables some tooling: any reader without the KMS client class and a permitted identity cannot open the file. That is the point, but it also applies to your compaction jobs, backfills and data quality tools.

Failure modes

  • RPC TLS assumed on. spark.ssl.enabled=true gives HTTPS on the UI while RPC stays plaintext. Check for spark.ssl.rpc.enabled explicitly.
  • Unauthenticated cipher. AES RPC with the CTR default encrypts but allows tampering. Set GCM.
  • Version split. Clients on authEngineVersion 2 and a shuffle service on 1 fail to authenticate; shuffle fetches fail and stages retry until the job dies.
  • Shuffle service left behind. Executors with TLS cannot fetch from a NodeManager shuffle service without it, and a missing certificate authority in its truststore looks like a network error.
  • Key on the wire. I/O encryption without RPC encryption sends the disk key in plaintext.
  • Output assumed covered. I/O encryption does not encrypt tables, event logs or files your code writes.
  • KMS outage. Readers of column-encrypted Parquet fail when the KMS is down or a service account loses unwrap permission; alert on it like any dependency.
  • Expired certificates. Long-running streaming jobs outlive certificates; plan rotation, or use truststore reloading on standalone clusters.

What to do next

  1. Turn on spark.authenticate everywhere, delivering the secret from a file on Kubernetes.
  2. Enable spark.ssl.rpc.enabled with certificates from your CA, including every external shuffle service. Where TLS is not yet possible, use AES with AES/GCM/NoPadding and authEngineVersion 2.
  3. Enable spark.io.encryption.enabled on any cluster whose shuffle data is sensitive, and only together with RPC encryption.
  4. Serve the UI and History Server over HTTPS behind authentication, and encrypt the event log directory at the storage layer.
  5. Classify table columns, put sensitive ones under Parquet column keys with a production KmsClient, and test that an unauthorised identity is refused.
  6. Add the tcpdump, disk grep and log checks to a staging smoke test, and benchmark one shuffle-heavy job before and after.
Key takeaway: Spark has one encryption switch per path, and each switch is off by default. Authenticate first, then encrypt RPC with TLS through spark.ssl.rpc, which spark.ssl.enabled does not turn on, or with the legacy AES protocol using GCM and version 2. Encrypt local disks with I/O encryption, and never without RPC encryption, because the disk key travels over RPC. Put the UI behind HTTPS and authentication. For data that outlives the job, use storage encryption plus Parquet column keys held in a real KMS. Then prove each layer with a packet capture, a disk grep and a log check, because a setting you have not tested is a setting you do not have.