A Spark job that reads from a database or an object store needs credentials, and the easy way to pass them is a --conf flag or an option on spark.read. The easy way leaks. A Spark application is a distributed program whose configuration is copied to every executor, written to event logs, shown in a web UI and often echoed into notebook output. A password passed casually can end up in five places, some of which are retained for months and readable by people who never had access to the database.

This article follows a credential through a Spark application, shows where it leaks, explains what Spark's redaction does and does not protect and gives concrete delivery patterns for Kubernetes, YARN and managed platforms. It also covers Spark's own internal secret, the one that authenticates executors to the driver, which is a separate problem that is often left at its insecure default. Configuration names are checked against the Spark 4.2 documentation.

Advertisement

Two different problems

Spark has two kinds of secrets, and they need different treatment. The first is application credentials: database passwords, cloud access keys, API tokens and keystore passwords. They let your code reach external systems. The second is Spark's own authentication secret, set by spark.authenticate, which stops a stranger on the network from registering a fake executor or reading shuffle data. The first is your data-access boundary. The second protects the cluster's internal traffic, and it is off by default.

The best application credential is one that does not exist. On every major cloud, a pod or VM can carry an identity, such as GKE Workload Identity, IAM roles for service accounts on EKS or an instance profile, and the storage connectors pick up short-lived tokens automatically. Static secrets remain for databases, third-party APIs and on-premises systems, and those are what the rest of this article is about.

Where a credential travels and leaks

Where a credential travels in a Spark application, and where it leaksspark-submit--conf, args, envDriverSparkConf, plans, closuresExecutor 1tasks, JDBC / S3 clientsExecutor Ntasks, JDBC / S3 clientsProcess list / pod specargv and env are readableEvent log + historyconf and plans persistedSpark UIEnvironment, SQL tabsLogs and notebooksprints, stack tracesBetter: identity, not secretsworkload identity, instance roles, delegation tokensOtherwise: reference, not valuemounted files, credential providers, secret managerRedaction hides what matches a regex in some outputs. It is a backstop, not a delivery mechanism.
A secret passed as configuration is copied to the driver and every executor and persisted by the event log. Each red box is a place it can be read later by someone other than the job.
SurfaceHow the secret gets thereWho can read it
Process list, pod specPassed on the command line or as a literal env varAnyone who can run ps or read pods in the namespace
Spark UI Environment tabAny spark.* conf, system property or env varAnyone who can reach the UI
Event log, History ServerConf is written into the application start eventAnyone with read access to the log directory, for its whole retention
SQL tab and explain outputOptions passed to a data source appear in plansUI users, logs that print plans
MetastoreCREATE TABLE ... OPTIONS with a password is stored as table propertiesAnyone who can describe the table
Notebooks and logsprint(), df.show() of a config table, exception messagesNotebook viewers, log pipelines

The metastore row deserves attention. A JDBC table defined with a password in its options stores that password with the table definition, where it outlives the job that created it. Use a view over a runtime read, or a connector that takes a reference to a secret, instead.

Advertisement

What redaction does, and its limits

Spark redacts configuration values whose key or value matches spark.redaction.regex, whose default is (?i)secret|password|token|access[.]?key. Matching entries are hidden in the Environment tab and in logs such as YARN and event logs. A second setting, spark.redaction.string.regex, has no default. When set, it replaces matching parts of strings Spark produces, currently the output of SQL explain commands.

Three limits matter. The regex is matched against keys and values, so a credential under a key such as spark.myapp.dbcred, whose value looks like random text, is not hidden unless you extend the regex. It applies to Spark's own outputs, not to your print statements, exceptions or data. And the value is still present in memory and passed to executors; redaction changes what is displayed, not who holds the secret. Extend the pattern cluster-wide in spark-defaults.conf so individual jobs cannot forget it:

# spark-defaults.conf
spark.redaction.regex         (?i)secret|pass|token|access[.]?key|cred|apikey|jdbc:.*@
spark.redaction.string.regex  (?i)(password|pwd)=[^;&\s]+

Test it: submit a job with a canary value such as spark.myapp.password=CANARY-7f3a, then search the UI, the event log file and the driver log for the canary. That check belongs in CI for your platform images.

Delivery on Kubernetes: mounted Secrets and secretKeyRef

Spark on Kubernetes can attach Kubernetes Secrets to the driver and executor pods without the value ever appearing in the Spark configuration. Two property families do this. spark.kubernetes.driver.secrets.[SecretName]=<mount path> and its executor twin mount a Secret as files. spark.kubernetes.driver.secretKeyRef.[EnvName]=name:key and its executor twin expose one key of a Secret as an environment variable. The Secret must be in the same namespace as the pods.

kubectl create secret generic orders-db -n analytics \
  --from-literal=username=etl --from-literal=password='...'

spark-submit \
  --master k8s://https://api.cluster:6443 --deploy-mode cluster \
  --conf spark.kubernetes.namespace=analytics \
  --conf spark.kubernetes.driver.secrets.orders-db=/etc/secrets/orders-db \
  --conf spark.kubernetes.executor.secrets.orders-db=/etc/secrets/orders-db \
  --conf spark.kubernetes.driver.secretKeyRef.DB_USER=orders-db:username \
  local:///opt/app/job.py

Prefer files to environment variables. Environment variables are inherited by child processes, show up in debug dumps and are captured by some crash reporters. A mounted file can be read once into memory. The Kubernetes RBAC question also matters: anyone who can read Secrets or exec into pods in the namespace can read the credential, so give Spark jobs their own namespace and service account. The deployment side is covered in Spark on Kubernetes.

Hadoop credential providers

On YARN and anywhere Hadoop libraries are in play, the Hadoop credential provider API stores secrets in an encrypted keystore, such as a JCEKS file on HDFS, and components look them up by alias. Spark's SSL password settings (keyPassword, keyStorePassword, trustStorePassword under each spark.ssl namespace) can be resolved this way, and the S3A connector can too. The path is set with hadoop.security.credential.provider.path in the Hadoop configuration or in SparkConf.

hadoop credential create spark.ssl.keyPassword -value '...' \
  -provider jceks://hdfs@nn1.example.com:9001/user/spark/ssl.jceks

spark-submit \
  --conf spark.hadoop.hadoop.security.credential.provider.path=jceks://hdfs@nn1.example.com:9001/user/spark/ssl.jceks \
  ...

The keystore file itself must be protected by HDFS permissions, and its keystore password is a secret in turn. The provider moves the problem to one well-guarded file rather than removing it. On Kerberised clusters, delegation tokens are the equivalent of workload identity: the job holds a renewable token, not a password. Spark on Kubernetes can read existing tokens from a Secret named by spark.kubernetes.kerberos.tokenSecret.name and spark.kubernetes.kerberos.tokenSecret.itemKey. The Kerberos side is covered in Hadoop Kerberos.

Fetching at runtime from a secret manager

The most flexible pattern is to pass only a reference, such as a secret name, and fetch the value at runtime from a manager like Google Secret Manager, AWS Secrets Manager or Vault, using the pod's workload identity. The reference is harmless in the UI and logs. The subtle part is executors: code inside mapPartitions runs on executors, and if every task calls the manager you will hit rate limits on a large job. Cache one client and one value per Python worker process, and refresh on expiry. The cache must live in an importable module shipped with --py-files: functions and globals defined in the driver script are pickled into each task by value, so a cache declared there starts empty in every task.

# secret_cache.py, shipped with --py-files; module state persists in each Python worker
import time
_client, _cache = None, {}

def get_secret(ref, ttl=300):
    global _client
    hit = _cache.get(ref)
    if hit and time.time() - hit[1] < ttl:
        return hit[0]
    if _client is None:
        from google.cloud import secretmanager
        _client = secretmanager.SecretManagerServiceClient()
    value = _client.access_secret_version(name=ref).payload.data.decode()
    _cache[ref] = (value, time.time())
    return value

# job.py
import os
from pyspark.sql import SparkSession

SECRET_REF = os.environ["ORDERS_DB_SECRET_REF"]      # a name, not a value

def write_partition(rows):
    import psycopg2
    from secret_cache import get_secret               # imported, so shared across tasks
    conn = psycopg2.connect(host="orders-db", user="etl", password=get_secret(SECRET_REF))
    try:
        with conn, conn.cursor() as cur:
            for r in rows:
                cur.execute("INSERT INTO audit VALUES (%s, %s)", (r.id, r.status))
    finally:
        conn.close()

spark = SparkSession.builder.getOrCreate()
spark.read.parquet("gs://lake/orders/").foreachPartition(write_partition)

Two details matter. The secret is fetched on the executor, not captured from the driver, so it is never serialised into task closures. And the environment variable holds only the reference, so seeing it is not enough to connect. For the built-in JDBC source, which takes the password as a read option on the driver, fetch on the driver just before the call and do not store it in SparkConf or in a table definition.

Securing Spark&#x27;s own secret and traffic

With spark.authenticate=true, Spark's internal connections are authenticated with a shared secret. On YARN, Spark generates it per application. On Kubernetes, Spark also generates a unique secret per application, but propagates it to executors through environment variables, which anyone who can read pod specs can see. The documented alternative is to mount the secret as a file with spark.authenticate.secret.file (or the .driver.file and .executor.file variants), available since 3.0.

SettingDefaultRecommendation
spark.authenticatefalsetrue on any shared network
spark.network.crypto.enabledfalsetrue, for AES-based RPC encryption
spark.network.crypto.cipherAES/CTR/NoPaddingAES/GCM/NoPadding for authenticated encryption (4.0.0, 3.5.2, 3.4.4+)
spark.network.crypto.authEngineVersion12, as the docs recommend (4.0.0+)
spark.io.encryption.enabledfalsetrue when shuffle and spill files may hold sensitive data
spark.acls.enablefalsetrue, with view and modify ACLs; front the UI with spark.ui.filters

The UI matters because the Environment and SQL tabs show exactly what this article worries about. Spark 4.x documents org.apache.spark.ui.JWSFilter for spark.ui.filters, which accepts a signed JWT in the Authorization header; otherwise put the UI and History Server behind an authenticating proxy. The Spark UI guide covers what each tab exposes.

Rotation and long-running jobs

A batch job reads a secret once and exits, so rotation only needs the next run to see the new value. Streaming jobs run for weeks. A Kubernetes Secret mounted as a volume, not through subPath, is eventually refreshed in the pod when the Secret changes, but an already-open connection pool keeps using the old password until it reconnects. Build for this: re-read the file or re-fetch from the manager on authentication failure, retry once, then fail the task so Spark reschedules it. Rotate with two valid credentials overlapping, so running executors never see both old and new rejected. The overlap pattern is covered in secrets rotation.

Worked example: moving one job off a plaintext password

A nightly job was submitted with --conf spark.myapp.dbpass=... and read the value in code. A review found the password in the History Server, retained for 90 days, and in the Environment tab, because the key did not match the redaction regex. The fix took four steps.

  1. Rotate the password immediately; it had been readable for months, so treat it as compromised.
  2. Store the new password in the secret manager and give the job's Kubernetes service account read access to that one secret only.
  3. Change the job to take ORDERS_DB_SECRET_REF and fetch at runtime, as above; remove the conf key.
  4. Extend spark.redaction.regex platform-wide, add the canary test to CI and purge old event logs that contain the old value.

The job now has no credential in its configuration, the History Server shows only a reference, and a rotation needs no redeploy.

Failure modes

  • Secret in a conf key the regex does not match: visible in the UI and event logs. Extend the regex and use canary tests.
  • Secret in a table definition: persists in the metastore after the job is gone. Audit table properties.
  • Per-task secret-manager calls: throttling fails the job under load. Cache in a shipped module, per worker process.
  • Kubernetes Secret in the wrong namespace: pods stay pending with mount errors. Keep Secrets beside the job.
  • Stale credential after rotation in a streaming job: authentication errors until restart. Reconnect and re-fetch on auth failure.
  • Default internal auth on a shared cluster: fake executors or shuffle reads are possible. Enable spark.authenticate and RPC encryption.

What to do next

  1. Grep your event log directory and History Server for known secret patterns and a canary value; rotate anything you find.
  2. Move cloud storage access to workload identity; remove static access keys from job configuration.
  3. Deliver the remaining credentials as mounted files, credential-provider aliases or secret-manager references, never as conf values.
  4. Extend spark.redaction.regex cluster-wide and add a CI canary test for the UI, event log and driver log.
  5. Turn on spark.authenticate, spark.network.crypto.enabled with AES/GCM and, where needed, spark.io.encryption.enabled.
  6. Put the Spark UI and History Server behind authentication and restrict the event log directory's read permissions.
Key takeaway: In Spark, a secret passed as configuration is copied to every executor, persisted in event logs and displayed in the UI, and redaction only hides names that match a regex. Prefer workload identity so there is no secret at all. Otherwise pass a reference and resolve it late: Kubernetes Secrets mounted as files, Hadoop credential providers, or a secret manager fetched once per worker process and refreshed on expiry. Separately, turn on Spark's own authentication and AES-GCM RPC encryption, protect the UI and History Server, and test the whole setup with a canary secret.