Rotating a secret sounds like one action: change the password. In Kubernetes it is a chain of copies. The value lives in a secret store, a controller copies it into a Kubernetes Secret, the kubelet copies that into a file in each pod, the application copies the file into memory, and a connection pool authenticated with it minutes or hours ago. Each copy updates on its own schedule, and some never update at all. Most rotation outages come from revoking the old credential while some copy still holds it.

This article follows that chain hop by hop, shows how to size the window in which old and new credentials must both work, and covers the tools that automate each hop. For rotation inside one cloud service, see AWS Secrets Manager rotation; for the wider picture of storing secrets, see Secrets management. Here the focus is the last mile: getting a new value into running pods without dropping requests.

Advertisement

Rotation is a propagation problem

Think of a rotation as an event that has to propagate through a pipeline. The moment you issue credential B, only the secret store knows it. The moment you revoke credential A, every consumer still on A starts failing. Between those two moments is the overlap window, and the one rule of safe rotation is that it must be longer than the slowest propagation path to any consumer.

That turns rotation into a measurement problem. You need to know, for each hop, the worst-case delay, and whether the hop updates at all. The diagram shows the common paths.

How a rotated credential travels from the secret store to a live connectionSecret storeVault / cloud managerSync controllerESO: refreshIntervalKubernetes Secretstored in etcdpollupdatekubeletwatch + sync loopwatchPod volume..data symlink swapatomic writeApp processre-reads fileConnection poolnew conns use BCSI driverpolls store, 2m defaultmount refreshEnv varfrozen at startpod start onlyOverlap windowold and new credential both valid until every hop has caught upEvery arrow adds delay; the old credential must outlive the slowest path.
The propagation chain. Delays add up along each path; the environment variable path never refreshes in a running container.

How the kubelet updates a mounted Secret

When a pod mounts a Secret as a volume, the kubelet writes each key as a file. It does this with an atomic writer: the files live in a timestamped directory, a symlink called ..data points at it, and each visible key is a symlink through ..data. On update the kubelet writes a new timestamped directory and swaps the ..data symlink in one rename, so a reader never sees half of an update.

The Kubernetes documentation describes the update as eventually consistent. The kubelet learns about changes through the strategy in its configMapAndSecretChangeDetectionStrategy setting (default Watch; Cache and Get are the alternatives), and the total delay can be as long as the kubelet sync period (syncFrequency, default one minute) plus the cache propagation delay. Measure it on your own clusters rather than assuming it.

Three ways of consuming a Secret do not follow updates at all:

  • Environment variables. They are set when the container starts and never change. A new value needs a new container.
  • subPath mounts. A subPath bind-mounts one file at pod start, so it keeps pointing at the old version after the symlink swap. The documentation states plainly that a container using a Secret as a subPath volume mount does not receive updates.
  • Immutable Secrets. A Secret marked immutable: true cannot be changed; rotating means creating a new Secret with a new name and rolling the pods onto it. That is explicit and versioned, but it is a restart strategy.
spec:
  containers:
    - name: api
      image: registry.example.com/payments-api:4.2.0
      volumeMounts:
        - name: db-creds
          mountPath: /var/run/secrets/db   # whole directory: receives updates
          readOnly: true
        # - mountPath: /etc/app/password   # subPath mount: NEVER receives updates
        #   subPath: password
  volumes:
    - name: db-creds
      secret:
        secretName: payments-db
Advertisement

Choosing a delivery mechanism

Several tools fill the gap between an external store and a pod. They differ in which hops they replace and how each hop is triggered.

MechanismHow the value gets into the podRotation behaviourWatch for
Native Secret, volume mountkubelet writes filesFiles update after kubelet delaySomething must update the Secret itself
External Secrets Operator (ESO)Controller copies store values into a Kubernetes SecretRefreshes on refreshInterval with the default Periodic policy; zero means fetch oncePolling cost against the store's rate limits
Secrets Store CSI driverDriver mounts values from the store directlyOpt-in via --enable-secret-rotation; --rotation-poll-interval defaults to 2mSynced Secrets used as env vars still need a restart
Vault Agent injectorSidecar renders templates into a shared volumeSidecar re-renders as leases and values change; agent-inject-command- can signal the appSidecar resource cost per pod
App reads the store directlySDK call at runtimeImmediate; app owns cachingStore becomes a hard runtime dependency

ESO is the common choice when teams want plain Kubernetes Secrets and GitOps-friendly manifests. Note that the Secret it produces sits in etcd, so envelope encryption at rest matters. A minimal ExternalSecret looks like this; the refreshInterval you choose is the first delay in your budget:

apiVersion: external-secrets.io/v1   # v1beta1 on older ESO releases
kind: ExternalSecret
metadata:
  name: payments-db
  namespace: payments
spec:
  refreshInterval: 15m          # set it explicitly; this is your hop-1 delay
  secretStoreRef:
    kind: ClusterSecretStore
    name: prod-secrets
  target:
    name: payments-db           # the Kubernetes Secret ESO owns
    creationPolicy: Owner
  data:
    - secretKey: password
      remoteRef:
        key: prod/payments/db
        property: password

The dual-credential pattern

The overlap window only exists if the backend accepts two credentials at once. Some systems allow two keys natively; databases usually need a little design. The robust pattern is two login identities that share one set of privileges, rotated alternately:

-- Two login roles share one group role that owns the privileges.
CREATE ROLE payments_rw NOLOGIN;
GRANT SELECT, INSERT, UPDATE, DELETE ON ALL TABLES IN SCHEMA payments TO payments_rw;
CREATE ROLE payments_app_a LOGIN PASSWORD '...' IN ROLE payments_rw;
CREATE ROLE payments_app_b LOGIN PASSWORD '...' IN ROLE payments_rw;

-- Rotation N: the inactive role (b) gets a fresh password, then is published.
ALTER ROLE payments_app_b PASSWORD 'new-random-value';
-- ...publish {user: payments_app_b, password: ...} to the store, wait the overlap...
-- Only after no session uses role a any more:
SELECT count(*) FROM pg_stat_activity WHERE usename = 'payments_app_a';  -- expect 0
ALTER ROLE payments_app_a PASSWORD NULL;  -- or a random value nobody knows

Rotating the inactive identity means the active one is never touched during propagation, so nothing fails while the new value travels. A useful PostgreSQL property helps here: a password is checked only when a session authenticates. Changing or removing it does not disconnect sessions that already logged in, which is why the check against pg_stat_activity is needed before you consider role A retired. The same reasoning applies to any system that authenticates at connect time: revocation only bites on the next connection.

If a system truly allows one credential at a time, you cannot get zero downtime from rotation alone. You either accept a short error spike and retry through it, or put a proxy in front that holds the backend credential while clients rotate theirs.

Making the application pick up the new value

Updating the file is not enough if the process read it once at startup. There are two honest options: reload in process, or restart.

In-process reload. Watch for the ..data swap and re-read. Watch the directory or the symlink target, not the key file: the visible file is a symlink, and some file-watching libraries miss changes that happen by renaming a directory behind it. Polling the symlink target every few seconds is simple and reliable. Feed the new value only into new connections; let the pool retire old ones through its maximum connection lifetime.

import logging, os, threading, time
import psycopg

log = logging.getLogger(__name__)
CRED_DIR = "/var/run/secrets/db"

class Credentials:
    """Re-read the mounted secret when the kubelet swaps the ..data symlink."""
    def __init__(self, path=CRED_DIR, poll=10):
        self.path, self.poll = path, poll
        self._lock = threading.Lock()
        self._stamp, self._password = None, None
        self._load()
        threading.Thread(target=self._watch, daemon=True).start()

    def _stamp_now(self):
        # ..data is the symlink the kubelet replaces atomically on every update
        return os.readlink(os.path.join(self.path, "..data"))

    def _load(self):
        with open(os.path.join(self.path, "password")) as f:
            pw = f.read().strip()
        with self._lock:
            self._stamp, self._password = self._stamp_now(), pw

    def _watch(self):
        while True:
            time.sleep(self.poll)
            try:
                if self._stamp_now() != self._stamp:
                    self._load()
                    log.info("db credential reloaded")   # never log the value
            except OSError as e:
                log.warning("credential reload failed: %s", e)  # keep the old one

    @property
    def password(self):
        with self._lock:
            return self._password

creds = Credentials()

def connect():
    # called by the pool for every NEW connection; existing ones keep working
    return psycopg.connect(host="db", user="payments_app", password=creds.password)

Restart on change. If the app cannot reload, roll the pods when the Secret changes. Stakater Reloader watches Secrets and ConfigMaps and triggers a rolling update of workloads annotated with reloader.stakater.com/auto: "true" (or secret.reloader.stakater.com/reload: "payments-db" to name specific Secrets). Its default strategy adds an environment variable to the pod template; an annotation strategy exists for GitOps setups that would otherwise see drift. For a manual roll, kubectl rollout restart deployment/payments-api does the same. Restarts are slower and need a PodDisruptionBudget and good readiness probes, but they also cover env vars and subPath mounts, which reload cannot.

Worked example: sizing the overlap

A payments API with 40 replicas uses ESO against a cloud secret manager and reloads in process. The pool's maximum connection lifetime is 30 minutes. The worst-case delays are:

HopWorst caseSource
Store to Kubernetes Secret15 minESO refreshInterval
Secret to pod fileabout 2 minkubelet sync plus cache; measured
File to process10 sreload poll interval
Process to every connection30 minpool max lifetime
Totalabout 47 min

Doubling for safety gives an overlap of at least 95 minutes, so the runbook revokes the old identity no sooner than two hours after publishing the new one, and only after the pg_stat_activity count for the old role is zero. If the overlap were unacceptable, the cheapest lever is the first hop: you can force an ExternalSecret to sync right away instead of waiting for the interval; ESO documents a force-sync annotation for this, so check the exact key for your version. The second cheapest is a shorter pool lifetime.

The other rotation: the encryption-at-rest key

Secrets are stored unencrypted in etcd unless the API server encrypts them. With an EncryptionConfiguration, the first provider in the list encrypts new writes and every listed key can decrypt. Rotating that key follows the same overlap logic. Add the new key as the second entry and restart every API server, so all of them can read it. Move it to first and restart again, so new writes use it. Rewrite every Secret so it is re-encrypted, as the Kubernetes documentation suggests with kubectl get secrets --all-namespaces -o json | kubectl replace -f -. Only then remove the old key. Remove it early and the API server can no longer read Secrets written under it. With the KMS v2 provider the key lives in an external KMS and the provider tracks key versions, though retiring an old one still means rewriting Secrets.

# kube-apiserver --encryption-provider-config=/etc/kubernetes/enc.yaml
apiVersion: apiserver.config.k8s.io/v1
kind: EncryptionConfiguration
resources:
  - resources: ["secrets"]
    providers:
      - aescbc:                 # FIRST provider encrypts new writes
          keys:
            - name: key-2026-10
              secret: <base64 32-byte key>
            - name: key-2026-04 # still listed so old data can be decrypted
              secret: <base64 32-byte key>
      - identity: {}

Failure modes

  • Revoking before propagation finishes. The classic outage. Fix: measured overlap, plus a check that no session still uses the old identity.
  • Credentials in environment variables. They silently stay old until the next deploy. Fix: mount files, or wire up restart on change.
  • subPath mounts. A convenient way to place one file in an existing directory, and a guaranteed rotation bug. Fix: mount the directory, or point the app at the mount path.
  • Reload that crashes on a partial read. If the app reads the file while it is missing or empty during a bad sync, it should keep the previous value and alert, not fail.
  • Store rate limits. Hundreds of ExternalSecrets with a short refresh interval can exhaust a cloud API quota, and then nothing refreshes. Fix: longer intervals plus force-sync during planned rotations.

Operating it

  • Track credential age per secret and alert when it exceeds policy; rotation that silently stopped looks exactly like rotation that works.
  • Watch ESO's ExternalSecret status conditions and sync error metrics; for the CSI driver, the SecretProviderClassPodStatus resources show which secret version each pod has mounted.
  • Log a credential version identifier in the application on reload so you can prove which pods have moved.
  • On the backend, record authentications per identity. The old identity's count reaching zero is your real signal that revocation is safe.

Trade-offs

Short credential lifetimes shrink the value of a leak but multiply the number of propagations, and with them the chances of an outage. Dynamic, per-pod credentials (for example from Vault's database secrets engine) remove the shared-password problem completely but add a hard runtime dependency on the secret service. In-process reload avoids restarts but must be written and tested in every application; restart on change works for any app but turns every rotation into a deployment. Most teams do best with files plus in-process reload for their own services, Reloader for third-party software that cannot reload, and dual identities on every backend that allows them. For the cluster basics behind all of this, see Kubernetes introduction.

What to do next

  1. Inventory every Secret consumer and mark how it consumes the value: volume, subPath, env var or SDK.
  2. Move rotated credentials out of env vars and subPath mounts, or add restart on change for those workloads.
  3. Give every backend that allows it two credentials or two identities, and rotate them alternately.
  4. Run a staged rotation, time every hop, and write the measured overlap into the runbook.
  5. Add a revocation gate: no old credential is disabled while the backend still sees sessions or authentications using it.
  6. Alert on credential age and on sync errors from ESO or the CSI driver.
  7. Confirm Secrets are encrypted at rest, and document how the encryption key itself is rotated.
Key takeaway: Secret rotation in Kubernetes is a propagation problem. The new value has to travel from the store, through a sync controller, the Kubernetes Secret and the kubelet, into a file, into the process and finally into every connection. Each hop has its own delay, and env vars and subPath mounts never update at all. Make it safe by letting two credentials be valid at once, measuring each hop so the overlap is longer than the slowest path, reloading in process or restarting on change, and revoking the old credential only when the backend shows nobody is using it.