Most writing about secrets management is aimed at the platform team: which vault to buy, how to audit it, who approves access. A developer meets the problem from the other end. Your service needs a database password, an API key for a payment provider and a signing key for webhooks, and somebody has told you that the value lives in a secret manager. What remains is a set of concrete coding decisions: how the process proves who it is, when it fetches the value, how long it keeps it, what happens when the value changes underneath a running connection pool, and how you run the same code on a laptop without copying production keys into a file.
This article follows that runtime path: the bootstrap problem, a provider interface with a refreshing cache, dynamic database credentials from HashiCorp Vault, local development and testing. The companion piece Secret Management, in depth covers inventory, leak channels and incident response; this one is about the code you write.
Secret zero: authenticate with identity, not a stored key
Every secret store has the same chicken-and-egg question. To read a secret, the application must authenticate. If it authenticates with a token or password, that credential is itself a secret, and it has to be delivered somehow. This is usually called secret zero. Skipping the question leaves a long-lived vault token in an environment variable, which moves the problem without solving it.
The fix is to authenticate with something the platform asserts about the workload rather than something the workload stores. AWS instance profiles and EKS pod identities, Google Cloud service accounts with GKE Workload Identity, and Azure managed identities all hand the SDK short-lived credentials. In Kubernetes generally, the kubelet projects a short-lived service account token into the pod that systems such as Vault can verify.
The property all of these share is that nothing long-lived is written to disk or configuration. The credential is minted on demand, expires within hours or minutes, and is bound to a workload the platform can name. Your code should therefore never accept a secret-store token as configuration in production. It should accept a role name and let the SDK or a login call turn identity into a token. For Vault with the Kubernetes auth method, that login is one HTTP call:
import requests
def vault_login_k8s(vault_addr: str, role: str) -> tuple[str, int]:
# Projected service account token, rotated by the kubelet.
with open("/var/run/secrets/kubernetes.io/serviceaccount/token") as f:
jwt = f.read()
r = requests.post(f"{vault_addr}/v1/auth/kubernetes/login",
json={"role": role, "jwt": jwt}, timeout=5)
r.raise_for_status()
auth = r.json()["auth"]
return auth["client_token"], auth["lease_duration"] # token, secondsThe Vault role binds a namespace and service account to policies, so the token can only read what that workload needs. The function reads the token file on every login rather than caching it at start-up: the kubelet rotates projected tokens, and a copy held in memory for days will eventually be rejected.
One interface, several backends
Application code should not know which store a value came from. Business logic wants a database password, not a Vault response body. Put a narrow interface between the two, and have each backend implement it. Local development then becomes a configuration change, and caching and redaction live in one place.
from dataclasses import dataclass
from typing import Protocol
import time
@dataclass(frozen=True)
class Lease:
values: dict[str, str] # e.g. {"username": ..., "password": ...}
expires_at: float # epoch seconds; float("inf") for static secrets
lease_id: str | None = None
renewable: bool = False
class SecretSource(Protocol):
def fetch(self, name: str) -> Lease: ...
class EnvSource:
"""Local development only: reads APP_SECRET_<NAME>_<FIELD> variables."""
def fetch(self, name: str) -> Lease:
import os
prefix = f"APP_SECRET_{name.upper()}_"
vals = {k[len(prefix):].lower(): v for k, v in os.environ.items()
if k.startswith(prefix)}
if not vals:
raise KeyError(name)
return Lease(vals, float("inf"))
class FileSource:
"""Files mounted by a CSI driver or agent sidecar: /run/secrets/<name>/<field>."""
def __init__(self, root="/run/secrets"):
self.root = root
def fetch(self, name: str) -> Lease:
import pathlib
d = pathlib.Path(self.root, name)
vals = {f.name: f.read_text().strip() for f in d.iterdir() if f.is_file()}
return Lease(vals, time.time() + 300) # re-read every 5 minutesA fetch returns a lease, not a bare string, so code that treats every value as expiring handles rotation without a restart. The file source re-reads periodically instead of once. Mounted secrets are updated in place by the kubelet or an agent, and Secret Rotation in Kubernetes explains the propagation delay you are designing around.
Fetching, caching and refreshing
Calling a secret manager on every request is slow and makes the store a dependency of every request; calling it once at start-up makes rotation impossible. The middle ground is a cache with three behaviours: refresh ahead of expiry, add jitter so a fleet does not refresh in lockstep, and keep serving the last good value for a bounded time if the store is unreachable.
import random, threading, logging
log = logging.getLogger("secrets")
class CachedSecrets:
def __init__(self, source: SecretSource, refresh_fraction=0.67,
stale_grace=600.0):
self.source, self.frac, self.grace = source, refresh_fraction, stale_grace
self._cache: dict[str, tuple[Lease, float]] = {} # name -> (lease, refresh_at)
self._lock = threading.Lock()
def _schedule(self, lease: Lease) -> float:
if lease.expires_at == float("inf"):
return time.time() + 3600 # recheck static secrets hourly
ttl = max(lease.expires_at - time.time(), 1.0)
return time.time() + ttl * self.frac * random.uniform(0.9, 1.0)
def get(self, name: str) -> Lease:
with self._lock:
hit = self._cache.get(name)
if hit and time.time() < hit[1]:
return hit[0]
try:
lease = self.source.fetch(name)
except Exception as exc:
# Store down: serve the old lease until it truly expires plus grace.
if hit and time.time() < hit[0].expires_at + self.grace:
log.warning("secret refresh failed for %s: %s", name, type(exc).__name__)
return hit[0]
raise
self._cache[name] = (lease, self._schedule(lease))
return leaseRefreshing at roughly two thirds of the lease leaves a third of the lifetime to absorb retries. The jitter spreads a fleet's refreshes over a window instead of a spike. The stale grace lets a short store outage pass unnoticed, but only helps while the value is still valid upstream. The warning logs the name and exception type, never the value or a full exception message, which some SDKs fill with request bodies.
Worked example: dynamic database credentials
Static secrets are shared and long-lived. Dynamic secrets invert both properties: the secret store creates a fresh credential for each client when asked, with a lease, and deletes it when the lease ends. Vault's database secrets engine is the common example. An operator configures a connection to the database and a role whose creation statement is a template; each read of database/creds/<role> runs that template with a generated username and password and returns them with a lease.
Worked example. An orders service on Kubernetes needs PostgreSQL access. The operator creates a role with a creation statement along the lines of CREATE ROLE "{{name}}" WITH LOGIN PASSWORD '{{password}}' VALID UNTIL '{{expiration}}'; GRANT orders_rw TO "{{name}}"; and a default TTL of one hour with a max TTL of 24 hours. The service logs in with its service account, reads credentials, and gets back something like this response body:
{
"lease_id": "database/creds/orders-rw/AbCd...",
"lease_duration": 3600,
"renewable": true,
"data": { "username": "v-k8s-orders-rw-XyZ...", "password": "..." }
}A VaultDbSource implementing the interface above turns that into a Lease with expires_at = now + 3600. The cache will fetch a new credential at about 40 minutes. Now the important part: a connection pool opened with the old username keeps working only until the old user is dropped at lease end. The pool must therefore be rebuilt, not merely told the new password:
class RotatingPool:
def __init__(self, secrets: CachedSecrets, name: str, make_pool):
self.secrets, self.name, self.make_pool = secrets, name, make_pool
self._lease, self._pool = None, None
def get(self):
lease = self.secrets.get(self.name)
if lease is not self._lease: # new credential issued
old = self._pool
self._pool = self.make_pool(lease.values["username"],
lease.values["password"])
self._lease = lease
if old is not None:
# close_when_idle: a method on your own pool wrapper, not a library API
old.close_when_idle(max_wait=300) # let in-flight work finish
return self._poolThree numbers must line up: the refresh point (about 40 minutes), the drain time for the old pool (5 minutes here) and the lease end (60 minutes). The drain must finish before the lease ends, with margin. If your longest transaction or streaming query can exceed that margin, lengthen the TTL or shorten the drain. Renewal, through sys/leases/renew, extends a lease up to its max TTL without changing the username; it suits long-lived connections but caps out, so you still need the replacement path.
The payoff is that each pod has its own database user, so audit logs name the pod, a leaked credential dies within the hour, and revoking one workload does not break the others. The cost is user churn on the database and a hard dependency on Vault at refresh time. KMS envelope encryption applies the same short-lived-key principle to data keys.
Local development that mirrors production
Local development is where most real leaks start, because convenience wins. The goal is that a laptop runs exactly the same code path as production with a different SecretSource, and that the values it receives are development credentials that would be harmless if they leaked.
Three patterns work, in order of preference. First, point local runs at the same secret manager using the developer's own single sign-on identity and a development-only path; the provider code is identical to production. Second, use a CLI that injects secrets into a child process's environment at launch, such as the 1Password CLI's op run with a file of secret references, or aws-vault exec for AWS credentials; nothing is written to disk in plain text. Third, keep encrypted configuration in the repository with a tool such as SOPS, which encrypts values but leaves keys readable so diffs stay reviewable, and decrypt at run time with a key tied to the developer's identity.
# .env.refs is committed: it contains references, not values.
APP_SECRET_PAYMENTS_API_KEY=op://dev-vault/payments-sandbox/api-key
APP_SECRET_ORDERSDB_PASSWORD=op://dev-vault/orders-db-local/password
# Launch: the CLI resolves references and passes values to the child only.
op run --env-file=.env.refs -- python -m orders.serviceA plain .env file of real values is the fallback, acceptable only if it holds development credentials, is in .gitignore and is covered by a pre-commit scanner as described in Secrets Scanning Architecture. Environment variables are a weak transport anyway: child processes inherit them and crash reporters capture them, so prefer files or direct API reads in production.
Testing code that reads secrets
Code that reads secrets is easy to test if it depends on the interface rather than on a vendor client. A fake source lets you exercise the paths that matter and that rarely get tested: expiry, refresh, store outage and rotation.
class FakeSource:
def __init__(self):
self.version, self.fail = 0, False
def fetch(self, name):
if self.fail:
raise ConnectionError("store down")
self.version += 1
return Lease({"password": f"pw{self.version}"}, time.time() + 3)
def test_serves_stale_during_outage_then_rotates():
src = FakeSource()
cache = CachedSecrets(src, refresh_fraction=0.5, stale_grace=10)
assert cache.get("db").values["password"] == "pw1"
time.sleep(2) # past the refresh point
src.fail = True
assert cache.get("db").values["password"] == "pw1" # stale but valid
src.fail = False
time.sleep(1)
assert cache.get("db").values["password"] == "pw2" # rotatedAdd one more test that formats every log line produced during these scenarios and asserts that no secret value appears in it. That single assertion catches the most common regression, which is a new debug log or exception message that includes the response body.
Failure modes
- Fetch at import time. A module-level
PASSWORD = client.get(...)runs once, before logging and retries are configured, and can never rotate. Fetch lazily through the cache. - Thundering refresh. Hundreds of pods started together refresh together and hit the store's rate limit; some get errors exactly when their old lease is ending. Jitter and the stale grace prevent it.
- Pool outlives the lease. Connections opened with a dynamic credential fail with authentication errors at lease end, often in the middle of a batch job. Rebuild the pool before expiry and keep the drain shorter than the remaining lease.
- Token file cached forever. A projected service account token read once at start-up is rejected after it rotates, and the next Vault login fails. Re-read the file on every login.
- Secrets in error paths. SDK exceptions, HTTP client debug logs and error trackers that capture local variables leak values. Log names and exception types, and scrub event payloads.
Trade-offs
| Approach | Strength | Cost |
|---|---|---|
| Static secret, read once at start | Simplest code | Rotation needs a restart; long-lived leak window |
| Static secret, cached with refresh | Rotation without restart | Must handle two valid versions during rotation |
| Dynamic credentials with leases | Per-workload identity, short leak window, clean revocation | Store becomes a runtime dependency; pool rotation logic; DB user churn |
| Platform identity, no secret at all | Nothing to leak or rotate | Only works where the target accepts the platform's tokens |
What to do next
- List every secret your service reads and, for each, write down its source, its lifetime and whether the target could accept platform identity instead.
- Remove any secret-store token from configuration; authenticate with workload identity and a role name.
- Introduce a
SecretSourceinterface and route every read through a cache with refresh-ahead, jitter and bounded stale serving. - Move one database to dynamic credentials in a staging environment and test pool rotation with a deliberately short TTL, such as five minutes.
- Replace committed or shared
.envfiles with secret references resolved at launch, and use development-only credentials on laptops. - Add tests for expiry, outage and rotation, plus an assertion that no log line contains a secret value.
- Read Secret Management, in depth for the inventory and incident runbook that sit around this code.