Most teams that suffer a credential leak already own a secret manager. The password was in the vault; it also ended up in a debug log, a stack trace shipped to an error tracker, a container layer, a CI log or a laptop's shell history. Storage is the solved half of secret management. The hard half is everything that happens after a value leaves the store: how your code holds it, where it can escape, who can still use it after it does, and how fast you can make it worthless.
This article is about that second half. The platform architecture, meaning workload identity, dynamic credentials, caching and versioning, is covered in cloud secrets architecture, rotation mechanics in secret rotation in Kubernetes, and detection tooling in secrets scanning. Here we start from first principles, build an inventory, harden the code path, close the leak channels, and finish with an incident runbook and a checklist.
A secret is any value whose possession is sufficient to act as someone. The definition is about capability, not format: API keys, database passwords, refresh tokens, TLS private keys, webhook signing keys and session cookies are all secrets because holding the bytes is enough. Short-lived does not mean safe; a one-hour token is plenty of time to exfiltrate a table.
Three properties decide how much damage a leaked secret can do, and they are the properties you engineer:
Everything below either keeps a secret from escaping or shrinks scope, lifetime and time-to-revoke when one does. Assume something will escape and design so that the escape is boring.
You cannot protect secrets you have not listed. An inventory is a table with one row per credential, not per system, and it should answer five questions for every row: who owns it, what it can do, where it is consumed, how long it lives, and exactly how to revoke it. The revoke column is the one teams skip and the one that matters during an incident, when nobody remembers which console page or API call kills a particular key.
| Credential | Owner | Scope | Lifetime | Revoke path | Risk |
|---|---|---|---|---|---|
| orders DB password | payments team | read/write orders schema | static, rotated quarterly | rotate in store; restart 3 services | high: shared by 3 services |
| payment provider API key | payments team | charge and refund | static | provider dashboard, then redeploy | critical: money movement |
| CI deploy credential | platform team | push images, deploy prod | federated, 1 hour | remove trust policy | medium: short-lived |
| webhook signing key | integrations | forge inbound events | static | dual-key overlap, then retire old | high: forgery is silent |
| service-to-service token | platform team | one audience | minted per call, 5 min | disable issuer role | low |
Rows that combine wide scope, long life and a slow revoke path are your priorities. The orders password is shared, so the first fix is one credential per consumer, turning one painful rotation into three independent ones. The provider key moves money and never expires, so it earns a restricted key if the provider offers one and an alert on unusual refund volume. The CI credential is federated and short-lived, which is what good looks like. Start from your store's listing, then grep configuration and infrastructure code for credentials that are not in the store; those strays are usually the oldest and widest.
Inside the process, the most effective control is boring: give secrets their own type. A plain string is printed by every logger, serialised by every JSON encoder and captured by every error tracker that records local variables. A wrapper type that refuses to reveal itself makes the safe path the default and the unsafe path visible in code review:
class Secret:
"""Holds a credential; never reveals it through str, repr, format or pickle."""
__slots__ = ("_value",)
def __init__(self, value: str):
self._value = value
def reveal(self) -> str: # grep-able: every call site is reviewed
return self._value
def __repr__(self) -> str:
return "Secret(****)"
__str__ = __repr__
def __format__(self, spec) -> str:
return "****"
def __reduce__(self): # pickling or deep-copy to disk is a bug
raise TypeError("Secret values cannot be pickled")
def __eq__(self, other):
import hmac
return isinstance(other, Secret) and hmac.compare_digest(
self._value.encode(), other._value.encode())
__hash__ = None
db_password = Secret(read_mounted_file("/run/secrets/orders-db"))
log.info("connecting with %s", db_password) # logs: connecting with Secret(****)
conn = connect(host, user, db_password.reveal()) # the one reviewed call siteLibraries offer the same idea, for example pydantic's SecretStr with get_secret_value(). A search for reveal( lists every place the raw bytes exist, and that list should stay short. Comparing secrets, for example an inbound webhook signature, must use a constant-time comparison such as hmac.compare_digest so response timing does not leak how many leading bytes matched.
A wrapper does not wipe memory. Python strings are immutable and the JVM moves objects during collection, which is why Java APIs take passwords as char[] that callers can overwrite, a mitigation rather than a guarantee. Treat process memory as readable by anyone who can attach a debugger or read a core dump: disable core dumps for services holding high-value keys, restrict ptrace, and never expose debug endpoints that dump the environment or heap.
Once the type is in place, close the side channels one by one. Each has a characteristic bug and a characteristic fix.
CI systems often hold the most powerful credentials in an organisation, because they deploy to production, and they execute code that anyone with a pull request can change. Three rules cover most of the risk.
First, replace stored long-lived cloud keys with federation. Major CI providers can issue each job a signed identity token, which the cloud provider exchanges for short-lived credentials if its claims (repository, branch, environment) match a trust policy you wrote. Nothing long-lived sits in CI settings. Claim names vary by provider; the key decision is to pin the policy to specific repositories and branches, never a whole organisation.
Second, separate untrusted code from secrets. A workflow that builds pull requests from forks must run without deployment credentials; most CI systems withhold secrets from fork-triggered runs by default, and the dangerous configurations are the ones that deliberately run fork code in a privileged context. Treat any workflow that checks out contributor code and also has secrets or write tokens as a privilege escalation path, and review it like one.
Third, do not rely on log masking. CI systems replace registered secret values with asterisks, but masking matches exact strings: base64-encode the value or let a tool echo a derived URL and the mask misses it. The real controls are short lifetimes, narrow scope and never echoing credentials, including through shell tracing modes that print expanded commands.
On developer machines, credentials should be personal, scoped to development resources and issued by single sign-on rather than shared in chat. Keep .env files out of version control with an ignore rule and a pre-commit scanner; configuration that must live in a repository should be encrypted to a KMS key or public-key recipients, so decryption requires an identity you can revoke. The key hierarchy behind that is explained in KMS envelope encryption.
When a secret leaks, order matters more than speed at any single step. The common mistake is to start by deleting the commit or log line. That is not containment: the value has already been copied by whatever saw it, including scrapers that watch public repositories continuously.
Write the runbook per credential class before you need it: which store path, which services to restart, and which check confirms the old value is rejected.
An illustrative case. A developer debugging a payment timeout enables verbose HTTP logging in staging. The configuration loader falls back to the production provider key when the staging variable is missing, and that variable was renamed last month. Verbose logging prints the authorization header into a log platform readable by 300 engineers.
Three days later a log scanner flags the provider's live-key prefix. Following the runbook, on-call deploys a new restricted key with the documented dual-key overlap, then revokes the old key, eleven minutes after the alert. The provider's request log for the exposure window shows only the payments service's own address. The log platform's index is purged for the affected stream.
Cause analysis finds four independent failures, and each gets a fix: the fallback to a production key (configuration now fails closed when a variable is missing), the plain-string credential (now a secret type, so even verbose logging prints a mask), the missing header filter on the log pipeline (authorization headers are now masked at ingestion), and a staging environment that could reach production credentials at all (separate store paths and separate identities, so staging cannot read production values). Any one fix would have prevented the incident.
Short lifetimes cut the value of a leak but add a token service, refresh logic and a failure mode when the issuer is down, so cache tokens until shortly before expiry and alarm on refresh failures. Mounted files are harder to leak than environment variables but need reload-on-change; fetching at startup keeps secrets off disk but makes the store a startup dependency. Per-consumer credentials multiply the values to manage but make every revocation surgical. In almost every case, pay for narrower scope and faster revocation first, because those reduce the cost of every future mistake rather than preventing one specific one.
reveal() call sites a review checkpoint.