Secret Management, in depth: inventory, secret types in code, leak channels, CI pipelines and a revoke-first incident runbook

By Sandeep Belgavi · 2026-10-03 · Category: Security
Advertisement
Secret storeversioned, auditedDeliveryfile, SDK or sidecarProcess memorySecret[T] wrapperUseTLS, DB, API callLogsformat stringsExceptionsrepr in tracesTelemetryspan attributesDumps, URLscore, query stringBuild outputimages, CI logsSource and historycommits, forks, laptopsCI pipelinerunners, caches, artefactsInventoryowner, scope, TTL, revoke pathThe store is rarely the weak point. Secrets leak sideways, out of the process and the delivery path.
Where a secret travels after it leaves the store, and the side channels it leaks through. The platform half (store, identity, rotation) is solved infrastructure; the red row is where most real incidents start.

Most teams that suffer a credential leak already own a secret manager. The password was in the vault; it also ended up in a debug log, a stack trace shipped to an error tracker, a container layer, a CI log or a laptop's shell history. Storage is the solved half of secret management. The hard half is everything that happens after a value leaves the store: how your code holds it, where it can escape, who can still use it after it does, and how fast you can make it worthless.

This article is about that second half. The platform architecture, meaning workload identity, dynamic credentials, caching and versioning, is covered in cloud secrets architecture, rotation mechanics in secret rotation in Kubernetes, and detection tooling in secrets scanning. Here we start from first principles, build an inventory, harden the code path, close the leak channels, and finish with an incident runbook and a checklist.

What makes a value a secret

A secret is any value whose possession is sufficient to act as someone. The definition is about capability, not format: API keys, database passwords, refresh tokens, TLS private keys, webhook signing keys and session cookies are all secrets because holding the bytes is enough. Short-lived does not mean safe; a one-hour token is plenty of time to exfiltrate a table.

Three properties decide how much damage a leaked secret can do, and they are the properties you engineer:

Everything below either keeps a secret from escaping or shrinks scope, lifetime and time-to-revoke when one does. Assume something will escape and design so that the escape is boring.

Advertisement

Start with an inventory

You cannot protect secrets you have not listed. An inventory is a table with one row per credential, not per system, and it should answer five questions for every row: who owns it, what it can do, where it is consumed, how long it lives, and exactly how to revoke it. The revoke column is the one teams skip and the one that matters during an incident, when nobody remembers which console page or API call kills a particular key.

CredentialOwnerScopeLifetimeRevoke pathRisk
orders DB passwordpayments teamread/write orders schemastatic, rotated quarterlyrotate in store; restart 3 serviceshigh: shared by 3 services
payment provider API keypayments teamcharge and refundstaticprovider dashboard, then redeploycritical: money movement
CI deploy credentialplatform teampush images, deploy prodfederated, 1 hourremove trust policymedium: short-lived
webhook signing keyintegrationsforge inbound eventsstaticdual-key overlap, then retire oldhigh: forgery is silent
service-to-service tokenplatform teamone audienceminted per call, 5 mindisable issuer rolelow

Rows that combine wide scope, long life and a slow revoke path are your priorities. The orders password is shared, so the first fix is one credential per consumer, turning one painful rotation into three independent ones. The provider key moves money and never expires, so it earns a restricted key if the provider offers one and an alert on unusual refund volume. The CI credential is federated and short-lived, which is what good looks like. Start from your store's listing, then grep configuration and infrastructure code for credentials that are not in the store; those strays are usually the oldest and widest.

Holding secrets in code

Inside the process, the most effective control is boring: give secrets their own type. A plain string is printed by every logger, serialised by every JSON encoder and captured by every error tracker that records local variables. A wrapper type that refuses to reveal itself makes the safe path the default and the unsafe path visible in code review:

class Secret:
    """Holds a credential; never reveals it through str, repr, format or pickle."""
    __slots__ = ("_value",)

    def __init__(self, value: str):
        self._value = value

    def reveal(self) -> str:          # grep-able: every call site is reviewed
        return self._value

    def __repr__(self) -> str:
        return "Secret(****)"
    __str__ = __repr__

    def __format__(self, spec) -> str:
        return "****"

    def __reduce__(self):             # pickling or deep-copy to disk is a bug
        raise TypeError("Secret values cannot be pickled")

    def __eq__(self, other):
        import hmac
        return isinstance(other, Secret) and hmac.compare_digest(
            self._value.encode(), other._value.encode())
    __hash__ = None

db_password = Secret(read_mounted_file("/run/secrets/orders-db"))
log.info("connecting with %s", db_password)      # logs: connecting with Secret(****)
conn = connect(host, user, db_password.reveal())  # the one reviewed call site

Libraries offer the same idea, for example pydantic's SecretStr with get_secret_value(). A search for reveal( lists every place the raw bytes exist, and that list should stay short. Comparing secrets, for example an inbound webhook signature, must use a constant-time comparison such as hmac.compare_digest so response timing does not leak how many leading bytes matched.

A wrapper does not wipe memory. Python strings are immutable and the JVM moves objects during collection, which is why Java APIs take passwords as char[] that callers can overwrite, a mitigation rather than a guarantee. Treat process memory as readable by anyone who can attach a debugger or read a core dump: disable core dumps for services holding high-value keys, restrict ptrace, and never expose debug endpoints that dump the environment or heap.

Closing the leak channels

Once the type is in place, close the side channels one by one. Each has a characteristic bug and a characteristic fix.

CI pipelines and developer machines

CI systems often hold the most powerful credentials in an organisation, because they deploy to production, and they execute code that anyone with a pull request can change. Three rules cover most of the risk.

First, replace stored long-lived cloud keys with federation. Major CI providers can issue each job a signed identity token, which the cloud provider exchanges for short-lived credentials if its claims (repository, branch, environment) match a trust policy you wrote. Nothing long-lived sits in CI settings. Claim names vary by provider; the key decision is to pin the policy to specific repositories and branches, never a whole organisation.

Second, separate untrusted code from secrets. A workflow that builds pull requests from forks must run without deployment credentials; most CI systems withhold secrets from fork-triggered runs by default, and the dangerous configurations are the ones that deliberately run fork code in a privileged context. Treat any workflow that checks out contributor code and also has secrets or write tokens as a privilege escalation path, and review it like one.

Third, do not rely on log masking. CI systems replace registered secret values with asterisks, but masking matches exact strings: base64-encode the value or let a tool echo a derived URL and the mask misses it. The real controls are short lifetimes, narrow scope and never echoing credentials, including through shell tracing modes that print expanded commands.

On developer machines, credentials should be personal, scoped to development resources and issued by single sign-on rather than shared in chat. Keep .env files out of version control with an ignore rule and a pre-commit scanner; configuration that must live in a repository should be encrypted to a KMS key or public-key recipients, so decryption requires an identity you can revoke. The key hierarchy behind that is explained in KMS envelope encryption.

When a secret leaks: revoke first

When a secret leaks, order matters more than speed at any single step. The common mistake is to start by deleting the commit or log line. That is not containment: the value has already been copied by whatever saw it, including scrapers that watch public repositories continuously.

  1. Revoke or rotate first. Make the leaked value worthless. If revocation will cause an outage because the credential is shared, accept a short, planned outage for a critical credential rather than leaving it live; this is exactly the trade the inventory's revoke column prepares you for.
  2. Scope the exposure. Establish when the value was first exposed, where (public or internal), and what it could do. The time window and the scope bound what you must investigate.
  3. Review usage. Pull the provider's or store's audit logs for the credential across the exposure window. Look for unfamiliar source addresses, unusual operations and volume spikes. No suspicious use is a finding to record, not an assumption.
  4. Clean up. Now remove the value from history, logs, tickets and images, and purge caches and mirrors where you can. Accept that copies outside your control remain; step one is what made them harmless.
  5. Fix the cause. Ask why the value could leak and why it was valuable: add the missing type, filter or scanner rule, and shrink scope or lifetime so the next leak matters less.

Write the runbook per credential class before you need it: which store path, which services to restart, and which check confirms the old value is rejected.

Worked example: a live key in staging logs

An illustrative case. A developer debugging a payment timeout enables verbose HTTP logging in staging. The configuration loader falls back to the production provider key when the staging variable is missing, and that variable was renamed last month. Verbose logging prints the authorization header into a log platform readable by 300 engineers.

Three days later a log scanner flags the provider's live-key prefix. Following the runbook, on-call deploys a new restricted key with the documented dual-key overlap, then revokes the old key, eleven minutes after the alert. The provider's request log for the exposure window shows only the payments service's own address. The log platform's index is purged for the affected stream.

Cause analysis finds four independent failures, and each gets a fix: the fallback to a production key (configuration now fails closed when a variable is missing), the plain-string credential (now a secret type, so even verbose logging prints a mask), the missing header filter on the log pipeline (authorization headers are now masked at ingestion), and a staging environment that could reach production credentials at all (separate store paths and separate identities, so staging cannot read production values). Any one fix would have prevented the incident.

Failure modes

Trade-offs

Short lifetimes cut the value of a leak but add a token service, refresh logic and a failure mode when the issuer is down, so cache tokens until shortly before expiry and alarm on refresh failures. Mounted files are harder to leak than environment variables but need reload-on-change; fetching at startup keeps secrets off disk but makes the store a startup dependency. Per-consumer credentials multiply the values to manage but make every revocation surgical. In almost every case, pay for narrower scope and faster revocation first, because those reduce the cost of every future mistake rather than preventing one specific one.

What to do next

  1. Build the inventory: one row per credential with owner, scope, lifetime, consumers and an exact revoke path; fix the wide, long-lived, slow-to-revoke rows first.
  2. Introduce a secret type in each service language and make reveal() call sites a review checkpoint.
  3. Add a header and key-pattern redaction filter to your log pipeline and an allow-list of recorded headers to your tracer.
  4. Make configuration fail closed when a secret is missing, and give each environment its own store path and identity.
  5. Replace long-lived cloud keys in CI with federated short-lived credentials pinned to repository and branch; strip secrets from fork-triggered runs.
  6. Use build secret mounts, scan images as well as repositories, and disable core dumps on services holding high-value keys.
  7. Write a revoke-first runbook per credential class and rehearse it once a quarter against a real, low-risk key.
Key takeaway: Storing secrets in a vault is the easy half. Most leaks happen after the value leaves the store, through logs, exceptions, telemetry, environment variables, images and CI. Inventory every credential with an owner and a revoke path, give secrets their own type in code, redact at the pipeline as a backstop, use federated short-lived credentials in CI, and when something leaks, revoke first and clean up second.