An ephemeral credential is one that is minted on demand, scoped to a task, and expires on its own within minutes or hours. Cloud session tokens from a security token service, database users created per lease by a secrets broker, and Kubernetes projected service-account tokens are all examples. The opposite is a static secret: an API key or password created once and kept until someone remembers to rotate it.

LLM workloads make the difference matter. Model servers pull weights from object storage, retrieval pipelines read databases, fine-tuning jobs write checkpoints for a day or more, and agents call tools on behalf of users. Every one of those needs credentials, and every one sits next to logs, traces and context windows that leak. A short lifetime turns a leak from "rotate everything and hope" into "it expires in fifteen minutes".

This article is about the lifecycle mechanics: choosing a TTL, refreshing without outages, jobs that outlive their tokens, revocation, and the failure modes of each. How a workload proves its identity in the first place is covered in Workload Identity for LLM Services, in depth; brokering credentials for agents is covered in Secret Management for Agents, in depth.

What a short lifetime buys

The value of a short lifetime is easy to state as arithmetic. If a credential is valid for T and leaks at a random moment, the attacker gets on average T/2 of use and at worst T. A static key leaked into a log that nobody reads is valid until rotation, which in many organisations means months.

Lifetime is only one of four properties that bound damage, and it is worth keeping them separate:

  • Lifetime: how long a leaked credential works.
  • Scope: what it can do while it works. A one-hour token with admin rights is worse than a one-day token that can only read one bucket prefix.
  • Binding: whether it can be used from anywhere. Some systems can bind a token to a client certificate or a proof-of-possession key, so a copied bearer string is useless alone.
  • Revocability: whether you can kill it before it expires.
An ephemeral credential's life: issue, use, refresh ahead, expireWorkloadpod / job / agent runIssuerSTS, Vault, token APIResourcebucket, DB, model API1. prove identity2. credential + expiry3. use (scoped)issueduse the cached credentialrefresh windowskew + gracerefresh at ~2/3 TTLexpiryExposure window if leaked = time remaining until expiry (or until revoked)TTL bounds the damage of a leak; it does not replace revocation, scope or auditShort TTL makes refresh a hot path: an issuer outage becomes your outage unless you plan for it.
Figure: the workload proves identity, receives a credential with an expiry, uses it, and refreshes ahead of expiry. The exposure window is what remains of the TTL.

Ephemeral credentials also remove a whole class of operational pain. There is no rotation project, because every credential rotates itself. There is no long-lived secret to store in a vault, a CI variable or a container image. The secret that remains is the bootstrap identity, and on modern platforms that is itself a short-lived, platform-issued token.

Issuers and their limits

Lifetimes and limits differ by issuer, and the limits are hard: ask for more and the call fails rather than silently shortening. These were checked against vendor documentation on 2026-10-05; confirm them for your environment.

IssuerTypical lifetimeLimits worth knowing
AWS STS AssumeRole1 hour by defaultDurationSeconds from 900 s up to the role's maximum session duration (1 to 12 hours). Role chaining caps the session at 1 hour, and asking for more fails.
Google Cloud service-account access tokens1 hourLonger lifetimes up to 12 hours need an organisation policy that allows lifetime extension.
Kubernetes projected service-account tokenSet by expirationSeconds, default 1 hourMinimum 10 minutes. The kubelet refreshes the file when it is over 80% of its TTL or older than 24 hours; your code must re-read it.
HashiCorp Vault dynamic secretsPer role default_ttl / max_ttlEach credential is a lease; it can be renewed up to max_ttl and revoked early by lease ID or prefix.

Vault's database engine is a good illustration of a truly dynamic credential: the broker creates a real database user for each lease and drops it when the lease ends.

# operator: define a role that mints read-only users for the RAG service
vault write database/roles/rag-reader \
    db_name=vectors \
    creation_statements="CREATE ROLE \"{{name}}\" WITH LOGIN PASSWORD '{{password}}' VALID UNTIL '{{expiration}}';
                         GRANT SELECT ON ALL TABLES IN SCHEMA embeddings TO \"{{name}}\";" \
    default_ttl=1h max_ttl=8h

# workload: obtain a lease
vault read database/creds/rag-reader
#   lease_id        database/creds/rag-reader/AbC123...
#   lease_duration  1h
#   lease_renewable true
#   password        ...
#   username        v-rag-reader-...

vault lease renew database/creds/rag-reader/AbC123...        # extend, up to max_ttl
vault lease revoke -prefix database/creds/rag-reader          # incident: kill every lease

Notice the VALID UNTIL clause: the database itself enforces expiry, so even if the broker is down when the lease ends, the password stops working.

Choosing a TTL

Picking a TTL is a balance. Too long and the exposure window grows; too short and refresh traffic, latency and issuer dependency grow. A useful lower bound is:

TTL  >=  issue_latency + longest_single_operation + clock_skew + refresh_margin

The longest single operation is the one that uses a credential once at the start and cannot re-authenticate in the middle. For LLM systems this is where surprises live:

  • Weight downloads. A 140 GB checkpoint pulled from object storage at a few hundred MB per second takes several minutes. Most SDKs sign each request or part separately, so they pick up refreshed credentials between parts, but a single long presigned URL does not. A presigned URL signed with temporary credentials stops working when those credentials expire, even if the URL's own expiry is later.
  • Long streams. A streaming completion or a long-lived database cursor authenticated once at connect time usually survives token expiry, because the check happened at connect. Do not rely on that as a security property, and do not rely on it surviving a reconnect.
  • Training and batch jobs. A fine-tuning run that checkpoints to storage every hour for thirty hours will outlive any reasonable token. The job must refresh, not ask for a thirty-hour credential.
  • Agent runs. An agent that pauses for human approval may resume after its token expired. Treat resumption as a fresh credential request, re-checking authorisation.

Refreshing without outages

Most refresh bugs come from hand-rolled caching. The rules are: refresh ahead of expiry, add jitter so a fleet does not stampede the issuer, let only one caller refresh at a time, and keep using a still-valid credential if a refresh fails. The class below implements those rules around any fetch function that returns a secret and an absolute expiry.

import random, threading, time
from dataclasses import dataclass

@dataclass
class Cred:
    secret: object
    expires_at: float          # epoch seconds, from the issuer's response

class RefreshingCredential:
    def __init__(self, fetch, refresh_fraction=0.66, skew=30.0, jitter=0.1):
        self._fetch = fetch              # () -> Cred; must raise on failure
        self._frac = refresh_fraction
        self._skew = skew                # treat expiry as this many seconds earlier
        self._jitter = jitter
        self._lock = threading.Lock()
        self._cred = None
        self._refresh_at = 0.0

    def _schedule(self, cred):
        now = time.time()
        life = max(0.0, cred.expires_at - self._skew - now)
        j = 1.0 - random.uniform(0, self._jitter)
        self._refresh_at = now + life * self._frac * j

    def get(self):
        now = time.time()
        cred = self._cred
        if cred and now < self._refresh_at:
            return cred.secret                       # fast path, no lock
        with self._lock:                             # single flight
            cred = self._cred
            if cred and time.time() < self._refresh_at:
                return cred.secret
            try:
                new = self._fetch()
                self._cred = new
                self._schedule(new)
                return new.secret
            except Exception:
                if cred and time.time() < cred.expires_at - self._skew:
                    self._refresh_at = time.time() + 15   # retry soon, serve stale-but-valid
                    return cred.secret
                raise                                # truly expired: fail closed

Two details matter. The expiry comes from the issuer's response, not from the TTL you asked for, because issuers may shorten it. And the failure path serves the old credential only while it is still valid; once it has expired the call fails closed rather than retrying with a dead token in a loop. For a Kubernetes projected token the fetch function is just "re-read the file and parse its exp claim".

Rules specific to LLM systems

Four rules are specific to LLM systems.

  1. Credentials never enter the context window. The model asks for an action; a tool executor outside the model attaches the credential. If a token appears in a prompt, a tool result or a trace, short lifetime limits the damage but the design is still wrong. See Capability Tokens, in depth for scoping each tool call more tightly than a cloud role can.
  2. One credential per run, not per service. An agent run or a batch job should get its own session with a session name or tag that identifies the run. Audit logs then answer "which run did this", and you can revoke one run without touching the others.
  3. Narrow at issue time. STS accepts an inline session policy that can only reduce the role's permissions. A retrieval run for one tenant can be issued a session limited to that tenant's prefix, even though the role could read every tenant.
  4. Scrub before you log. Error messages from SDKs sometimes include signed URLs or headers. Redact known token shapes at the logger, because a credential that expires in an hour still works for that hour.

Revocation is a separate feature

Short lifetime is not revocation. When you detect a leak you want the credential dead now, and how you do that depends on the issuer.

  • Vault leases: revoke by lease ID or prefix. Because the database user is dropped, this is real, immediate revocation.
  • AWS role sessions: issued session tokens cannot be individually deleted. The console action to revoke active sessions attaches a deny policy to the role conditioned on aws:TokenIssueTime, so every session issued before that moment is refused. Legitimate workloads must then re-assume the role, which they will do automatically if they refresh correctly.
  • Self-contained JWTs: a signed token is valid until its exp unless the resource checks something else, such as an introspection endpoint, a deny list of token IDs, or a rotated signing key. If you need early revocation for JWTs, design it in; you cannot add it during the incident.

Rehearse it. A quarterly drill that revokes a test workload's sessions and confirms both that the attacker's copy fails and that the real workload recovers within a minute is the only reliable proof that your refresh logic works. Fold it into LLM Incident Response, in depth.

Failure modes

FailureSymptomFix
Fleet refreshes at the same instantIssuer throttling, latency spikes every TTLJittered refresh, as in the class above
Clock skewFresh tokens rejected as not yet valid or already expiredNTP everywhere; a skew margin in the cache
Issuer outageEvery workload fails once cached tokens expireServe stale-but-valid; size TTL to cover a plausible outage
Requested TTL above limitHard error at startup, e.g. a chained role asked for 2 hoursRead limits from config; alert on issue errors
Token file read oncePod works for an hour, then 401sRe-read projected token files; never cache at import
Long presigned URLDownload fails part-way after credential expirySign with longer-lived identity or download in parts
Token in logs or promptsValid credential visible to anyone with log accessRedaction at the logger; executor-side injection

Worked example: a RAG service off static secrets

Worked example. A RAG service reads a Postgres vector store and a document bucket. Before: a static database password and an access key, both in a Kubernetes secret, both two years old. After:

  1. The pod's projected service-account token, with a 1-hour expiry, authenticates it to Vault and to the cloud provider.
  2. Vault issues a rag-reader database lease with default_ttl=1h; the service renews it at about 40 minutes and gets a fresh lease when max_ttl is reached.
  3. The cloud SDK assumes a read-only role for 1 hour with a session name of the pod name. A per-tenant batch job adds an inline session policy for its tenant's prefix.
  4. Both credentials sit behind the refresh class; dashboards track issue latency, issue errors and seconds-to-expiry of the oldest credential in use.
  5. The incident runbook names the two revocation commands and was rehearsed once.

The worst-case leak window fell from "until someone notices" to one hour, and the only long-lived secret left is the platform's own signing key, which the team never handles. For the rest of the static-secret inventory, see Secrets Management for LLM Apps, in depth.

What to do next

  1. Inventory every credential your LLM services use and mark each as static or ephemeral.
  2. For each static one, name the issuer that could replace it and the identity that would authenticate to that issuer.
  3. Measure your longest single operation per workload, and set TTLs from the formula above rather than by habit.
  4. Wrap every credential in one shared refresh helper with jitter, single flight and fail-closed expiry.
  5. Give each agent run or batch job its own session name and, where possible, a narrowed session policy.
  6. Add redaction for token shapes at the logger and confirm no tool result is passed to the model with a credential in it.
  7. Write and rehearse the revocation runbook for each issuer, including recovery of legitimate workloads.
Key takeaway: Ephemeral credentials bound a leak by time, but lifetime is one of four controls alongside scope, binding and revocability. Know each issuer's hard limits, set TTLs from your longest single operation, refresh ahead with jitter and fail closed at expiry, give each run its own narrowed session, keep tokens out of prompts and logs, and rehearse revocation before you need it.