An ephemeral credential is one that is minted on demand, scoped to a task, and expires on its own within minutes or hours. Cloud session tokens from a security token service, database users created per lease by a secrets broker, and Kubernetes projected service-account tokens are all examples. The opposite is a static secret: an API key or password created once and kept until someone remembers to rotate it.
LLM workloads make the difference matter. Model servers pull weights from object storage, retrieval pipelines read databases, fine-tuning jobs write checkpoints for a day or more, and agents call tools on behalf of users. Every one of those needs credentials, and every one sits next to logs, traces and context windows that leak. A short lifetime turns a leak from "rotate everything and hope" into "it expires in fifteen minutes".
This article is about the lifecycle mechanics: choosing a TTL, refreshing without outages, jobs that outlive their tokens, revocation, and the failure modes of each. How a workload proves its identity in the first place is covered in Workload Identity for LLM Services, in depth; brokering credentials for agents is covered in Secret Management for Agents, in depth.
What a short lifetime buys
The value of a short lifetime is easy to state as arithmetic. If a credential is valid for T and leaks at a random moment, the attacker gets on average T/2 of use and at worst T. A static key leaked into a log that nobody reads is valid until rotation, which in many organisations means months.
Lifetime is only one of four properties that bound damage, and it is worth keeping them separate:
- Lifetime: how long a leaked credential works.
- Scope: what it can do while it works. A one-hour token with admin rights is worse than a one-day token that can only read one bucket prefix.
- Binding: whether it can be used from anywhere. Some systems can bind a token to a client certificate or a proof-of-possession key, so a copied bearer string is useless alone.
- Revocability: whether you can kill it before it expires.
Ephemeral credentials also remove a whole class of operational pain. There is no rotation project, because every credential rotates itself. There is no long-lived secret to store in a vault, a CI variable or a container image. The secret that remains is the bootstrap identity, and on modern platforms that is itself a short-lived, platform-issued token.
Issuers and their limits
Lifetimes and limits differ by issuer, and the limits are hard: ask for more and the call fails rather than silently shortening. These were checked against vendor documentation on 2026-10-05; confirm them for your environment.
| Issuer | Typical lifetime | Limits worth knowing |
|---|---|---|
AWS STS AssumeRole | 1 hour by default | DurationSeconds from 900 s up to the role's maximum session duration (1 to 12 hours). Role chaining caps the session at 1 hour, and asking for more fails. |
| Google Cloud service-account access tokens | 1 hour | Longer lifetimes up to 12 hours need an organisation policy that allows lifetime extension. |
| Kubernetes projected service-account token | Set by expirationSeconds, default 1 hour | Minimum 10 minutes. The kubelet refreshes the file when it is over 80% of its TTL or older than 24 hours; your code must re-read it. |
| HashiCorp Vault dynamic secrets | Per role default_ttl / max_ttl | Each credential is a lease; it can be renewed up to max_ttl and revoked early by lease ID or prefix. |
Vault's database engine is a good illustration of a truly dynamic credential: the broker creates a real database user for each lease and drops it when the lease ends.
# operator: define a role that mints read-only users for the RAG service
vault write database/roles/rag-reader \
db_name=vectors \
creation_statements="CREATE ROLE \"{{name}}\" WITH LOGIN PASSWORD '{{password}}' VALID UNTIL '{{expiration}}';
GRANT SELECT ON ALL TABLES IN SCHEMA embeddings TO \"{{name}}\";" \
default_ttl=1h max_ttl=8h
# workload: obtain a lease
vault read database/creds/rag-reader
# lease_id database/creds/rag-reader/AbC123...
# lease_duration 1h
# lease_renewable true
# password ...
# username v-rag-reader-...
vault lease renew database/creds/rag-reader/AbC123... # extend, up to max_ttl
vault lease revoke -prefix database/creds/rag-reader # incident: kill every leaseNotice the VALID UNTIL clause: the database itself enforces expiry, so even if the broker is down when the lease ends, the password stops working.
Choosing a TTL
Picking a TTL is a balance. Too long and the exposure window grows; too short and refresh traffic, latency and issuer dependency grow. A useful lower bound is:
TTL >= issue_latency + longest_single_operation + clock_skew + refresh_marginThe longest single operation is the one that uses a credential once at the start and cannot re-authenticate in the middle. For LLM systems this is where surprises live:
- Weight downloads. A 140 GB checkpoint pulled from object storage at a few hundred MB per second takes several minutes. Most SDKs sign each request or part separately, so they pick up refreshed credentials between parts, but a single long presigned URL does not. A presigned URL signed with temporary credentials stops working when those credentials expire, even if the URL's own expiry is later.
- Long streams. A streaming completion or a long-lived database cursor authenticated once at connect time usually survives token expiry, because the check happened at connect. Do not rely on that as a security property, and do not rely on it surviving a reconnect.
- Training and batch jobs. A fine-tuning run that checkpoints to storage every hour for thirty hours will outlive any reasonable token. The job must refresh, not ask for a thirty-hour credential.
- Agent runs. An agent that pauses for human approval may resume after its token expired. Treat resumption as a fresh credential request, re-checking authorisation.
Refreshing without outages
Most refresh bugs come from hand-rolled caching. The rules are: refresh ahead of expiry, add jitter so a fleet does not stampede the issuer, let only one caller refresh at a time, and keep using a still-valid credential if a refresh fails. The class below implements those rules around any fetch function that returns a secret and an absolute expiry.
import random, threading, time
from dataclasses import dataclass
@dataclass
class Cred:
secret: object
expires_at: float # epoch seconds, from the issuer's response
class RefreshingCredential:
def __init__(self, fetch, refresh_fraction=0.66, skew=30.0, jitter=0.1):
self._fetch = fetch # () -> Cred; must raise on failure
self._frac = refresh_fraction
self._skew = skew # treat expiry as this many seconds earlier
self._jitter = jitter
self._lock = threading.Lock()
self._cred = None
self._refresh_at = 0.0
def _schedule(self, cred):
now = time.time()
life = max(0.0, cred.expires_at - self._skew - now)
j = 1.0 - random.uniform(0, self._jitter)
self._refresh_at = now + life * self._frac * j
def get(self):
now = time.time()
cred = self._cred
if cred and now < self._refresh_at:
return cred.secret # fast path, no lock
with self._lock: # single flight
cred = self._cred
if cred and time.time() < self._refresh_at:
return cred.secret
try:
new = self._fetch()
self._cred = new
self._schedule(new)
return new.secret
except Exception:
if cred and time.time() < cred.expires_at - self._skew:
self._refresh_at = time.time() + 15 # retry soon, serve stale-but-valid
return cred.secret
raise # truly expired: fail closedTwo details matter. The expiry comes from the issuer's response, not from the TTL you asked for, because issuers may shorten it. And the failure path serves the old credential only while it is still valid; once it has expired the call fails closed rather than retrying with a dead token in a loop. For a Kubernetes projected token the fetch function is just "re-read the file and parse its exp claim".
Rules specific to LLM systems
Four rules are specific to LLM systems.
- Credentials never enter the context window. The model asks for an action; a tool executor outside the model attaches the credential. If a token appears in a prompt, a tool result or a trace, short lifetime limits the damage but the design is still wrong. See Capability Tokens, in depth for scoping each tool call more tightly than a cloud role can.
- One credential per run, not per service. An agent run or a batch job should get its own session with a session name or tag that identifies the run. Audit logs then answer "which run did this", and you can revoke one run without touching the others.
- Narrow at issue time. STS accepts an inline session policy that can only reduce the role's permissions. A retrieval run for one tenant can be issued a session limited to that tenant's prefix, even though the role could read every tenant.
- Scrub before you log. Error messages from SDKs sometimes include signed URLs or headers. Redact known token shapes at the logger, because a credential that expires in an hour still works for that hour.
Revocation is a separate feature
Short lifetime is not revocation. When you detect a leak you want the credential dead now, and how you do that depends on the issuer.
- Vault leases: revoke by lease ID or prefix. Because the database user is dropped, this is real, immediate revocation.
- AWS role sessions: issued session tokens cannot be individually deleted. The console action to revoke active sessions attaches a deny policy to the role conditioned on
aws:TokenIssueTime, so every session issued before that moment is refused. Legitimate workloads must then re-assume the role, which they will do automatically if they refresh correctly. - Self-contained JWTs: a signed token is valid until its
expunless the resource checks something else, such as an introspection endpoint, a deny list of token IDs, or a rotated signing key. If you need early revocation for JWTs, design it in; you cannot add it during the incident.
Rehearse it. A quarterly drill that revokes a test workload's sessions and confirms both that the attacker's copy fails and that the real workload recovers within a minute is the only reliable proof that your refresh logic works. Fold it into LLM Incident Response, in depth.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
| Fleet refreshes at the same instant | Issuer throttling, latency spikes every TTL | Jittered refresh, as in the class above |
| Clock skew | Fresh tokens rejected as not yet valid or already expired | NTP everywhere; a skew margin in the cache |
| Issuer outage | Every workload fails once cached tokens expire | Serve stale-but-valid; size TTL to cover a plausible outage |
| Requested TTL above limit | Hard error at startup, e.g. a chained role asked for 2 hours | Read limits from config; alert on issue errors |
| Token file read once | Pod works for an hour, then 401s | Re-read projected token files; never cache at import |
| Long presigned URL | Download fails part-way after credential expiry | Sign with longer-lived identity or download in parts |
| Token in logs or prompts | Valid credential visible to anyone with log access | Redaction at the logger; executor-side injection |
Worked example: a RAG service off static secrets
Worked example. A RAG service reads a Postgres vector store and a document bucket. Before: a static database password and an access key, both in a Kubernetes secret, both two years old. After:
- The pod's projected service-account token, with a 1-hour expiry, authenticates it to Vault and to the cloud provider.
- Vault issues a
rag-readerdatabase lease withdefault_ttl=1h; the service renews it at about 40 minutes and gets a fresh lease when max_ttl is reached. - The cloud SDK assumes a read-only role for 1 hour with a session name of the pod name. A per-tenant batch job adds an inline session policy for its tenant's prefix.
- Both credentials sit behind the refresh class; dashboards track issue latency, issue errors and seconds-to-expiry of the oldest credential in use.
- The incident runbook names the two revocation commands and was rehearsed once.
The worst-case leak window fell from "until someone notices" to one hour, and the only long-lived secret left is the platform's own signing key, which the team never handles. For the rest of the static-secret inventory, see Secrets Management for LLM Apps, in depth.
What to do next
- Inventory every credential your LLM services use and mark each as static or ephemeral.
- For each static one, name the issuer that could replace it and the identity that would authenticate to that issuer.
- Measure your longest single operation per workload, and set TTLs from the formula above rather than by habit.
- Wrap every credential in one shared refresh helper with jitter, single flight and fail-closed expiry.
- Give each agent run or batch job its own session name and, where possible, a narrowed session policy.
- Add redaction for token shapes at the logger and confirm no tool result is passed to the model with a credential in it.
- Write and rehearse the revocation runbook for each issuer, including recovery of legitimate workloads.