An LLM application holds more credentials than it looks like it should. There is the platform API key for each model provider, often one per environment. There may be keys your customers bring for their own provider accounts. Then come the tokens the agent's tools use, the vector database password, webhook signing secrets and the key for your tracing vendor. Each one is a bill someone else can run up, or data someone else can read.
The architecture for keeping credentials out of the model's context window, with secret references and resolution at the tool boundary, is covered in secrets management architecture. This article is about the lifecycle around it: knowing what you hold, fetching and caching it safely, rotating without downtime, storing customer keys with envelope encryption, keeping keys out of logs and traces, and what to do in the first hour after a leak. The examples use AWS Secrets Manager and KMS because their APIs are concrete; Vault and the GCP and Azure equivalents follow the same shapes.
The secrets an LLM app actually holds
You cannot rotate what you have not listed. Start with an inventory that records, for every secret: what it unlocks, who owns it, where it is stored, which workloads read it, how it is rotated and what a leak would cost.
| Secret | Typical blast radius | Rotation reality |
|---|---|---|
| Platform model-provider key | Spend on your account; access to stored files or fine-tunes on that account | Usually created and revoked by hand in the provider console; plan for overlap |
| Customer-supplied provider key (BYOK) | The customer's spend and data; your reputation | The customer rotates it; you must accept updates cleanly |
| Tool and OAuth tokens | Whatever the scope allows the agent to do | Short-lived if issued per call; see the architecture article |
| Vector DB and database credentials | Your retrieval corpus, possibly every tenant's documents | Dynamic credentials where the store supports them |
| Webhook signing secrets | Forged events into your system | Dual-secret verification during rotation |
| Observability vendor keys | Write access to traces that may contain prompts | Rarely rotated; often the oldest key you hold |
The last row is the one teams forget. Tracing keys tend to be long-lived, broadly scoped and copied into every service, and the traces they write contain full prompts.
Architecture: one way in, filtered ways out
Two rules shape the design. First, there is one way in: a running service gets a credential from the secret store or by decrypting with KMS, authenticated by its workload identity, never from a baked-in environment file. Workload identity covers how to remove that last bootstrap secret. Second, every way out is filtered: logs, traces, error messages, eval datasets and support tickets pass through redaction, because something will eventually try to print a header.
Platform keys: one per service and environment
Do not share one provider key across the company. Create one key per service, per environment and, where the provider supports project or workspace scoping, per project. Separate keys buy you three things: attribution, wherever the provider's usage reporting breaks spend down by key or project; containment, because you can revoke one key without taking down every service; and limits, because spend caps and rate limits set on a narrow key bound the damage from a leak.
Name keys after the workload that owns them, for example support-agent-prod, and record the owner in the inventory. Store each as its own secret, so access policies can grant a service exactly the keys it needs. A developer laptop should never hold a production key; give developers a separate key with a low spend cap.
Fetching and caching
Fetching the secret on every model call makes the secret store a hard dependency of every request and adds latency. Fetching once at startup means a rotated key is not picked up until the next deploy. The middle path is an in-process cache with a short TTL, a bounded stale-if-error window and a forced refresh when the provider says the key is wrong.
import threading, time, logging
import boto3
log = logging.getLogger("secrets")
class SecretCache:
def __init__(self, secret_id, ttl=300, max_stale=900, min_force_interval=30):
self._sm = boto3.client("secretsmanager")
self._id, self._ttl, self._max_stale = secret_id, ttl, max_stale
self._min_force = min_force_interval
self._value, self._fetched, self._lock = None, 0.0, threading.Lock()
def get(self, force=False):
with self._lock:
now = time.monotonic()
age = now - self._fetched
if self._value and age < self._ttl and not (force and age >= self._min_force):
return self._value
try:
resp = self._sm.get_secret_value(SecretId=self._id, VersionStage="AWSCURRENT")
self._value, self._fetched = resp["SecretString"], now
except Exception:
if self._value and age < self._max_stale:
log.warning("refresh failed for %s; serving cached value", self._id)
return self._value
raise
return self._value
provider_key = SecretCache("llm/support-agent/prod/provider-key")
def call_model(http, url, payload):
key = provider_key.get()
r = http.post(url, json=payload, headers={"Authorization": f"Bearer {key}"}, timeout=60)
if r.status_code == 401: # rotated or revoked: refresh once
fresh = provider_key.get(force=True)
if fresh != key:
r = http.post(url, json=payload, headers={"Authorization": f"Bearer {fresh}"}, timeout=60)
r.raise_for_status()
return r.json()The header name and scheme are illustrative; providers differ, so use their SDK where you can and pass the key in its client constructor. Three details matter. The minimum force interval stops a burst of 401s from becoming a burst of secret-store calls. The stale window keeps the app working through a short secret-store outage but not forever. And the cache never logs the value, only the secret's name.
Rotation without downtime
Rotation without downtime needs two valid keys at once. Secrets Manager tracks versions with staging labels: AWSCURRENT is what readers get by default, AWSPREVIOUS is the version before it, and AWSPENDING marks a version being prepared by a rotation function. Calling put_secret_value without naming stages makes the new version AWSCURRENT and moves the old one to AWSPREVIOUS, which is exactly the overlap a manual provider-key rotation needs.
- Create a new key in the provider console for the same workload. Both keys now work.
- Store it: put_secret_value on the secret. Readers fetching AWSCURRENT now get the new key.
- Wait at least one cache TTL plus your longest request, so every process has refreshed. With a 300-second TTL, ten minutes is comfortable.
- Confirm the old key is idle, using the provider's per-key usage view if it has one, or your own request logs tagged with the secret version.
- Revoke the old key at the provider. Any straggler gets a 401, forces a refresh, and recovers.
Where a provider offers an API to create and revoke keys, the same steps fit a Secrets Manager rotation function, which runs createSecret, setSecret, testSecret and finishSecret steps. Most teams rotate provider keys by hand on a schedule, quarterly for example, and immediately on any suspected exposure. Rehearse the manual version before you need it at 3 a.m.
Customer-supplied keys and envelope encryption
If customers bring their own provider keys, you are storing other people's credentials, and a database dump must not hand them over. Use envelope encryption: KMS generates a data key, you encrypt the customer's key locally with it, and you store only the ciphertext plus the KMS-wrapped data key. The encryption context binds each record to its tenant, so a wrapped key copied onto another tenant's row will not decrypt, and every decrypt is logged with that context.
import os
import boto3
from cryptography.hazmat.primitives.ciphers.aead import AESGCM
kms = boto3.client("kms")
KEY_ID = "alias/byok-provider-keys"
def _ctx(tenant_id, provider):
return {"tenant": tenant_id, "provider": provider, "purpose": "llm-provider-key"}
def seal(tenant_id, provider, api_key):
dk = kms.generate_data_key(KeyId=KEY_ID, KeySpec="AES_256",
EncryptionContext=_ctx(tenant_id, provider))
nonce = os.urandom(12)
aad = f"{tenant_id}|{provider}".encode()
ct = AESGCM(dk["Plaintext"]).encrypt(nonce, api_key.encode(), aad)
return {"wrapped_key": dk["CiphertextBlob"], "nonce": nonce,
"ciphertext": ct, "last4": api_key[-4:]}
def unseal(tenant_id, provider, row):
dk = kms.decrypt(CiphertextBlob=row["wrapped_key"],
EncryptionContext=_ctx(tenant_id, provider))
aad = f"{tenant_id}|{provider}".encode()
return AESGCM(dk["Plaintext"]).decrypt(row["nonce"], row["ciphertext"], aad).decode()Validate a customer key with one cheap provider call when it is submitted, store last4 so the UI can show which key is active without decrypting it, and never return the full key through any API. Decrypt just before the outbound call and keep the plaintext in a local variable rather than a long-lived cache. Python cannot reliably wipe memory, so keep the window short instead. Restrict kms:Decrypt to the service that makes provider calls, not the web tier that accepts the key.
Keeping keys out of logs and traces
Most real leaks are not attacks; they are logging. HTTP client debug logging prints request headers. Exception messages include the failing request. Tracing integrations record request attributes. Prompt logs capture a user who pasted a key into the chat. Each copy lands in a store with its own retention and access list.
Defend in layers. Turn off header-level debug logging in production. Configure SDKs and tracing to record the prompt and response, not transport headers. And put a redaction filter in front of every sink as a backstop:
import logging, re
PATTERNS = [
re.compile(r"(?i)(authorization|x-api-key|api[_-]?key)(['\"]?\s*[:=]\s*['\"]?)(bearer\s+)?[A-Za-z0-9._\-]{12,}"),
re.compile(r"\bsk-[A-Za-z0-9_\-]{16,}"), # one common key prefix; add your providers' formats
]
class RedactSecrets(logging.Filter):
def filter(self, record):
msg = record.getMessage()
for pat in PATTERNS:
msg = pat.sub("[REDACTED]", msg)
record.msg, record.args = msg, None
return True
for handler in logging.getLogger().handlers:
handler.addFilter(RedactSecrets())Regexes miss things, so pair the filter with a real detector on stored data. Secret scanners covers detectors, verification and scanning prompts and tool results, and audit logging for LLM systems covers what to keep and for how long.
Worked example: the first hour after a leak
A support engineer notices a production provider key in a trace: a timeout exception had serialised the full outbound request, headers included. The first hour looks like this.
- Minute 0 to 10: create a replacement key, put_secret_value, and wait for caches to refresh. Because rotation was rehearsed, nothing breaks.
- Minute 10 to 15: revoke the leaked key. Revoke first if there is any sign of abuse, accepting a short outage.
- Minute 15 to 30: check provider usage for the leaked key since the trace's timestamp. Unexplained spend or calls from unknown addresses turn this from an exposure into an incident.
- Minute 30 to 60: find every copy. Search trace, log and eval stores, and any dataset exported from them, for the key and its last characters; purge or restrict them.
- Afterwards: fix the cause, here the exception serialiser and a missing redaction filter on the trace exporter, and add a test that fails if a header reaches a log.
Write this up as a playbook entry; incident response for LLM systems covers the wider process. Note that revocation is what stops abuse. Purging copies only reduces future exposure, which is why rotation speed matters more than search speed.
Failure modes
- Rotation that breaks production. The old key is revoked before caches refresh. Always overlap by at least one TTL.
- A refresh storm. Every request forces a secret-store call after a 401. Rate-limit forced refreshes.
- Keys in eval and fine-tuning data. Datasets built from logs inherit their secrets. Scan before export.
- Over-broad decrypt rights. Any service that can call kms:Decrypt on the BYOK key can read every customer key.
- Stale inventory. A key nobody owns is a key nobody rotates.
Trade-offs
Shorter cache TTLs pick up rotations faster but add secret-store load and a dependency on its availability; a few minutes is a good default. Per-service keys multiply the number of things to rotate, but turn every leak from a company-wide event into a single-service one. Envelope encryption costs a KMS call per decrypt, which you can reduce by caching data keys briefly, at the price of a wider window if a process is compromised. Choose deliberately and record the choice in the inventory.
What to do next
- Build the inventory: every secret, owner, store, readers, rotation method and blast radius.
- Split shared provider keys into one key per service and environment, with spend caps where available.
- Replace startup-only secret loading with a TTL cache, a stale-if-error window and a rate-limited 401 refresh.
- Rehearse a provider-key rotation end to end in staging and time it.
- Store customer keys with KMS envelope encryption bound to the tenant by encryption context.
- Disable header debug logging in production and add a redaction filter to every log and trace sink.
- Write the leak playbook: rotate, revoke, check usage, find copies, fix the cause.