An LLM provider API key is a bearer credential with a price tag. Anyone who holds the string can spend your money, read whatever your account can read (fine-tuned models, stored files, batch outputs) and burn through your rate limits, and nothing in the request proves who sent it. That combination makes provider keys some of the most actively hunted secrets on the internet. Attackers scan public repositories, mobile app bundles and leaked environment files for them, and stolen access to hosted models is resold through reverse proxies. Security researchers at Sysdig named the pattern LLMjacking in 2024, after finding stolen cloud credentials being used to run up model bills for the account owner.
This article is about the keys themselves: where they may live, who may hold them, how one leak is contained to one tenant and one budget, and what to do in the first hour after a leak. General secret storage and zero-downtime rotation are covered in Secrets Management for LLM Apps; here we build the provider-key layer on top of it: a gateway that holds the real keys and hands out revocable virtual keys, with budgets, scoping, detection and an incident runbook.
How provider keys leak
Start with how provider keys actually leak, because each path needs a different control.
| Leak path | Typical cause | Control |
|---|---|---|
| Source control | Key pasted into a notebook, test or config and pushed | Pre-commit and server-side secret scanning; keys never in code |
| Client bundles | Browser or mobile app calls the provider directly | Clients never get provider keys; they call your backend |
| Logs and traces | Request headers or the whole client config logged on error | Redact the authorization header at the logger and the tracer |
| Prompts and agent context | Key in an environment variable an agent can read and echo | Agents call tools through a broker; no key in the process the model can inspect |
| CI and build caches | Key exported for integration tests, cached in an artifact | Short-lived CI credentials; a separate low-budget test key |
| Shared dev keys | One key in a team chat, copied into ten laptops | Per-developer keys with small budgets, or the gateway |
Two properties make the damage larger than for most secrets. First, the cost is immediate and unbounded: a stolen key used for bulk generation spends real money per token, and the bill arrives before your monthly review. Second, keys are usually account-wide by default. A key created for one chatbot can often call every model, read uploaded files and start fine-tuning jobs. Least privilege therefore has to be designed in, because the default gives every key the whole account.
The scanning is fast. Treat any key that reached a public repository as compromised within minutes, whatever you do afterwards. Some providers participate in GitHub secret scanning: Anthropic, for example, documents that keys found in public GitHub repositories are reported to it and deactivated automatically, with an email to the owner. That is a safety net, not a control.
One holder: the gateway architecture
The architecture that removes most of these paths is a single LLM gateway that is the only component holding provider keys. Every other caller authenticates to the gateway with something you control and can revoke individually: a user session, a workload identity (mTLS certificate, OIDC token from your platform) or a virtual key you issued. The gateway maps the caller to a tenant, checks that tenant's budget and model allow-list, attaches the real provider key and forwards the request.
The data flow for one request is short. The caller presents its credential; the gateway hashes it, looks up the tenant record, checks the model against the allow-list, reserves an estimated cost against the budget, fetches the provider key from an in-memory cache (populated from the secret store with a short TTL), forwards the call, then settles the reservation with the actual token counts from the response and writes one usage record. The provider key never appears in a response, a log line or the caller's environment.
Cloud-hosted model endpoints change the picture for the provider leg. When the model is reached through your cloud account, the gateway can authenticate with its own IAM role and short-lived credentials instead of a static string, which is strictly better: there is nothing long-lived to steal from the gateway's configuration. LLMjacking attacks against hosted models typically start from leaked long-lived cloud access keys, so the same discipline applies one level down.
Virtual keys with budgets
A virtual key is a random token you mint, store only as a hash and attach to a tenant record with limits. It gives you everything the provider key lacks: per-caller identity, per-caller budgets and instant individual revocation. The core of a gateway check fits in a page.
import hashlib, hmac, secrets, time
from dataclasses import dataclass, field
PEPPER = b"loaded-from-secret-store" # server-side secret, not in the DB
def mint_virtual_key() -> tuple[str, str]:
raw = "vk_" + secrets.token_urlsafe(32) # shown to the caller once
digest = hmac.new(PEPPER, raw.encode(), hashlib.sha256).hexdigest()
return raw, digest # store only the digest
@dataclass
class Tenant:
name: str
models: set[str]
monthly_budget_usd: float
spent_usd: float = 0.0
reserved_usd: float = 0.0
revoked: bool = False
expires_at: float = field(default_factory=lambda: time.time() + 90 * 86400)
TENANTS: dict[str, Tenant] = {} # digest -> tenant (a DB in practice)
def authorize(raw_key: str, model: str, est_cost: float) -> Tenant:
digest = hmac.new(PEPPER, raw_key.encode(), hashlib.sha256).hexdigest()
t = TENANTS.get(digest)
if t is None or t.revoked or time.time() > t.expires_at:
raise PermissionError("unknown, revoked or expired key")
if model not in t.models:
raise PermissionError(f"model {model} not allowed for {t.name}")
if t.spent_usd + t.reserved_usd + est_cost > t.monthly_budget_usd:
raise PermissionError("budget exhausted") # HTTP 429 or 402 to the caller
t.reserved_usd += est_cost # atomic in a real store
return t
def settle(t: Tenant, est_cost: float, actual_cost: float) -> None:
t.reserved_usd -= est_cost
t.spent_usd += actual_costThree details matter. Hash with a keyed HMAC rather than a bare SHA-256 so a database dump alone cannot be used to test guesses. Reserve before the call and settle after it, using the maximum output tokens for the estimate, so a burst of concurrent requests cannot overshoot the budget. And give every virtual key an expiry, so forgotten keys die on their own. The reservation pattern is the same one used for token-weighted limits in Rate Limiting for LLM Endpoints; budgets are rate limits measured in money over a month instead of tokens over a minute.
Scoping at the provider
The gateway is your main scoping tool, but use what the providers offer as a second layer. Most major providers now let you split an organisation into projects or workspaces, issue keys bound to one of them, and in some cases restrict a key's permissions or set usage limits per project or workspace. The exact controls differ by provider and change often, and some are only available on certain account types, so check the current console rather than assuming. The design principles are stable:
- One provider key per environment and blast radius. Production, staging and development get separate keys in separate projects, so a leaked dev key cannot touch production data or production rate limits.
- Two live keys per production slot. Keep a primary and a secondary so you can revoke one immediately and keep serving on the other while you rotate. The gateway reads both versions from the secret store and fails over on an authentication error.
- Provider-side spend caps where available. A cap enforced by the provider still holds if the gateway is bypassed or buggy. Set it above normal peak, well below a painful bill.
- Least-privilege permissions. If a key only needs inference, do not let it manage files, fine-tunes or other keys. Administrative keys that can create keys belong to a human break-glass process, never to an application.
Clients, agents and logs
The most common design mistake is shipping a provider key in a client. A browser bundle, a mobile app binary or a desktop app is fully readable by its user; obfuscation and string splitting buy minutes. If the app needs model access, the app calls your backend with the user's session, and the backend (or the gateway) calls the provider. That also gives you per-user quotas, abuse detection and the ability to change providers without shipping a new app.
The same rule applies to agents. An agent process that has the key in its environment can be prompted into reading and printing it, or into sending it to a tool. Run the model-facing agent without provider credentials and let a broker attach them outside the model's reach, as described in Egress Control for Agents. If a tool needs a third-party key, the tool executor holds it, not the prompt.
Logging is the quiet leak path. HTTP client debug logging prints headers; exception handlers dump client configuration; tracing libraries capture request attributes. Add a redaction filter that removes the authorization and API-key headers at the logging layer, then test it by grepping a staging log stream for your key prefix. Run the same patterns in a secret scanner over logs and prompts, as covered in Secret Scanners, in depth.
Detecting a leaked key
Prevention will eventually fail, so measure every key's behaviour against its own baseline. The gateway usage log already contains what you need: tenant, model, input and output tokens, cost, source address and time. A few cheap signals catch most abuse:
- Spend per tenant per hour above a multiple of its trailing seven-day hourly p99.
- A model the tenant has never used, especially the most expensive one.
- Requests from source networks or regions the tenant has never used.
- A sudden jump in output-to-input token ratio, typical of bulk generation for resale.
- Provider-side usage that does not match gateway usage: proof that someone is calling the provider with your key without passing through the gateway.
The last signal is the one that catches a leaked provider key rather than a leaked virtual key. Reconcile daily: pull usage per provider key or project from the provider's usage reporting and compare it with the gateway's ledger. Any surplus beyond a small tolerance is unexplained spend.
def check_reconciliation(provider_usd: float, gateway_usd: float,
tolerance: float = 0.03) -> str:
"""Daily: provider-reported spend vs the gateway's own ledger for one key."""
if gateway_usd == 0:
return "PAGE: key used with no gateway traffic" if provider_usd > 1.0 else "ok"
surplus = provider_usd - gateway_usd
if surplus > tolerance * gateway_usd:
return f"PAGE: {surplus:.2f} USD unexplained ({surplus / gateway_usd:.0%})"
return "ok"Worked example: the gateway records 412.80 USD for the production key yesterday; the provider reports 1,903.55 USD. The surplus is 1,490.75 USD, 361% of gateway spend, so the check pages. Because the gateway is the only legitimate caller, the provider key itself has leaked, and the runbook below starts.
The leak runbook
Write the leak runbook before you need it, and rehearse it. The order matters: stop the spending first, investigate second.
- Revoke. For a virtual key, mark it revoked in the gateway; the next request fails. For a provider key, switch the gateway to the secondary key, confirm traffic is healthy, then revoke the leaked key in the provider console. Do not wait to find the leak source.
- Contain. Lower provider-side spend caps temporarily and check that no new keys, fine-tuning jobs or uploaded files were created with the leaked credential if it had those permissions.
- Rotate the spare. Mint a new secondary so you are back to two live keys.
- Find the source. Search repositories and history, CI logs, application logs, client bundles and chat tools for the key prefix and its fingerprint. Removing it from the latest commit is not enough; it is still in history and in every clone.
- Quantify. Use provider usage reports to bound the cost and time window, and contact the provider about fraudulent usage if it is material.
- Fix the path. Add the scanner rule, logger redaction or architecture change that closes the leak route, and record it in the post-incident review.
Trade-offs
| Option | Gains | Costs |
|---|---|---|
| Keys in each service's environment | Simple, no extra hop | N copies to leak and rotate; no per-caller budgets |
| Central gateway with virtual keys | One place for keys, budgets, audit, failover | Extra hop and an availability dependency; must be highly available |
| Cloud-hosted models via IAM | No static key, short-lived credentials | Tied to one cloud; IAM misconfiguration becomes the new risk |
| Provider-side projects and caps only | No infrastructure to run | Coarse; no per-user identity or business-level budgets |
Most teams end up combining them: a gateway for identity and budgets, provider projects and caps as a backstop, and IAM-based access wherever a cloud endpoint is available. Run at least two gateway instances; its availability, not its latency, is the real cost.
Failure modes
- The gateway becomes a single point of failure. One instance, a slow secret store call per request, no cache: an outage of the store takes down every model call.
- Budgets checked after the call. Concurrent requests all pass the check and overshoot; reserve first.
- Virtual keys that never expire. Departed contractors keep working credentials for years.
- Revocation not tested. The first real revocation discovers that nothing reads the secondary key and production goes down.
What to do next
- Inventory every provider key: owner, project, permissions, where it is stored, who can read it.
- Remove keys from clients, notebooks and agent environments; route those calls through a backend.
- Put a gateway in front of providers with virtual keys, model allow-lists, reservations and expiry.
- Keep two production provider keys and rehearse failover plus revocation in staging.
- Set provider-side spend caps and least-privilege permissions where the provider supports them.
- Add logger redaction and secret scanning for repos, CI, logs and prompts.
- Build daily provider-versus-gateway reconciliation and hourly per-tenant spend alerts.