Most writing about API keys and LLMs is from the consumer side: you hold a provider's key and must not leak it. That side is covered in API key management for LLM providers. This article is about the other side. You run an LLM service, whether a public API, an internal platform or a model gateway for other teams, and you issue keys to callers. Every key you hand out is a bearer credential that spends your GPUs directly, so design choices that look cosmetic, such as how the key string is shaped, decide how fast you can detect a leak, how cheaply you can reject garbage, and how much one stolen key can cost before anyone notices.

We will design a key format from first principles, store keys so that a database dump does not hand out working credentials, build the request path that validates them, attach scopes and token budgets, rotate without downtime, and wire up leak reporting. A worked example follows one leaked key from a public commit to revocation.

Why LLM keys are worth stealing

Three properties make LLM API keys worth more to an attacker than typical SaaS keys. They convert directly into compute that can be resold: stolen keys are used to run unrelated workloads or offered through proxy services. Usage is expensive per request, so a stolen key can burn thousands of dollars of inference in hours. And abuse is attributed to the key's owner: harmful generations, policy violations and rate-limit bans land on your customer, not the thief.

The issuer's goals follow: make keys easy to recognise when they leak, impossible to recover from your own storage, cheap to reject when wrong, narrow in what they can do, bounded in what they can spend, and fast to revoke.

Designing the key string

A good key string has four parts: a recognisable prefix, an identifier you can index, a random secret, and a checksum. GitHub's 2021 token redesign is a useful public reference. Its tokens start with a type prefix such as ghp_ for personal access tokens; the underscore was chosen because it is not a Base64 character, so random strings such as SHAs cannot accidentally look like tokens, and GitHub anticipated the prefix alone would bring secret-scanning false positives down to about 0.5 percent. The last six characters hold a CRC32 checksum encoded in Base62, so a candidate can be validated offline without a database query.

Apply the same ideas to your own service. A prefix such as csk_live_ or csk_test_ says whose key it is and which environment it unlocks, so a scanner, a log filter or a human can spot it. A short public key id lets you look the record up directly. The secret should carry at least 128 bits, and more is cheap: 32 Base62 characters carry about 190 bits. The checksum covers everything before it.

import secrets, zlib, hmac, hashlib

B62 = "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz"

def b62(n, width):
    out = ""
    while n:
        n, r = divmod(n, 62)
        out = B62[r] + out
    return out.rjust(width, "0")

def rand62(length):
    return "".join(secrets.choice(B62) for _ in range(length))

def mint(env="live"):
    key_id = rand62(10)                        # public, indexed
    secret = rand62(32)                        # ~190 bits, never stored
    body = f"csk_{env}_{key_id}_{secret}"
    return body + b62(zlib.crc32(body.encode()), 6)

def well_formed(key):
    body, check = key[:-6], key[-6:]
    return (key.startswith(("csk_live_", "csk_test_"))
            and len(key) == len("csk_live_") + 10 + 1 + 32 + 6
            and b62(zlib.crc32(body.encode()), 6) == check)

The checksum is not a security control; anyone can compute CRC32. It is a noise filter. It lets the edge reject typos, truncated pastes and random strings without touching the key store, and lets a secret scanner discard look-alikes before reporting them.

Storing keys you cannot recover

Never store the key. Store a hash of the secret and the metadata around it. Because the secret is long and random, a fast hash is enough: there is no dictionary to brute force, unlike a human password, so bcrypt or Argon2 add tens of milliseconds to every API call and buy nothing. Use HMAC-SHA-256 keyed with a server-side pepper held in a KMS or secrets manager. A leaked database then yields neither keys nor anything that can be checked offline without the pepper.

CREATE TABLE api_keys (
  key_id        text PRIMARY KEY,          -- the public 10-char id
  org_id        text NOT NULL,
  secret_hmac   bytea NOT NULL,            -- HMAC-SHA-256(pepper, secret)
  pepper_ver    smallint NOT NULL,         -- supports pepper rotation
  scopes        text[] NOT NULL,           -- e.g. {chat:write, embed:write}
  models        text[],                    -- null = org default
  monthly_cap_usd numeric,                 -- per-key spend ceiling
  expires_at    timestamptz,
  state         text NOT NULL DEFAULT 'active',  -- active|revoked
  created_by    text NOT NULL,
  last_used_at  timestamptz,
  last_used_ip  inet
);
def verify(key, store, peppers):
    if not well_formed(key):
        return None                               # no DB access at all
    key_id, secret = key.split("_")[2], key.split("_")[3][:32]
    rec = store.get(key_id)                       # cached lookup by id
    if rec is None or rec.state != "active":
        return None
    mac = hmac.new(peppers[rec.pepper_ver], secret.encode(), hashlib.sha256).digest()
    return rec if hmac.compare_digest(mac, rec.secret_hmac) else None

Show the full key exactly once, at creation, and display only the prefix and id afterwards. Record who created each key and when it was last used; those two columns answer most incident questions.

The validation path

Every inference request crosses the authentication path, so it has to be fast and it has to fail closed. Order the checks from cheapest to most expensive. Parse and checksum at the edge. Look the key id up in a short-lived in-process or shared cache. Go to the store only on a cache miss. Then compare the HMAC in constant time, apply scopes, and only then reserve budget and forward to the model.

Request path for an issued key: cheap rejections first, the database lastClientBearer keyEdge parseprefix + checksumvalidAuth cachekey id -> recordmissKey storehash, scopes, statebad checksum401, no DB hitcheap rejectionrecordVerify + policyhash compare, scopeBudget metertokens, spendmodelRevocation buspush invalidationsThe checksum lets the edge drop typos and scanner noise without a lookup; revocations are pushed, not waited out.
Malformed keys die at the edge, valid ones are resolved by id through a cache, and revocations are pushed to every cache instead of waiting for TTLs to expire.

The cache is where revocation latency hides. With a five-minute TTL, a revoked key keeps working for up to five minutes on every node that cached it, which at LLM prices can be a lot of tokens. Publish revocations on a bus that every node subscribes to and evicts on, and keep the TTL as a backstop. Cache negative results briefly too, so a flood of requests with a well-formed but unknown key does not become a flood of database queries.

Scopes, budgets and short-lived tokens

A key should carry the narrowest authority its caller needs. For an LLM service the useful dimensions are:

  • Operations: chat, embeddings, fine-tuning, file upload and key administration are different risks. A key that only embeds cannot exfiltrate uploaded files.
  • Models: restrict expensive or sensitive models to keys that need them.
  • Token and request rates: requests per minute and tokens per minute per key, with the org limit above it. How to derive these from GPU capacity is covered in Rate limiting for LLM endpoints.
  • Spend ceilings: a monthly or daily dollar cap per key turns a leak from unbounded loss into a known worst case. Enforce it by reserving the request's maximum possible cost (prompt tokens plus max output tokens) before running and settling after.
  • Network and time: optional IP allowlists for server-to-server callers, and expiry dates for keys issued to contractors, demos or CI.

Browsers and mobile apps should never hold a long-lived key. Have the customer's backend mint a short-lived, narrowly scoped token from its real key and pass that to the client, so a key scraped from a page expires in minutes.

Rotation without a flag day

Rotation fails when it requires a flag day. Allow at least two active keys per application so a customer can create the new key, deploy it, watch the old key's last_used_at stop advancing, and then revoke the old one. Surface last-used time and source in the dashboard; it is the evidence customers need to revoke with confidence. Offer an optional automatic expiry with warning emails rather than forcing short lifetimes on everyone, because forced rotation of server keys mostly produces outages. When you rotate your own pepper, store its version per row and rehash on next successful use.

Make the rotation API scriptable as well as clickable. A customer with fifty services will not rotate through a dashboard, so expose create, list with last-used data, and revoke as API operations under a separate administration scope that inference keys never carry. That separation matters: if any leaked inference key could mint new keys, revoking it would not end the incident, because the attacker would already hold a fresh one.

Finding leaked keys

Leaks are found in three places: public code, your own traffic and reports from others. For public code, GitHub runs a secret scanning partner program: a provider registers patterns for its token formats and an endpoint, and GitHub sends matches found in public repositories to that endpoint, signed so the provider can verify the report came from GitHub. Your distinctive prefix and checksum are what make the pattern precise enough to qualify. Handle reports idempotently: verify the signature, check the checksum, look the key up, and act according to policy, which for live keys is usually immediate revocation plus notification.

In your own traffic, alert on changes that a legitimate customer rarely makes suddenly: a new country or autonomous system, a jump in tokens per hour beyond the key's history, a shift to the most expensive model, or many concurrent streams from a key that used to send one at a time. Tools that scan your customers' repositories and logs are covered in Secret scanners.

Worked example: one leaked key, two timelines

At 14:02 a developer at a customer pushes a notebook with a live key to a public repository. At 14:03 the scanning partner report arrives. The handler verifies the signature, sees a valid checksum and an active key, marks it revoked, and publishes a revocation event. By 14:03:05 every gateway node has evicted the key; the next request with it gets 401. An email and dashboard banner tell the customer which key, where it was found and when it was last used legitimately.

Now replay the same leak with the defaults many services start with: an unprefixed hex key, no partner registration, a ten-minute auth cache and no spend cap. Nobody is told. A scraper finds the key within the hour and runs a large model continuously. The customer learns from a bill days later. The difference between the two timelines is entirely in issuer-side design.

Failure modes

FailureConsequenceMitigation
Keys stored in plaintext or reversiblyDatabase leak is a credential leakHMAC with KMS-held pepper
Slow password hash per requestLatency and CPU on every callFast keyed hash; secret is high entropy
Long auth cache, no pushRevoked keys work for minutesRevocation bus plus short TTL
One key per org, all scopesAny leak exposes everythingPer-app keys with scopes and models
No spend ceilingUnbounded loss from a leakPer-key caps with reservation
Generic key formatScanners cannot find leaksPrefix, id and checksum
Single active keyRotation needs downtimeTwo keys plus last-used telemetry

Trade-offs

Each control has a price. Prefixes reveal which vendor a key belongs to, which helps an attacker triage a dump; that is outweighed by detection. Push-based revocation adds an event system to operate. Spend reservations can reject a request that would have fit because its maximum output was reserved; let callers set max tokens to reduce that. Automatic revocation on a partner report can break a production integration if the report is about a key that was meant to be public, which should never happen with secret keys but does; notify loudly and make re-issuing easy.

What to do next

  1. Define a key format with environment prefix, public id, at least 128 bits of secret and a checksum; validate it at the edge.
  2. Migrate storage to keyed hashes with a versioned pepper in your KMS, and show keys only once.
  3. Add scopes for operations and models, per-key token rates, and a per-key spend ceiling with reservation.
  4. Allow two active keys per application and expose last-used time and source.
  5. Push revocations to every auth cache and measure time from revoke to first 401.
  6. Register your format with secret scanning partner programs and build an idempotent report handler.
  7. Alert on per-key anomalies: new networks, token spikes, model shifts, concurrency jumps.
  8. Review the consumer side in Secrets management for LLM apps so your own outbound provider keys get the same care.
Key takeaway: An issued LLM key is money in bearer form. Make it recognisable with a prefix and checksum, store only a keyed hash, reject bad keys before any lookup, narrow every key by operation, model, rate and spend, let customers rotate with two live keys, and push revocations everywhere within seconds of a leak report.