When people say a model was stolen they usually mean one of two very different things. Either someone used the public API to train a cheaper model that imitates it, or someone walked off with the weights. Between those sits a third, less familiar case: the API's own output format giving away exact pieces of the model's architecture and parameters.
This article covers the second and third cases in detail and the first one briefly, because model extraction through the API already has its own deep dive on surrogate training, detection and deterrence. Here you will learn why a logit vector is a leak, how to measure your own endpoint, and which custody controls matter for checkpoints.
What can be stolen, and through which door
Model stealing is not one attack. The table separates the targets, because a control that stops one does nothing for another. Rate limits slow distillation but do nothing about an over-permissive checkpoint bucket.
| Target | Door | What the attacker gets | Who owns the defence |
|---|---|---|---|
| Behaviour (functional clone) | Generated text at scale | A student model trained on your outputs | API product and abuse teams |
| Architecture facts | Full or partial logit vectors | Hidden size, identity of the base model | API design |
| Exact final layer | Logprobs plus logit bias | The output projection matrix, up to a symmetry | API design |
| Full weights | Storage, hosts, people | Everything, including future fine-tuning | Security and infrastructure |
Why a logit vector is a leak: the softmax bottleneck
A decoder-only transformer ends the same way every time. The last layer produces a hidden vector h of size d, the hidden size, and an output projection W of shape V by d turns it into V logits, one per vocabulary token. V is large, often 32,000 to more than 200,000. d is much smaller, typically a few thousand.
That size gap is the whole story. Every logit vector the model can ever produce is W times some h, so all of them lie in a subspace of dimension at most d inside a V-dimensional space. Collect more than d full logit vectors from random prompts, stack them into a matrix and compute its singular values: about d of them are large and the rest collapse to numerical noise. The position of that drop is the hidden size. Log-probabilities are logits minus a per-prompt constant, so they span at most d + 1 dimensions, and the same trick works.
Two 2024 papers turned this into practical attacks on production APIs. Carlini and colleagues, in Stealing Part of a Production Language Model, recovered the hidden size and the output projection matrix (up to an unavoidable symmetry) of OpenAI's ada and babbage models for under $20, confirming hidden sizes of 1024 and 2048. They also recovered the hidden size of gpt-3.5-turbo and estimated that the full projection would cost under $2,000 in queries. Their key enabler was that the API combined logprobs with a logit bias parameter, so an attacker could push chosen tokens into the visible top-k and reconstruct full vectors piece by piece. Finlayson, Ren and Swayamdipta, in Logits of API-Protected LLMs Leak Proprietary Information, used the same low-rank structure to estimate gpt-3.5-turbo's embedding size at about 4096. They also showed that one full output can identify which model produced it. The affected vendors deployed mitigations after disclosure.
For a defender, the point is that this is a property of the architecture, not one vendor's bug. Any endpoint that returns enough of the distribution exposes it, and the cost falls as each query reveals more tokens.
Audit your own endpoint before someone else does
You do not need to attack anyone to know your exposure. Run the measurement on a model you own and you will see exactly what a full-logit API would reveal. The script below samples random prompts, takes the final-position logits and finds the largest drop between consecutive singular values.
# Audit what full logit vectors from a model you own reveal. Runs locally.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
name = "your-org/your-model" # a checkpoint you control
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForCausalLM.from_pretrained(name, torch_dtype=torch.float32).eval()
d = model.config.hidden_size # the number the attack would recover
n = d + 256 # need more samples than the hidden size
rows = []
with torch.no_grad():
for i in range(n):
ids = torch.randint(0, tok.vocab_size, (1, 16)) # random prompts are enough
logits = model(ids).logits[0, -1] # length V
rows.append(logits)
L = torch.stack(rows) # shape (n, V)
s = torch.linalg.svdvals(L)
ratio = s[:-1] / s[1:]
knee = int(torch.argmax(ratio[: n - 1])) + 1 # largest drop between singular values
print("configured hidden size:", d, " numerical rank (knee):", knee)
# Expect knee == d (d + 1 for log-probabilities).On any standard decoder the knee lands on the configured hidden size, which makes the point concrete for reviewers. Then repeat the exercise against your own staging API with the options your customers actually have: top-k logprobs, logit bias, prompt echo, the embeddings endpoint. For each option, ask how many queries it takes to rebuild one full vector. If both logprobs and logit bias can be combined freely, you have the configuration the Carlini paper exploited.
Treat the result as an input to a risk decision. For a fine-tune of a well-known open-weight base, the hidden size is already public; what stays sensitive is the fine-tuned output layer and which base you use. For a model trained from scratch, the hidden size is a real secret, because it narrows down parameter count and compute.
An output-interface policy
Each API feature should go through an explicit decision. The table lists the common ones, what each reveals, and the usual mitigation.
| Feature | What it reveals | Typical policy |
|---|---|---|
| Full logprob vector | The whole subspace, quickly | Do not offer it on any public tier |
| Top-k logprobs | k coordinates per position | Cap k, and count logprob tokens against quota |
| Logit bias | Lets the caller choose which coordinates become visible | Never combine with logprobs on the same request |
| Prompt echo with logprobs | Logprobs for every input position, many vectors per call | Disable |
| Embeddings endpoint | A different layer's geometry | Separate quotas and monitoring |
The policy belongs in the gateway, not in client libraries, and it should be tiered. Researchers and enterprise customers can get more, under contract and with tighter monitoring. The sketch below shows the shape. Field names follow the common chat-completions style. Map them onto whatever your API actually exposes.
# Gateway policy: decide what a request may see, per tier, before it reaches the model.
from dataclasses import dataclass
@dataclass
class Tier:
max_top_logprobs: int # 0 = no logprobs at all
allow_logit_bias: bool
allow_bias_with_logprobs: bool
daily_token_cap: int
TIERS = {
"free": Tier(0, False, False, 200_000),
"standard": Tier(5, True, False, 5_000_000),
"research": Tier(20, True, False, 50_000_000), # still never both together
}
def enforce(req: dict, tier: Tier, used_today: int) -> dict:
k = int(req.get("top_logprobs") or 0)
if k > tier.max_top_logprobs:
raise PermissionError(f"top_logprobs > {tier.max_top_logprobs}")
bias = req.get("logit_bias") or {}
if bias and not tier.allow_logit_bias:
raise PermissionError("logit_bias not available on this tier")
if bias and k and not tier.allow_bias_with_logprobs:
raise PermissionError("logit_bias cannot be combined with logprobs")
if used_today >= tier.daily_token_cap:
raise PermissionError("daily token cap reached")
return {**req, "echo": False} # never return prompt-token logprobsAdding noise to logprobs or rounding them reduces precision but does not change the low-rank structure. Both raise the query count rather than closing the leak, and both degrade the feature for honest users. Use them next to the policy, not instead of it.
Functional cloning, briefly
Distillation needs only generated text, so output restrictions do not stop it. Detection of coverage-shaped query patterns and contractual terms carry most of the defence, covered in the extraction deep dive.
Two LLM-specific additions are worth making. Do not return hidden reasoning or intermediate tool traces unless the product needs them, because they are high-value training signal. And keep an evaluation set of canary prompts with distinctive answers. If a competitor's model reproduces them, you have evidence that is far stronger than general similarity. Output provenance techniques, discussed in AI provenance and watermarking, serve the same purpose.
Weight custody: the larger prize
None of the above matters if the checkpoint can be copied. Weights are the most valuable artifact a model team produces. They are also large, frequently moved and readable by many systems. RAND's report Securing AI Model Weights catalogues 38 distinct attack vectors and defines five security levels, SL1 to SL5, matched to attackers ranging from opportunistic criminals to the most capable nation states. The useful habit for any team: decide which attacker you are defending against and check every hop of the artifact path against it.
- Checkpoint store. A dedicated bucket or account for checkpoints, no public access settings anywhere in that account, object reads limited to the training and promotion roles, and encryption with a customer-managed key whose administrators are not the same people as its users. See secure inference for the serving-side half of this.
- Registry. Promote models by signed manifest (shard hashes, producing job, approver), and have inference hosts verify it before loading.
- Inference fleet. Weights sit in plaintext on serving hosts, so restrict shell access, disable core dumps and memory-dumping debug endpoints, and treat a compromised host as a weight-loss incident. LLM infrastructure security covers host hardening.
- Egress. Checkpoints are hundreds of gigabytes; egress limits and large-transfer alerts turn a silent copy into a noisy one.
Read-volume monitoring is cheap and catches both insiders and stolen credentials. A daily job comparing each principal's checkpoint reads with its own history flags the event that matters: a full model's worth of bytes read by someone who never read one before.
# Weight-custody alert: bytes read from checkpoint prefixes, per principal, versus its own baseline.
from collections import defaultdict
from statistics import median
def read_volume_alerts(access_log_rows, history, factor=5.0, floor_gb=50):
today = defaultdict(float)
for r in access_log_rows:
if r["op"] == "GET" and r["key"].startswith("checkpoints/"):
today[r["principal"]] += r["bytes"] / 1e9
alerts = []
for who, gb in today.items():
base = median(history.get(who, [0.0])) or 0.0
if gb >= floor_gb and gb > factor * max(base, 1.0):
alerts.append((who, round(gb, 1), round(base, 1)))
return alerts
Worked example: reviewing a fine-tuned model API
A team serves a domain fine-tune of an open-weight 8B model. It has a hidden size of 4096 and a vocabulary of about 128,000. The draft API offers top-20 logprobs, logit bias and prompt echo on every tier.
Step one is the audit. The local script finds the knee at 4096, but that number is already public for this base, so it reveals nothing new. The sensitive asset is the fine-tuned output behaviour, and full vectors would let a third party fingerprint the base and track every retrain. Step two: with echo on, every long prompt yields hundreds of logprob positions per call, and logit bias lets a caller choose which tokens become visible. Step three is the decision. The free tier gets no logprobs at all. Paid tiers get top-5, with logit bias allowed but never on the same request as logprobs, and echo is removed everywhere. A research tier gets top-20 under contract, with no bias combination. Step four, custody: the CI role that builds the inference image can read every checkpoint prefix, and so can three former contractors. The fix narrows the CI role, revokes the stale accounts and pages on read volume.
The custody findings were the bigger risk, which is the usual result of this review.
Failure modes
- Policy in the SDK instead of the gateway. Anyone who calls the raw HTTP endpoint skips it.
- Per-key limits only. An attacker spreads queries over many free accounts. Aggregate by organisation, payment instrument and network origin.
- Encrypted bucket, broad key policy. Encryption at rest means nothing if every role that can read the objects can also use the key.
Trade-offs
| Decision | More open | More closed |
|---|---|---|
| Logprob depth | Better calibration, research and tooling for customers | Fewer coordinates leaked per query |
| Logit bias | Constrained generation and classification use cases | Removes the attacker's way of choosing which tokens become visible |
| Checkpoint access | Fast iteration, any engineer can evaluate | Fewer people who can lose the weights |
What to do next
- Write down which of the four targets in the first table you are defending, and who owns each.
- Run the singular-value audit on your model and record what the knee reveals that is not already public.
- Inventory every endpoint that returns logprobs, accepts logit bias or echoes prompts, including batch and evaluation routes.
- Move the output policy into the gateway, tier it, and forbid logit bias on any request that also returns logprobs.
- List every principal that can read checkpoint objects or use the checkpoint key, and remove the ones without a current need.
- Sign promoted model manifests and make inference hosts verify them before loading.
- Turn on storage access logs for checkpoint prefixes and alert on unusual read volume per principal.
- Keep a canary prompt set with distinctive answers so you can test suspected clones.