When people say a model was stolen they usually mean one of two very different things. Either someone used the public API to train a cheaper model that imitates it, or someone walked off with the weights. Between those sits a third, less familiar case: the API's own output format giving away exact pieces of the model's architecture and parameters.

This article covers the second and third cases in detail and the first one briefly, because model extraction through the API already has its own deep dive on surrogate training, detection and deterrence. Here you will learn why a logit vector is a leak, how to measure your own endpoint, and which custody controls matter for checkpoints.

Advertisement

What can be stolen, and through which door

Model stealing is not one attack. The table separates the targets, because a control that stops one does nothing for another. Rate limits slow distillation but do nothing about an over-permissive checkpoint bucket.

TargetDoorWhat the attacker getsWho owns the defence
Behaviour (functional clone)Generated text at scaleA student model trained on your outputsAPI product and abuse teams
Architecture factsFull or partial logit vectorsHidden size, identity of the base modelAPI design
Exact final layerLogprobs plus logit biasThe output projection matrix, up to a symmetryAPI design
Full weightsStorage, hosts, peopleEverything, including future fine-tuningSecurity and infrastructure
Three ways a model leaves the building, and where each is defended1. Through the output interfaceClient queriesprompts, optionsAPI gatewaypolicy, quotasModelh = f(prompt)Logits = W hrank at most dLeaks: generated text (distillation), logprobs and logit bias (hidden size, final layer), timing2. Through the infrastructureTraining clustercheckpointsCheckpoint storebucket + KMSModel registrysigned artifactsInference fleetweights in memoryLeaks: over-broad read access, exported copies, debug dumps, compromised serving hosts3. Through people and processInsiders and contractorslegitimate access, wrong useThird partieseval vendors, partnersSupply chainbuild and CI credentialsDefences: output policy at the gateway (1), custody controls on every hop (2),least privilege, two-person export and read-volume alerts for all three
The output interface, the artifact path and the people around them are three separate attack surfaces with separate owners.

Why a logit vector is a leak: the softmax bottleneck

A decoder-only transformer ends the same way every time. The last layer produces a hidden vector h of size d, the hidden size, and an output projection W of shape V by d turns it into V logits, one per vocabulary token. V is large, often 32,000 to more than 200,000. d is much smaller, typically a few thousand.

That size gap is the whole story. Every logit vector the model can ever produce is W times some h, so all of them lie in a subspace of dimension at most d inside a V-dimensional space. Collect more than d full logit vectors from random prompts, stack them into a matrix and compute its singular values: about d of them are large and the rest collapse to numerical noise. The position of that drop is the hidden size. Log-probabilities are logits minus a per-prompt constant, so they span at most d + 1 dimensions, and the same trick works.

Two 2024 papers turned this into practical attacks on production APIs. Carlini and colleagues, in Stealing Part of a Production Language Model, recovered the hidden size and the output projection matrix (up to an unavoidable symmetry) of OpenAI's ada and babbage models for under $20, confirming hidden sizes of 1024 and 2048. They also recovered the hidden size of gpt-3.5-turbo and estimated that the full projection would cost under $2,000 in queries. Their key enabler was that the API combined logprobs with a logit bias parameter, so an attacker could push chosen tokens into the visible top-k and reconstruct full vectors piece by piece. Finlayson, Ren and Swayamdipta, in Logits of API-Protected LLMs Leak Proprietary Information, used the same low-rank structure to estimate gpt-3.5-turbo's embedding size at about 4096. They also showed that one full output can identify which model produced it. The affected vendors deployed mitigations after disclosure.

For a defender, the point is that this is a property of the architecture, not one vendor's bug. Any endpoint that returns enough of the distribution exposes it, and the cost falls as each query reveals more tokens.

Advertisement

Audit your own endpoint before someone else does

You do not need to attack anyone to know your exposure. Run the measurement on a model you own and you will see exactly what a full-logit API would reveal. The script below samples random prompts, takes the final-position logits and finds the largest drop between consecutive singular values.

# Audit what full logit vectors from a model you own reveal. Runs locally.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

name = "your-org/your-model"          # a checkpoint you control
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForCausalLM.from_pretrained(name, torch_dtype=torch.float32).eval()

d = model.config.hidden_size           # the number the attack would recover
n = d + 256                            # need more samples than the hidden size
rows = []
with torch.no_grad():
    for i in range(n):
        ids = torch.randint(0, tok.vocab_size, (1, 16))   # random prompts are enough
        logits = model(ids).logits[0, -1]                  # length V
        rows.append(logits)
L = torch.stack(rows)                                      # shape (n, V)
s = torch.linalg.svdvals(L)
ratio = s[:-1] / s[1:]
knee = int(torch.argmax(ratio[: n - 1])) + 1               # largest drop between singular values
print("configured hidden size:", d, " numerical rank (knee):", knee)
# Expect knee == d (d + 1 for log-probabilities).

On any standard decoder the knee lands on the configured hidden size, which makes the point concrete for reviewers. Then repeat the exercise against your own staging API with the options your customers actually have: top-k logprobs, logit bias, prompt echo, the embeddings endpoint. For each option, ask how many queries it takes to rebuild one full vector. If both logprobs and logit bias can be combined freely, you have the configuration the Carlini paper exploited.

Treat the result as an input to a risk decision. For a fine-tune of a well-known open-weight base, the hidden size is already public; what stays sensitive is the fine-tuned output layer and which base you use. For a model trained from scratch, the hidden size is a real secret, because it narrows down parameter count and compute.

An output-interface policy

Each API feature should go through an explicit decision. The table lists the common ones, what each reveals, and the usual mitigation.

FeatureWhat it revealsTypical policy
Full logprob vectorThe whole subspace, quicklyDo not offer it on any public tier
Top-k logprobsk coordinates per positionCap k, and count logprob tokens against quota
Logit biasLets the caller choose which coordinates become visibleNever combine with logprobs on the same request
Prompt echo with logprobsLogprobs for every input position, many vectors per callDisable
Embeddings endpointA different layer's geometrySeparate quotas and monitoring

The policy belongs in the gateway, not in client libraries, and it should be tiered. Researchers and enterprise customers can get more, under contract and with tighter monitoring. The sketch below shows the shape. Field names follow the common chat-completions style. Map them onto whatever your API actually exposes.

# Gateway policy: decide what a request may see, per tier, before it reaches the model.
from dataclasses import dataclass

@dataclass
class Tier:
    max_top_logprobs: int        # 0 = no logprobs at all
    allow_logit_bias: bool
    allow_bias_with_logprobs: bool
    daily_token_cap: int

TIERS = {
    "free":       Tier(0, False, False, 200_000),
    "standard":   Tier(5, True,  False, 5_000_000),
    "research":   Tier(20, True, False, 50_000_000),   # still never both together
}

def enforce(req: dict, tier: Tier, used_today: int) -> dict:
    k = int(req.get("top_logprobs") or 0)
    if k > tier.max_top_logprobs:
        raise PermissionError(f"top_logprobs > {tier.max_top_logprobs}")
    bias = req.get("logit_bias") or {}
    if bias and not tier.allow_logit_bias:
        raise PermissionError("logit_bias not available on this tier")
    if bias and k and not tier.allow_bias_with_logprobs:
        raise PermissionError("logit_bias cannot be combined with logprobs")
    if used_today >= tier.daily_token_cap:
        raise PermissionError("daily token cap reached")
    return {**req, "echo": False}             # never return prompt-token logprobs

Adding noise to logprobs or rounding them reduces precision but does not change the low-rank structure. Both raise the query count rather than closing the leak, and both degrade the feature for honest users. Use them next to the policy, not instead of it.

Functional cloning, briefly

Distillation needs only generated text, so output restrictions do not stop it. Detection of coverage-shaped query patterns and contractual terms carry most of the defence, covered in the extraction deep dive.

Two LLM-specific additions are worth making. Do not return hidden reasoning or intermediate tool traces unless the product needs them, because they are high-value training signal. And keep an evaluation set of canary prompts with distinctive answers. If a competitor's model reproduces them, you have evidence that is far stronger than general similarity. Output provenance techniques, discussed in AI provenance and watermarking, serve the same purpose.

Weight custody: the larger prize

None of the above matters if the checkpoint can be copied. Weights are the most valuable artifact a model team produces. They are also large, frequently moved and readable by many systems. RAND's report Securing AI Model Weights catalogues 38 distinct attack vectors and defines five security levels, SL1 to SL5, matched to attackers ranging from opportunistic criminals to the most capable nation states. The useful habit for any team: decide which attacker you are defending against and check every hop of the artifact path against it.

  • Checkpoint store. A dedicated bucket or account for checkpoints, no public access settings anywhere in that account, object reads limited to the training and promotion roles, and encryption with a customer-managed key whose administrators are not the same people as its users. See secure inference for the serving-side half of this.
  • Registry. Promote models by signed manifest (shard hashes, producing job, approver), and have inference hosts verify it before loading.
  • Inference fleet. Weights sit in plaintext on serving hosts, so restrict shell access, disable core dumps and memory-dumping debug endpoints, and treat a compromised host as a weight-loss incident. LLM infrastructure security covers host hardening.
  • Egress. Checkpoints are hundreds of gigabytes; egress limits and large-transfer alerts turn a silent copy into a noisy one.

Read-volume monitoring is cheap and catches both insiders and stolen credentials. A daily job comparing each principal's checkpoint reads with its own history flags the event that matters: a full model's worth of bytes read by someone who never read one before.

# Weight-custody alert: bytes read from checkpoint prefixes, per principal, versus its own baseline.
from collections import defaultdict
from statistics import median

def read_volume_alerts(access_log_rows, history, factor=5.0, floor_gb=50):
    today = defaultdict(float)
    for r in access_log_rows:
        if r["op"] == "GET" and r["key"].startswith("checkpoints/"):
            today[r["principal"]] += r["bytes"] / 1e9
    alerts = []
    for who, gb in today.items():
        base = median(history.get(who, [0.0])) or 0.0
        if gb >= floor_gb and gb > factor * max(base, 1.0):
            alerts.append((who, round(gb, 1), round(base, 1)))
    return alerts

Worked example: reviewing a fine-tuned model API

A team serves a domain fine-tune of an open-weight 8B model. It has a hidden size of 4096 and a vocabulary of about 128,000. The draft API offers top-20 logprobs, logit bias and prompt echo on every tier.

Step one is the audit. The local script finds the knee at 4096, but that number is already public for this base, so it reveals nothing new. The sensitive asset is the fine-tuned output behaviour, and full vectors would let a third party fingerprint the base and track every retrain. Step two: with echo on, every long prompt yields hundreds of logprob positions per call, and logit bias lets a caller choose which tokens become visible. Step three is the decision. The free tier gets no logprobs at all. Paid tiers get top-5, with logit bias allowed but never on the same request as logprobs, and echo is removed everywhere. A research tier gets top-20 under contract, with no bias combination. Step four, custody: the CI role that builds the inference image can read every checkpoint prefix, and so can three former contractors. The fix narrows the CI role, revokes the stale accounts and pages on read volume.

The custody findings were the bigger risk, which is the usual result of this review.

Failure modes

  • Policy in the SDK instead of the gateway. Anyone who calls the raw HTTP endpoint skips it.
  • Per-key limits only. An attacker spreads queries over many free accounts. Aggregate by organisation, payment instrument and network origin.
  • Encrypted bucket, broad key policy. Encryption at rest means nothing if every role that can read the objects can also use the key.

Trade-offs

DecisionMore openMore closed
Logprob depthBetter calibration, research and tooling for customersFewer coordinates leaked per query
Logit biasConstrained generation and classification use casesRemoves the attacker's way of choosing which tokens become visible
Checkpoint accessFast iteration, any engineer can evaluateFewer people who can lose the weights

What to do next

  1. Write down which of the four targets in the first table you are defending, and who owns each.
  2. Run the singular-value audit on your model and record what the knee reveals that is not already public.
  3. Inventory every endpoint that returns logprobs, accepts logit bias or echoes prompts, including batch and evaluation routes.
  4. Move the output policy into the gateway, tier it, and forbid logit bias on any request that also returns logprobs.
  5. List every principal that can read checkpoint objects or use the checkpoint key, and remove the ones without a current need.
  6. Sign promoted model manifests and make inference hosts verify them before loading.
  7. Turn on storage access logs for checkpoint prefixes and alert on unusual read volume per principal.
  8. Keep a canary prompt set with distinctive answers so you can test suspected clones.
Key takeaway: Model stealing has three doors, and each needs its own lock. Distillation goes through generated text and is handled by detection and contracts. Architecture and final-layer theft go through the softmax bottleneck, so whether an endpoint leaks is decided by the output interface: cap logprobs, never combine them with logit bias, drop echo, and measure the leak on your own model. Full-weight theft goes through storage, hosts and people, and it is by far the costliest loss. That makes checkpoint custody, signed promotion and read-volume alerts the highest-value controls most teams are missing.