Most abuse of a language model API has one thing in common: somebody other than the paying customer is spending money that the customer, or you, will be billed for. A leaked key is resold behind a proxy. A free trial is opened two hundred times with throwaway emails. A script pins every request to your most expensive model with the longest output it can get. An agent loop retries itself until a budget runs out. These attacks often look harmless when you inspect a single request, but they all change how much is being spent, how fast, and on what.
Cost-based abuse detection uses that fact directly. You price every request as it settles, add the spend up per entity, and look for spend that is too fast, has the wrong shape, or is tied to entities that should not be connected. It complements content classifiers and rate limits, because its main signal is the one the attacker is trying to maximise. This article builds the pricer, the spend ledger, four detectors with working code, a leaked-key incident, failure modes and a checklist.
Why cost is the signal
Request counts are a poor proxy for harm in an LLM service, because what one request costs can vary by about four orders of magnitude. A short classification call on a small model might cost a hundredth of a cent. A long-context call to a large model that asks for the maximum output, with reasoning enabled, can cost tens of cents. A limit of 60 requests per minute therefore allows anything from a few cents to several dollars per minute, and attackers who are paid in your compute know which end of that range to aim for. The rate limiting deep dive covers how to make limits token-weighted. Detection is the layer above that: limits say how much is allowed, detection asks whether what is happening is normal for this entity.
Most money-driven abuse falls into a small number of families, and each one leaves a different mark on the spend data:
| Family | What the attacker gains | Cost signature |
|---|---|---|
| Stolen key resale (often called LLMjacking) | Free model access sold on to others | Sudden step up in spend, new source networks, all-hours traffic, model mix swings to premium |
| Trial and free-tier farming | Free credit, many times over | Many young accounts, each spending close to its credit cap, linked by card, device or network |
| Denial of wallet | Harm to you or a customer | Output-heavy requests at max_tokens, cache-busting prefixes, long contexts with low value |
| Distillation and scraping | Training data from your outputs | High, steady output volume, very varied prompts, no follow-up turns |
| Runaway automation | Nothing (often a bug, sometimes induced) | Spend climbs while cost per successful outcome explodes |
Pricing requests and building the spend ledger
Detection is only as good as the cost figure it uses. Provider invoices and billing exports arrive hours or days late, so they are useful for reconciliation but not for detection. Instead, price every request at the gateway as soon as its token counts are final, using a versioned price table that you keep yourself. Record the input tokens, the cached input tokens, the output tokens (including hidden reasoning tokens where the model bills them), the model, and the price table version, so you can reprice past events when prices change.
Then aggregate along every dimension an attacker could hide behind: per key for a leaked credential, per account and organisation for a compromised customer, and per payment instrument, device fingerprint and network (ASN) for farming, where each account looks small. Keep 5-minute, 1-hour and 24-hour windows. The LLM FinOps supply chain article covers keeping this ledger accurate; for security it must also be written inline, because a detector on yesterday's export only reports damage already done.
Four detectors
Rate against the entity's own baseline. Spend varies a lot between customers, so a single global threshold either misses small accounts or pages you about every large one. Each entity needs its own baseline. Spend is heavy-tailed, so model the logarithm of hourly spend with an exponentially weighted mean and variance. Set a floor on the variance so a very steady account does not alert on a 10% wobble, and freeze the baseline while a score is high so the attack does not teach the model that attack-level spend is normal.
Burn horizon. Divide the remaining credit or budget by the current spend rate. If an account will empty a $500 prepaid balance in three hours when it normally takes three months, that is worth acting on even before the anomaly score is confirmed, and the explanation is one the customer understands straight away.
Cost shape. Total spend can stay flat while its make-up changes. Track the share of spend on each model, the ratio of output to input tokens, the cache-hit share of input tokens, the share of requests that reach max_tokens, and how spend is spread across hours of the day. A proxy reselling a key shows almost no prompt-cache reuse (its users do not share your customer's system prompt), runs around the clock, and favours the most capable model. Distillation shows a high output-to-input ratio and almost no multi-turn sessions.
Cost per outcome. Where the product has a success event (a task completed, a document processed), divide spend by successes. Runaway agents and amplification attacks push this ratio up even when traffic looks normal.
Linkage. Farming defeats per-account detectors on purpose. Join accounts that share a payment fingerprint, a device or browser fingerprint, or a narrow network range, and score the cluster's total spend against what a real customer of that age would spend. The abuse detection architecture article describes the entity graph and risk scoring in general. The cost-specific point is that clusters should be ranked by spend, so review time goes where the money is.
The detectors in code
The core is small. The pricer turns token counts into money, and the baseline keeps two numbers per entity. This is the version we tested; the prices are illustrative and not any provider's real list.
from dataclasses import dataclass
import math
PRICE = { # illustrative USD per million tokens; load the real, versioned table
"small": {"in": 0.20, "cached_in": 0.05, "out": 0.80},
"large": {"in": 3.00, "cached_in": 0.75, "out": 15.00},
}
def request_cost(model, tokens_in, tokens_cached, tokens_out):
p = PRICE[model]
fresh = tokens_in - tokens_cached
return (fresh * p["in"] + tokens_cached * p["cached_in"]
+ tokens_out * p["out"]) / 1e6
@dataclass
class SpendBaseline:
"""EWMA of log hourly spend for one entity, with a variance floor."""
alpha: float = 0.05 # roughly a 20-hour memory
mean: float = 0.0
var: float = 0.25
n: int = 0
min_var: float = 0.04
def score(self, spend):
x = math.log1p(spend * 100) # cents, log-compressed
z = (x - self.mean) / math.sqrt(max(self.var, self.min_var))
return x, z
def update(self, x, z, z_freeze=4.0):
if self.n >= 24 and z > z_freeze: # do not learn the attack
return
a = self.alpha if self.n >= 24 else 1.0 / (self.n + 1)
d = x - self.mean
self.mean += a * d
self.var = (1 - a) * (self.var + a * d * d)
self.n += 1
def hourly_check(b, spend, credit_left, burn_hours_alert=6.0):
x, z = b.score(spend)
alerts = []
if b.n >= 24 and z > 4.0:
alerts.append(f"spend z={z:.1f}")
if spend > 0 and credit_left / spend < burn_hours_alert:
alerts.append(f"credit exhausted in {credit_left / spend:.1f} h")
b.update(x, z)
return alertsFor example, a request with 12,000 input tokens (8,000 of them cached) and 900 output tokens on the illustrative large model costs $0.0315: $0.012 for fresh input, $0.006 for cached input and $0.0135 for output. Output is priced five times input, so max-output requests are the cheapest way for an attacker to burn your money. Cost-shape features are a single aggregation over the ledger:
SELECT entity_id,
SUM(cost_usd) AS spend_1h,
SUM(CASE WHEN model = 'large' THEN cost_usd END)
/ NULLIF(SUM(cost_usd), 0) AS premium_share,
SUM(tokens_out) / NULLIF(SUM(tokens_in), 0) AS out_in_ratio,
SUM(tokens_cached) / NULLIF(SUM(tokens_in), 0) AS cache_share,
AVG(CASE WHEN hit_max_tokens THEN 1.0 ELSE 0.0 END) AS max_tokens_share,
COUNT(DISTINCT src_asn) AS asn_count
FROM spend_ledger
WHERE ts >= now() - INTERVAL '1 hour'
GROUP BY entity_id;Compare each feature with the entity's own 28-day percentiles, not fixed constants: a cache share falling from 0.7 to 0.02 matters only for an account that caches.
Worked example: a leaked key at 2 a.m.
A small analytics customer usually spends about $1.20 an hour, with a standard deviation of about a fifth of that, mostly on the small model, with a 70% cache share from a fixed system prompt, and almost nothing at night. Their key is committed to a public repository at 01:40. At 02:00 the key is spending $38 an hour, all of it on the large model, the cache share is close to zero, requests come from eleven networks the account has never used, and the prepaid balance is $120.
After 72 hours of normal history, the baseline code above returns a z-score of about 17.5 for the first attack hour, far above the threshold of 4, and a burn horizon of 3.2 hours. Because the update freezes at high z, the baseline mean stays where it was for every following attack hour instead of drifting up. The shape detector agrees on its own: premium share went from about 5% to 100%, and the cache share fell from 0.7 to near zero. With two independent detectors in agreement, the policy goes straight to a soft cap: the key is held to three times its baseline hourly spend, the owner gets an email and an in-console banner with the evidence, and a one-click rotation link. The account itself is not suspended, so the customer's production traffic, which uses a different key, keeps running.
Without the detector, the hard cap would have stopped the loss only when the $120 balance ran out, about three hours later, leaving a dead integration at 05:00. With it, the loss was about one hour of attack spend.
From score to action
Responses should be graduated, as with any abuse system, but cost signals make the levels easier to set, because the harm is measured in money and can be bounded:
- Observe. Score and log. Entities with under 24 hours of history use cohort priors (plan, region, signup channel).
- Notify. Tell the owner, with spend rate, burn horizon, model and key.
- Soft cap. Hold the key or cluster to a multiple of its baseline, with an error that names the cause.
- Step up. Require re-verification or key rotation to lift the cap.
- Suspend. Only for confirmed farming or clearly stolen keys; keep it reversible and record who approved it.
Hard budget caps sit outside this ladder, enforced at admission as in the LLM denial-of-service article. They must hold even when every detector is wrong.
Failure modes
- Stale price table. A model's price changes, or a new model launches, and the pricer still uses old numbers or prices the new model at zero. Fail closed: an unknown model gets the highest price in its family, and an alert fires when billed and computed spend drift more than a few percent apart.
- Slow-ramp poisoning. An attacker who adds 15% a day stays under any EWMA z threshold, because the baseline follows them up. Add a long-window check (this week against the 28-day median) and an absolute ceiling per plan.
- Missing reasoning tokens. If hidden reasoning tokens are billed but not counted in the gateway, both cost and output ratio are understated for exactly the requests that cost the most.
- Legitimate spikes. Product launches, backfills and month-end batch jobs look like attacks. Let customers declare expected events with a budget and an end time, and lower the alert weight inside that window instead of turning it off.
- Shared keys and seasonality. One key per company hides which workload changed, so push customers toward one key per workload. A weekday-only business makes every Monday look like a spike; keep hour-of-week baselines.
- Alert fatigue. A z threshold of 3 on hourly data for 100,000 entities produces hundreds of false alerts per hour. Require two independent detectors before acting automatically, and rank the review queue by dollars at risk.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Inline pricing vs billing export | Detection within minutes | You maintain a price table and reconcile it |
| Per-entity baseline vs global threshold | Sensitive for small accounts, quiet for big ones | State per entity; cold start needs cohort priors |
| Freeze baseline when z is high | The attack does not become normal | A genuine step change needs a manual reset or a declared event |
| Soft cap vs suspension | Stops the loss without breaking the customer | Some loss continues up to the cap |
| Linkage clustering | Finds farming that per-account rules miss | Fingerprinting has privacy and false-match costs |
| Cost per outcome | Catches runaway loops and amplification | Needs a success event from the product |
What to do next
- Price every request inline from a versioned table that includes cached and reasoning tokens. Reconcile it against the invoice daily.
- Write a spend ledger keyed by key, account, organisation, payment instrument and network, with 5-minute, 1-hour and 24-hour windows.
- Deploy the log-spend EWMA baseline with a variance floor and freeze-on-alert, and use cohort priors for new entities.
- Add a burn-horizon alert on prepaid and budgeted accounts.
- Compute cost-shape features (premium share, output/input ratio, cache share, max-tokens share, network count) and compare them with each entity's own percentiles.
- Cluster accounts by shared payment, device and network fingerprints, and rank the clusters by spend.
- Map detector agreement to the notify, soft-cap, step-up and suspend ladder, and keep hard caps independent of it.
- Replay one leaked-key and one farming scenario against staging every quarter, and check time to detection and dollars lost.