An LLM API key is a line of credit. Whoever holds it can spend your GPU budget, read whatever your retrieval layer exposes, and probe your model at your expense. Most of the incidents that hurt LLM operators are not clever prompt attacks; they are ordinary usage that is wrong in shape: a key leaked into a public repository and resold, an agent stuck in a retry loop burning output tokens all night, a customer scraping completions to train a competing model, or a compromised account suddenly calling the most expensive model from a new country. Each of these shows up first in usage telemetry, long before anyone reads a transcript.

This article is about the detectors that find those shapes: what to log, how to build a baseline for each key and tenant, why robust statistics beat the textbook mean and standard deviation, where exponentially weighted averages fail, which behavioural features catch what volume misses, and how to prove a detector works by injecting synthetic incidents. What happens after a detection, the risk scores and the graduated enforcement ladder, is covered in LLM abuse detection architecture; here we stay with the statistics and the plumbing that feed it.

What to log

Anomaly detection is only as good as the event it reads. Emit one structured record per model call at the gateway, not inside the model server, so that every route, retry and cached response is counted the same way. A useful minimum, where api_key_id is an identifier and never the secret itself, and prompt_fingerprint is a similarity hash rather than the text:

{
  "ts": "2026-10-06T14:03:11.402Z",
  "api_key_id": "key_7f3a",
  "tenant_id": "acme",
  "principal": "svc-support-bot",
  "model": "large-v3",
  "input_tokens": 1830, "output_tokens": 412, "cached_input_tokens": 1200,
  "status": 200, "latency_ms": 2140, "stream": true,
  "src_ip_asn": 16509, "src_country": "DE", "user_agent_family": "python-httpx",
  "prompt_fingerprint": "simhash:9c41e07a5b",
  "tool_calls": 2, "moderation_flag": null
}

Three choices here matter. First, log identifiers and hashes, not prompt text: the detection pipeline should not become a second copy of every customer conversation, and the retention rules in audit logging for LLM systems apply to it. A similarity hash such as SimHash over normalised prompt text lets you measure how varied a key's prompts are without storing them. Second, record input and output tokens separately, because cost and abuse live mostly in output tokens and in long uncached inputs. Third, record the network origin as an autonomous system number and country rather than a raw IP, which is enough to see a key move from a cloud provider in one region to a residential proxy network elsewhere.

Aggregate those events into per-entity time buckets. Hourly buckets per API key, per tenant and per principal are a good default: fine enough to catch an overnight runaway within an hour, coarse enough that each bucket has a stable count. For each bucket keep request count, input and output tokens, distinct prompt fingerprints, distinct countries and ASNs, error rate, and the share of calls to each model.

The pipeline

Usage anomaly detection: telemetry in, triaged alerts outLLM gatewayone event per callEvent streamkey, tenant, tokens, modelAggregatorper entity per hourFeature storehistory, baselinesVolume detectorsseasonal median + MAD, EWMABehaviour detectorsdiversity, geo, model mixBudget detectorscumulative spend vs capfeaturesCorrelatordedupe, join detectors, attach evidencePage on-callstolen key, runaway costAbuse queueextraction, policy probingOwner noticeexpected growth, bugs
Events leave the gateway, are rolled up per entity and hour, scored by three detector families, correlated into one alert per incident, and routed by type.

Keep three families of detectors and do not merge them into one model too early. Volume detectors ask whether this entity is using much more than it normally does at this time. Behaviour detectors ask whether the usage looks different in kind, even at normal volume. Budget detectors ask whether cumulative spend is on track to break a limit, which catches slow drift that per-bucket tests never see. The correlator joins them, so a stolen key that triggers all three becomes one page with three pieces of evidence instead of three pages.

Robust baselines

LLM usage is spiky and seasonal: a support bot peaks at 14:00 on weekdays and idles on Sunday night. Comparing this hour with the last hour therefore fires constantly. Compare it instead with the same hour of the same weekday over the past several weeks. This is the seasonal baseline that infrastructure monitoring uses too, described in metrics anomaly detection.

The second decision is the statistic. The usual z-score, (x - mean) / standard deviation, is fragile because past incidents live in the history. One earlier runaway inflates both the mean and the deviation so much that the next runaway looks normal. The median and the median absolute deviation (MAD) ignore a few extreme points. Multiplying MAD by 1.4826 puts it on the same scale as a standard deviation for normally distributed data, so thresholds read the same way.

import statistics

def robust_z(history, x, floor=1.0):
    """How many robust standard deviations x sits above the median of history."""
    med = statistics.median(history)
    mad = statistics.median(abs(h - med) for h in history)
    scale = max(1.4826 * mad, floor)      # 1.4826 makes MAD match sigma for normal data
    return (x - med) / scale

class EWMA:
    """Exponentially weighted mean and variance; flags points far from the running mean."""
    def __init__(self, alpha=0.1, warmup=24):
        self.alpha, self.warmup, self.n = alpha, warmup, 0
        self.mean, self.var = 0.0, 0.0

    def score(self, x):
        self.n += 1
        if self.n == 1:
            self.mean = x
            return 0.0
        sd = max(self.var ** 0.5, 1e-9)
        z = (x - self.mean) / sd if self.n > self.warmup else 0.0
        d = x - self.mean
        self.mean += self.alpha * d
        self.var = (1 - self.alpha) * (self.var + self.alpha * d * d)
        return z

The floor argument matters more than it looks. A key with perfectly flat usage has a MAD of zero, and any change would score infinity. Set the floor in the unit you are measuring, for example a few thousand tokens, so that tiny keys do not page anyone for going from 3 calls to 9.

Worked example

Take one API key whose output tokens in the 14:00 hour on the last eight Tuesdays were, in thousands: 41, 38, 45, 52, 40, 47, 39, 44. The median is 42.5 and the MAD is 3.0, giving a robust scale of about 4.45 thousand tokens.

Today at 14:00 (k tokens)Robust zReading
583.48busy day; log it, no alert
757.31notify owner: new feature or a bug
41082.63page: runaway loop or stolen key

Now suppose one of those eight weeks contained an earlier incident, so the history is 41, 38, 45, 52, 40, 47, 39, 300. With mean and standard deviation (75.2 and 90.9) the 410 hour scores only 3.68 and slips under a threshold of 4. The robust score is 61.88. This is the most common way naive detectors go blind: every incident you fail to remove from history makes the next one harder to see.

The EWMA detector has the opposite weakness. It adapts quickly, which is good for keys with no long history, but adaptation is exactly what a careful abuser exploits. Feeding it 48 hours of noisy usage around 100 and then a ramp that adds 12 per hour, reaching 244 after twelve hours, the highest z it ever reports is 3.9: the running mean chases the ramp and the alarm never fires at 4. A frozen baseline catches the same ramp: scoring each ramp hour with robust_z against the first 48 hours gives 3.14 in the first hour and 6.29 in the second. That is why you run both, and why budget detectors exist at all.

Behavioural signals

Volume is not enough, because several important patterns sit at normal volume:

  • Credential theft. A key that has called from one ASN in one country for months starts calling from a residential proxy range in several countries, often with a different client library. Score the share of traffic from ASNs and countries never seen for that key in the last 30 days.
  • Model extraction and distillation scraping. The prompts become unusually diverse and systematic: many distinct fingerprints, near-template prompts with one slot varied, high output-to-input ratio, sustained near the rate limit. Track distinct fingerprints per thousand calls and the entropy of the fingerprint distribution; see model extraction for the attack itself.
  • Policy probing. A burst of moderation flags or refusals from one principal, spread over paraphrases of the same request, is a jailbreak search in progress.
  • Model-mix shift. A key that used a small model for classification suddenly sends everything to the largest model. Cost per call jumps even when call counts do not.
  • Runaway agents. Repeated near-identical prompts with growing context, rising tool-call counts and a falling cache-hit ratio usually mean a loop, not an attacker.

For multivariate shapes, an isolation forest over a dozen per-key hourly features is a reasonable second stage, but treat its score as a ranking for human review rather than a paging signal: it has no notion of which direction is bad, so a key that suddenly becomes very quiet scores as anomalous too.

Proving the detectors work

You cannot measure recall from real incidents alone, because there are too few and you only know about the ones you caught. Instead, replay last month's real aggregated telemetry and inject synthetic incidents into it: a 20x spike on a random key, a two-day ramp, a country switch, a fingerprint-diversity jump at constant volume. Then count how many the detectors find and how fast.

def inject_and_score(series_by_key, detector, scenarios, threshold=4.0):
    """series_by_key: {key: [hourly values]}; scenarios: functions that mutate a copy."""
    found, delays, false_alarms, clean_hours = 0, [], 0, 0
    for key, series in series_by_key.items():
        for make_incident in scenarios:
            mutated, start = make_incident(list(series))
            hits = [t for t in range(start, len(mutated))
                    if detector(mutated[:t], mutated[t]) > threshold]
            if hits:
                found += 1
                delays.append(hits[0] - start)
        clean_hits = sum(detector(series[:t], series[t]) > threshold
                         for t in range(168, len(series)))
        false_alarms += clean_hits
        clean_hours += max(len(series) - 168, 0)
    total = len(series_by_key) * len(scenarios)
    key_days = max(clean_hours / 24, 1e-9)
    return {"recall": found / total, "median_delay_h": sorted(delays)[len(delays) // 2]
            if delays else None, "false_alarms_per_1k_key_days": 1000 * false_alarms / key_days}

Report three numbers per detector and scenario: recall, median time to detection, and false alarms per thousand key-days on clean data. The last one drives everything else. With a hundred thousand active keys, a false-positive rate of one in a thousand key-hours is 2,400 alerts a day, which nobody will read. Tune thresholds against the alert budget your team can actually triage, then see what recall that buys.

Failure modes

  • Cold start. New keys have no history. Score them against a cohort baseline (same plan, same model, first week of life) until they have four weeks of their own.
  • Expected growth looks like attack. A customer launch is a genuine 10x. Let owners pre-register expected changes and suppress volume alerts, never behaviour alerts, for the window.
  • Aggregation hides spread-out abuse. An attacker using fifty keys each at normal volume beats per-key detectors. Run the same detectors per tenant, per payment method and per ASN.
  • Detection without a fast lever. An alert that takes a ticket and two days to act on is a report. Pair the paging tier with an automatic, reversible action, such as dropping the key to a low rate limit; see rate limiting for LLM endpoints.
  • Pipeline lag. If aggregation runs hourly in batch, a runaway costs two hours before anyone looks. Keep a streaming budget check on the gateway path.

Trade-offs

Robust seasonal baselines are cheap, explainable and resistant to contamination, but they need weeks of history and react within the bucket size at best. EWMA reacts quickly with little history and loses slow ramps. Budget detectors are blunt but impossible to evade by pacing. Behavioural features catch theft and extraction at normal volume, at the cost of more state per key. Multivariate models find shapes nobody wrote a rule for, but cannot explain themselves to the on-call engineer. The practical answer is a small ensemble with one clear owner per alert type.

What to do next

  1. Add a structured per-call usage event at the gateway with key, tenant, model, token counts, ASN, country and a prompt similarity hash.
  2. Build hourly per-key and per-tenant aggregates with at least eight weeks of retention.
  3. Implement the robust seasonal z-score with a unit-aware floor and run it in shadow for two weeks, logging what it would have alerted on.
  4. Add a cumulative daily budget check on the gateway path for every key.
  5. Add two behavioural features: new-ASN share and distinct fingerprints per thousand calls.
  6. Build the injection harness and record recall, delay and false alarms per detector.
  7. Set thresholds from your triage budget, wire the paging tier to an automatic rate-limit drop, and review every page in a weekly retrospective.
Key takeaway: Detect LLM usage anomalies from structured gateway telemetry, not transcripts. Compare each key and tenant with its own same-hour history using the median and MAD, back that up with fast EWMA, budget and behavioural detectors, and prove recall by injecting synthetic incidents into replayed data before you trust any threshold.