Prompt injection is an attack on your users through content they did not write. Abuse is different: the person sending requests is the adversary. They may be generating content your policies forbid, running jailbreak campaigns, harvesting outputs to train a copy of your model, farming free-tier accounts, running up cost on a stolen API key, or using your model as a component in spam and fraud. Content classifiers alone do not stop this, because they judge one request at a time, and abuse is a pattern across many requests, accounts and days.
This article describes an abuse detection architecture from first principles: what to detect, which entities to track, the split between fast inline checks and slower asynchronous detection, how to turn signals into a decaying risk score, how to act in graduated steps, how to keep false positives from destroying trust, and how to run the system. The code is a working sketch of the scoring core with a worked example.
What counts as abuse
| Class | What it looks like | Primary signals |
|---|---|---|
| Policy-violating generation | repeated requests for prohibited content | input and output classifier hits per entity |
| Jailbreak campaigns | templated prompts iterated until one works | template similarity, refusal-then-rephrase loops |
| Extraction and distillation | high-volume, diverse prompts, outputs stored | volume, prompt diversity, long outputs, low reuse |
| Account and key abuse | stolen keys, free-tier farming, resale | new geographies, sign-up velocity, shared payment or device |
| Resource abuse | cost exhaustion, maximum-length requests | cost per entity against baseline, token ratios |
| Downstream misuse | spam, phishing or fraud text at scale | near-duplicate outputs, templated personalisation |
| Agent and tool misuse | using tools to reach systems the user should not | tool-call patterns, denied permission checks |
Each class needs different evidence and a different response, so write the taxonomy down before building detectors. It becomes the label set for reviewers, the key for runbooks, and the dimension along which you report precision. Model extraction is covered in more depth in model extraction attacks.
The architecture: an inline path and an asynchronous path
Abuse detection splits into two paths with very different budgets. The inline path runs on every request and must add only milliseconds: authenticate the key, read the entity's current enforcement state, apply quotas and rate limits, run fast input classifiers, call the model, and run output classifiers before returning or streaming. The inline path makes per-request decisions and does not try to understand patterns.
The asynchronous path consumes an event log of every request's metadata and classifier results, computes features over time windows for each entity, runs detectors, maintains risk scores and decides actions. Its output is written back to an enforcement store that the inline path reads. This separation is the core design decision: expensive analysis such as clustering prompts across accounts or scoring with a large model never sits in the request path, and the inline path stays simple enough to be reliable.
Entities: who or what you are scoring
Abuse is attributed to entities, and there is more than one kind: API key, user account, organisation, payment instrument, device or client fingerprint, IP address and network (autonomous system). They form a graph. One organisation owns many keys; one payment card may fund many accounts; one residential proxy network may front many unrelated users.
Score each entity type separately and propagate carefully along edges. Evidence against a key should raise the owning account's score; evidence against several accounts sharing one card should raise the card's score and, through it, any new account that uses it. Propagate a fraction of the score rather than all of it, and never propagate from IP addresses alone to accounts: carrier-grade NAT, corporate proxies and university networks put thousands of innocent users behind one address. IP and network signals are useful as features, not as identities. The data model is also a tenancy boundary, so read tenant isolation for LLM systems before letting one customer's signals affect another's.
Features and detectors
Features are computed per entity over sliding windows, typically one hour, one day and one week, with a baseline for comparison. Useful families are:
- Volume and cost: requests, input and output tokens, and spend, compared with the entity's own history and with its cohort (new free-tier accounts, established enterprise keys).
- Content outcomes: rates of input and output classifier hits and of model refusals. A single hit means little; a sustained rate, or a refusal followed by rephrased attempts, means a lot. See content moderation for LLMs for the classifiers themselves.
- Similarity: near-duplicate prompts within an account (MinHash or embedding clusters) and similarity to known jailbreak templates, which catches variants that exact signatures miss.
- Account context: account age, verification level, sign-up velocity from the same payment instrument or device, and sudden changes in geography or client.
- Shape: ratio of output to input tokens, prompt diversity across topics, and time-of-day regularity typical of scripts.
Detectors come in three kinds and you need all of them. Rules encode known patterns and are easy to explain. Supervised models learn from reviewer labels and catch combinations rules miss. Unsupervised methods, such as clustering and anomaly scores, find new campaigns nobody has labelled yet. Every detector emits a named signal with a reference to its evidence, never just a number.
Turning signals into a risk score
Signals arrive at different times, for different reasons, with different strength. A per-entity score that adds weighted evidence and decays over time handles this simply: recent, repeated evidence pushes the score up; an account that stops misbehaving drifts back down without anyone clearing it by hand. The ladder maps score ranges to actions.
import time
from dataclasses import dataclass, field
HALF_LIFE_S = 6 * 3600 # evidence loses half its weight every six hours
WEIGHTS = {
"jailbreak_template_match": 15,
"output_policy_block": 10,
"input_policy_block": 5,
"near_duplicate_burst": 8,
"cost_spike": 12,
"new_account_high_volume": 10,
}
LADDER = [(20, "log"), (40, "warn"), (60, "throttle"),
(80, "restrict"), (95, "suspend_pending_review")]
@dataclass
class EntityRisk:
score: float = 0.0
updated: float = field(default_factory=time.time)
evidence: list = field(default_factory=list)
def _decay(self, now):
self.score *= 0.5 ** ((now - self.updated) / HALF_LIFE_S)
self.updated = now
def add(self, signal, now, ref):
self._decay(now)
self.score = min(100.0, self.score + WEIGHTS[signal])
self.evidence.append((now, signal, ref)) # keep why, not just how much
def current(self, now):
self._decay(now)
return self.score
def action_for(score):
action = "allow"
for threshold, name in LADDER:
if score >= threshold:
action = name
return actionThe weights and half-life are starting points to be tuned on labelled data, not recommendations. The evidence list matters as much as the score: every enforcement action must be explainable to a reviewer and, often, to the customer.
Worked example
An account created yesterday sends three prompts that match known jailbreak templates within a few minutes. Each adds 15 points; decay over minutes is negligible, so the score reaches 45 and the ladder action becomes warn, which shows the user a policy notice. One of the attempts produces an output that the output classifier blocks, adding 10 for 55, still a warning. The account's hourly spend then jumps far above its cohort baseline, adding 12 for 67, which crosses 60 and moves the account to throttle: lower rate limits and a smaller maximum output length. Nothing reaches suspension without a human, because the top rung is defined as a review queue, not an automatic ban.
Now suppose the behaviour stops. With a six-hour half-life, after twelve hours the score is 67 times 0.25, about 17, below the first threshold, and the throttle lifts. If the same payment card funds three more new accounts that start the same pattern, the card's score rises through propagation, and the fourth account starts at warn on its first request. That is how the system follows a campaign rather than a single account.
Graduated enforcement
Actions should escalate in steps, each proportionate and reversible: log only; warn the user; add friction such as re-verification; throttle rates, lengths or concurrency; restrict access to specific models, tools or features; suspend pending review; and terminate after review. Prefer actions that raise the attacker's cost without destroying a legitimate user's work. Throttling a compromised key limits damage while the owner is contacted; banning it immediately may break a customer's production system over a leaked credential.
Enforcement decisions are written to a small, fast store keyed by entity, which the inline path reads on every request. Record every action with its evidence, the policy version and whether a human approved it, in logs designed for audit, as described in audit logging for LLM systems. Provide an appeal path and measure how often appeals succeed. Some content carries legal duties regardless of your own policy; in the United States, for example, providers must report apparent child sexual abuse material to NCMEC under 18 U.S.C. 2258A. Map those obligations with counsel and build them into runbooks rather than improvising.
False positives and the base-rate problem
Abuse is rare, which makes even accurate detectors produce mostly false alarms. Take one million active accounts of which 0.1 percent, or 1,000, are abusive. A detector that catches 99 percent of them and wrongly flags only 1 percent of legitimate accounts flags 990 abusers and 9,990 legitimate accounts. Precision is 990 out of 10,980, about 9 percent: eleven of every twelve flags are wrong.
The architecture answers this in three ways. Combine independent signals into a score so no single detector triggers a strong action. Tie action severity to precision: low-precision signals may only log or warn, and only high-precision combinations may throttle or restrict. And measure precision per rung by sending samples of each action to reviewers, running new detectors in shadow mode before they act, and keeping a small holdout to estimate what was missed.
Failure modes
- Adaptation. Attackers rotate templates, spread load across accounts and slow down to stay under thresholds. Static signatures decay within weeks; similarity and clustering age better.
- Shared infrastructure. Scoring IPs as identities punishes everyone behind a proxy or NAT.
- Feedback loops. Training only on what the current system flagged teaches models to find more of the same and nothing new; sample unflagged traffic for review.
- Reviewer overload. A queue that grows faster than it is worked means suspensions linger for innocent users; cap the rate of review-requiring actions.
- Enforcement lag. Hourly batch scoring lets a stolen key spend for an hour; add streaming counters for cost and volume.
- Privacy drift. Retaining full prompts for detection expands what you hold about users. Store features and hashes where possible, restrict access, and set retention limits.
Operating the system
Run abuse detection as a product with owners, an on-call rotation and service levels: time from signal to action, review queue age, appeal turnaround, and precision per rung. Write a runbook for each taxonomy class that says what evidence confirms it, which actions are allowed, who approves them, and what is communicated to the customer. Review weights and thresholds monthly against labelled outcomes, and re-run detectors against recent red-team output. Keep abuse detection as one layer among several, as described in defense in depth for LLM applications.
What to do next
- Write a taxonomy of abuse classes for your product and use it as the reviewer label set.
- Log request metadata and classifier results per request to an event stream keyed by key, account, organisation and payment instrument.
- Stand up the enforcement store and make the inline path read it on every request.
- Implement a decaying per-entity score with named, evidence-backed signals, starting in shadow mode.
- Define the enforcement ladder, and require human review for suspension and termination.
- Compute expected precision from your base rate before switching any detector from shadow to action.
- Add streaming cost and volume counters so stolen keys are throttled in minutes.
- Publish an appeal path and track appeal outcomes as a precision metric.