A classic web service has roughly uniform request costs, so limiting requests per second limits load. An LLM endpoint does not. A one-line question with a short answer might occupy a GPU for a fraction of a second; a request that fills the context window and asks for the maximum output can hold memory and batch capacity for a minute and cost a thousand times more. An attacker who understands that asymmetry does not need a botnet. A few well-shaped requests per minute can starve every other tenant, and a leaked API key can empty a budget without degrading anything at all.
The OWASP Top 10 for LLM Applications 2025 captures this as LLM10, Unbounded Consumption, broadened from the Model Denial of Service entry of the earlier list to include denial of wallet and model extraction by volume. This page explains the cost model from first principles, lists the attack classes that exploit it, and builds a layered defence you can implement: token-based admission with reservation and settlement, engine limits, output and agent budgets, and metering that feeds abuse detection.
First principles: what one request costs
Serving a request has two phases. Prefill processes every prompt token in parallel to build the key-value cache; its compute grows with prompt length, and the attention part grows with its square, so it is compute-bound and fast per token. Decode then generates one token per step for each sequence in the batch; each step reads the model weights and the sequence's cache, so it is memory-bandwidth-bound and occupies a batch slot for as long as the output continues. Throughout, the sequence holds KV-cache memory proportional to prompt plus output tokens.
The cache is usually the binding resource. Per token it costs 2 (keys and values) × layers × KV heads × head dimension × bytes per value. For an illustrative model with 32 layers, 8 KV heads of dimension 128 and 16-bit values, that is 2 × 32 × 8 × 128 × 2 = 131,072 bytes, 128 KiB per token. One 32,768-token sequence then needs 4 GiB of cache. If the engine can spare 40 GiB for cache after the weights, ten such sequences fill it, and every other request waits or is preempted. The numbers change with the model; the shape does not. Paged KV cache explains how engines allocate and preempt that memory.
So the cost of a request is roughly a function of input tokens, output tokens, and for reasoning models hidden reasoning tokens, weighted by model size. That function, not the request count, is what an attacker maximises and what your limits must bound.
The attack classes
| Attack | Resource exhausted | Why request limits miss it | Primary control |
|---|---|---|---|
| Long-input flood | Prefill compute, KV-cache memory | Few requests, each near the context limit | Input token cap per tier; token-based rate limit |
| Maximum-output generation | Decode time and batch slots | Output length is chosen by the model and the client's max_tokens | Server-side output cap; reserve and settle |
| Reasoning inflation | Hidden reasoning tokens | Tokens are spent before any visible output | Reasoning budget where supported; cost-based limits |
| Injected decoys in retrieved content | Reasoning tokens of other users' requests | Attacker never calls your API directly | Treat retrieved text as untrusted; per-request budgets |
| Agent loops and tool fan-out | Model calls, tool quotas, downstream APIs | Each call is small; the loop is unbounded | Step, tool-call and wall-clock budgets per task |
| Concurrency flood and slow readers | Queue slots, connections, KV cache held by idle streams | Rate is legal; duration is the attack | Per-tenant concurrency; idle-stream timeout; cancel on disconnect |
| Prefix-cache busting | Cache hit rate, so prefill cost for everyone | Each request looks normal | Per-tenant cache partitioning or accounting; monitor hit rate |
| Stolen or leaked keys | Money (denial of wallet) | Traffic is authenticated | Spend caps, anomaly alerts, key scoping and rotation |
Two research results show how far this goes. Shumailov and colleagues' sponge examples (IEEE EuroS&P 2021) searched for inputs that maximise energy and latency of neural networks, including language models, showing that input content alone can move cost. The OverThink paper (arXiv 2502.02542, 2025) injects decoy problems, such as a Markov decision process to solve, into content a reasoning model retrieves; the authors report reasoning token counts inflated by up to 46 times on one benchmark while the visible answer stays correct, and found common prompt-injection filters largely ineffective. The attacker never needs an account on your service: they need their text in your retrieval corpus or a web page your agent reads.
The architecture: bound every resource at the layer that sees it
The edge handles volumetric network attacks exactly as for any service; see cloud DDoS protection. Identity attaches every request to a tenant and tier, because every later limit is per tenant; anonymous access to a model endpoint is an open budget. Cost admission estimates the request's token cost before it touches a GPU and reserves that many tokens from the tenant's budget. The concurrency layer caps simultaneous requests per tenant and gives every queued request a deadline, so a flood from one tenant queues behind itself rather than in front of everyone else. The engine enforces hard ceilings. The output guard caps generation and stops work when the client goes away. Agent budgets bound the loops that sit above single calls. Metering records what actually happened and feeds both settlement and abuse detection.
Token-based admission: reserve, then settle
A request-count limiter cannot see cost. A token bucket whose unit is weighted tokens can. The subtlety is that output length is unknown at admission, so reserve the worst case the request is allowed, then refund the difference when it finishes. Without settlement a tenant who always asks for large max_tokens and uses little is throttled unfairly; without reservation, one burst of long requests is admitted before any cost is recorded.
import time, threading
class TokenBudget:
"""Per-tenant bucket in weighted tokens: refills continuously, reserves up front, settles after."""
def __init__(self, rate_per_s, burst):
self.rate, self.burst = rate_per_s, burst
self.level, self.t = burst, time.monotonic()
self.lock = threading.Lock()
def _refill(self):
now = time.monotonic()
self.level = min(self.burst, self.level + (now - self.t) * self.rate)
self.t = now
def reserve(self, cost):
with self.lock:
self._refill()
if cost > self.level:
return False # reject or queue with a deadline; never admit on credit
self.level -= cost
return True
def settle(self, reserved, actual):
with self.lock:
self.level = min(self.burst, self.level + reserved - actual)
OUTPUT_WEIGHT = 4 # decode tokens cost more than prefill tokens on most deployments; measure yours
def admit(tenant, prompt_tokens, requested_max, limits):
max_out = min(requested_max, limits.max_output_tokens) # never trust the client's number
if prompt_tokens > limits.max_input_tokens:
raise ValueError("input too long for tier")
cost = prompt_tokens + OUTPUT_WEIGHT * max_out
if not tenant.budget.reserve(cost):
raise RuntimeError("429: token budget exhausted")
if not tenant.concurrency.acquire(timeout=limits.queue_deadline_s):
tenant.budget.settle(cost, 0)
raise RuntimeError("429: too many concurrent requests")
return cost, max_outAfter generation, call settle with the actual weighted cost and release the concurrency slot, in a finally block so an error or disconnect cannot leak a slot. Count tokens with the model's own tokenizer at the gateway; character-based estimates are easy to game with text that tokenises badly. Distributed rate limiter architecture covers keeping these buckets consistent across gateway replicas.
Engine limits and the serving scheduler
The inference engine is the last line and should never trust the gateway alone. In vLLM, --max-model-len caps prompt plus output length per sequence, --max-num-seqs caps sequences scheduled in one iteration, --max-num-batched-tokens caps tokens processed per iteration, and --gpu-memory-utilization sets how much GPU memory the engine may use. --enable-chunked-prefill splits long prompts into chunks bounded by the batched-token budget, so one huge prompt cannot stall the decode steps of everyone else, and --enable-prefix-caching reuses cache blocks for shared prefixes, which makes cache hit rate a resource worth monitoring.
Understand what happens at saturation. With continuous batching, when cache memory runs out the scheduler preempts sequences and must later recompute or swap them. An attacker filling the cache with long sequences does not just take their own share; they cause preemption churn that wastes work for everyone. Admission above the engine, as in admission control for LLM serving, should keep the cache below the point where preemption becomes frequent.
Output, reasoning and agent budgets
Cap output tokens on the server per tier, whatever the client sends. Stream responses, and when the client disconnects, abort the generation in the engine; an orphaned stream generating to nobody is the cheapest possible attack. Set an idle timeout for clients that open a stream and read slowly, since the sequence holds cache the whole time.
For reasoning models, use the provider's or engine's reasoning budget control where one exists, and in every case count hidden reasoning tokens in the tenant's budget; billing sees them even when the user does not. For agents, bound the task, not only the call: maximum model steps, maximum tool calls, maximum total tokens and a wall-clock deadline per task, enforced by the orchestrator. A loop that calls the model with the same failing tool result forty times is a denial of service you caused yourself, and the same budget stops an attacker who induces it through injected content.
Worked example: one tenant, three requests a minute
A tenant on a paid tier has 120,000 weighted tokens per minute, a burst of 120,000, a 16,000-token input cap and a 2,000-token output cap. An attacker holding that key sends one request every 20 seconds with a 15,000-token prompt and max_tokens of 100,000.
Without the design above: three requests per minute passes any request-rate limit; each asks for 100,000 output tokens, and the engine's context limit is the only brake. Each sequence can hold cache for minutes, and a handful of such keys fill the cache for the whole cluster.
With it: the output is clamped to 2,000, so the reserved cost is 15,000 + 4 × 2,000 = 23,000. Three a minute reserve 69,000 of the 120,000, so they are admitted, but each is now bounded at 17,000 tokens of sequence and holds cache for seconds rather than minutes. Sending faster only drains the burst: sustained, the budget admits about five such requests per minute. If the model stops after 300 tokens, settlement refunds 4 × 1,700 = 6,800. The attacker's maximum damage is now bounded by the tier they pay for, and other tenants' queue deadlines are unaffected because the concurrency cap keeps this tenant from monopolising slots. The metering layer records a tenant whose inputs are consistently near the cap with near-zero cache hits, a pattern worth a risk score.
Detection and monitoring
Watch per-tenant weighted tokens per second, queue wait, time to first token at the 99th percentile, KV-cache utilisation, preemption count, prefix-cache hit rate and cost per request. Alert on the distribution, not the mean: an attack shows up as a tenant whose cost per request sits at the policy ceiling. Signals like these feed abuse detection, which turns them into graduated throttling and blocking. For denial of wallet, set hard spend caps per key and per project and alert at fractions of them; a cap that only emails someone after the money is gone is not a control.
Failure modes and trade-offs
- Limits only at the public gateway. Internal agents and batch jobs call the model directly and bypass them. Put budgets where the model is called.
- Client retries amplify overload. A 429 without Retry-After, answered by aggressive retries, turns a spike into a storm. Send Retry-After and use jittered backoff in your own clients.
- Reserving without settling punishes honest users; settling without reserving admits bursts before cost is known. Do both.
- One global queue lets one tenant's flood delay everyone. Queue per tenant or per tier with weighted fair scheduling.
- Tight caps hurt legitimate long-context use. Tier them, and let trusted tenants raise them through a reviewed process rather than removing them.
What to do next
- Measure your cost model: GPU seconds per input token and per output token on your model and hardware, and set the output weight from data.
- Clamp max_tokens and input length on the server per tier and add a test that a client value above the cap is reduced.
- Replace request-count limits with weighted-token budgets that reserve up front and settle in a finally block.
- Add per-tenant concurrency caps and queue deadlines; verify generation is aborted when a client disconnects.
- Set engine ceilings such as max-model-len and max-num-seqs explicitly instead of accepting defaults.
- Give every agent task step, tool-call, token and wall-clock budgets, and set hard spend caps with alerts per key.