A classic web service has roughly uniform request costs, so limiting requests per second limits load. An LLM endpoint does not. A one-line question with a short answer might occupy a GPU for a fraction of a second; a request that fills the context window and asks for the maximum output can hold memory and batch capacity for a minute and cost a thousand times more. An attacker who understands that asymmetry does not need a botnet. A few well-shaped requests per minute can starve every other tenant, and a leaked API key can empty a budget without degrading anything at all.

The OWASP Top 10 for LLM Applications 2025 captures this as LLM10, Unbounded Consumption, broadened from the Model Denial of Service entry of the earlier list to include denial of wallet and model extraction by volume. This page explains the cost model from first principles, lists the attack classes that exploit it, and builds a layered defence you can implement: token-based admission with reservation and settlement, engine limits, output and agent budgets, and metering that feeds abuse detection.

Advertisement

First principles: what one request costs

Serving a request has two phases. Prefill processes every prompt token in parallel to build the key-value cache; its compute grows with prompt length, and the attention part grows with its square, so it is compute-bound and fast per token. Decode then generates one token per step for each sequence in the batch; each step reads the model weights and the sequence's cache, so it is memory-bandwidth-bound and occupies a batch slot for as long as the output continues. Throughout, the sequence holds KV-cache memory proportional to prompt plus output tokens.

The cache is usually the binding resource. Per token it costs 2 (keys and values) × layers × KV heads × head dimension × bytes per value. For an illustrative model with 32 layers, 8 KV heads of dimension 128 and 16-bit values, that is 2 × 32 × 8 × 128 × 2 = 131,072 bytes, 128 KiB per token. One 32,768-token sequence then needs 4 GiB of cache. If the engine can spare 40 GiB for cache after the weights, ten such sequences fill it, and every other request waits or is preempted. The numbers change with the model; the shape does not. Paged KV cache explains how engines allocate and preempt that memory.

So the cost of a request is roughly a function of input tokens, output tokens, and for reasoning models hidden reasoning tokens, weighted by model size. That function, not the request count, is what an attacker maximises and what your limits must bound.

The attack classes

AttackResource exhaustedWhy request limits miss itPrimary control
Long-input floodPrefill compute, KV-cache memoryFew requests, each near the context limitInput token cap per tier; token-based rate limit
Maximum-output generationDecode time and batch slotsOutput length is chosen by the model and the client's max_tokensServer-side output cap; reserve and settle
Reasoning inflationHidden reasoning tokensTokens are spent before any visible outputReasoning budget where supported; cost-based limits
Injected decoys in retrieved contentReasoning tokens of other users' requestsAttacker never calls your API directlyTreat retrieved text as untrusted; per-request budgets
Agent loops and tool fan-outModel calls, tool quotas, downstream APIsEach call is small; the loop is unboundedStep, tool-call and wall-clock budgets per task
Concurrency flood and slow readersQueue slots, connections, KV cache held by idle streamsRate is legal; duration is the attackPer-tenant concurrency; idle-stream timeout; cancel on disconnect
Prefix-cache bustingCache hit rate, so prefill cost for everyoneEach request looks normalPer-tenant cache partitioning or accounting; monitor hit rate
Stolen or leaked keysMoney (denial of wallet)Traffic is authenticatedSpend caps, anomaly alerts, key scoping and rotation

Two research results show how far this goes. Shumailov and colleagues' sponge examples (IEEE EuroS&P 2021) searched for inputs that maximise energy and latency of neural networks, including language models, showing that input content alone can move cost. The OverThink paper (arXiv 2502.02542, 2025) injects decoy problems, such as a Markov decision process to solve, into content a reasoning model retrieves; the authors report reasoning token counts inflated by up to 46 times on one benchmark while the visible answer stays correct, and found common prompt-injection filters largely ineffective. The attacker never needs an account on your service: they need their text in your retrieval corpus or a web page your agent reads.

Advertisement

The architecture: bound every resource at the layer that sees it

Layered defence: each layer bounds a different resource, and metering feeds back into every layerEdgeDDoS, WAF, TLSIdentityAPI key, tenant, tierCost admissionestimate, reserve tokensConcurrency + queueper tenant, deadlinesInference enginemax-model-len, max-num-seqsOutput guardmax tokens, cancel on disconnectAgent budgetsteps, tools, wall clockMetering and settlementactual tokens, GPU seconds, cost per tenantAbuse signalsrisk score, throttle, blockfeedbackRequests per minute alone bounds none of these: one request can cost a thousand times another
Each layer bounds a different resource. Volumetric traffic stops at the edge, cost at admission, duration and memory at the engine and output guard, loops at the agent budget; metering closes the loop.

The edge handles volumetric network attacks exactly as for any service; see cloud DDoS protection. Identity attaches every request to a tenant and tier, because every later limit is per tenant; anonymous access to a model endpoint is an open budget. Cost admission estimates the request's token cost before it touches a GPU and reserves that many tokens from the tenant's budget. The concurrency layer caps simultaneous requests per tenant and gives every queued request a deadline, so a flood from one tenant queues behind itself rather than in front of everyone else. The engine enforces hard ceilings. The output guard caps generation and stops work when the client goes away. Agent budgets bound the loops that sit above single calls. Metering records what actually happened and feeds both settlement and abuse detection.

Token-based admission: reserve, then settle

A request-count limiter cannot see cost. A token bucket whose unit is weighted tokens can. The subtlety is that output length is unknown at admission, so reserve the worst case the request is allowed, then refund the difference when it finishes. Without settlement a tenant who always asks for large max_tokens and uses little is throttled unfairly; without reservation, one burst of long requests is admitted before any cost is recorded.

import time, threading

class TokenBudget:
    """Per-tenant bucket in weighted tokens: refills continuously, reserves up front, settles after."""
    def __init__(self, rate_per_s, burst):
        self.rate, self.burst = rate_per_s, burst
        self.level, self.t = burst, time.monotonic()
        self.lock = threading.Lock()

    def _refill(self):
        now = time.monotonic()
        self.level = min(self.burst, self.level + (now - self.t) * self.rate)
        self.t = now

    def reserve(self, cost):
        with self.lock:
            self._refill()
            if cost > self.level:
                return False              # reject or queue with a deadline; never admit on credit
            self.level -= cost
            return True

    def settle(self, reserved, actual):
        with self.lock:
            self.level = min(self.burst, self.level + reserved - actual)

OUTPUT_WEIGHT = 4        # decode tokens cost more than prefill tokens on most deployments; measure yours

def admit(tenant, prompt_tokens, requested_max, limits):
    max_out = min(requested_max, limits.max_output_tokens)      # never trust the client's number
    if prompt_tokens > limits.max_input_tokens:
        raise ValueError("input too long for tier")
    cost = prompt_tokens + OUTPUT_WEIGHT * max_out
    if not tenant.budget.reserve(cost):
        raise RuntimeError("429: token budget exhausted")
    if not tenant.concurrency.acquire(timeout=limits.queue_deadline_s):
        tenant.budget.settle(cost, 0)
        raise RuntimeError("429: too many concurrent requests")
    return cost, max_out

After generation, call settle with the actual weighted cost and release the concurrency slot, in a finally block so an error or disconnect cannot leak a slot. Count tokens with the model's own tokenizer at the gateway; character-based estimates are easy to game with text that tokenises badly. Distributed rate limiter architecture covers keeping these buckets consistent across gateway replicas.

Engine limits and the serving scheduler

The inference engine is the last line and should never trust the gateway alone. In vLLM, --max-model-len caps prompt plus output length per sequence, --max-num-seqs caps sequences scheduled in one iteration, --max-num-batched-tokens caps tokens processed per iteration, and --gpu-memory-utilization sets how much GPU memory the engine may use. --enable-chunked-prefill splits long prompts into chunks bounded by the batched-token budget, so one huge prompt cannot stall the decode steps of everyone else, and --enable-prefix-caching reuses cache blocks for shared prefixes, which makes cache hit rate a resource worth monitoring.

Understand what happens at saturation. With continuous batching, when cache memory runs out the scheduler preempts sequences and must later recompute or swap them. An attacker filling the cache with long sequences does not just take their own share; they cause preemption churn that wastes work for everyone. Admission above the engine, as in admission control for LLM serving, should keep the cache below the point where preemption becomes frequent.

Output, reasoning and agent budgets

Cap output tokens on the server per tier, whatever the client sends. Stream responses, and when the client disconnects, abort the generation in the engine; an orphaned stream generating to nobody is the cheapest possible attack. Set an idle timeout for clients that open a stream and read slowly, since the sequence holds cache the whole time.

For reasoning models, use the provider's or engine's reasoning budget control where one exists, and in every case count hidden reasoning tokens in the tenant's budget; billing sees them even when the user does not. For agents, bound the task, not only the call: maximum model steps, maximum tool calls, maximum total tokens and a wall-clock deadline per task, enforced by the orchestrator. A loop that calls the model with the same failing tool result forty times is a denial of service you caused yourself, and the same budget stops an attacker who induces it through injected content.

Worked example: one tenant, three requests a minute

A tenant on a paid tier has 120,000 weighted tokens per minute, a burst of 120,000, a 16,000-token input cap and a 2,000-token output cap. An attacker holding that key sends one request every 20 seconds with a 15,000-token prompt and max_tokens of 100,000.

Without the design above: three requests per minute passes any request-rate limit; each asks for 100,000 output tokens, and the engine's context limit is the only brake. Each sequence can hold cache for minutes, and a handful of such keys fill the cache for the whole cluster.

With it: the output is clamped to 2,000, so the reserved cost is 15,000 + 4 × 2,000 = 23,000. Three a minute reserve 69,000 of the 120,000, so they are admitted, but each is now bounded at 17,000 tokens of sequence and holds cache for seconds rather than minutes. Sending faster only drains the burst: sustained, the budget admits about five such requests per minute. If the model stops after 300 tokens, settlement refunds 4 × 1,700 = 6,800. The attacker's maximum damage is now bounded by the tier they pay for, and other tenants' queue deadlines are unaffected because the concurrency cap keeps this tenant from monopolising slots. The metering layer records a tenant whose inputs are consistently near the cap with near-zero cache hits, a pattern worth a risk score.

Detection and monitoring

Watch per-tenant weighted tokens per second, queue wait, time to first token at the 99th percentile, KV-cache utilisation, preemption count, prefix-cache hit rate and cost per request. Alert on the distribution, not the mean: an attack shows up as a tenant whose cost per request sits at the policy ceiling. Signals like these feed abuse detection, which turns them into graduated throttling and blocking. For denial of wallet, set hard spend caps per key and per project and alert at fractions of them; a cap that only emails someone after the money is gone is not a control.

Failure modes and trade-offs

  • Limits only at the public gateway. Internal agents and batch jobs call the model directly and bypass them. Put budgets where the model is called.
  • Client retries amplify overload. A 429 without Retry-After, answered by aggressive retries, turns a spike into a storm. Send Retry-After and use jittered backoff in your own clients.
  • Reserving without settling punishes honest users; settling without reserving admits bursts before cost is known. Do both.
  • One global queue lets one tenant's flood delay everyone. Queue per tenant or per tier with weighted fair scheduling.
  • Tight caps hurt legitimate long-context use. Tier them, and let trusted tenants raise them through a reviewed process rather than removing them.

What to do next

  1. Measure your cost model: GPU seconds per input token and per output token on your model and hardware, and set the output weight from data.
  2. Clamp max_tokens and input length on the server per tier and add a test that a client value above the cap is reduced.
  3. Replace request-count limits with weighted-token budgets that reserve up front and settle in a finally block.
  4. Add per-tenant concurrency caps and queue deadlines; verify generation is aborted when a client disconnects.
  5. Set engine ceilings such as max-model-len and max-num-seqs explicitly instead of accepting defaults.
  6. Give every agent task step, tool-call, token and wall-clock budgets, and set hard spend caps with alerts per key.
Key takeaway: LLM denial of service exploits the gap between request count and request cost. Bound cost where it is incurred: identity and per-tenant weighted-token budgets that reserve and settle, concurrency caps with deadlines, server-side input and output clamps, engine ceilings, cancellation on disconnect, and task-level budgets for agents and reasoning. Meter everything and feed it back into abuse detection and spend caps, because the cheapest attack is a legitimate-looking key spending your money.