A request limit such as 60 per minute works for most APIs because their requests cost about the same. LLM requests do not. A 40-token question with a 20-token answer and a 100,000-token document with a 4,000-token summary are both one request, yet the second uses thousands of times more GPU time and holds a slot in the KV cache for far longer. An endpoint limited only by request count is either too strict for normal use or wide open to anyone who sends large requests.
This article is about designing the limits themselves. It covers which dimensions to limit, how to derive the numbers from measured GPU capacity, how to weight input, cached and output tokens, how to cap concurrent streams, how to share capacity fairly between tenants, and what the client side must do. The attack side, including reasoning inflation and slow-reader floods, is covered in LLM denial of service. The generic limiter algorithms, GCRA and composite limits in Redis, are in the rate limiter architecture deep dive.
Four limits, four resources
Each limit protects one resource. Enforce several at once, because each fails differently when used alone:
| Limit | Protects | Fails alone because |
|---|---|---|
| Requests per minute | gateway, auth, per-request overhead | ignores request size |
| Weighted tokens per minute | GPU compute time | ignores how long a stream holds memory |
| Concurrent streams | KV cache and batch slots | ignores how much compute each stream uses |
| Spend per day or month | the tenant's budget and yours | too coarse to protect latency |
Key every limit by the account or organisation as well as by API key. Otherwise a customer who creates twenty keys gets twenty times the quota. Keep a per-key limit beneath the organisation limit so one leaked key cannot spend the whole budget.
Deriving limits from GPU capacity
Limits should come from measurement. Load-test one replica with your real prompt and output length mix, at your latency targets for time to first token and time per output token. Record two numbers: prefill throughput in input tokens per second, and total decode throughput in output tokens per second across all running streams. The figures below are illustrative. Use your own.
| Quantity | Value | Note |
|---|---|---|
| Prefill throughput per replica | 20,000 tokens/s | measured at the latency target |
| Decode throughput per replica | 2,500 tokens/s | all streams combined |
| Output token weight | 8 | 20,000 / 2,500 |
| Replica capacity | 20,000 units/s | 1 unit = 1 uncached input token |
| Fleet of 8 replicas | 9.6M units/min | 8 x 20,000 x 60 |
| Sellable after 30% headroom | about 6.7M units/min | for spikes, failover and retries |
| Concurrent sequences per replica | 64 | from KV-cache size at your context mix |
Decode tokens usually cost more GPU time than prefill tokens. Prefill processes many tokens in one parallel pass, while decode produces one token per step per stream and is limited by memory bandwidth. Your measured ratio sets the output weight. If the engine has prefix caching, cached input tokens skip most of the prefill work, so give them a weight below 1. Measure the saving instead of guessing it, and check whether caching works across tenants or only within one.
Hidden reasoning tokens are output tokens. Count them at the output weight even though the user never sees them. Otherwise a reasoning model consumes far more capacity than your limits record.
Per-tenant limits follow from the sellable total. The sum of guaranteed rates must fit inside it. Burst allowances can exceed it, because tenants rarely peak together, but they need the fair queue described below as a backstop. Concurrency works the same way: 8 replicas at 64 sequences each gives 512 slots, and per-tenant stream caps should leave room for the rest of the fleet.
Admission at the gateway
Output length is unknown when a request arrives. The standard fix is to reserve the maximum and refund the difference at the end, explained in detail in the denial-of-service article. Here is how it fits with the weights and the stream cap at a gateway. The bucket is in-process for clarity. In production it sits in a shared store, as in the limiter deep dive.
import asyncio, math, time
W_IN, W_CACHED, W_OUT = 1.0, 0.25, 8.0 # from your own measurements
class Bucket:
def __init__(self, units_per_min, burst):
self.rate, self.cap = units_per_min / 60.0, burst
self.level, self.t = burst, time.monotonic()
def _refill(self):
now = time.monotonic()
self.level = min(self.cap, self.level + (now - self.t) * self.rate)
self.t = now
def try_take(self, n):
"""Take n units, or return the seconds until n will be available."""
self._refill()
if self.level >= n:
self.level -= n
return 0.0
return (n - self.level) / self.rate
def give_back(self, n):
self._refill()
self.level = min(self.cap, self.level + n)
async def handle(tenant, req, engine, client):
n_in, n_cached = count_tokens(req) # the model's own tokenizer
max_out = min(req.max_tokens, MODEL_MAX_OUT)
reserve = W_IN * (n_in - n_cached) + W_CACHED * n_cached + W_OUT * max_out
if reserve > tenant.bucket.cap:
return error(400, "request exceeds burst allowance; lower max_tokens")
wait = tenant.bucket.try_take(reserve)
if wait:
return error(429, "token rate limit", retry_after=math.ceil(wait))
if tenant.streams.locked(): # asyncio.Semaphore(max_streams)
tenant.bucket.give_back(reserve)
return error(429, "concurrent stream limit", retry_after=1)
out = 0
async with tenant.streams:
try:
async for chunk in engine.stream(req, max_tokens=max_out):
if await client.disconnected():
break # closing the stream must abort generation
out += chunk.n_tokens
await client.send(chunk)
finally:
tenant.bucket.give_back(W_OUT * (max_out - out))Three details matter. A request whose reservation exceeds the bucket's capacity can never succeed, so it gets a 400 telling the caller what to change, not a 429 that invites retries. The stream semaphore is released in a finally path, so a crash or disconnect cannot leak a slot. Breaking out of the loop only helps if the engine client aborts the sequence when the stream is closed. Test that by disconnecting mid-stream and confirming the engine's running-request count drops. Otherwise abandoned streams keep decoding and hold KV cache for nobody.
Fair sharing under contention
Hard per-tenant limits waste capacity at night and still allow overload when everyone bursts at once. Put a fair queue between admission and the engine. When the fleet has spare capacity, requests pass straight through. When it is saturated, deficit round robin hands out capacity in proportion to tenant weight, measured in weighted tokens, not in requests:
from collections import deque
def drr(queues, weight, base=50_000):
"""queues: tenant -> deque of (units, request). Yields requests in fair order."""
deficit = {t: 0.0 for t in queues}
while any(queues.values()):
for t, q in queues.items():
if not q:
deficit[t] = 0.0 # idle tenants do not bank credit
continue
deficit[t] += base * weight[t]
while q and q[0][0] <= deficit[t]:
units, req = q.popleft()
deficit[t] -= units
yield reqAdd priority classes on top. Interactive traffic goes before batch, and batch is shed first, with 503 and a longer Retry-After, when queue wait passes its target. Many teams run batch traffic on a separate endpoint with a lower price and looser latency, so it fills idle capacity without competing with chat.
Telling clients, and pacing on the client
Rejections have to tell the client what to do. Return 429 Too Many Requests, defined in RFC 6585, when the caller exceeded its own quota. Return 503 when the service is overloaded regardless of the caller. Include Retry-After, defined in RFC 9110 as either seconds or an HTTP date, computed from the bucket and not a constant. The IETF HTTPAPI working group's RateLimit and RateLimit-Policy header fields let servers advertise quotas and remaining allowance. As of draft 11 (May 2026) they are still an Internet-Draft, and LLM providers use their own vendor-specific header names, so read each provider's documentation rather than assuming one.
On the client side, the pattern is to pace before sending and back off after a rejection:
async def call(client, req, pacer, attempts=6):
for i in range(attempts):
await pacer.acquire(estimate_units(req)) # local bucket set a little under your quota
r = await client.post(req)
if r.status == 429 or r.status >= 500:
ra = r.headers.get("Retry-After", "")
delay = float(ra) if ra.isdigit() else min(60, 2 ** i)
await asyncio.sleep(delay * random.uniform(1.0, 1.5)) # jitter
continue
return r # 2xx, or a 4xx you must fix, not retry
raise RuntimeError("rate limited after retries")Never retry a 400 such as context too long, because it will fail the same way. Set max_tokens close to the output you really need, since the reservation is charged at that size until the request finishes.
Worked example: a batch summarisation job
A tenant on a 300,000 units per minute tier, with a 600,000 unit burst and 16 streams, wants to summarise 10,000 documents. Each has 6,000 input tokens and needs about 500 output tokens.
- Real cost per document: 6,000 plus 8 times 500, or 10,000 units. The whole job is 100M units, which takes about 5.6 hours at 300,000 per minute. That is the floor and no client trick beats it.
- Their first attempt sets
max_tokensto 4,096 and fires 200 requests in parallel. Each reserves 6,000 plus 8 times 4,096, about 38,800 units. The burst admits 15 requests and the rest get 429s. Throughput then depends on how fast refunds return, and retries without jitter arrive in waves. - The fix: set
max_tokensto 800, which makes the reservation 12,400 units, run 16 workers to match the stream cap, and pace at 29 documents per minute, just under the 30 the tier allows. The job finishes close to the 5.6-hour floor with almost no 429s. - If it has to finish faster, the answer is a higher tier or the batch endpoint, not more parallelism.
Failure modes
- Limits keyed only by API key. Many keys multiply the quota. Aggregate to the organisation.
- Wrong tokenizer at the gateway. Counting with a different model's tokenizer can be off by tens of percent. Use the serving model's tokenizer, or settle against the engine's reported usage.
- Reservations that never refund. A gateway crash mid-stream leaves reserved units consumed. Store reservations as leases with a time-to-live and reconcile them from engine usage logs.
- Gateway limits with an unbounded engine queue. The gateway admits within quota, but the engine queue still grows during a replica failure. Bound the engine queue too and return 503 when it is full.
- Limiter store outage. Decide in advance whether to fail open with conservative local per-node caps, or fail closed. Failing closed on a cache outage turns a small incident into a full one.
- Synchronised retries. Constant
Retry-Aftervalues make every client return at the same second. Compute them from the bucket and add jitter.
Running it in production
Watch per tenant: share of requests rejected by each limit type, weighted units used against quota, reservation-to-actual ratio, and stream-cap hits. Watch for the fleet: queue wait, KV-cache use and shed counts by priority. A tenant whose reservation is ten times its actual usage needs a note about max_tokens, not a bigger quota. Feed usage anomalies into abuse detection, since a key suddenly hitting every limit is often a stolen key. Re-derive the capacity table after every model, engine or hardware change, because the output weight moves with all three.
What to do next
- Load-test one replica with your real traffic mix and record prefill and decode throughput at your latency targets.
- Compute the output weight and the cached-input weight, and write the capacity table with explicit headroom.
- Implement the four limits keyed by organisation and key, with reserve-and-refund and a stream semaphore released in a
finallyblock. - Verify that client disconnects abort generation in the engine.
- Add a weighted fair queue and a batch priority class that is shed first.
- Return 429 or 503 with a computed
Retry-After, publish a client pacing example, and set tiers with quota system design and LLM cost analysis.