A request limit such as 60 per minute works for most APIs because their requests cost about the same. LLM requests do not. A 40-token question with a 20-token answer and a 100,000-token document with a 4,000-token summary are both one request, yet the second uses thousands of times more GPU time and holds a slot in the KV cache for far longer. An endpoint limited only by request count is either too strict for normal use or wide open to anyone who sends large requests.

This article is about designing the limits themselves. It covers which dimensions to limit, how to derive the numbers from measured GPU capacity, how to weight input, cached and output tokens, how to cap concurrent streams, how to share capacity fairly between tenants, and what the client side must do. The attack side, including reasoning inflation and slow-reader floods, is covered in LLM denial of service. The generic limiter algorithms, GCRA and composite limits in Redis, are in the rate limiter architecture deep dive.

Four limits, four resources

Each limit protects one resource. Enforce several at once, because each fails differently when used alone:

LimitProtectsFails alone because
Requests per minutegateway, auth, per-request overheadignores request size
Weighted tokens per minuteGPU compute timeignores how long a stream holds memory
Concurrent streamsKV cache and batch slotsignores how much compute each stream uses
Spend per day or monththe tenant's budget and yourstoo coarse to protect latency
One request, four limits: each checks a different resourceClientpaced, honours 4291. Request rateper key and per org2. Weighted tokensreserve, then refund3. Concurrent streamssemaphore per tenant4. Spend per dayhard cap, alert at 80%Fair queueweighted per tenantadmittedEngine replicasbatching schedulerKV cachebounds concurrencyUsage meteractual tokens outon finishrefundRejections: 429 with Retry-After, or 503 if the fleet is overloadedLimits are derived from measured replica throughput and KV-cache headroom, not picked by hand.
Admission checks four independent limits, then a fair queue shares the fleet. Actual output tokens are metered on completion and the unused reservation is refunded.

Key every limit by the account or organisation as well as by API key. Otherwise a customer who creates twenty keys gets twenty times the quota. Keep a per-key limit beneath the organisation limit so one leaked key cannot spend the whole budget.

Deriving limits from GPU capacity

Limits should come from measurement. Load-test one replica with your real prompt and output length mix, at your latency targets for time to first token and time per output token. Record two numbers: prefill throughput in input tokens per second, and total decode throughput in output tokens per second across all running streams. The figures below are illustrative. Use your own.

QuantityValueNote
Prefill throughput per replica20,000 tokens/smeasured at the latency target
Decode throughput per replica2,500 tokens/sall streams combined
Output token weight820,000 / 2,500
Replica capacity20,000 units/s1 unit = 1 uncached input token
Fleet of 8 replicas9.6M units/min8 x 20,000 x 60
Sellable after 30% headroomabout 6.7M units/minfor spikes, failover and retries
Concurrent sequences per replica64from KV-cache size at your context mix

Decode tokens usually cost more GPU time than prefill tokens. Prefill processes many tokens in one parallel pass, while decode produces one token per step per stream and is limited by memory bandwidth. Your measured ratio sets the output weight. If the engine has prefix caching, cached input tokens skip most of the prefill work, so give them a weight below 1. Measure the saving instead of guessing it, and check whether caching works across tenants or only within one.

Hidden reasoning tokens are output tokens. Count them at the output weight even though the user never sees them. Otherwise a reasoning model consumes far more capacity than your limits record.

Per-tenant limits follow from the sellable total. The sum of guaranteed rates must fit inside it. Burst allowances can exceed it, because tenants rarely peak together, but they need the fair queue described below as a backstop. Concurrency works the same way: 8 replicas at 64 sequences each gives 512 slots, and per-tenant stream caps should leave room for the rest of the fleet.

Admission at the gateway

Output length is unknown when a request arrives. The standard fix is to reserve the maximum and refund the difference at the end, explained in detail in the denial-of-service article. Here is how it fits with the weights and the stream cap at a gateway. The bucket is in-process for clarity. In production it sits in a shared store, as in the limiter deep dive.

import asyncio, math, time

W_IN, W_CACHED, W_OUT = 1.0, 0.25, 8.0      # from your own measurements

class Bucket:
    def __init__(self, units_per_min, burst):
        self.rate, self.cap = units_per_min / 60.0, burst
        self.level, self.t = burst, time.monotonic()
    def _refill(self):
        now = time.monotonic()
        self.level = min(self.cap, self.level + (now - self.t) * self.rate)
        self.t = now
    def try_take(self, n):
        """Take n units, or return the seconds until n will be available."""
        self._refill()
        if self.level >= n:
            self.level -= n
            return 0.0
        return (n - self.level) / self.rate
    def give_back(self, n):
        self._refill()
        self.level = min(self.cap, self.level + n)

async def handle(tenant, req, engine, client):
    n_in, n_cached = count_tokens(req)           # the model's own tokenizer
    max_out = min(req.max_tokens, MODEL_MAX_OUT)
    reserve = W_IN * (n_in - n_cached) + W_CACHED * n_cached + W_OUT * max_out
    if reserve > tenant.bucket.cap:
        return error(400, "request exceeds burst allowance; lower max_tokens")
    wait = tenant.bucket.try_take(reserve)
    if wait:
        return error(429, "token rate limit", retry_after=math.ceil(wait))
    if tenant.streams.locked():                  # asyncio.Semaphore(max_streams)
        tenant.bucket.give_back(reserve)
        return error(429, "concurrent stream limit", retry_after=1)
    out = 0
    async with tenant.streams:
        try:
            async for chunk in engine.stream(req, max_tokens=max_out):
                if await client.disconnected():
                    break                        # closing the stream must abort generation
                out += chunk.n_tokens
                await client.send(chunk)
        finally:
            tenant.bucket.give_back(W_OUT * (max_out - out))

Three details matter. A request whose reservation exceeds the bucket's capacity can never succeed, so it gets a 400 telling the caller what to change, not a 429 that invites retries. The stream semaphore is released in a finally path, so a crash or disconnect cannot leak a slot. Breaking out of the loop only helps if the engine client aborts the sequence when the stream is closed. Test that by disconnecting mid-stream and confirming the engine's running-request count drops. Otherwise abandoned streams keep decoding and hold KV cache for nobody.

Fair sharing under contention

Hard per-tenant limits waste capacity at night and still allow overload when everyone bursts at once. Put a fair queue between admission and the engine. When the fleet has spare capacity, requests pass straight through. When it is saturated, deficit round robin hands out capacity in proportion to tenant weight, measured in weighted tokens, not in requests:

from collections import deque

def drr(queues, weight, base=50_000):
    """queues: tenant -> deque of (units, request). Yields requests in fair order."""
    deficit = {t: 0.0 for t in queues}
    while any(queues.values()):
        for t, q in queues.items():
            if not q:
                deficit[t] = 0.0                 # idle tenants do not bank credit
                continue
            deficit[t] += base * weight[t]
            while q and q[0][0] <= deficit[t]:
                units, req = q.popleft()
                deficit[t] -= units
                yield req

Add priority classes on top. Interactive traffic goes before batch, and batch is shed first, with 503 and a longer Retry-After, when queue wait passes its target. Many teams run batch traffic on a separate endpoint with a lower price and looser latency, so it fills idle capacity without competing with chat.

Telling clients, and pacing on the client

Rejections have to tell the client what to do. Return 429 Too Many Requests, defined in RFC 6585, when the caller exceeded its own quota. Return 503 when the service is overloaded regardless of the caller. Include Retry-After, defined in RFC 9110 as either seconds or an HTTP date, computed from the bucket and not a constant. The IETF HTTPAPI working group's RateLimit and RateLimit-Policy header fields let servers advertise quotas and remaining allowance. As of draft 11 (May 2026) they are still an Internet-Draft, and LLM providers use their own vendor-specific header names, so read each provider's documentation rather than assuming one.

On the client side, the pattern is to pace before sending and back off after a rejection:

async def call(client, req, pacer, attempts=6):
    for i in range(attempts):
        await pacer.acquire(estimate_units(req))      # local bucket set a little under your quota
        r = await client.post(req)
        if r.status == 429 or r.status >= 500:
            ra = r.headers.get("Retry-After", "")
            delay = float(ra) if ra.isdigit() else min(60, 2 ** i)
            await asyncio.sleep(delay * random.uniform(1.0, 1.5))   # jitter
            continue
        return r                                      # 2xx, or a 4xx you must fix, not retry
    raise RuntimeError("rate limited after retries")

Never retry a 400 such as context too long, because it will fail the same way. Set max_tokens close to the output you really need, since the reservation is charged at that size until the request finishes.

Worked example: a batch summarisation job

A tenant on a 300,000 units per minute tier, with a 600,000 unit burst and 16 streams, wants to summarise 10,000 documents. Each has 6,000 input tokens and needs about 500 output tokens.

  • Real cost per document: 6,000 plus 8 times 500, or 10,000 units. The whole job is 100M units, which takes about 5.6 hours at 300,000 per minute. That is the floor and no client trick beats it.
  • Their first attempt sets max_tokens to 4,096 and fires 200 requests in parallel. Each reserves 6,000 plus 8 times 4,096, about 38,800 units. The burst admits 15 requests and the rest get 429s. Throughput then depends on how fast refunds return, and retries without jitter arrive in waves.
  • The fix: set max_tokens to 800, which makes the reservation 12,400 units, run 16 workers to match the stream cap, and pace at 29 documents per minute, just under the 30 the tier allows. The job finishes close to the 5.6-hour floor with almost no 429s.
  • If it has to finish faster, the answer is a higher tier or the batch endpoint, not more parallelism.

Failure modes

  • Limits keyed only by API key. Many keys multiply the quota. Aggregate to the organisation.
  • Wrong tokenizer at the gateway. Counting with a different model's tokenizer can be off by tens of percent. Use the serving model's tokenizer, or settle against the engine's reported usage.
  • Reservations that never refund. A gateway crash mid-stream leaves reserved units consumed. Store reservations as leases with a time-to-live and reconcile them from engine usage logs.
  • Gateway limits with an unbounded engine queue. The gateway admits within quota, but the engine queue still grows during a replica failure. Bound the engine queue too and return 503 when it is full.
  • Limiter store outage. Decide in advance whether to fail open with conservative local per-node caps, or fail closed. Failing closed on a cache outage turns a small incident into a full one.
  • Synchronised retries. Constant Retry-After values make every client return at the same second. Compute them from the bucket and add jitter.

Running it in production

Watch per tenant: share of requests rejected by each limit type, weighted units used against quota, reservation-to-actual ratio, and stream-cap hits. Watch for the fleet: queue wait, KV-cache use and shed counts by priority. A tenant whose reservation is ten times its actual usage needs a note about max_tokens, not a bigger quota. Feed usage anomalies into abuse detection, since a key suddenly hitting every limit is often a stolen key. Re-derive the capacity table after every model, engine or hardware change, because the output weight moves with all three.

What to do next

  1. Load-test one replica with your real traffic mix and record prefill and decode throughput at your latency targets.
  2. Compute the output weight and the cached-input weight, and write the capacity table with explicit headroom.
  3. Implement the four limits keyed by organisation and key, with reserve-and-refund and a stream semaphore released in a finally block.
  4. Verify that client disconnects abort generation in the engine.
  5. Add a weighted fair queue and a batch priority class that is shed first.
  6. Return 429 or 503 with a computed Retry-After, publish a client pacing example, and set tiers with quota system design and LLM cost analysis.
Key takeaway: LLM requests differ in cost by orders of magnitude, so request counts alone cannot protect an endpoint. Limit requests, weighted tokens, concurrent streams and spend together, keyed by organisation. Derive the numbers from measured prefill and decode throughput and KV-cache headroom. Reserve at max_tokens and refund on completion, make disconnects abort generation, share saturated capacity with a weighted fair queue, and give clients a computed Retry-After and a pacing pattern to follow.