A model API is the slowest, most expensive and most variable dependency most services have ever had. One request can take 300 milliseconds or three minutes. It can cost a fraction of a cent or several dollars. It streams for a while and then may stall. And the provider, or your own inference cluster, runs admission control of its own that answers with HTTP 429 when you exceed a quota measured in tokens, not just requests. The protective patterns that keep ordinary RPC dependencies from taking a service down, including timeouts, retries, circuit breakers and bounded queues, all still apply, but almost every default value and several of the underlying assumptions are wrong for this kind of dependency.
This article is about the protection layer in front of one model endpoint: how to admit, queue, limit, break and stream so that a degrading model does not exhaust your threads, connections, budget or users' patience, and so that your own clients do not make the overload worse. Routing among several providers is a separate problem, and the generic breaker state machine is covered in Circuit Breaker Pattern; here we concentrate on what changes when the dependency generates tokens.
Why model calls break the usual assumptions
Start with Little's law: requests in flight equal the arrival rate multiplied by the average time each spends in the system. A service receiving 40 generation requests per second with a mean latency of 8 seconds holds about 320 requests in flight. If the model slows to 16 seconds, which happens routinely under provider load, the same arrival rate holds 640. Nothing failed and no error was returned, but every connection pool, worker thread and buffer sized for 320 is now oversubscribed, and the first symptom is usually your own service timing out on unrelated endpoints.
Three properties make this worse for model APIs than for a database. First, service time depends on output length, so the raw latency distribution is wide even when the endpoint is healthy, and a latency threshold that flags a sick database flags every long answer. Second, the unit of capacity is the token: the endpoint's real limit is prefill work for input tokens plus decode work for output tokens, as explained for self-hosted servers in Continuous Batching for LLM Serving, so counting requests treats a 50-token classification and a 100,000-token document summary as equal. Third, failure is gradual: before an endpoint returns errors it usually gets slower to first token, then slower per token, then begins to stall mid-stream. A protection layer that only counts errors reacts last.
Token-weighted admission: reserve, then reconcile
The first gate is admission, and it should charge requests in the unit the endpoint is limited by. Before the call you know the input token count exactly, if you tokenize locally, and you know the maximum output you asked for. Reserve input plus maximum output from a token bucket that refills at your quota rate. When the response completes, the usage report tells you how many output tokens were actually generated; refund the difference. Without the refund, a service that always asks for 4,000 output tokens but typically uses 300 would admit about a tenth of the traffic its quota allows.
When the endpoint answers 429 with a Retry-After value, treat it as authoritative backpressure: pause admission for that route until the indicated time and shrink the local refill rate a little, instead of letting every waiting request discover the limit separately. Rejections here should be fast: a request that waits to be rejected still holds a connection.
import threading, time
class TokenBucket:
"""Admission in tokens: reserve estimate up front, refund unused output after."""
def __init__(self, tokens_per_min, burst):
self.rate = tokens_per_min / 60.0
self.capacity = burst
self.level = burst
self.paused_until = 0.0
self.t = time.monotonic()
self.lock = threading.Lock()
def _refill(self, now):
self.level = min(self.capacity, self.level + (now - self.t) * self.rate)
self.t = now
def try_reserve(self, n):
with self.lock:
now = time.monotonic()
if now < self.paused_until:
return False
self._refill(now)
if self.level >= n:
self.level -= n
return True
return False
def reconcile(self, reserved, used):
with self.lock:
self.level = min(self.capacity, self.level + max(0, reserved - used))
def pushback(self, retry_after_s):
with self.lock:
self.paused_until = max(self.paused_until, time.monotonic() + retry_after_s)
self.rate *= 0.9 # recover slowly via a separate timer, not shown
Adaptive concurrency instead of a fixed pool size
The token bucket protects your quota; it does not protect you from latency. For that you need a limit on requests in flight, and a fixed number is always wrong: too high when the endpoint is slow, too low when it is fast. An adaptive limiter measures latency and adjusts the limit, following the same idea as TCP congestion control. Envoy's adaptive concurrency filter uses a gradient controller that periodically measures a minimum round-trip time under very low concurrency and compares recent samples against it; Netflix's open-source concurrency-limits library offers AIMD, Vegas and gradient-style limiters built on the same principle.
For model calls, feed the limiter a latency signal that does not depend on answer length. Time to first token works well for interactive traffic, because it reflects queueing and prefill on the endpoint. Inter-token time, the average gap between streamed tokens, reflects decode pressure. Total latency divided by output tokens is a reasonable fallback for non-streaming calls. The simplest controller that works is additive increase, multiplicative decrease:
import time
class AIMDLimiter:
def __init__(self, initial=32, lo=4, hi=512, ttft_target_s=1.5, cut_every_s=10):
self.limit, self.lo, self.hi = initial, lo, hi
self.target, self.cut_every_s = ttft_target_s, cut_every_s
self.in_flight, self.last_cut = 0, 0.0
def acquire(self):
if self.in_flight >= self.limit:
return False
self.in_flight += 1
return True
def release(self, ttft_s=None, overloaded=False):
self.in_flight -= 1
now = time.monotonic()
if overloaded or (ttft_s is not None and ttft_s > self.target):
if now - self.last_cut >= self.cut_every_s: # at most one cut per window
self.limit = max(self.lo, int(self.limit * 0.9))
self.last_cut = now
elif self.in_flight >= self.limit - 1:
self.limit = min(self.hi, self.limit + 1) # probe up only when saturatedCut at most once per window, or a burst of slow completions collapses the limit to its floor in a second. Only raise it when saturated, and keep one limiter per model and region.
Queues that respect deadlines
Requests that cannot start immediately wait in a queue, and the queue is where overload quietly becomes outage. An unbounded queue converts a short spike into minutes of latency for everyone behind it, and by the time a request reaches the front its caller has often given up. Every queue in front of a model call should be bounded in length and should check deadlines on the way in and on the way out. On the way in, reject if the expected wait plus the expected service time exceeds the caller's remaining deadline. On the way out, drop requests whose deadline has already passed rather than spending tokens on an answer nobody will read. Pass the remaining deadline down with the request, as described for tool calls in Tool-Calling Reliability: Timeouts, Idempotency, and Compensation.
Separate queues by priority. Interactive user traffic, agent steps inside a user-visible task, and background work such as batch summarisation and evaluation runs should not share one line. Give each lane its own share of the concurrency limit, let interactive traffic borrow from background when background is idle, and shed background first when the limit shrinks. This is the model-API version of the general practice in Load shedding: dropping work to survive overload.
Circuit breakers tuned for model endpoints
A breaker exists to stop sending work to a dependency that cannot handle it, both to fail fast for your callers and to give the dependency room to recover. For a model endpoint, the definition of failure needs to be wider than errors and narrower than everything that went wrong:
| Outcome | Counts toward breaker? | Why |
|---|---|---|
| 5xx, connection reset, timeout | Yes | Endpoint unhealthy or overloaded |
| Time to first token above a slow-call threshold | Yes, as a slow call | Degradation appears here first |
| Stream stall: no token for N seconds mid-stream | Yes | Common failure shape for long generations |
| 429 with Retry-After | No; pause admission instead | The endpoint is working; you are over quota |
| 400 or request validation error | No; alert | Your bug, not the endpoint's |
| Valid response that fails your schema | No; track separately | Model or prompt problem, not availability |
Open the breaker when the failure rate or the slow-call rate over a sliding window exceeds its threshold, with a minimum number of calls in the window so that three failures at 3 a.m. do not trip it. While open, reject immediately and serve the degraded path. After a cool-down, move to half-open and let a small number of real requests through. Do not probe with a tiny synthetic prompt: an endpoint that answers "ping" quickly may still time out on real 30,000-token requests, and the breaker will flap.
Retries that do not amplify an outage
Retries are the most common way a partial outage becomes a total one. If each of three layers retries three times, one user request becomes up to 27 model calls, arriving exactly when the endpoint is least able to serve them. Retry at one layer only, preferably the one closest to the model call that knows the error class, and cap retries with a retry budget: retries may be at most a fixed fraction, say 10 percent, of first attempts over a recent window.
gRPC's retry throttling is a well-specified version of this idea: each client keeps a token count per server that starts at maxTokens, every failed call subtracts 1, every success adds tokenRatio, and retries are disabled while the count is at or below half of maxTokens. The Google SRE book's client-side adaptive throttling goes one step further and rejects new requests locally, with probability max(0, (requests - K × accepts) / (requests + 1)) over the last couple of minutes, so that a backend already spending its capacity on rejections is spared the traffic entirely. Either mechanism translates directly to a model client. Also note what never to retry: a request whose stream already delivered tool-call fragments or user-visible text, unless your handling of partial output is designed for it.
Backpressure on the way out: streaming to slow consumers
Backpressure also runs the other direction. When you relay a token stream to a browser or to another service, the reader may be slower than the model, or may have gone away. If the relay buffers without bound, a mobile client on a poor connection makes your process hold the entire answer in memory; multiply by thousands of streams and memory becomes the outage. Give each stream a bounded buffer. When it fills, stop reading from the upstream connection so transport-level flow control pushes back, and if the client has not read for longer than a stall timeout, close both sides.
Most importantly, propagate cancellation. When the user closes the tab or an agent abandons a branch, cancel the upstream request; generation otherwise continues, and is usually billed. Without that, you pay for tokens nobody reads and hold concurrency slots everyone else needs. Measure it: count output tokens generated after the client disconnected, per route. It should be near zero.
Worked example: a slow Tuesday
A support product calls a hosted model with a quota of 8 million tokens per minute. Normal traffic is 25 requests per second averaging 3,000 input and 400 output tokens, about 5.1 million tokens per minute, and requests reserve 1,000 output tokens at admission, about 6 million. Median time to first token is 0.7 seconds. At 14:05 the provider's time to first token rises to 4 seconds without errors. The AIMD limiter, targeting 1.5 seconds and cutting by 10 percent at most every 10 seconds, takes its limit from 180 to about 106 in under a minute. Interactive requests keep their lane; the queue for the background summarisation lane grows, and its deadline check starts rejecting summarisation jobs, which are retried by the batch scheduler at 15:00.
At 14:12 the provider begins to stall some streams mid-answer. The slow-call and stall rate for the route crosses 30 percent over a 60-second window with more than 200 calls, and the breaker opens. The chat interface switches to its degraded mode, answering from retrieved help-centre articles with a notice, while half-open probes with realistic prompts run every 30 seconds. Retries stay within a 10 percent budget throughout. At 14:31 probes succeed, the breaker closes and the limiter climbs back over several minutes. Nothing in the service ran out of memory or connections, which was the goal.
Failure modes and what to watch
- Retry storms from several layers each retrying. Audit every client and SDK for built-in retries and disable all but one.
- Quota starvation because reservations are not refunded on exceptions and cancellations. Refund on every exit path.
- Dashboards per route: admitted and rejected tokens, in-flight versus limit, queue wait per lane, breaker state, time to first token and inter-token percentiles, retry ratio, and 429s.
What to do next
- Compute in-flight requests with Little's law at normal and doubled latency, and check every pool and buffer against the doubled number.
- Add token-weighted admission with reservation of maximum output and refund on completion; honour Retry-After by pausing admission.
- Replace fixed concurrency with an adaptive limiter per route, driven by time to first token or inter-token time.
- Bound every queue, check deadlines on enqueue and dequeue, and split interactive from background lanes.
- Configure breakers to count timeouts, slow calls and stream stalls, not 400s or 429s, with realistic half-open probes.
- Keep retries at one layer, with a retry budget or gRPC-style throttling.
- Bound per-stream buffers and cancel upstream calls when clients disconnect; alert on tokens generated after disconnect.