LiteLLM is an open-source Python library and proxy server that puts one OpenAI-shaped API in front of many model backends: hosted providers, cloud platforms and your own OpenAI-compatible servers such as vLLM. To an application team it is a convenience. To whoever runs GPUs it is a control point. Every request that reaches your inference fleet passes through one process that authenticates the caller, charges a budget, picks a replica, retries or fails over, and writes down what it cost.

This article treats LiteLLM as GPU-serving infrastructure: the proxy's architecture, a config for two vLLM pools and a hosted fallback, how routing, retries and cooldowns interact with replica capacity, per-team budgets, and the lessons of the March 2026 PyPI compromise. Configuration keys follow the LiteLLM documentation as of October 2026; check the docs for your pinned version before copying any key.

SDK and proxy: two products in one package

There are two products in one package, and it pays to keep them apart.

The SDK is a Python function, litellm.completion(), that accepts OpenAI-style messages and a provider-prefixed model string such as openai/<model>, translates the request into the provider's wire format and converts the response back. It runs in your process with no shared state, which suits a single Python service.

The proxy is a server, started with litellm --config config.yaml, that listens on port 4000 by default and exposes OpenAI-compatible routes, so any language can call it. It adds virtual API keys, spend tracking in Postgres, rate limits shared across proxy replicas through Redis, an admin UI, and a router that load-balances one public model name across many deployments.

The key idea is the split between a model name and a deployment. Clients ask for chat-default. The config maps that name to several concrete deployments: replica sets of a self-hosted model, or a hosted model in some region. Clients never learn where a request went, so you can add GPUs, drain a pool or move traffic to a provider without shipping client code.

Architecture: one request through the proxy

LiteLLM proxy in front of self-hosted vLLM pools and a hosted providerApps and agentsOpenAI SDK, any languageBatch jobsevals, backfillsBearer sk-...LiteLLM proxy (port 4000)1. Virtual key auth2. Budget and rpm/tpm check3. Router: pick deployment4. Translate, call, log costfallbackvLLM pool A4 replicas, zone 1vLLM pool B2 replicas, zone 2Hosted providerpay per tokenPostgreskeys, spend logsRedisshared rpm, cacheOne model_name maps to many deployments.Clients only ever see the model_name.
The proxy sits between every caller and every backend. Postgres holds keys and spend; Redis shares counters between proxy replicas.

Follow one request. The caller sends an OpenAI-format chat request with a bearer token. The proxy resolves it as a virtual key, checks the key's allowed models, budget and rpm/tpm limits, then hands the request to the router, which picks one deployment of the model name, skipping any in cooldown. The adapter calls the backend, and as tokens stream back the proxy counts and prices them against the key and team.

Postgres holds keys, teams, budgets and spend logs, so virtual keys need it. Redis is how several proxy replicas share rate-limit counters and router state. Without it each replica enforces limits alone, so a 100 rpm limit across three replicas quietly becomes up to 300 rpm.

A config for two GPU pools and a hosted fallback

Below is a config for the setup in the diagram. Two vLLM pools serve the same open-weights model, and a hosted model is the overflow and fallback target. vLLM speaks the OpenAI protocol, so its deployments use the openai/ prefix plus an api_base. Secrets are read from the environment with the os.environ/NAME syntax rather than written into the file. Model ids in angle brackets are placeholders.

model_list:
  # Pool A: four vLLM replicas behind one internal load balancer, zone 1
  - model_name: chat-default
    litellm_params:
      model: openai/<served-model-id>
      api_base: http://vllm-pool-a.internal:8000/v1
      api_key: "os.environ/VLLM_KEY"
      rpm: 1900            # what pool A can sustain at the TTFT target
  # Pool B: two replicas, zone 2
  - model_name: chat-default
    litellm_params:
      model: openai/<served-model-id>
      api_base: http://vllm-pool-b.internal:8000/v1
      api_key: "os.environ/VLLM_KEY"
      rpm: 950
  # Hosted overflow, exposed under its own name
  - model_name: chat-hosted
    litellm_params:
      model: <provider>/<hosted-model-id>
      api_key: "os.environ/PROVIDER_API_KEY"

litellm_settings:
  num_retries: 2
  request_timeout: 60
  allowed_fails: 3
  cooldown_time: 30
  fallbacks: [{"chat-default": ["chat-hosted"]}]

router_settings:
  routing_strategy: simple-shuffle
  context_window_fallbacks: [{"chat-default": ["chat-hosted"]}]
  enable_pre_call_checks: true

general_settings:
  master_key: "os.environ/LITELLM_MASTER_KEY"
  database_url: "os.environ/DATABASE_URL"

Each pool is one deployment pointing at its own internal load balancer, not one per GPU replica, so the config survives autoscaling and cache-aware routing stays inside the pool. The rpm values are measured capacity. The hosted model has its own name, so spend logs show which traffic overflowed.

Routing strategies and what they cannot see

The router decides which deployment of a model name receives each call. The routing docs list these strategies; simple-shuffle is the default and the one the docs recommend.

routing_strategyChooses byFits GPU pools whenWatch out for
simple-shuffleRandom pick, weighted by rpm, tpm or weight if setPools sit behind their own balancersUniform if you omit the weights; blind to queue depth
least-busyFewest in-flight requests seen by this proxyOne proxy replica, uneven request lengthsEach proxy replica only sees its own in-flight count
latency-based-routingRecent response latencyPools in different regionsLong generations distort latency unless you look at time to first token
usage-based-routing-v2Lowest token usage in the current minuteHosted quotas are the binding limitEqualises load, not capacity share; needs Redis across proxies
cost-based-routingCheapest healthy deploymentMixing self-hosted and paid modelsWill drain the cheap pool to saturation first

For self-hosted GPUs the important point is what LiteLLM cannot see. It does not know a replica's KV-cache occupancy, its prefix-cache hits or its batch size. Those signals live in vLLM's metrics. So let LiteLLM choose between pools and let a cache-aware router or an admission controller inside each pool choose the replica. LLM routing strategies covers that inner layer, and admission control for LLM serving covers shedding load before KV pressure forces preemption.

Retries, cooldowns and fallbacks

What happens when a deployment errors: retry, cool down, then fall backRequestmodel: chat-defaultPick deploymentskip cooled-down onesCall itwithin request_timeoutokStream to clientlog tokens and costerrorCount failureover allowed_fails: cool downnum_retries leftretries spentFallback groupsame steps, next model_nameContext too long or policy refusal taketheir own fallback lists, not the generic one.
Retry inside the model group, cool down a deployment that keeps failing, and only then cross to a fallback group.

Four settings decide what a backend failure costs. num_retries is LiteLLM's own retry loop within a model group. allowed_fails is how many failures per minute a deployment may have before it is put into cooldown, and cooldown_time is how long it stays out of rotation. fallbacks names another model group to try once retries are spent. Separate lists, context_window_fallbacks and content_policy_fallbacks, handle prompts that are too long and provider refusals. Those errors will not succeed on a sibling replica, so retrying them is waste.

The documentation stresses that num_retries is not the provider SDK's max_retries. They multiply. If the client SDK retries twice, LiteLLM retries twice and the provider SDK retries twice, one user click can become 3 x 3 x 3 = 27 attempts against a backend that is already failing. Choose one layer to own retries, normally LiteLLM because it knows about the other deployments, and set the others to zero.

With streaming, a retry is only safe before the first token reaches the client; after that it would print a second, different answer. Test what your pinned version does when a stream dies mid-generation.

Virtual keys, budgets and cost

With master_key and database_url set, the admin calls /key/generate to mint virtual keys. Each key can be limited to certain models, given a spending cap that resets on a schedule, and rate limited:

curl -X POST http://litellm.internal:4000/key/generate \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"models": ["chat-default"],
       "max_budget": 200,
       "budget_duration": "30d",
       "tpm_limit": 400000,
       "rpm_limit": 600,
       "metadata": {"team": "support-bot", "owner": "oncall-support"}}'

Notice what is missing from this key: chat-hosted. The support bot can only reach the self-hosted pools directly. It gets the paid model only through the configured fallback, so overflow spend is a platform decision rather than something any team can opt into. Hosted calls are priced from LiteLLM's model cost map. Self-hosted deployments cost whatever you tell LiteLLM they cost, which is zero unless you set per-token prices for them. Set an internal price derived from GPU-hours per million tokens so that chargeback reflects the fleet. LLM spend alerts shows how to alert on spend burn rates once these numbers exist.

Worked example: losing a zone at peak

Size the config above for a peak of 40 requests per second. Load tests show one replica holds the p95 time-to-first-token target up to 8 requests per second for this workload, with a mean of 600 prompt tokens and 250 output tokens.

ScenarioCapacity (req/s)Offered (req/s)Outcome
Both pools healthy6 x 8 = 4840Fits with 17% headroom
Pool B down (zone loss)4 x 8 = 32408 req/s must overflow
Pool A loses one replica5 x 8 = 4040Zero headroom; tail latency climbs

The rpm fields matter here. The routing docs say simple-shuffle weights its random pick by a deployment's rpm or tpm when set, and picks uniformly when they are not. Omit them and pool B, with a third of the GPUs, gets half the traffic: 20 requests per second against 16 it can carry. With 1,900 and 950 (32 and 16 per second, times 60, rounded down) the split is about 2:1.

When zone 2 fails, pool B returns connection errors, is cooled down after more than three failures in a minute, and every request lands on pool A, 8 requests per second over capacity. Pool A's admission controller should reject the excess fast with a 429 so LiteLLM's fallback sends it to chat-hosted. At 850 tokens per request that is about 24.5 million hosted tokens per hour. Multiply by your provider's price, and put the hourly cost of losing a zone in the runbook before the outage.

The March 2026 supply-chain compromise

On 24 March 2026, LiteLLM versions 1.82.7 and 1.82.8 were published to PyPI by attackers who had gained publishing access. LiteLLM's incident notes say 1.82.7 hid a payload in litellm/proxy/proxy_server.py and 1.82.8 added a litellm_init.pth file. Python runs executable lines in .pth files under site-packages at interpreter start-up, so 1.82.8 could run on any Python start in that environment without import litellm. Security write-ups describe the payload harvesting environment variables, cloud credentials, SSH keys and Kubernetes tokens. PyPI quarantined the project and the maintainers removed the two versions. If you installed either, treat every secret that environment could read as stolen and rotate it.

A gateway is a perfect target. It holds every provider key, the master key and the database URL, and it can reach every GPU pool. Run it like the security boundary it is:

  • Pin exact versions with hashes (pip install --require-hashes, or a locked image digest) and promote upgrades through staging. Never install latest in production.
  • Alert on new .pth files in site-packages in your image scan.
  • Restrict the proxy's outbound traffic to the provider endpoints and internal pools it needs. That would have blocked the exfiltration call.
  • Store provider keys in a secret manager scoped to the proxy's identity, not in a shared environment that CI and notebooks also load.

Failure modes

  • Limits that multiply. Several proxy replicas without Redis each enforce the full rpm and tpm limits, so the real limit is N times the configured one.
  • Retry storms. Retries at the client, LiteLLM and the provider SDK multiply during exactly the outage you are trying to ride out.
  • Cooldown flapping. A short cooldown_time on a pool that is overloaded rather than broken sends it back into rotation just in time to fail again. Fix capacity, not the timer.
  • Silent fallback spend. A broken self-hosted pool can run up a large hosted bill without any error reaching users. Alert on fallback rate as well as error rate.
  • Postgres as a hidden dependency. Key checks and spend logging touch the database. Test in staging what your version does when Postgres is unreachable, and decide fail-open or fail-closed before an outage decides for you.

Trade-offs

ChoiceGainCost
SDK in-processNo extra hop, no new serviceNo shared budgets or keys; every language needs its own client
ProxyOne policy point for all callers and backendsA service to run, scale, patch and secure
LiteLLM vs a dedicated gatewayPython-native, broad provider coverageFewer enterprise traffic controls than Kong AI Gateway-style products
Pool-level deploymentsStable config, cache-aware routing stays in the poolLiteLLM cannot cool down a single bad replica
Hosted fallbackRides out zone lossVariable cost and different model behaviour during incidents

LiteLLM makes switching cheap but does not make two models behave the same; see multi-provider LLM strategy before you depend on a fallback model.

What to do next

  1. Install a pinned, hash-locked LiteLLM in a test environment and run the proxy with a one-deployment config pointing at a vLLM server.
  2. Load-test each pool to find the request rate where p95 time to first token breaks your target, and write those numbers into the rpm fields.
  3. Decide which layer owns retries, set num_retries there and zero elsewhere, then confirm the attempt count with a deliberately failing backend.
  4. Add Redis before you run a second proxy replica, and check that a 60 rpm key really stops at 60 when traffic is spread across replicas.
  5. Mint per-team virtual keys with model lists, budgets and owners. Give self-hosted deployments an internal per-token price.
  6. Kill one pool in staging and measure fallback latency, fallback rate and hosted spend per hour. Put all three in the runbook.
  7. Add egress restrictions, .pth scanning and a secret-rotation drill to the proxy's deployment checklist.
Key takeaway: LiteLLM's proxy turns many backends into one OpenAI-shaped API. That makes it the place where keys, budgets, routing and failover are enforced for your GPU fleet. Map public model names to pool-level deployments, size rpm from load tests, let one layer own retries, add Redis before scaling out the proxy, and price self-hosted tokens so spend is real. Above all, run it as a security boundary: pinned and hash-locked, with restricted egress, because it holds every key you have.