DeepInfra is an inference provider. It hosts open-weight models, including chat LLMs, embedding and reranking models, image, speech and vision models, on its own GPU fleet and sells access per token, or per GPU-hour for dedicated deployments. For most teams it is a drop-in replacement for the OpenAI API: change the base URL and the model name, and existing client code runs.

Providers like this look interchangeable, and their docs list similar features, so it is worth knowing where they differ. This page covers what decides whether DeepInfra works for your traffic: the four ways to run a model, the per-model concurrency limit and the throughput ceiling it implies, a client that respects that limit, the custom deployment model and its limits, the Batch API, and a break-even calculation between per-token and per-hour pricing. For the general pattern of routing across several providers, see Multi-Provider LLM Strategy.

Four surfaces over one fleet

DeepInfra: one account, four ways to run a modelYour serviceOpenAI SDK or plain HTTPClient limitersemaphore + backoff on 429Serverless, OpenAI-compatible/v1/openai, per token, 200 concurrent per modelNative inference API/v1/inference/{model}, all model typesBatch APIJSONL, 24h window, 20% below real-timeCustom deployment/deploy/llm, per GPU-hour, up to 4 GPUsShared GPU fleetA100 to B300Pick the surface by traffic shape: steady and heavy goes dedicated, bursty goes serverless, offline goes batch.
Four surfaces over one GPU fleet. The client-side limiter matters because the serverless limit counts concurrent requests per model, not requests per minute.

The serverless OpenAI-compatible API at https://api.deepinfra.com/v1/openai covers chat, completions and embeddings. The native API at https://api.deepinfra.com/v1/inference/{model} is plain HTTP, reaches every model type, and returns token counts with each response. The Batch API follows OpenAI's file-and-batch workflow at a discount. Custom deployments run a Hugging Face model you choose on dedicated A100, H100, H200, B200 or B300 GPUs, billed per GPU-hour. LoRA adapters can be uploaded on supported base models and are called by name, like any other model.

Calling the API

With the OpenAI SDK, only the key and base URL change. Keep the model ID in configuration: DeepInfra's catalogue changes as new open-weight releases arrive and old ones are retired. The IDs below come from DeepInfra's own documentation at the time of writing. Check the model list before you pin one.

import os
from openai import OpenAI

client = OpenAI(api_key=os.environ["DEEPINFRA_API_KEY"],
                base_url="https://api.deepinfra.com/v1/openai")

chat_model = os.environ.get("CHAT_MODEL", "deepseek-ai/DeepSeek-V4-Flash-0731")
resp = client.chat.completions.create(
    model=chat_model,
    messages=[{"role": "user", "content": "Summarize RAID 10 in two sentences."}],
    max_tokens=200,
)
print(resp.choices[0].message.content, resp.usage)

emb = client.embeddings.create(model="Qwen/Qwen3-Embedding-8B",
                               input=["first text", "second text"],
                               encoding_format="float")
print(len(emb.data), len(emb.data[0].embedding))

The native API is useful when you do not want an SDK, or for model types the OpenAI schema does not cover:

curl "https://api.deepinfra.com/v1/inference/deepseek-ai/DeepSeek-V4-Flash-0731" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $DEEPINFRA_API_KEY" \
  -d '{"input": "The capital of France is", "max_new_tokens": 20, "stream": false}'

Log the usage block from every response. Billing is per token, and the usage block is the only per-request record you will have to reconcile against the invoice.

The concurrency limit is a throughput ceiling

DeepInfra's default limit is 200 concurrent requests per model per account. It is not a requests-per-minute limit. Requests over it get HTTP 429 with a Rate limited message, and you can ask for an increase from the dashboard. A concurrency limit turns into a throughput ceiling through Little's law: concurrency = throughput × latency. So the most requests per second you can sustain is 200 divided by the average request duration.

Average request durationMax sustained throughputPer minute
1 s (short embedding or classification)200 req/s12,000
8 s (chat, ~500 output tokens)25 req/s1,500
60 s (long generation, agents)3.3 req/s200

Work through the middle row. A support assistant generates about 500 tokens per answer, and at the observed decode speed each request takes 8 seconds end to end. At 200 concurrent, the ceiling is 25 requests per second. If peak traffic is 40 requests per second, a streaming client keeps 320 requests open and 120 of them come back as 429s. A faster network does not fix this. The options are to shorten the requests (lower max_tokens, a faster model), split traffic across two models, ask for a higher limit, or move the workload to a dedicated deployment.

Enforce the limit in the client instead of finding it through errors. A semaphore set a bit below the limit, plus jittered backoff for the 429s that still get through, keeps tail latency predictable:

import asyncio, random
from openai import AsyncOpenAI, RateLimitError

client = AsyncOpenAI(api_key=KEY, base_url="https://api.deepinfra.com/v1/openai",
                     max_retries=0)              # we own retries
gate = asyncio.Semaphore(180)                     # headroom under 200 per model

async def complete(messages, model, attempts=5):
    for i in range(attempts):
        async with gate:
            try:
                return await client.chat.completions.create(
                    model=model, messages=messages, max_tokens=500)
            except RateLimitError:
                pass                              # release the slot before sleeping
        await asyncio.sleep(min(30, 0.5 * 2 ** i) * random.uniform(0.5, 1.5))
    raise RuntimeError("rate limited after retries")

Because the limit applies per model, one semaphore per model ID is the correct shape. A single global semaphore throttles you for no reason when you call several models, because each model has its own 200 slots and the global gate would cap them all together.

Custom deployments

A custom deployment runs a model from a Hugging Face repository on GPUs reserved for you. It is created with one call:

curl -X POST https://api.deepinfra.com/deploy/llm \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $DEEPINFRA_API_KEY" \
  -d '{"model_name": "support-bot", "gpu": "A100-80GB", "num_gpus": 2,
       "max_batch_size": 64, "hf": {"repo": "your-org/your-model"},
       "settings": {"min_instances": 1, "max_instances": 2}}'

You call it through the normal OpenAI endpoint, with the model set to YOUR_USERNAME/support-bot. You can change the scaling settings later with PUT /deploy/DEPLOY_ID. The documentation lists the limits you should design around. Deployments are billed per GPU-hour while instances run, not per token. min_instances: 0 scales to zero and saves money, at the price of a cold start on the next request. There is a default limit of 4 GPUs per user, such as four 1-GPU instances or one 4-GPU instance, which can be raised by contacting DeepInfra. GPU capacity during scale-up is not guaranteed. Quantization is not currently supported for custom deployments, so size GPU memory for the full-precision weights plus KV cache.

max_batch_size is the main tuning knob. Raising it lets the server batch more concurrent decodes per step, which increases total tokens per second but also the latency of each token. Set it from a load test, not a guess.

The Batch API

Offline work such as classifying a corpus, embedding a document store or generating evaluation answers does not need an interactive endpoint. The Batch API takes a JSONL file in which each line has a unique custom_id, "method": "POST", a url equal to the batch endpoint, and the same body you would send in real time. All requests in a file must use one model. Supported endpoints are /v1/chat/completions, /v1/completions and /v1/embeddings. The completion window is 24 hours, and the price is 20 percent below the real-time price for the same model. A file holds at most 50,000 requests and 200 MB, and an account can run 100 batches at once.

f = client.files.create(file=open("requests.jsonl", "rb"), purpose="batch")
batch = client.batches.create(input_file_id=f.id,
                              endpoint="/v1/chat/completions",
                              completion_window="24h")
# poll client.batches.retrieve(batch.id) until a terminal status, then
# download output_file_id and error_file_id; join on custom_id, not on line order

Output order is not guaranteed, so join results back on custom_id. Batch jobs also do not count against your interactive traffic, which makes batch the easiest way to keep a nightly job from taking concurrency slots from live users.

Serverless or dedicated: the break-even

Serverless costs tokens x price_per_token. A deployment costs gpus x hours x price_per_gpu_hour whether or not it is busy. The break-even point is the token volume at which they are equal: tokens_per_hour = gpus x gpu_hour_price / token_price. Take hypothetical numbers, not DeepInfra's current prices, which you should read from its pricing page: $0.30 per million tokens blended, and $2.00 per GPU-hour. Two GPUs break even at 2 × 2.00 / 0.30 = 13.3 million tokens per hour, about 3,700 tokens per second sustained, around the clock.

A deployment wins only if your traffic is steady at that level, the model is not in the serverless catalogue, or you need its isolation and fixed configuration. Bursty traffic almost always favours serverless, because an idle GPU-hour costs the same as a busy one. Run the calculation with your measured tokens per second at peak and off-peak, and include the 20 percent batch discount for any offline share. The Together AI page applies the same model to another provider, and comparing the two on your own traffic is a quick way to decide.

Measuring and operating it

Four numbers tell you whether the integration is healthy: time to first token, output tokens per second for each request, the 429 rate, and cost per request from the usage block. Time to first token reflects queueing and prefill on the provider side. Tokens per second reflects decode speed and how busy the shared fleet is. Measure both by streaming, because a non-streaming call only gives you the total.

import time

def timed_stream(client, model, messages):
    t0 = time.perf_counter()
    first, chunks = None, 0
    stream = client.chat.completions.create(
        model=model, messages=messages, stream=True,
        stream_options={"include_usage": True})
    usage = None
    for ev in stream:
        if ev.choices and ev.choices[0].delta.content:
            first = first or time.perf_counter()
            chunks += 1
        if ev.usage:
            usage = ev.usage
    end = time.perf_counter()
    ttft = (first or end) - t0
    out_tokens = usage.completion_tokens if usage else chunks
    return {"ttft_s": ttft, "tok_per_s": out_tokens / max(end - (first or end), 1e-6),
            "usage": usage}

Check that the provider returns the final usage event when you ask for it with stream_options. If it does not, fall back to counting chunks, which is an approximation. Export these figures per model as histograms, and alert on the 95th percentile of time to first token and on the 429 rate, not on averages. Then compare the summed usage against the billing usage endpoint once a week. A gap means a code path is missing from your logs, often a retry. Finally, run a fixed evaluation prompt set every day against each model you depend on. That catches changes in model version or serving configuration that the API would not report.

Failure modes

  • Treating 429 as an outage. It is the concurrency limit. Retrying immediately makes it worse. Limit in the client and back off with jitter.
  • Pinning a retired model. Catalogue entries change. Keep model IDs in config and alert on model-not-found errors.
  • Assuming provider equivalence. The same open model on two providers can differ in precision, context limit and tokenizer handling. Run your evaluation set before switching; provider failover covers how to do that safely.
  • Scale-to-zero on a latency path. min_instances: 0 gives the first user after an idle period a cold start.
  • Assuming scale-up capacity. Scale-up capacity is not guaranteed. Keep a serverless fallback for overflow.
  • Joining batch results by line number. Output order is not guaranteed; join on custom_id.

Trade-offs

DeepInfra's strengths are a broad open-weight catalogue, OpenAI compatibility, per-token pricing with no minimum, and a cheaper batch path. Its constraints are a per-model concurrency limit you have to plan for, a small default GPU cap on dedicated deployments, and no quantization on custom deployments for now. Running the model yourself with vLLM gives full control of precision, batching and placement, but you take on capacity planning, on-call and the cost of idle GPUs. A common pattern is serverless DeepInfra for bursty and experimental traffic, batch for offline work, and self-hosted or dedicated capacity only for the steady base load.

What to do next

  1. Point an existing OpenAI client at the DeepInfra base URL with a model ID from config, and log the usage block from every response.
  2. Measure your average request duration and compute 200 divided by it. Compare that with your peak requests per second.
  3. Add a per-model semaphore below the limit, plus jittered backoff on 429.
  4. Move any offline workload to the Batch API and join results on custom_id.
  5. Run the break-even calculation with current prices before creating a deployment, and keep a serverless fallback if you do.
  6. Run your evaluation set against the exact model and provider before sending production traffic.
Key takeaway: DeepInfra serves open-weight models through an OpenAI-compatible API, a native HTTP API, a discounted 24-hour Batch API and per-GPU-hour custom deployments. Its default limit is 200 concurrent requests per model, so sustained throughput is 200 divided by request duration. Enforce that in the client, send offline work to batch, and pay for dedicated GPUs only when steady traffic passes the break-even.