Picking an LLM provider looks like a leaderboard question and is really a procurement and engineering question. The model that tops a public benchmark may fail your extraction task, cost three times more per resolved ticket than a smaller one, sit in a region your data cannot leave, or be retired a year after you build on it. Teams that choose well do it the same way they would choose a database: requirements first, hard constraints next, then measurement on their own workload, a cost model, a contract review and a plan for leaving.
This guide is deliberately vendor-neutral. Prices, model names, context limits and rate limits change every few months, so any table of them would be out of date before you read it. Instead it gives you the method, the code for an evaluation and latency harness, the arithmetic for cost per task, the questions to put to every vendor, and the failure modes that catch teams after launch. Runtime routing between models, once you have more than one, is covered in LLM model routing.
What you are actually choosing
Three things are often bundled under the word provider. The model is the weights and their behaviour. The platform is who serves it: the model developer's own API, a cloud provider's managed model service, an inference company hosting open-weight models, or your own GPUs. The contract is the terms: data retention, use of your data for training, regional processing, service levels, support and deprecation notice. The same model can be available from several platforms with different latency, quotas, regions and terms, and the same platform can offer models from several developers.
Separate the three in your evaluation. You might choose a model for quality, a platform for data residency and existing cloud commitments, and negotiate the contract for retention and notice periods. Open-weight models add a fourth option, self-hosting, which trades per-token prices for GPU costs, operations work and full control over versions.
The selection pipeline
The order matters. Hard filters are cheap to apply and remove candidates that would otherwise consume weeks of evaluation. Evaluation on your own data is expensive and should run only on a shortlist. The cost model needs the evaluation's token counts, so it comes after. The contract review comes last because it is easiest to negotiate when you know what you need and have a credible alternative.
Step 1: requirements
Write down every task the model will perform, with an example input, the expected output and how you would judge it. Classify each task by difficulty and volume, because a few hard tasks and many easy ones usually points to two models rather than one. Then record the constraints that apply across tasks:
- Latency: time to first token for interactive use, total time for batch use, and the percentile you care about.
- Volume: requests per minute at peak, average input and output tokens, and growth over a year.
- Context: the longest input you must handle, measured in the provider's tokens, not words.
- Capabilities: tool calling, structured output against a JSON schema, vision or audio input, streaming, batch processing, prompt caching, fine-tuning.
- Data: what the prompts contain, such as personal data, health data, source code or customer documents, and where it is allowed to be processed.
- Availability: what happens to your product when the model is down for an hour.
Step 2: hard filters
Apply filters that are pass or fail before any quality testing. Ask each candidate in writing, and keep the answers with the contract file:
- Is customer data used for training by default, and can that be turned off contractually?
- How long are prompts and outputs retained, for abuse monitoring or logs, and is a zero-retention option available for your use case?
- Which regions process the data, including for any fallback capacity, and can processing be pinned to a region?
- Which certifications and agreements are available, such as SOC 2 reports, ISO 27001, or a business associate agreement for health data?
- Are model versions pinnable to a dated snapshot, and how much notice is given before a version is retired?
- What service level is offered in writing, and what are the remedies?
A no on any requirement that your legal or security team treats as mandatory removes the candidate, however good its model. Expect that the same model behind different platforms gives different answers here.
Step 3: an evaluation harness on your data
Public benchmarks tell you about general capability. They do not tell you whether a model follows your output schema, handles your domain vocabulary, or refuses your legitimate requests. Build a small harness with an adapter per provider, a fixed dataset of real tasks and scorers you trust. Start with 100 to 300 examples drawn from production or realistic drafts, including the awkward cases.
from dataclasses import dataclass
import json, statistics, time
@dataclass
class Result:
text: str
input_tokens: int
output_tokens: int
ttft_s: float
total_s: float
class Adapter: # one subclass per provider SDK or HTTP API
name = "base"
def complete(self, system: str, user: str, schema: dict | None) -> Result: ...
def score(task, output: str) -> float:
if task["kind"] == "extract": # exact fields, deterministic
try:
got = json.loads(output)
except ValueError:
return 0.0
want = task["expected"]
return sum(got.get(k) == v for k, v in want.items()) / len(want)
return judge(task, output) # rubric-graded, spot-checked by humans
def run(adapter, dataset, system):
rows = []
for task in dataset:
r = adapter.complete(system, task["input"], task.get("schema"))
rows.append({"id": task["id"], "score": score(task, r.text),
"in": r.input_tokens, "out": r.output_tokens,
"ttft": r.ttft_s, "total": r.total_s})
return {"provider": adapter.name,
"mean_score": statistics.mean(x["score"] for x in rows),
"p95_total_s": sorted(x["total"] for x in rows)[int(0.95 * len(rows)) - 1],
"rows": rows}Rules that keep the comparison honest: tune the prompt for each model within a fixed time budget rather than using one prompt for all, because prompts do not transfer perfectly; use deterministic scorers wherever possible; if you use a model as a judge, do not let it judge its own provider's outputs alone, and have humans check a sample; and record token counts from each provider's response, since tokenizers differ. The approach generalises into eval-driven development, which is how you keep these evals running after the decision.
Step 4: cost per task, not price per token
Price lists quote rates per million input and output tokens, but the number that matters is cost per successful task. Two models with the same rates can differ in cost by the number of tokens their tokenizer produces for your text, how verbose their answers are, how often they need a retry, and whether you can use prompt caching or batch discounts. Compute it from the harness rows:
def cost_per_success(rows, rate_in, rate_out, cached_frac=0.0, cache_rate=0.0,
retry_rate=0.0, success_threshold=0.8):
# rates are per token; cache_rate is the discounted rate for cached input tokens
total = 0.0
for r in rows:
cached = r["in"] * cached_frac
fresh = r["in"] - cached
total += fresh * rate_in + cached * cache_rate + r["out"] * rate_out
total *= (1 + retry_rate)
successes = sum(r["score"] >= success_threshold for r in rows)
return total / max(successes, 1)A worked example with illustrative round-number rates (check current price lists, which change often): a support-summarisation task averages 3,000 input and 400 output tokens. Model A charges 3 dollars per million input tokens and 15 per million output, and succeeds on 92 percent of tasks. Model B charges 0.5 and 2, and succeeds on 78 percent. Per call, A costs 0.009 plus 0.006, 1.5 cents; B costs 0.0015 plus 0.0008, 0.23 cents. Per success, A is about 1.6 cents and B about 0.3 cents. B is cheaper, but 14 more failures per hundred might each cost a human escalation worth far more than a cent. Put the cost of a failure into the model, and the answer can flip. If 2,500 of the 3,000 input tokens are a stable system prompt and the provider offers caching at a discount, the input term can shrink substantially; check whether caching is automatic or explicit and how long cached prefixes live.
Step 5: latency and limits under load
Measure latency with streaming enabled, from the region where your servers run, at your expected concurrency, for at least a day so you see busy hours. Record time to first token, output tokens per second and total time at p50 and p95, plus the rate of 429 and 5xx responses. Rate limits are usually expressed in requests and tokens per minute and tied to a usage tier or a committed-capacity contract, so ask what limit you get on day one, how it rises, and whether provisioned or reserved throughput is available for steady loads.
Treat the provider's status history as data. Ask for incident history over the past year, and design for the outage you will eventually have: a timeout and retry policy with backoff, a fallback model, and a product that degrades gracefully, as described in multi-provider routing and graceful degradation.
Step 6: score the trade-offs
Put the results in one table and weight the criteria before you look at totals, so the weights are not chosen to justify a favourite. A typical weighting for a customer-facing feature might give quality 40 percent, cost per success 20, latency 15, reliability and limits 15 and contract terms 10, but your weights should come from your requirements. Score each candidate from 1 to 5 per criterion with the evidence written next to the score. Plotting quality against cost per success, as in quality-cost-latency frontiers, often shows that two candidates dominate the rest and the real choice is between them.
Avoiding lock-in
Lock-in in LLMs is mostly in prompts, tool schemas and response parsing, not in data. Keep one internal interface for completions, tools and structured output, with a thin adapter per provider, and route all traffic through it, ideally behind a gateway that also handles keys, quotas, logging and fallback; see LLM gateway architecture. Store prompts as versioned files with their eval scores, not as strings scattered through code. Avoid depending on a provider-only feature unless it is worth the switching cost, and when you do, isolate it behind your own function.
Failure modes after launch
- Choosing from a leaderboard and discovering in production that the model ignores your output schema one time in fifty.
- Using an unpinned model alias that the provider updates, so behaviour changes without a deploy. Pin dated versions where offered and re-run evals before moving.
- Missing a deprecation notice and having weeks to migrate a feature whose prompts were tuned for the retired model.
- Comparing price per token across providers with different tokenizers and output verbosity.
- Hitting rate limits on launch day because the account was still on a low usage tier.
- Sending regulated data to a platform before the contract terms for retention and region were confirmed.
- Having no second provider tested, so an outage becomes a product outage.
What to do next
- List every LLM task with examples, volumes, latency targets and a scoring method.
- Send the hard-filter questions to each candidate platform and drop any that fail a mandatory requirement.
- Build the harness with one adapter per shortlisted provider and 100 to 300 real tasks.
- Tune prompts per model within a fixed budget, run the evals, and spot-check judged scores by hand.
- Compute cost per successful task with current rates, caching, retries and the cost of a failure.
- Load test from your production region at expected concurrency for a full day.
- Score candidates against weights fixed in advance, and record the evidence for each score.
- Negotiate retention, region, version pinning and deprecation notice, then put a gateway in front and keep a second provider tested.
- Schedule a re-evaluation every quarter and on every model version change.