Model selection is usually argued from a leaderboard: one model scores a few points higher on a public benchmark, so it is assumed to be the better choice. In production the question is different. A team is choosing a configuration, meaning a model plus a prompt, a reasoning budget, a retrieval setup and serving parameters, and that choice spends money and adds latency on every request. A configuration that is two points better on quality but four times more expensive and twice as slow at the 95th percentile may be the wrong choice, or the only acceptable one, depending on what a correct answer is worth and what the latency budget allows.
The tool that makes this trade-off explicit is the Pareto frontier: the set of configurations for which no alternative is at least as good on every axis and strictly better on one. This article covers how to define the three axes so they are measurable, how to collect the numbers on your own workload, how to compute the frontier without being fooled by noise, how to pick a point on it with explicit constraints, and how to keep the analysis current as prices and models change.
Three axes, defined precisely
Vague axes produce meaningless frontiers, so each one needs an operational definition before any measurement starts. Quality is the pass rate, or mean rubric score, on a versioned golden dataset drawn from real traffic, scored by the same graders for every configuration. Cost is dollars per completed task, not per token: it includes input, cached input and output tokens, reasoning tokens (billed as output on most APIs), retries, tool-call round trips and any follow-up turns the configuration needs to finish the task. Latency is end-to-end time as the user experiences it, reported at the percentile the product promises, measured at the concurrency the system will actually run.
Per-token prices are misleading on their own because configurations use very different numbers of tokens for the same task. A cheaper model that needs a longer prompt with more examples, or that fails and retries twenty percent of the time, can cost more per completed task than a pricier model that succeeds on the first attempt. Likewise, time to first token matters for streaming chat but not for a batch extraction job, where only total completion time counts. Choose the latency statistic that matches the product surface and write it down.
- Quality: pass rate on golden dataset version v, with a 95% confidence interval, overall and for each critical slice.
- Cost: mean dollars per completed task, including every attempt, at current list or contracted prices.
- Latency: end-to-end p95 (or p99) at target concurrency; add time to first token for streaming surfaces.
Measuring each configuration
Run every candidate through the same harness on the same cases, and record raw per-case results rather than summaries: pass or fail, token counts by type, attempt count, wall-clock latency and any error. Replay latency measurements under a load generator at the intended concurrency rather than sequentially, because many serving stacks, hosted and self-hosted, degrade at the tail as load rises.
Compute cost from token counts and a price table held separately from the measurements. That separation matters: prices change more often than models, and a frontier built from recorded dollar figures goes stale the day a provider cuts a price, while one built from token counts can be refreshed in seconds.
from dataclasses import dataclass
@dataclass(frozen=True)
class Price: # USD per million tokens
input: float
cached_input: float
output: float # reasoning tokens are usually billed here too
def cost_per_task(attempts, price):
"""attempts: every call made for one task, including retries and tool round trips."""
total = 0.0
for a in attempts:
uncached = a.input_tokens - a.cached_input_tokens
total += (uncached * price.input
+ a.cached_input_tokens * price.cached_input
+ a.output_tokens * price.output) / 1_000_000
return totalFor self-hosted models the price table is replaced by a throughput calculation. The effective price per million output tokens is the hourly cost of the serving hardware divided by the tokens it sustains per hour at the latency target. An instance costing ten dollars an hour that sustains 2,500 output tokens per second produces nine million tokens an hour, or about $1.11 per million. The same instance run at a smaller batch size to meet a tighter latency target might sustain half that throughput and cost twice as much per token, which is exactly the coupling between cost and latency that the frontier is meant to expose.
Computing the frontier
With one row per configuration, the frontier is a simple dominance filter. Configuration A dominates B if A is at least as good on every axis and strictly better on at least one. Any dominated configuration can be discarded, because there is some alternative that is no worse in any respect. What remains is the set of real trade-offs, and it is usually much smaller than the candidate list: of a dozen configurations, often three to five survive.
@dataclass(frozen=True)
class Config:
name: str
quality: float # pass rate on the golden set
cost: float # USD per completed task
p95_ms: float # end-to-end p95 at target concurrency
def dominates(a, b):
no_worse = a.quality >= b.quality and a.cost <= b.cost and a.p95_ms <= b.p95_ms
better = a.quality > b.quality or a.cost < b.cost or a.p95_ms < b.p95_ms
return no_worse and better
def pareto_front(configs):
return [c for c in configs
if not any(dominates(o, c) for o in configs if o is not c)]Plot the result with cost on a logarithmic x-axis, because candidate costs routinely span two orders of magnitude, and quality on the y-axis, with latency encoded by marker or used as a pre-filter. The frontier traces the best quality available at each spend level. Its shape is informative: a steep early segment means small spending increases buy large quality gains, and a flat late segment means the most expensive configurations are buying very little.
Uncertainty: dominance you can trust
Quality numbers from a few hundred cases carry real sampling error. At 500 cases and a pass rate near 80%, the 95% confidence interval is roughly plus or minus 3.5 points, so a configuration that scores 81% has not been shown to beat one that scores 79%. Naive dominance on point estimates will happily declare winners that a re-run with a different sample would reverse, and the frontier will jitter every time it is recomputed.
Two practices fix this. First, compare configurations on the same cases with a paired test, which removes case difficulty from the variance and is far more sensitive than comparing two independent intervals. Second, treat quality differences inside the interval as ties when testing dominance: A dominates B on quality only if the paired interval for A minus B excludes zero.
import random
def paired_bootstrap(a_pass, b_pass, iters=2000, seed=0):
"""a_pass, b_pass: per-case 0/1 results on the same cases.
Returns a 95% interval for pass_rate(a) - pass_rate(b)."""
rng = random.Random(seed)
n = len(a_pass)
diffs = []
for _ in range(iters):
idx = [rng.randrange(n) for _ in range(n)]
diffs.append(sum(a_pass[i] - b_pass[i] for i in idx) / n)
diffs.sort()
return diffs[int(0.025 * iters)], diffs[int(0.975 * iters) - 1]
Choosing a point on the frontier
The frontier removes bad choices; it does not make the final one. That needs constraints and a value model. Hard constraints come first: the latency SLO, any quality floor that product or compliance requires on critical slices, and the per-task budget. Filter the frontier to configurations that satisfy them. If nothing survives, the constraints are incompatible with the available technology, and that is an important finding to take back to the product owner rather than something to paper over by quietly relaxing one.
Among feasible configurations, choose by expected net value. Estimate what a successful task is worth and what a failure costs; failures are rarely free, since they create support tickets, human review or lost users. The net value of a configuration is then its pass rate times the value of success, minus the failure rate times the cost of failure, minus the cost of running it. This makes the trade-off explicit and auditable: someone can disagree with the value estimate, but the reasoning is visible.
def pick(front, value_success, cost_failure, p95_slo_ms, budget):
feasible = [c for c in front if c.p95_ms <= p95_slo_ms and c.cost <= budget]
if not feasible:
return None # constraints conflict: escalate, don't silently relax
def net(c):
return c.quality * value_success - (1 - c.quality) * cost_failure - c.cost
return max(feasible, key=net)When the value of a success is large relative to inference cost, as in legal review, code changes or customer-facing answers with liability attached, the choice slides toward the high-quality end of the frontier, and cost barely matters. When value is small and volume is huge, such as classification or tagging at scale, the choice slides toward the cheap end, and a point of quality is rarely worth a doubling of cost. Run the calculation per workload; one company-wide default model is almost always wrong for some of them.
Configurations, not just models
The candidate list should include variations of each model, not only different models. Many of the cheapest improvements come from configuration changes that move a point up or to the left without changing the model at all, and they often dominate a switch to a larger model.
- Prompt and examples: shorter prompts cut cost on every call; a few well-chosen examples can buy quality more cheaply than a bigger model.
- Reasoning budget: where a model exposes a thinking or effort setting, each level is a separate point with its own cost and latency.
- Output limits and format: structured output and tight length limits reduce output tokens, which are usually the most expensive.
- Prompt caching: a stable prefix billed at the cached rate can change the cost ranking of two configurations outright.
- Retrieval depth: fewer, better-ranked passages reduce input tokens and sometimes improve faithfulness.
- Serving choices: for self-hosted models, quantization and batch size move both cost and latency.
Cascades and routing on the frontier
A cascade sends every request to a cheap configuration first and escalates to an expensive one only when a check fails, such as a verifier, a confidence threshold or a schema validation. Its expected cost is the cheap cost plus the escalation rate times the expensive cost, and its quality depends on how well the check separates good cheap answers from bad ones. A good cascade can land above the frontier line connecting its two components, reaching near-large quality at near-small cost, which is why cascades belong in the candidate set as configurations in their own right.
Latency is where cascades surprise people. Escalated requests pay for both calls in sequence, so if more than about five percent of requests escalate, the p95 falls inside the escalated population and roughly equals the small model's latency plus the large model's. A cascade that looks excellent on mean latency can therefore break a p95 SLO. Measure cascades end to end under load like any other configuration, and track the escalation rate in production, because it drifts with the traffic mix and moves the cascade along the cost axis as it does.
Latency under real load
A frontier measured with one request at a time describes a system nobody runs. Hosted APIs vary by time of day, region and account tier, and self-hosted servers trade per-request latency for throughput as batch sizes grow. Measure at the concurrency and prompt-length distribution you expect in production, over long enough windows to include the provider's normal variation, and record the percentile you actually promise. A configuration can sit on the frontier at the median and fall off it at the 99th percentile.
Keeping the frontier current
A frontier is a snapshot. Prices fall, new models ship, providers change default behavior, and the golden dataset gets a new version. Make the analysis a scheduled job: re-run the candidate set whenever the dataset version, a price or the candidate list changes, recompute costs from stored token counts when only prices move, and publish the frontier with the dataset version, price table date and load profile attached. Keep the production configuration on every plot and alert when something dominates it.
Failure modes
- Public benchmark as quality axis. Leaderboard scores measure someone else's task distribution and are exposed to contamination; only your own golden set predicts your production quality.
- Per-token pricing as cost axis. Ignoring prompt length, retries and reasoning tokens misranks configurations by large factors.
- Sequential latency tests. Single-request timings hide tail behavior under load and flatter cascades and small batch sizes.
- Point-estimate dominance. Declaring winners inside the confidence interval produces a frontier that reshuffles on every run.
- Aggregate-only quality. A configuration can win overall while failing a critical slice; apply slice floors as hard constraints.