Total cost of ownership asks a different question from cost per token. Cost per token is a snapshot: what one unit of work costs on one configuration today. TCO is a forecast: what it costs to run a language model workload over its life, typically three years, including the demand you have not seen yet, the people who keep it running, the commitments you sign, the hardware you might own, and the model you will replace halfway through. Most self-hosting decisions that go wrong were right on a per-token spreadsheet and wrong on a TCO one.
This article builds a three-year TCO model for an LLM serving workload, compares a hosted API, on-demand cloud GPUs, committed cloud GPUs and owned hardware, and shows which inputs actually move the answer. For the per-token mechanics behind it see LLM cost analysis; for the fully loaded cost of an owned GPU-hour see GPU infrastructure cost. Every price below is a labelled assumption; replace each with your own quotes.
The cost stack
The cost stack has more layers than the GPU line item. Compute is the obvious one: API tokens, cloud GPU-hours, or purchased servers. Owned hardware adds power multiplied by facility overhead, space, spares and support contracts, minus whatever the hardware is worth when you retire it. Every self-hosted option adds engineers: on-call, upgrades of the serving stack, capacity planning and incident response. All options carry evaluation work, because each model change needs a regression suite run before it ships, and migration work, because the model you serve today will not be the one you serve in eighteen months.
Then there are the costs of mismatch between capacity and demand. Self-hosted capacity is bought for peak traffic plus headroom, while demand averages well below peak, so you pay for idle time. Commitments are bought before demand is known, so they either strand money or force overflow onto expensive on-demand capacity. An API has no mismatch cost at all, which is a large part of why it wins at small scale.
How serving software turns GPUs into capacity
The single most important number in a self-hosted TCO is how many requests one GPU serves per second at your latency target. It comes from how the serving software uses the hardware. Prefill, which processes the prompt, is dominated by matrix multiplications and is limited by compute. Decode, which generates one token at a time, must read the model's weights and the growing key-value cache for every step, so it is limited by memory bandwidth. A single request in decode leaves most of the arithmetic units idle.
Batching fixes that by decoding many requests in one pass over the weights, and continuous batching keeps the batch full as requests finish. Batch size is capped by memory: every concurrent request holds its key-value cache, which grows with context length, and paged KV cache management reduces the waste. Longer prompts, longer outputs, tighter time-to-first-token targets and bigger models all lower requests per GPU-second. That is why the number must come from a load test of your model, your prompt mix and your SLO, never from a vendor's peak tokens-per-second figure.
The model as code
The model below runs month by month for 36 months. Demand compounds. The API option pays per token with an annual price decline. On-demand cloud pays for average need times an autoscaling inefficiency factor. Committed cloud buys a fixed fleet sized to the month-12 peak and bursts to on-demand for the excess. Owned hardware buys whole eight-GPU servers three months ahead of need and recovers a resale fraction at the end. Each month is discounted at 8 percent a year.
import math
from dataclasses import dataclass
@dataclass
class Workload:
req_per_day: float = 4_000_000; growth_per_month: float = 0.05
in_tokens: int = 1_500; out_tokens: int = 300; peak_to_avg: float = 2.5
@dataclass
class Prices: # every value is an assumption to replace with your quotes
api_in_per_m: float = 0.50; api_out_per_m: float = 1.50; api_decline_per_year: float = 0.25
gpu_on_demand_hr: float = 4.00; gpu_committed_hr: float = 2.50
gpu_capex: float = 40_000; gpu_residual: float = 0.20; owned_opex_hr: float = 0.60
fte_year: float = 250_000
@dataclass
class Serving:
req_per_gpu_s: float = 2.0; headroom: float = 0.20; min_gpus: int = 4
autoscale_factor: float = 1.4; ops_fte_self: float = 1.5; ops_fte_api: float = 0.25
HOURS, DISCOUNT = 730, 0.08
def demand(w, m):
return w.req_per_day * (1 + w.growth_per_month) ** m
def gpus_for_peak(w, s, m):
peak_rps = demand(w, m) / 86_400 * w.peak_to_avg
return max(s.min_gpus, math.ceil(peak_rps * (1 + s.headroom) / s.req_per_gpu_s))
def monthly(option, w, p, s, m, commit_gpus=0, owned_gpus=0):
if option == "api":
tok = demand(w, m) * 30.4 * (w.in_tokens * p.api_in_per_m + w.out_tokens * p.api_out_per_m) / 1e6
return tok * (1 - p.api_decline_per_year) ** (m / 12) + s.ops_fte_api * p.fte_year / 12
people = s.ops_fte_self * p.fte_year / 12
if option == "on_demand":
avg_gpus = demand(w, m) / 86_400 / s.req_per_gpu_s
return max(s.min_gpus, avg_gpus * s.autoscale_factor) * HOURS * p.gpu_on_demand_hr + people
if option == "committed":
overflow = max(0, gpus_for_peak(w, s, m) - commit_gpus) * HOURS * 0.35 # bursts near peak only
return commit_gpus * HOURS * p.gpu_committed_hr + overflow * p.gpu_on_demand_hr + people
return owned_gpus * HOURS * p.owned_opex_hr + people # owned
def tco(option, w=Workload(), p=Prices(), s=Serving(), months=36):
npv, owned, capex = 0.0, 0, 0.0
commit = gpus_for_peak(w, s, 12)
for m in range(months):
buy_cost = 0.0
if option == "owned":
need = gpus_for_peak(w, s, min(m + 3, months - 1))
if need > owned:
buy = math.ceil((need - owned) / 8) * 8
buy_cost, owned = buy * p.gpu_capex, owned + buy
capex += buy_cost
cost = monthly(option, w, p, s, m, commit, owned) + buy_cost
npv += cost / (1 + DISCOUNT) ** (m / 12)
if option == "owned":
npv -= capex * p.gpu_residual / (1 + DISCOUNT) ** 3
return npvThe model is deliberately small. Every line maps to a decision someone in your organisation can change, which is the point: TCO is a tool for arguing about inputs, not a number to defend.
Worked example
Take a support assistant at 4 million requests a day, growing 5 percent a month (about 5.5 times over three years), with 1,500 input and 300 output tokens per request, a 2.5 times peak-to-average ratio, and a load test showing 2 requests per GPU-second at the latency target. At month 0 the peak needs 70 GPUs with headroom; by month 35 it needs 384. Running the model gives these three-year net present costs:
| Scenario | Hosted API | On-demand | Committed | Owned |
|---|---|---|---|---|
| Base case | $7.55M | $8.87M | $10.55M | $15.42M |
| Throughput 4 req/s per GPU | $7.55M | $4.94M | $5.80M | $8.25M |
| Throughput 1 req/s per GPU | $7.55M | $16.74M | $20.08M | $29.74M |
| No growth | $3.41M | $4.06M | $5.13M | $4.45M |
| Peak-to-average 1.5 | $7.55M | $8.87M | $6.74M | $9.72M |
| API price decline 50% a year | $4.23M | $8.87M | $10.55M | $15.42M |
Read the table as a map of sensitivities, not a verdict. In the base case the API is cheapest, because self-hosting pays for idle capacity and people. Doubling per-GPU throughput, which serving-stack work such as better batching, quantisation or a smaller distilled model can deliver, makes on-demand self-hosting the cheapest option by a third. Halving it makes every self-hosted option more than twice as expensive as the API. Ownership loses under fast growth because it buys ahead of demand and sizes for peak. With flat demand it beats committed cloud ($4.45M against $5.13M), though on-demand and the API still come out cheaper at this scale. A flatter traffic shape helps commitments most, because fewer committed GPUs sit idle off-peak.
A more useful output than any single row is the break-even throughput: the requests per GPU-second at which a self-hosted option matches the API. Sweep req_per_gpu_s and the base case crosses over at about 2.4 for on-demand cloud and about 2.9 for committed cloud. That turns the decision into an engineering question with a target. If the serving team believes it can reach 2.4 requests per GPU-second at the latency target, with evidence from a load test rather than hope, on-demand self-hosting is worth piloting; if the best measured figure is 2.0 and no change on the roadmap moves it by a fifth, stay on the API and revisit when the model, the prompt shape or the traffic changes. Run the same sweep over growth and peak-to-average ratio, and hand each break-even to the person who owns that input.
What moves the answer
Rank the inputs by how much one plausible change moves the answer, and spend measurement effort in that order. For the self-hosted options in this model, requests per GPU-second and demand growth lead, and the peak-to-average ratio matters mainly for committed and owned capacity. For the API option, the price trend is the input that matters, since a steep annual decline cut its total by more than 40 percent. Hardware price negotiations move a self-hosted TCO by tens of percent; throughput and utilization move it by multiples. GPU cost optimization lists the engineering levers in detail.
| Cost often left out | Why it matters | How to include it |
|---|---|---|
| Engineers on call | Self-hosting needs 24x7 coverage once users depend on it | Count FTEs, not hours; a rota needs several people |
| Evaluation runs | Every model or serving change needs a regression suite | GPU-hours per eval times releases per year |
| Model migration | Models are typically replaced within one or two years | Engineer-months per migration, plus double-running |
| Egress and storage | Logs, checkpoints and cross-region traffic | Per-GB prices times measured volume |
| Compliance | Data residency, audits, contractual terms | Legal and security review time per option |
| Stranded commitments | Demand fails to arrive or the model gets smaller | Run the model with growth at half the forecast |
Failure modes
The failures repeat across organisations. The throughput figure comes from a benchmark with short prompts and no latency target, so production needs three times the GPUs. Capacity is sized for average load, and the first busy Monday breaches the SLO. The API bill is extrapolated linearly while prices fall, overstating it. People are left out because they already work there, although their time is what self-hosting consumes. A three-year commitment is signed for a model that is replaced by a smaller one in month nine, stranding most of it. And no one re-runs the model, so a decision made on last year's inputs survives long after the inputs changed.
Each failure has the same remedy: tie every input to a measurement with an owner and a date, and re-run the model each quarter with fresh measurements, as a scheduled review rather than a one-off justification.
Trade-offs
The API buys elasticity, zero idle cost and the provider's model improvements, and costs control: data leaves your boundary, rate limits and deprecations are someone else's decision, and per-token prices can rise as well as fall. On-demand cloud keeps control and elasticity but pays the highest hourly rate and may not have capacity when you need it. Commitments trade flexibility for rate. Ownership gives the lowest marginal cost and complete control when demand is large, steady and predictable, and exposes you to obsolescence, lead times and the operational burden of a fleet.
Many teams end up hybrid: a committed or owned base sized near the off-peak floor, on-demand or API capacity for peaks and new models, and the option to move traffic between them. The TCO model above extends to that easily by splitting demand between options month by month.
What to do next
- Load-test your model with your real prompt mix and record requests per GPU-second at your SLO.
- Pull 90 days of traffic and compute average demand, peak-to-average ratio and monthly growth.
- Collect written quotes for API tokens, on-demand and committed GPU-hours, and servers if you might own.
- Count the people each option needs, including on-call rotation, and price them at loaded cost.
- Run the model for all four options, then run it again with throughput halved and growth halved.
- Pick the option that wins across the plausible range, not the one that wins only in the base case.
- Schedule a quarterly re-run with an owner for each input.