Total cost of ownership asks a different question from cost per token. Cost per token is a snapshot: what one unit of work costs on one configuration today. TCO is a forecast: what it costs to run a language model workload over its life, typically three years, including the demand you have not seen yet, the people who keep it running, the commitments you sign, the hardware you might own, and the model you will replace halfway through. Most self-hosting decisions that go wrong were right on a per-token spreadsheet and wrong on a TCO one.

This article builds a three-year TCO model for an LLM serving workload, compares a hosted API, on-demand cloud GPUs, committed cloud GPUs and owned hardware, and shows which inputs actually move the answer. For the per-token mechanics behind it see LLM cost analysis; for the fully loaded cost of an owned GPU-hour see GPU infrastructure cost. Every price below is a labelled assumption; replace each with your own quotes.

The cost stack

Three-year TCO: every option is demand x unit cost, plus what the option drags inDemand curverequests/day, growth, peakToken shapeinput and output per requestMeasured throughputreq/s per GPU at your SLOCapacity planGPUs for peak + headroomAPI meteringtokens x price, declineOn-demand cloudautoscaled GPU-hoursCommitted cloudfixed GPUs + overflowOwnedcapex steps + opex - resaleHosted APIper-token billMonthly cost, discounted to NPVplus people, evaluation, migration, egress, complianceThroughput and utilization sit upstream of every self-hosted option; prices are inputs, not answers.
The TCO pipeline: demand and measured throughput feed a capacity plan for self-hosted options and a token bill for the API; each option's monthly cost is then discounted.

The cost stack has more layers than the GPU line item. Compute is the obvious one: API tokens, cloud GPU-hours, or purchased servers. Owned hardware adds power multiplied by facility overhead, space, spares and support contracts, minus whatever the hardware is worth when you retire it. Every self-hosted option adds engineers: on-call, upgrades of the serving stack, capacity planning and incident response. All options carry evaluation work, because each model change needs a regression suite run before it ships, and migration work, because the model you serve today will not be the one you serve in eighteen months.

Then there are the costs of mismatch between capacity and demand. Self-hosted capacity is bought for peak traffic plus headroom, while demand averages well below peak, so you pay for idle time. Commitments are bought before demand is known, so they either strand money or force overflow onto expensive on-demand capacity. An API has no mismatch cost at all, which is a large part of why it wins at small scale.

How serving software turns GPUs into capacity

The single most important number in a self-hosted TCO is how many requests one GPU serves per second at your latency target. It comes from how the serving software uses the hardware. Prefill, which processes the prompt, is dominated by matrix multiplications and is limited by compute. Decode, which generates one token at a time, must read the model's weights and the growing key-value cache for every step, so it is limited by memory bandwidth. A single request in decode leaves most of the arithmetic units idle.

Batching fixes that by decoding many requests in one pass over the weights, and continuous batching keeps the batch full as requests finish. Batch size is capped by memory: every concurrent request holds its key-value cache, which grows with context length, and paged KV cache management reduces the waste. Longer prompts, longer outputs, tighter time-to-first-token targets and bigger models all lower requests per GPU-second. That is why the number must come from a load test of your model, your prompt mix and your SLO, never from a vendor's peak tokens-per-second figure.

The model as code

The model below runs month by month for 36 months. Demand compounds. The API option pays per token with an annual price decline. On-demand cloud pays for average need times an autoscaling inefficiency factor. Committed cloud buys a fixed fleet sized to the month-12 peak and bursts to on-demand for the excess. Owned hardware buys whole eight-GPU servers three months ahead of need and recovers a resale fraction at the end. Each month is discounted at 8 percent a year.

import math
from dataclasses import dataclass

@dataclass
class Workload:
    req_per_day: float = 4_000_000; growth_per_month: float = 0.05
    in_tokens: int = 1_500; out_tokens: int = 300; peak_to_avg: float = 2.5

@dataclass
class Prices:   # every value is an assumption to replace with your quotes
    api_in_per_m: float = 0.50; api_out_per_m: float = 1.50; api_decline_per_year: float = 0.25
    gpu_on_demand_hr: float = 4.00; gpu_committed_hr: float = 2.50
    gpu_capex: float = 40_000; gpu_residual: float = 0.20; owned_opex_hr: float = 0.60
    fte_year: float = 250_000

@dataclass
class Serving:
    req_per_gpu_s: float = 2.0; headroom: float = 0.20; min_gpus: int = 4
    autoscale_factor: float = 1.4; ops_fte_self: float = 1.5; ops_fte_api: float = 0.25

HOURS, DISCOUNT = 730, 0.08

def demand(w, m):
    return w.req_per_day * (1 + w.growth_per_month) ** m

def gpus_for_peak(w, s, m):
    peak_rps = demand(w, m) / 86_400 * w.peak_to_avg
    return max(s.min_gpus, math.ceil(peak_rps * (1 + s.headroom) / s.req_per_gpu_s))

def monthly(option, w, p, s, m, commit_gpus=0, owned_gpus=0):
    if option == "api":
        tok = demand(w, m) * 30.4 * (w.in_tokens * p.api_in_per_m + w.out_tokens * p.api_out_per_m) / 1e6
        return tok * (1 - p.api_decline_per_year) ** (m / 12) + s.ops_fte_api * p.fte_year / 12
    people = s.ops_fte_self * p.fte_year / 12
    if option == "on_demand":
        avg_gpus = demand(w, m) / 86_400 / s.req_per_gpu_s
        return max(s.min_gpus, avg_gpus * s.autoscale_factor) * HOURS * p.gpu_on_demand_hr + people
    if option == "committed":
        overflow = max(0, gpus_for_peak(w, s, m) - commit_gpus) * HOURS * 0.35   # bursts near peak only
        return commit_gpus * HOURS * p.gpu_committed_hr + overflow * p.gpu_on_demand_hr + people
    return owned_gpus * HOURS * p.owned_opex_hr + people                          # owned

def tco(option, w=Workload(), p=Prices(), s=Serving(), months=36):
    npv, owned, capex = 0.0, 0, 0.0
    commit = gpus_for_peak(w, s, 12)
    for m in range(months):
        buy_cost = 0.0
        if option == "owned":
            need = gpus_for_peak(w, s, min(m + 3, months - 1))
            if need > owned:
                buy = math.ceil((need - owned) / 8) * 8
                buy_cost, owned = buy * p.gpu_capex, owned + buy
                capex += buy_cost
        cost = monthly(option, w, p, s, m, commit, owned) + buy_cost
        npv += cost / (1 + DISCOUNT) ** (m / 12)
    if option == "owned":
        npv -= capex * p.gpu_residual / (1 + DISCOUNT) ** 3
    return npv

The model is deliberately small. Every line maps to a decision someone in your organisation can change, which is the point: TCO is a tool for arguing about inputs, not a number to defend.

Worked example

Take a support assistant at 4 million requests a day, growing 5 percent a month (about 5.5 times over three years), with 1,500 input and 300 output tokens per request, a 2.5 times peak-to-average ratio, and a load test showing 2 requests per GPU-second at the latency target. At month 0 the peak needs 70 GPUs with headroom; by month 35 it needs 384. Running the model gives these three-year net present costs:

ScenarioHosted APIOn-demandCommittedOwned
Base case$7.55M$8.87M$10.55M$15.42M
Throughput 4 req/s per GPU$7.55M$4.94M$5.80M$8.25M
Throughput 1 req/s per GPU$7.55M$16.74M$20.08M$29.74M
No growth$3.41M$4.06M$5.13M$4.45M
Peak-to-average 1.5$7.55M$8.87M$6.74M$9.72M
API price decline 50% a year$4.23M$8.87M$10.55M$15.42M

Read the table as a map of sensitivities, not a verdict. In the base case the API is cheapest, because self-hosting pays for idle capacity and people. Doubling per-GPU throughput, which serving-stack work such as better batching, quantisation or a smaller distilled model can deliver, makes on-demand self-hosting the cheapest option by a third. Halving it makes every self-hosted option more than twice as expensive as the API. Ownership loses under fast growth because it buys ahead of demand and sizes for peak. With flat demand it beats committed cloud ($4.45M against $5.13M), though on-demand and the API still come out cheaper at this scale. A flatter traffic shape helps commitments most, because fewer committed GPUs sit idle off-peak.

A more useful output than any single row is the break-even throughput: the requests per GPU-second at which a self-hosted option matches the API. Sweep req_per_gpu_s and the base case crosses over at about 2.4 for on-demand cloud and about 2.9 for committed cloud. That turns the decision into an engineering question with a target. If the serving team believes it can reach 2.4 requests per GPU-second at the latency target, with evidence from a load test rather than hope, on-demand self-hosting is worth piloting; if the best measured figure is 2.0 and no change on the roadmap moves it by a fifth, stay on the API and revisit when the model, the prompt shape or the traffic changes. Run the same sweep over growth and peak-to-average ratio, and hand each break-even to the person who owns that input.

What moves the answer

Rank the inputs by how much one plausible change moves the answer, and spend measurement effort in that order. For the self-hosted options in this model, requests per GPU-second and demand growth lead, and the peak-to-average ratio matters mainly for committed and owned capacity. For the API option, the price trend is the input that matters, since a steep annual decline cut its total by more than 40 percent. Hardware price negotiations move a self-hosted TCO by tens of percent; throughput and utilization move it by multiples. GPU cost optimization lists the engineering levers in detail.

Cost often left outWhy it mattersHow to include it
Engineers on callSelf-hosting needs 24x7 coverage once users depend on itCount FTEs, not hours; a rota needs several people
Evaluation runsEvery model or serving change needs a regression suiteGPU-hours per eval times releases per year
Model migrationModels are typically replaced within one or two yearsEngineer-months per migration, plus double-running
Egress and storageLogs, checkpoints and cross-region trafficPer-GB prices times measured volume
ComplianceData residency, audits, contractual termsLegal and security review time per option
Stranded commitmentsDemand fails to arrive or the model gets smallerRun the model with growth at half the forecast

Failure modes

The failures repeat across organisations. The throughput figure comes from a benchmark with short prompts and no latency target, so production needs three times the GPUs. Capacity is sized for average load, and the first busy Monday breaches the SLO. The API bill is extrapolated linearly while prices fall, overstating it. People are left out because they already work there, although their time is what self-hosting consumes. A three-year commitment is signed for a model that is replaced by a smaller one in month nine, stranding most of it. And no one re-runs the model, so a decision made on last year's inputs survives long after the inputs changed.

Each failure has the same remedy: tie every input to a measurement with an owner and a date, and re-run the model each quarter with fresh measurements, as a scheduled review rather than a one-off justification.

Trade-offs

The API buys elasticity, zero idle cost and the provider's model improvements, and costs control: data leaves your boundary, rate limits and deprecations are someone else's decision, and per-token prices can rise as well as fall. On-demand cloud keeps control and elasticity but pays the highest hourly rate and may not have capacity when you need it. Commitments trade flexibility for rate. Ownership gives the lowest marginal cost and complete control when demand is large, steady and predictable, and exposes you to obsolescence, lead times and the operational burden of a fleet.

Many teams end up hybrid: a committed or owned base sized near the off-peak floor, on-demand or API capacity for peaks and new models, and the option to move traffic between them. The TCO model above extends to that easily by splitting demand between options month by month.

What to do next

  1. Load-test your model with your real prompt mix and record requests per GPU-second at your SLO.
  2. Pull 90 days of traffic and compute average demand, peak-to-average ratio and monthly growth.
  3. Collect written quotes for API tokens, on-demand and committed GPU-hours, and servers if you might own.
  4. Count the people each option needs, including on-call rotation, and price them at loaded cost.
  5. Run the model for all four options, then run it again with throughput halved and growth halved.
  6. Pick the option that wins across the plausible range, not the one that wins only in the base case.
  7. Schedule a quarterly re-run with an owner for each input.
Key takeaway: LLM TCO is a three-year forecast, not a per-token snapshot. Model demand growth, peak traffic, measured requests per GPU-second, people, evaluation and migration for every option, discount the monthly costs, and test the answer against halved throughput and growth. Throughput and utilization usually move the result far more than prices, so measure them first and re-run the model every quarter.