An LLM price list looks like one number per model, but the bill is computed from five or six meters that move independently. Input and output tokens are priced differently. Cached input is priced differently again, with a surcharge to write the cache and a discount to read it. Reasoning tokens are billed even though you never see them. Batch and flex tiers trade latency for price, some prices depend on prompt length, and some depend on where inference runs. Committed capacity replaces all of that with a flat hourly charge that you pay whether or not you use it.
This page explains each pricing model from first principles, writes the bill as code, and works through one month of a real-shaped workload to show which levers move the total. The cost of serving a token on your own GPUs, and the self-host versus API break-even, are covered in LLM cost analysis. This page is about the prices providers charge and how to read them. Provider facts quoted here were checked against Anthropic's, Google's and OpenAI's pricing documentation on 2026-10-03. Prices change often, so the worked example uses a hypothetical rate card, and you should re-check the live pages before quoting a number.
The meters on every bill
Start with why input and output are priced differently, because every other rule builds on it. Input tokens are processed in prefill, where the whole prompt goes through the model in parallel and the GPU is limited by compute. Output tokens are generated one at a time in decode. Each step reads all the weights and the key-value cache, so the GPU is limited by memory bandwidth and serves many requests at once to stay busy. One output token ties up far more hardware time than one input token, so output is priced at a multiple of input. On the cards checked for this page the ratio is 5x for Anthropic's models and 8x for Gemini 2.5 Pro on shorter prompts, so read your provider's card. In practice, verbose answers cost more than long prompts.
Reasoning models add a meter you cannot see. Their hidden thinking tokens are generated in decode exactly like visible output, and they are billed as output. Google's Gemini pricing page states that its output prices include thinking tokens. A request that returns a 400-token answer after 1,200 tokens of reasoning is billed for 1,600 output tokens. Agent frameworks add another quiet meter. When you pass tools, Anthropic's API adds a tool-use system prompt of a few hundred tokens per request, and the exact count depends on the model and the tool_choice setting. Your tool schemas are billed as input on every call too. Server-side tools may carry their own fee: Anthropic lists web search at $10 per 1,000 searches, plus the tokens of the results.
Prompt caching: a surcharge, a discount and a break-even
Prompt caching is the meter with the biggest effect on most bills, and the one most often misread. Recomputing the key-value state for a long, unchanging prefix (system prompt, tool definitions, a policy document) on every request is wasted prefill. Providers can store that state and reuse it, so they charge less to read it. Anthropic's pricing documentation gives the multipliers relative to base input: a 5-minute cache write costs 1.25x, a 1-hour cache write costs 2x, and a cache read costs 0.1x, with lower read multipliers on some newer models. The same page says these multipliers stack with the Batch API discount and with data-residency pricing. Google prices cached input at a reduced rate plus a separate storage charge per million tokens per hour, so a cache that sits idle still costs money.
The write surcharge creates a break-even hit rate. With hit rate h and the 5-minute multipliers, the cost per prefix token, relative to sending it uncached, is h × 0.1 + (1 − h) × 1.25. That is below 1 only when h is above about 21.7%. For the 1-hour cache it is h × 0.1 + (1 − h) × 2, which breaks even near 52.6%. Below those hit rates, caching raises the bill. Hit rate is decided by your traffic, not by the provider. It depends on how many requests share the exact same prefix within the cache lifetime, and on whether anything dynamic, such as a timestamp or user name, was placed before the cache breakpoint. For the mechanics of placing breakpoints, see prompt caching.
Batch, flex, length and geography modifiers
Asynchronous tiers sell idle capacity. A provider has to provision for peak traffic, so off-peak GPUs are cheap to fill with work that can wait. Anthropic's Batch API charges 50% of standard rates on both input and output. Google's batch mode halves Gemini prices too. OpenAI's flex processing, enabled with service_tier="flex" on Responses or Chat Completions requests, is billed at Batch API rates. In return it is slower and sometimes unavailable. When capacity runs short it returns HTTP 429, OpenAI says you are not charged for those requests, and it recommends raising client timeouts to 15 minutes. The trade is the same everywhere: a discount for giving up latency guarantees.
Length tiers and geography are the other modifiers. Google prices Gemini 2.5 Pro in two bands split at a 200,000-token prompt, and above that threshold both input and output rates are higher. So one long request can cost far more than its token count suggests. Anthropic took the opposite route for its recent models: Claude 4.6 and later include the full 1M-token context window at standard pricing. Geography can carry a premium too. On Anthropic's API, pinning inference to the US with inference_geo applies a 1.1x multiplier to every token category. For Claude 4.5 and later models on Bedrock and Google Cloud, regional endpoints carry a 10% premium over global ones.
One more trap affects comparisons rather than bills. Tokenizers differ between providers and between model generations. Anthropic notes that its newer tokenizer produces about 30% more tokens for the same text. Comparing two price cards per million tokens is therefore not comparing cost per task. Measure the token counts of your own prompts on each candidate model before you compare prices.
The bill as code
The bill as code makes all of this testable. The model below covers the per-request meters and the batch modifier. The rate card is hypothetical, chosen to resemble a mid-tier model, and the cache multipliers are Anthropic's 5-minute ones. Swap in your provider's live numbers before you rely on the output.
from dataclasses import dataclass
@dataclass
class RateCard:
input_per_m: float
output_per_m: float
cache_write_mult: float = 1.25
cache_read_mult: float = 0.10
batch_mult: float = 0.50
def request_cost(rc, prefix, variable, output, reasoning=0, hit_rate=0.0, batch=False):
m = 1e-6
hit = prefix * hit_rate * rc.input_per_m * rc.cache_read_mult * m
miss = prefix * (1 - hit_rate) * rc.input_per_m * (rc.cache_write_mult if hit_rate else 1.0) * m
var = variable * rc.input_per_m * m
out = (output + reasoning) * rc.output_per_m * m
total = hit + miss + var + out
return total * (rc.batch_mult if batch else 1.0)
rc = RateCard(input_per_m=3.00, output_per_m=15.00) # hypotheticalTwo simplifications are worth naming. Every miss is treated as a cache write, which is right when the prefix is marked cacheable. And hit_rate=0 means caching is off, not caching that never hits. The second case would pay the write surcharge on every request.
Worked example: one million support requests
Take a support assistant serving one million requests a month. Each request has a 6,000-token stable prefix (instructions, tool schemas and policy text), 1,500 tokens of variable input (the conversation and retrieved snippets), and a 400-token answer. Running the model above gives these monthly totals:
| Scenario | Monthly cost | What changed |
|---|---|---|
| A. No caching | $28,500 | Every prefix token billed at base input |
| B. Caching at 95% hit rate | $13,335 | Prefix mostly read at 0.1x; writes at 1.25x on misses |
| C. B on a reasoning model, 1,200 thinking tokens | $31,335 | Hidden output at 15/M outweighs every saving in B |
| D. B with 30% of traffic moved to batch | $11,335 | Overnight summaries and evaluations at half price |
| Caching at a 20% hit rate | $28,860 | Below the 21.7% break-even; worse than A |
Three lessons generalise. First, caching more than halves this bill, because the stable prefix is 80% of the input. Your ratio of prefix to variable input sets the ceiling on what caching can save. Second, reasoning can wipe out every input-side saving. Here 1,200 thinking tokens add $18,000 a month, more than the whole cached bill. Reasoning effort is a pricing decision, and it should be set per route, not globally. Third, a badly placed cache breakpoint is worse than none. At a 20% hit rate the bill rises above the uncached baseline, and nothing in the API response flags it. You only see it by measuring hit rate.
Committed capacity and the utilisation test
Committed capacity is a different pricing model. Instead of paying per token, you pay a flat rate per hour for a reserved slice of throughput, usually for a fixed term. Each cloud platform sells this under its own name and its own throughput unit. You gain predictable latency and protection from rate limits. You pay for every hour whether traffic arrives or not, so the economics come down to utilisation. If a unit costs C per hour and the same tokens would cost P per hour on demand at full load, the commitment pays off only when average utilisation exceeds C / P.
def commitment_break_even(flat_per_hour, on_demand_at_full_load_per_hour):
"""Average utilisation above which reserved capacity is cheaper."""
return flat_per_hour / on_demand_at_full_load_per_hour
# hypothetical: unit costs 20/hour; the tokens it can serve would cost 32/hour on demand
print(commitment_break_even(20, 32)) # 0.625 -> needs 62.5% average utilisationDaily traffic cycles make 62.5% average utilisation hard to reach for interactive traffic alone. The usual pattern is to size the commitment to the overnight trough, send the peak to on-demand, and backfill the idle reserved hours with batch-style work you would otherwise send to a batch tier. The same utilisation logic, applied to GPUs you own, is in GPU infrastructure cost.
Failure modes
- Comparing price per million tokens across tokenizers. The same document can produce about 30% more tokens on one model than on another. Compare cost per task on your own prompts.
- Ignoring reasoning tokens. A cheap-looking model with a high reasoning effort can cost more per answer than a pricier one that answers directly.
- Caching below break-even. A timestamp, request ID or user name ahead of the cache breakpoint drops the hit rate toward zero while you keep paying the write surcharge.
- Crossing a length tier by accident. Retrieval that grows the prompt past a provider's threshold reprices the request at the higher band.
- Treating flex or batch 429s as outages. They are part of the deal. Back off and retry, or fall back to standard, and budget for that fallback.
- Unused commitments. Reserved capacity bought for a launch peak sits idle for the rest of the term. Size it to the trough.
- Tool overhead. Large tool schemas are re-billed on every agent step. Put them in the cached prefix.
What to do next
- Export a week of request logs with input, cached-read, cache-write, output and reasoning token counts split out. If your client does not record these usage fields, fix that first.
- Compute your stable-prefix fraction and your measured cache hit rate. Compare them with the 21.7% (5-minute) and 52.6% (1-hour) break-evens.
- Port the cost function above, load each candidate provider's live rate card, and price your real token mix, re-tokenized on each model.
- Tag every route as interactive or deferrable, and move the deferrable work (evaluations, summaries, backfills) to a batch or flex tier.
- Set reasoning effort per route, and alert when reasoning tokens per answer drift upward.
- Before buying committed capacity, plot hourly token demand and compute utilisation against the commitment's break-even.
- Add a cost-per-request panel next to latency and quality; LLM KPI dashboards shows the recording rules, and GPU cost optimisation orders the levers.