Every LLM product eventually faces the same question from finance: what does one answer cost? The honest answer is not a single number. A GPU-hour is the thing you pay for, but users consume tokens, and the number of tokens a GPU produces per hour swings by more than an order of magnitude depending on batch size, context length, the split between input and output, and how much of the day the GPU is actually busy.
This article builds that conversion from first principles. It explains why reading a prompt (prefill) and writing an answer (decode) cost very different amounts, derives both from datasheet numbers, puts them into a small cost model you can run, and works an example on an H100 with an 8-billion-parameter model. It then turns token cost into request cost, shows how idle time multiplies it, and sets up a break-even against paying an API per token. Every number below comes from running the model in this article; prices are assumptions, labelled as such, and you should replace them with your own.
The unit that matters: cost per token, split by phase
An inference request has two phases. In prefill the model reads the whole prompt at once and builds its key-value (KV) cache, the stored attention keys and values for every token so far. In decode it generates the answer one token at a time, each step reading the weights and the KV cache to produce the next token. Prefill processes thousands of tokens in parallel per weight read; decode processes one token per sequence per weight read. That asymmetry is why API price sheets charge less for input tokens than for output tokens, and why your own cost model must keep the two separate.
So the targets of a cost analysis are three numbers: dollars per million input tokens, dollars per million output tokens, and dollars per request for your actual traffic mix. A fourth, dollars per user per month, follows from requests per user. If someone gives you one blended figure, ask which mix it assumed.
The model in one picture
Model shape and GPU specs give a prefill rate and a decode rate; traffic shape and the latency target limit the batch. Token counts divided by the rates give GPU-seconds per request; multiply by price and divide by utilization, the fraction of paid time doing useful work, to get dollars per request.
Decode: a memory-bandwidth problem
During decode, every step must read every weight from GPU memory (HBM) once, whatever the batch size, and must also read each sequence's own KV cache. Each weight read feeds only a couple of multiply-adds per sequence in the batch, so at small batches the arithmetic units wait on memory. Step time is therefore close to bytes read divided by memory bandwidth.
Take an 8-billion-parameter dense model in BF16 (2 bytes per parameter), so 16 GB of weights, with an attention shape like Llama 3 8B: 32 layers, 8 KV heads, head dimension 128. The KV cache per token is 2 (K and V) times 32 times 8 times 128 times 2 bytes, which is 131,072 bytes or 128 KiB. A sequence holding 2,000 tokens carries 0.262 GB of KV cache. On an H100 SXM, with 80 GB of HBM at 3.35 TB/s, reading 16 GB of weights takes 16 / 3.35 = 4.78 ms. Dividing decimal gigabytes by terabytes per second gives milliseconds, a handy unit check.
At batch 1, a step reads 16.26 GB and takes 4.85 ms: about 206 tokens per second for the one user, and 206 for the GPU. At batch 64 the step reads 16 GB of weights plus 64 times 0.262 GB of KV, 32.8 GB in all, taking 9.78 ms. Each user now sees about 102 tokens per second, but the GPU produces 6,541. Batching is how decode gets cheap: the weight read is shared, and only the KV read grows with the batch.
Two limits cap the batch: weights plus every sequence's KV must fit in memory with headroom, and each user gets one token per step, so bigger batches stream slower. How much KV a deployment needs is covered in KV cache sizing.
Prefill: a compute problem
Prefill multiplies the whole prompt against each weight matrix in one pass, so each weight read is reused across thousands of tokens and the arithmetic units, not memory, set the pace. A dense transformer needs about 2 FLOPs per parameter per token for the forward pass, ignoring attention, which grows with context length and matters for very long prompts. For 8 billion parameters that is 16 GFLOP per token.
The H100 SXM's dense BF16 peak is about 989 TFLOPS. Real kernels reach a fraction of peak, expressed as model FLOPs utilization (MFU); 50 percent is a reasonable planning assumption for a well-tuned engine, but measure yours. At 50 percent, prefill runs at 989e12 times 0.5 divided by 16e9, about 30,900 tokens per second, which at an assumed $3.00 per GPU-hour is $0.027 per million input tokens at full utilization. Compare that to decode at batch 64, $0.127 per million at the ideal ceiling: output is roughly five times more expensive than input even before efficiency losses, and the gap widens at small batches.
The cost model as code
The model below encodes the arithmetic above. It is deliberately simple: dense models, one GPU, weights and KV read once per step, attention FLOPs ignored in prefill. Its job is to show which inputs dominate, not to replace measurement.
from dataclasses import dataclass
@dataclass
class Model:
params: float # parameters read every decode step (dense model)
bytes_per_param: float # 2 for BF16, 1 for FP8
layers: int
kv_heads: int
head_dim: int
kv_bytes: float = 2 # bytes per cached K or V element
def weight_bytes(self):
return self.params * self.bytes_per_param
def kv_bytes_per_token(self):
return 2 * self.layers * self.kv_heads * self.head_dim * self.kv_bytes # K and V
@dataclass
class Gpu:
hbm_bytes: float # capacity
hbm_bw: float # bytes/s, datasheet peak
flops: float # dense BF16 FLOP/s, datasheet peak
dollars_per_hour: float
def decode_step_seconds(m, g, batch, ctx, bw_eff):
"""Memory-bound decode: each step streams the weights once plus every sequence's KV cache."""
bytes_read = m.weight_bytes() + batch * ctx * m.kv_bytes_per_token()
return bytes_read / (g.hbm_bw * bw_eff)
def fits(m, g, batch, ctx, reserve=0.10):
return m.weight_bytes() + batch * ctx * m.kv_bytes_per_token() <= g.hbm_bytes * (1 - reserve)
def prefill_tokens_per_second(m, g, mfu):
"""Compute-bound prefill: about 2 FLOPs per parameter per token (attention ignored)."""
return g.flops * mfu / (2 * m.params)
def dollars_per_mtok(tokens_per_second, g, utilization):
return g.dollars_per_hour / (tokens_per_second * 3600 * utilization) * 1e6
model = Model(params=8.0e9, bytes_per_param=2, layers=32, kv_heads=8, head_dim=128)
h100 = Gpu(hbm_bytes=80e9, hbm_bw=3.35e12, flops=989e12, dollars_per_hour=3.00) # price assumed
ctx = 2000
for b in (1, 16, 64, 128, 200, 256):
t = decode_step_seconds(model, h100, b, ctx, bw_eff=1.0)
print(b, fits(model, h100, b, ctx), round(t * 1e3, 2), round(b / t),
round(dollars_per_mtok(b / t, h100, 1.0), 3))
prefill = prefill_tokens_per_second(model, h100, mfu=0.5)
decode = 64 / decode_step_seconds(model, h100, 64, ctx, bw_eff=0.6)
gpu_seconds = 2000 / prefill + 300 / decode # one request: 2,000 in, 300 out
for util in (1.0, 0.4):
print(util, gpu_seconds * h100.dollars_per_hour / 3600 / util)Two parameters carry the realism. bw_eff is the fraction of peak memory bandwidth your engine achieves during decode, and mfu the fraction of peak FLOPs during prefill. Both are assumptions until you benchmark your own engine, model and traffic; treat the outputs as ceilings when they are set to 1.0.
Worked example: what batching buys
Running the model with a 2,000-token average context, at peak bandwidth (a ceiling, not a forecast) and an assumed $3.00 per H100-hour, gives the following.
| Batch | Fits in 80 GB (10% reserve) | Step time | Per-user tok/s | GPU tok/s | Ceiling $/M output tok |
|---|---|---|---|---|---|
| 1 | yes | 4.85 ms | 206 | 206 | $4.045 |
| 16 | yes | 6.03 ms | 166 | 2,654 | $0.314 |
| 64 | yes | 9.78 ms | 102 | 6,541 | $0.127 |
| 128 | yes | 14.79 ms | 68 | 8,653 | $0.096 |
| 200 | yes | 20.43 ms | 49 | 9,791 | $0.085 |
| 256 | no | 24.81 ms | 40 | 10,319 | (does not fit) |
Three things stand out. Going from batch 1 to batch 64 cuts output cost by about 30 times, because the weight read is amortised. Beyond about 128, returns diminish: KV reads now dominate each step, so throughput grows slowly while each user's stream slows down. And capacity, not throughput, ends the curve: at 256 sequences of 2,000 tokens the weights plus KV no longer fit with headroom.
Note that 2,000 tokens is an average, and the KV a sequence reads grows by one token every step it decodes. A few long-context users consume KV capacity out of proportion to their numbers, which is why serving engines page KV memory and why a cost model built on the mean context underestimates the cost of the tail.
From token cost to request cost, and the utilization tax
Take a request with 2,000 input tokens and 300 output tokens, served at batch 64 with 60 percent of peak bandwidth (an assumption) and 50 percent MFU in prefill. Decode then runs at about 3,925 tokens per second. The request needs 2,000 / 30,906 = 0.065 GPU-seconds of prefill plus 300 / 3,925 = 0.076 GPU-seconds of decode, 0.141 GPU-seconds in all. At $3.00 per hour that is $0.000118 per request, or $118 per million requests, if the GPU is never idle.
It is always idle sometimes: traffic peaks daily and you provision for peak plus failover. If useful work fills 40 percent of paid hours, every number divides by 0.4: $0.000294 per request, $294 per million, and output tokens at about $0.53 per million rather than the $0.127 ceiling. Utilization usually moves the answer more than any kernel optimisation, which is why it tops GPU cost optimization.
Note the input share too: the 2,000-token prompt took almost half the GPU time. Prompt-heavy workloads such as retrieval-augmented generation are dominated by prefill, which prompt caching avoids repeating.
Self-host versus API: a break-even you can defend
Paying a hosted API per token has zero fixed cost and a steady marginal cost. Self-hosting on rented replicas that run all month is the reverse: a fixed floor, the replicas you need for availability plus the engineers who run them, and near-zero marginal cost until traffic needs another replica. The GPU time is already paid for, so do not add the per-request GPU cost on top. Break-even volume is fixed monthly cost divided by the API cost per request.
As an illustration only, assume API prices of $0.10 per million input tokens and $0.40 per million output tokens for a comparable model; these are placeholders, not quotes. The request above then costs $0.0002 + $0.00012 = $0.00032 on the API. Two H100s for redundancy cost 2 times $3.00 times 730 hours, $4,380 a month. Break-even is $4,380 / $0.00032, about 13.7 million requests a month, or 5.2 per second averaged over 730 hours. Two GPUs at 0.141 GPU-seconds per request serve about 14 per second, so that volume is about 37 percent utilization.
Now add people. At $5,000 a month of engineering time, break-even rises to about 29.3 million requests a month, 11.2 per second or 79 percent average utilization, so peaks need a third GPU, which raises the fixed cost again. At small volumes the people line decides; at large, steady volumes the hardware line does. Residency, quality and latency control are often the real reasons to self-host; say so. For the full cost of the GPU-hour itself, see GPU infrastructure cost.
Levers and how each moves the model
- Lower-precision weights. FP8 or INT4 cut the weight bytes read per decode step, helping most at small batches, and free memory for KV. Re-check quality.
- KV cache quantization. Fewer KV bytes per token raise the batch ceiling and cut the KV share of step time at large batches.
- Prefix caching. Shared prompts and documents skip prefill on a hit.
- Speculative decoding. A draft model proposes tokens the large model checks in one step, trading spare compute for fewer weight reads; best at small batches.
- Mixture-of-experts. Compute follows active parameters, but memory follows total parameters, and at large batches nearly every expert is read each step.
- Shorter outputs. Output tokens are the expensive ones; asking for concise answers is a cost control.
Where cost analyses go wrong
Uniform benchmark prompts hide the long tail that sets KV capacity and latency; replay real traffic or sample real length distributions. Costs follow the mean, but capacity follows the worst concurrent mix, so size replicas from peak concurrency and long-context share.
Different tokenizers split the same text into different token counts, so per-token prices across model families are not comparable; compare per request on your own prompts, and include hidden reasoning tokens. Timeouts, rejections and retries burn GPU time with no answer, so measure delivered answers per GPU-hour.
Finally, the latency target: batch 200 delivers about 49 tokens per second per user at the ceiling, less in practice. Faster streaming costs more per token, and that is the honest price of the requirement. See cost and latency budgets.
What to do next
- Pull a week of real traffic and record the distributions of input tokens, output tokens and concurrent sessions.
- Fill in the cost model with your model's shape, your GPU's datasheet figures and your real GPU-hour price.
- Benchmark your serving engine at several batch sizes to measure achieved bandwidth efficiency and prefill MFU, and replace the assumptions.
- Find the largest batch that meets your per-user tokens-per-second target and still fits KV for your long-context tail.
- Measure utilization over a full week and compute cost per request at that utilization, not at 100 percent.
- Run the break-even against API pricing for your own prompts, including engineering time and redundancy.
- Rank levers by expected saving: utilization first, then prefix caching, precision and output length.
- Re-run the model whenever the model, the engine or the traffic shape changes, and track cost per delivered answer as a metric.