The H200 is an H100 with a different memory system. The Hopper compute die, tensor-core peaks and power envelope are the same. What changes is 141 GB of HBM3e at 4.8 TB/s in place of 80 GB of HBM3 at 3.35 TB/s. Cloud providers and resellers charge a premium for it. The economic question is therefore narrow and answerable: for your workload, does the extra memory buy more tokens per dollar than the premium costs?

This article turns that question into a formula, a small cost model you can run, and a worked example on 70B and 8B models. It also shows why the answer is often much larger or much smaller than the 43 percent bandwidth increase suggests. The hardware itself, including KV-cache sizing and the NVL variant, is covered in the H200 deep dive; this page treats it as a purchasing decision.

What the premium buys

SXM partH100H200Ratio
HBM capacity80 GB HBM3141 GB HBM3e1.76x
HBM bandwidth3.35 TB/s4.8 TB/s1.43x
Dense BF16 / FP8 tensor peakabout 989 / 1,979 TFLOPSsame1.0x
Maximum powerup to 700 Wup to 700 W1.0x
NVLink per GPU900 GB/s900 GB/s1.0x

Three things follow. Anything limited by tensor-core throughput, such as large-batch prefill or dense training at a healthy batch size, runs at about the same speed on both. Anything limited by memory bandwidth, above all autoregressive decode, gets up to 1.43 times faster per step. And anything limited by memory capacity, such as how many sequences fit in the KV cache or whether the model fits on one GPU at all, can change shape entirely. That last effect is where the surprising answers come from. Software carries over unchanged: it is the same architecture, so CUDA kernels, FP8 recipes and serving engines need no port.

The unit that matters: cost per token

Compare the GPUs on cost per token, never on price or speed alone. For a deployment of g GPUs at price P per GPU-hour that sustains T tokens per second within your latency target, the cost per million tokens is g × P / (T × 3600) × 106. Divide both sides by g and the comparison becomes a break-even rule:

The H200 is cheaper per token when PH200 / PH100 is less than TH200 / TH100, with both throughputs measured per GPU.

Everything else in this article is about estimating the right-hand side honestly. The left-hand side comes from your quotes. Public on-demand listings in 2026 spread several-fold between hyperscalers and specialist clouds, and the H200 premium over the H100 varies by provider and contract length. Collect real quotes for the same term and region. The examples below use an illustrative $2.50 per GPU-hour for the H100 and $3.40 for the H200, a ratio of 1.36. They are inputs, not market data.

Why capacity can beat bandwidth

Each decode step reads every weight once and the KV cache of every sequence in the batch. In the memory-bound regime, step time is bytes read divided by effective bandwidth. When a server packs as many sequences as memory allows, almost all of memory is read every step. Step time then approaches usable memory divided by bandwidth, which is about 21 ms of raw transfer on an H100 and 28 ms on an H200. A full H200 therefore takes longer per step than a full H100, but each step produces a token for many more sequences.

So throughput per GPU scales with the batch that fits, and the batch is set by the memory left after the weights. That is a subtraction, not a ratio. A 70B model in FP8 needs about 71 GB of weights. On two H100s with tensor parallelism, after a working reserve, each GPU keeps about 37 GB for KV cache. On one H200 there are about 62 GB, and on two H200s about 98 GB per GPU. The bigger the model relative to the GPU, the more the extra 61 GB is worth. That is why the H200 advantage can exceed both the 1.43 bandwidth ratio and the 1.76 capacity ratio.

A decode cost model you can run

The model below turns that reasoning into numbers. It sizes the batch for the worst case, every sequence at full context. It then takes the slower of memory time and compute time per step and prices the result. Efficiencies are deliberately conservative placeholders. Replace them with measurements as soon as you have them.

from dataclasses import dataclass

@dataclass
class Gpu:
    name: str
    mem_gb: float        # HBM capacity per GPU
    bw_tbs: float        # HBM bandwidth per GPU, TB/s
    fp8_tflops: float    # dense FP8 tensor peak per GPU
    price_hr: float      # what YOU pay per GPU-hour

@dataclass
class Model:
    params_b: float      # billions of parameters
    weight_bytes: float  # bytes per parameter (1 = FP8, 2 = BF16)
    kv_bytes_tok: int    # KV-cache bytes per token, all layers

def decode_cost(gpu, model, gpus, ctx, reserve_gb=8, mem_eff=0.7, flop_eff=0.5):
    weights = model.params_b * 1e9 * model.weight_bytes
    free = gpus * (gpu.mem_gb - reserve_gb) * 1e9 - weights
    if free <= 0:
        return None                                    # does not fit
    batch = int(free // (ctx * model.kv_bytes_tok))    # worst case: every sequence at full context
    if batch == 0:
        return None
    bytes_step = weights + batch * ctx * model.kv_bytes_tok
    t_mem = bytes_step / (gpus * gpu.bw_tbs * 1e12 * mem_eff)
    t_flop = 2 * model.params_b * 1e9 * batch / (gpus * gpu.fp8_tflops * 1e12 * flop_eff)
    step = max(t_mem, t_flop)
    tok_s = batch / step
    usd_per_mtok = gpus * gpu.price_hr / (tok_s * 3600) * 1e6
    return dict(batch=batch, step_ms=step * 1e3, tok_s=tok_s,
                tok_s_per_gpu=tok_s / gpus, usd_per_mtok=usd_per_mtok)

H100 = Gpu("H100 SXM", 80, 3.35, 1979, price_hr=2.50)
H200 = Gpu("H200 SXM", 141, 4.8, 1979, price_hr=3.40)
L70 = Model(70.6, 1, 80 * 8 * 128 * 2 * 1)   # Llama 3.1 70B: 80 layers, 8 KV heads, dim 128, FP8 KV
L8 = Model(8.0, 2, 32 * 8 * 128 * 2 * 2)     # Llama 3.1 8B: 32 layers, BF16 weights and KV

The KV term is layers × KV heads × head dimension × 2 (keys and values) × bytes per value. That is 160 KiB per token for the 70B model with an FP8 cache, and 128 KiB for the 8B model in BF16.

Worked example: 70B and 8B

Model, contextDeploymentBatchTokens/s (per GPU)$ per M tokens
70B FP8, 4K1 x H10026510.65
70B FP8, 4K2 x H100, TP21093,556 (1,778)0.391
70B FP8, 4K1 x H200922,336 (2,336)0.404
70B FP8, 4K2 x H200, TP22917,355 (3,677)0.257
70B FP8, 16K2 x H100, TP227885 (443)1.569
70B FP8, 16K2 x H200, TP2721,834 (917)1.030
8B BF16, 4K1 x H1001043,3950.205
8B BF16, 4K1 x H2002175,5030.172

Read the per-GPU column against the break-even rule. For the 70B model, one H200 delivers 1.31 times the per-GPU throughput of an H100 pair. That is below the illustrative 1.36 price ratio, so it costs slightly more per token, though it removes tensor-parallel communication, which this model does not charge for. Two H200s deliver 2.07 times the per-GPU throughput of two H100s. They clear any plausible premium, cutting cost per token by about a third, because the H100 pair spends most of its memory on weights. A single H100 is hopeless for this model: two sequences fit. For the 8B model the ratio is 1.62. The H200 wins, but by less, because the weights were never the constraint.

Check the compute bound before trusting any of this. At batch 291 on two H200s, a step needs about 2 × 70.6 × 109 × 291 ≈ 4.1 × 1013 FLOPs. That is about 21 ms at half of dense FP8 peak, against 39.6 ms of memory time, so the step is still memory-bound. Attention arithmetic adds several percent at 4K context and more at long context. Measured results point the same way at smaller magnitude. In MLPerf Inference v4.0, NVIDIA reported the H200 about 45 percent ahead of the H100 on Llama 2 70B, and said part of that gain came from a custom thermal solution. That is one carefully tuned benchmark, not your traffic.

Llama 3.1 70B FP8 decode at 4K context: cost per million tokens (model output)2 x H100, TP2$0.3913,556 tok/s, batch 1091 x H200$0.4042,336 tok/s, batch 922 x H200, TP2$0.2577,355 tok/s, batch 291Illustrative prices: H100 $2.50 and H200 $3.40 per GPU-hour. Your quotes change the bars, not the method.Measure tokens/s per GPUat your SLO and context mixThroughput ratio TH200 over H100, per GPUBuy H200 if P < TP = price ratio per GPU-hour
Model output for the 70B case at 4K context, and the decision rule it feeds. Bars are outputs of decode_cost with the illustrative prices, not measurements.

What the model leaves out

  • Latency is not free. A full H200 step is longer: 39 ms against 31 ms in the model, so each user sees about 25 tokens per second instead of 33. If your service-level target caps per-user speed, cap the batch and recompute. Part of the H200 advantage then turns into lower latency rather than lower cost.
  • Worst-case KV sizing is pessimistic. Real requests are shorter than the maximum context, and paged KV caches pack them tightly. The batch grows on both GPUs, but more on the one with more spare memory.
  • Tensor-parallel communication is free in the model. It is not on real hardware, so configurations that avoid it, such as one H200 instead of two H100s, are undervalued.
  • Prefill is ignored. Prompt processing is compute-bound and runs at about the same speed on both. Workloads dominated by long prompts and short answers see a smaller H200 benefit.
  • The same efficiency is applied to both. Kernels tuned for one memory system may not reach the same fraction of peak on the other. Measure.

Training and fine-tuning economics

Training is mostly compute-bound, so the same step on the same parallel layout runs at nearly the same speed on both GPUs. At a 1.36 price ratio, paying for memory you do not use is a 36 percent cost increase. The H200 earns its premium in training only when the capacity lets you delete an overhead:

  • Turning off full activation recomputation. Recomputing the forward pass costs roughly a third more compute per step.
  • Shrinking tensor or pipeline parallelism, which removes communication and pipeline bubbles.
  • Larger micro-batches that raise kernel efficiency, especially at long sequence lengths.
  • Fitting a fine-tuning job, such as a 70B LoRA run, on fewer GPUs or a single node.

The rule is the same as for serving. Run one short job on each GPU with the best layout for each, measure tokens per second per GPU, and compare the ratio with the price ratio. A configuration that only fits on the H200 is not a reason to buy it unless it also wins on cost per token.

Owning instead of renting

If you buy rather than rent, replace the hourly price with amortised capital cost plus power, cooling, space and operations per GPU-hour, as built up in the LLM TCO model. Both parts share a 700 W ceiling, so facility cost per GPU is similar, and energy per token falls in proportion to the throughput gain. The comparison then hinges on the purchase premium and on utilisation. An owned H200 that sits idle half the day loses to a fully loaded H100. For the arithmetic of tokens and bandwidth on the H100 side, see the H100 article.

Failure modes

  • Comparing list prices, not cost per token. A 36 percent premium looks expensive until the throughput ratio is 2.
  • Benchmarking at a batch that fits both. If the test uses a batch the H100 can hold, you measure only the 1.43 bandwidth effect and miss the capacity effect.
  • Benchmarking offline throughput for an online service. Measure at your latency target and your real prompt and output length mix.
  • Mixing fleets without routing. If long-context traffic lands on H100s, the H200s sit underused. Route by context length and model size.
  • Ignoring the next generation. Prices for both parts move as newer GPUs ship. Re-run the break-even calculation at every renewal instead of locking in a single answer.

What to do next

  1. Collect real quotes for both GPUs for the same term, region and interconnect, and compute the price ratio.
  2. Fill in decode_cost with your model, KV dtype and context distribution, and find where the batch, not bandwidth, sets throughput.
  3. Benchmark your serving engine on one node of each at your latency target, at the largest batch each GPU can hold.
  4. Compare the measured per-GPU throughput ratio with the price ratio, and include the tensor-parallel layouts each GPU makes possible.
  5. Read decode math and KV cache sizing to sharpen the inputs, and record the decision with its assumptions so it can be re-run.
Key takeaway: The H200 sells memory, not compute, so judge it on cost per token. It wins when the H200 to H100 price ratio is below the per-GPU throughput ratio you measure at your latency target. For large models in decode, the capacity left after the weights sets the batch, and that ratio can reach 2. For small models or compute-bound training it sits near 1 to 1.6. Run the model, measure both GPUs at the batch each can hold, and re-check when prices move.