Ask five people what an H100 cost to rent in 2024 and you will hear numbers from about $2 to over $9 an hour, all of them honest. The H100 is the first GPU whose rental price has a public, multi-year history across several kinds of seller, and that history is a good teacher: most of the disagreement comes from comparing different products that happen to contain the same chip.

This page treats price as data. It walks through one published index, explains why its tiers diverge, shows how to turn a quoted dollars-per-GPU-hour into the number that matters (cost per useful hour of training or per million tokens served), and ends with a small model for deciding when to lock in a contract. For the story of scarcity, quotas and lead times, read H100 availability history first; this page assumes it and does not repeat it.

The data: one index, three tiers

Silicon Data publishes an H100 rental index (ticker SDH100RT) and a blog post summarising it by period and by seller type. The medians below are per GPU-hour as that post reports them; they are one vendor's aggregation, not a market-wide truth, but they are dated and tiered, which is what you need.

PeriodHyperscalerNeocloudMarketplace
Aug to Dec 2023$7.76 mediannot in indexnot in index
Jan to May 2024$7.92 mediannot in indexnot in index
Jun to Dec 2024$9.34 median$2.99 median (from Jul)$2.58 median (from Jun)
Jan to May 2025$8.96 median$3.50 median$2.29 median
Jun 2025$6.94$3.29$2.00
Jul to Dec 2025$6.26 median$3.33 median$1.95 median

Two other dated anchors come straight from providers. AWS announced price cuts of up to 45% for P5 (H100) On-Demand instances, effective 1 June 2025, consistent with the June 2025 step in the table (the index tier blends several hyperscalers, so it fell by less). And price moves are not one-way: in January 2026 AWS raised Capacity Block prices for its eight-GPU H200 instances (p5e.48xlarge and p5en.48xlarge) by about 15%, with steeper rises in some regions, as reported by DCD and The Register. Purchase prices for H100 boards and HGX systems were widely reported in the tens of thousands of dollars per GPU but were never published as list prices, so treat any single figure as anecdotal.

Why the tiers diverge

Read the table closely and three lessons appear. First, the index contained only hyperscaler prices until mid-2024. Anyone plotting a single "average" line across that boundary would see a cliff that is really a change in composition: cheaper sellers entered the sample. Always ask what a series contains before reading a trend from it.

Second, hyperscaler and neocloud prices barely moved together. Hyperscaler list prices are sticky and set by price-book decisions; neocloud and marketplace prices react to supply. The hyperscaler premium buys things a bare GPU does not: integration with the provider's storage, identity and networking, the ability to start and stop instances by the hour, enterprise support and contracts your procurement team already has. Most large customers also never pay list; savings plans and private pricing apply.

Third, the bottom tier is not the same product. Marketplace offers mix full eight-GPU HGX nodes with PCIe cards, single GPUs and machines with no high-speed fabric. A $2 GPU in a box with no InfiniBand or RoCE network is excellent for inference and fine-tuning and nearly useless for a 256-GPU pre-training job. Prices are only comparable within a configuration.

Same GPU, different prices: what a quoted H100 rate actually includesHyperscaler liston-demand, per instanceNeocloudon-demand or contractMarketplacespot-like, variable qualityOwnedcapex + power + staffNormalise: $ per GPU-hour, same commitment, same fabric, same storage and egressDivide by goodput: utilisation x (1 - failure loss) x MFU$ per useful hour, $ per million tokensMost "price history" disagreements are tier and unit mismatches, not market disagreements.
Normalising a quote. Compare like with like first, then divide by the fraction of paid hours that produce useful work.

Units that make quotes incomparable

Quotes come in many units, and confusing them is the most common error in price comparisons:

  • Per instance versus per GPU. Cloud price books list eight-GPU instances. Divide by eight, and check whether the quote includes the CPUs, RAM and local NVMe that come with it.
  • Commitment. On-demand, one-year and three-year prices for the same GPU can differ by a factor of two or more. A three-year price bought in 2023 and an on-demand price in 2025 are different financial instruments.
  • Capacity guarantees. Spot and interruptible capacity is cheaper because the provider can take it back. Reserved blocks guarantee start time; on-demand often does not.
  • Fabric and storage. Some providers bundle the cluster network and parallel file system; others bill them separately. Egress for checkpoints and datasets is billed per gigabyte.
  • Currency of the comparison. A 2023 quote paid upfront is worth more than the same dollar figure paid monthly in arrears.

From price to cost per useful hour and per token

The price you pay per GPU-hour is not your cost per unit of work. Three ratios sit between them. Utilisation is the fraction of paid hours the job is actually running (queueing, debugging, idle reservations). Failure loss is the fraction of running time thrown away by crashes and restarts since the last checkpoint. MFU is how much of the GPU's peak arithmetic the job uses while running. The script below turns a quote into dollars per useful hour and per million training tokens.

def effective_cost(price_gpu_hr, utilisation, failure_loss, mfu,
                   params, peak_flops=989e12):
    # peak_flops: H100 SXM dense BF16 (989 TFLOP/s, no sparsity).
    goodput = utilisation * (1 - failure_loss)
    per_useful_hr = price_gpu_hr / goodput
    tokens_per_gpu_s = mfu * peak_flops / (6 * params)       # 6ND training FLOPs
    tokens_per_paid_hr = tokens_per_gpu_s * 3600 * goodput
    return per_useful_hr, price_gpu_hr / tokens_per_paid_hr * 1e6

for name, price, util, loss, mfu in [
    ("hyperscaler on-demand", 6.26, 0.70, 0.05, 0.40),
    ("neocloud 1-yr",         3.33, 0.90, 0.08, 0.40),
    ("marketplace, no fabric",1.95, 0.85, 0.15, 0.25),
]:
    hr, mtok = effective_cost(price, util, loss, mfu, params=8e9)
    print(f"{name:24s} ${hr:5.2f}/useful hr  ${mtok:5.3f}/M tokens")

With these illustrative inputs for an 8B-parameter model the output is roughly $9.41, $4.02 and $2.70 per useful hour, and about $0.32, $0.14 and $0.15 per million training tokens. The headline price ratio between hyperscaler and marketplace was 3.2x; the cost-per-token ratio is about 2.2x, and the marketplace's 40% sticker discount against the neocloud turns into a slightly higher cost per token, because the box without a fabric loses MFU and more work to failures. Your inputs will differ. The point is to measure them, because they move the answer as much as the price does.

Notice also what the hyperscaler row assumed: 70% utilisation. That is typical of a team that holds on-demand instances through debugging sessions and weekends. The same team on a neocloud contract would face the identical utilisation problem, and the contract would make it worse because unused hours are still paid. Utilisation is the input most worth improving, and it is entirely within your control: queue work so that reserved GPUs are never idle, release on-demand capacity aggressively, and run evaluation and low-priority fine-tuning jobs as backfill.

Commit or wait: pricing a contract

The history also prices a decision every buyer faced: commit now or wait. A team that signed a multi-year contract at 2023 prices locked in a rate that the market undercut within about eighteen months. A team that waited could not train at all in 2023. Neither choice was wrong; they bought different things. You can make the trade explicit by comparing a fixed contract price with the expected spot path:

def contract_regret(contract_price, start_spot, monthly_decline, months, gpu_hours_per_month):
    # Positive result: dollars overpaid versus a market that declines geometrically.
    spot = start_spot
    regret = 0.0
    for _ in range(months):
        regret += (contract_price - spot) * gpu_hours_per_month
        spot *= (1 - monthly_decline)
    return regret

# 512 GPUs, 24-month term, 70% of hours used.
hours = 512 * 730 * 0.70
print(contract_regret(3.00, 3.30, 0.02, 24, hours))   # declining market
print(contract_regret(3.00, 3.30, 0.00, 24, hours))   # flat market

Run it with your own decline rate. At this scale a 2% monthly decline turns a contract that starts about 9% below spot into an overpayment of roughly $2.3 million over two years, while a flat market turns the same contract into a saving of about $1.9 million. What the contract really buys is certainty of capacity, so price that certainty explicitly: what does a month of delay cost your product? If the answer is more than the regret, sign. Compare this with the ownership route in bare metal versus cloud GPU.

The owner's view: price history as depreciation

For anyone who owns GPUs, the rental history is also a depreciation curve. The value of an owned H100 is roughly the rental income it can still earn, so when market rates fall, the payback period stretches. A simple model makes this concrete. The all-in cost per GPU below is an assumption for illustration, not a quoted price; substitute your own.

def payback_months(capex_per_gpu, rent_gpu_hr, utilisation, opex_gpu_hr):
    # opex: power, cooling, colocation and staff per GPU-hour, paid whether busy or not
    monthly = rent_gpu_hr * 730 * utilisation - opex_gpu_hr * 730
    return capex_per_gpu / monthly if monthly > 0 else float("inf")

for rent in (3.30, 1.95):
    print(rent, round(payback_months(30000, rent, 0.85, 0.50), 1))

At an assumed $30,000 per GPU and $0.50 per hour of operating cost, earning the neocloud median of $3.30 at 85% utilisation pays back in about 18 months; at the marketplace median of $1.95 it takes about 35 months, close to the point where the next generation makes the part hard to rent at all. The same arithmetic tells an internal platform team what to charge other teams, and when an old fleet is better sold or redeployed to inference than kept on training work.

What the next generation does to the old one

Each generation repeats the pattern. When the H200 and then Blackwell parts shipped, the H100 became the value option and its price fell; the newest part carried the scarcity premium. That makes the previous generation a good default for work that does not need the newest features: fine-tuning, inference of models that fit, and research runs. The question to ask is cost per token on your workload, which is how H200 versus H100 economics frames the comparison. For the longer view of how cost per training run is falling, see the training cost trajectory.

Failure modes

  • Comparing per-node and per-GPU quotes. An eight-GPU instance at $50 an hour is $6.25 per GPU, not $50.
  • Mixing tiers in one average. An index whose composition changes produces fake trends.
  • Survivorship in marketplace quotes. The cheapest listing is often the one nobody rented because it is unreliable; check interruption history and node health.
  • Ignoring the fabric. A cheap GPU without a cluster network forces a smaller job or pipeline-heavy parallelism and lowers MFU.
  • Ignoring idle reserved hours. A reservation used 60% of the time costs 1.67x its sticker price per used hour.
  • Assuming prices only fall. The 2026 H200 Capacity Block increase shows list prices can rise when demand shifts.

Trade-offs

ChoiceGets youCosts you
Hyperscaler on-demandElasticity, integration, procurement simplicityHighest list price, start not guaranteed
Reserved block or contractGuaranteed capacity and fabricPrice risk if the market falls
NeocloudLower price, large fabric-attached clustersVendor diligence, fewer managed services
Marketplace or spotLowest sticker priceInterruptions, mixed hardware, more ops time
Own the hardwareLowest marginal cost at high utilisationCapex, depreciation risk, staff

What to do next

  1. Write down your last three GPU quotes in one unit: dollars per GPU-hour, same commitment term, with fabric and storage noted.
  2. Measure utilisation of your reserved hours for a month; most teams overestimate it.
  3. Measure failure loss from your job logs: lost hours since last checkpoint per interruption.
  4. Plug price, utilisation, failure loss and MFU into the effective-cost script and rank providers on dollars per million tokens.
  5. For any contract longer than six months, run the regret model with a pessimistic decline rate and set a cost-of-delay figure beside it.
  6. Re-run the comparison every quarter and whenever a new GPU generation ships; the previous generation reprices quickly.
Key takeaway: Read H100 prices per tier: in one published index the hyperscaler median fell from about $9.34 per GPU-hour in late 2024 to $6.26 in late 2025, stepping down in June 2025; neocloud prices stayed roughly flat at $3 to $3.50; marketplace prices drifted from $2.58 to $1.95. Most disagreement in quoted prices is tier, unit and commitment mismatch. Normalise quotes, divide by utilisation, failure loss and MFU, compare on cost per token, and price contracts against both market decline and cost of delay.