Between 2023 and 2025 the NVIDIA H100 went from the scarcest piece of computing hardware on the market to something you could usually rent by the hour. That history is worth knowing not as trivia but because each new GPU generation repeats it: a launch, a scramble, a loosening and a price slide. Teams that understood the pattern the first time planned training runs and serving fleets around it; teams that did not lost months waiting for quota that was never going to turn into machines.

This article walks through the dated milestones that can be checked, the three phases they mark, the mechanisms behind them, and what changed for an engineer trying to get GPUs in each phase. It then gives a procurement model with code, a worked example and the failure modes that outlived the shortage. The supply chain itself is covered in GPU supply chain.

A dated timeline

Precise market data is scarce, so this timeline uses dated announcements and labels secondary reports as reports:

WhenEventWhy it mattered
March 2022H100 announced at GTCDemand forecasts were set before the LLM boom
October 2022US export rules restrict top data-centre GPUs to ChinaMarket splits; cut-down variants appear
Late 2022 to 2023LLM training wave beginsDemand jumps far beyond plans
26 July 2023AWS P5 instances generally availableFirst hyperscaler H100 GA
7 August 2023Azure ND H100 v5 generally availableSecond major cloud
29 August 2023Google says A3 will be GA the following monthAll three big clouds by autumn
September 2023TSMC warns advanced packaging stays tight for about 18 monthsSupply limit is packaging, not wafers
October 2023Export rules tightened againVariants built for the 2022 rules also restricted
1 November 2023EC2 Capacity Blocks for ML GA (P5, US East Ohio first)Short, dated reservations become a product
Late 2023Lead times reported at 40 to 52 weeksPeak of scarcity
2024Lead times reported falling to 8 to 12 weeks; H200 shipsLoosening; resale and neoclouds add supply
2025Blackwell rampsH100 becomes the previous generation
1 June 2025AWS cuts P5 on-demand prices by 44%Commodity pricing at a hyperscaler

Three phases

Phase 1, scarcity (2023). Supply was set by CoWoS packaging and HBM, demand by every lab and startup at once. Clouds rationed through quota approvals and sales conversations; neoclouds sold multi-year contracts with prepayment; on-demand H100s in popular regions were effectively unavailable. Large buyers got hardware by committing early and big. Everyone else waited or used A100s.

Phase 2, loosening (2024). Packaging capacity grew, early buyers found they had more than they could use and resold or subleased, and new providers brought clusters online. Products appeared for the middle of the market: reservations measured in days or weeks rather than years. Reported lead times dropped from quarters to weeks. Short training runs became plannable.

Phase 3, commodity (2025). With H200 and then Blackwell shipping, the H100 became the value option. Hyperscalers cut list prices, spot and on-demand capacity became normal in most regions, and the scarce resource moved to the new generation and to very large contiguous clusters with a fast fabric.

Three phases of H100 availability and how teams got capacity202320242025Scarcityallocation, long contractscloud quota rarely honouredlead times reported in quartersLooseningshort reservations appearcapacity blocks, neocloudslead times reported in weeksCommodityprice cuts, Blackwell rampon-demand usually workslarge clusters still bookedConstant across all three: quota is permission, not capacitydesign for InsufficientInstanceCapacity and preemption in every phaseSupply drivers: CoWoS, HBMDemand drivers: LLM training wave
The three phases. What changed was how capacity was obtained; what did not change was that quota approval never guaranteed machines.

Why the curve had this shape

Three forces explain the shape of the curve, and all three will recur.

Supply is set years ahead. An H100 is a large die joined to HBM stacks on a silicon interposer using TSMC's CoWoS packaging. Wafer capacity was not the limit; packaging lines and HBM output were, and both take a year or more to expand. TSMC's own guidance in September 2023 that the tightness would last about eighteen months is the clearest statement of that lag.

Demand arrived as a step. Forecasts made when the part was announced in 2022 assumed steady growth in data-centre AI. The LLM wave instead made every frontier lab, cloud and well-funded startup try to buy tens of thousands of GPUs in the same few quarters. A fixed supply line meeting a step in demand produces rationing, not just higher prices, because the vendor allocates to strategic customers first.

Hoarding amplifies both directions. When delivery is uncertain, buyers order more than they need and sign longer contracts than they want. When supply catches up, the same over-ordering becomes surplus, which resurfaces as resale, subleasing and new rental capacity. That is a large part of why the 2024 loosening was faster than packaging growth alone would suggest, and why prices fell further in 2025 as newer parts arrived.

Export controls added a fourth, regional force: they split the market into regions that could buy the full part and regions that could not, which shifted demand between products rather than changing total supply.

Quota is not capacity

The most expensive misunderstanding of 2023 was treating a quota increase as a promise. Quota is permission to request instances; capacity is whether the provider has machines free in that zone at that moment. Each cloud reports the difference with its own error: InsufficientInstanceCapacity on AWS, ZONE_RESOURCE_POOL_EXHAUSTED on Google Cloud, allocation failures on Azure. A launcher that treats these as fatal gives up; one that treats them as normal tries other zones, regions and instance types, and falls back to a reservation.

Capacity Blocks turned that fallback into an API. You ask for offerings by instance type, count, duration and date window, and buy one. The request requires a duration in hours; each block holds up to 64 instances:

import sys, boto3, datetime as dt

ec2 = boto3.client("ec2", region_name="us-east-2")
now = dt.datetime.now(dt.timezone.utc)

offers = ec2.describe_capacity_block_offerings(
    InstanceType="p5.48xlarge",
    InstanceCount=4,                       # 32 GPUs
    CapacityDurationHours=7 * 24,          # 1-day steps up to 14 days
    StartDateRange=now,
    EndDateRange=now + dt.timedelta(days=21),
)["CapacityBlockOfferings"]
if not offers:
    sys.exit("no offering in window; widen the dates or try another type or region")

best = min(offers, key=lambda o: (o["StartDate"], float(o["UpfrontFee"])))
print(best["CapacityBlockOfferingId"], best["AvailabilityZone"],
      best["StartDate"], best["UpfrontFee"], best["CurrencyCode"])
# ec2.purchase_capacity_block(CapacityBlockOfferingId=best["CapacityBlockOfferingId"],
#                             InstancePlatform="Linux/UNIX")   # commits money: review first

The lesson generalises beyond AWS: if your plan says when the run starts, you need an instrument that says the same thing. See reserved capacity for long commitments and spot instances for the opposite end.

A procurement model

A procurement decision is a trade between price, start date and the risk of not getting the machines at all. A short model makes the trade explicit. For each option, estimate the probability capacity is there when you need it and the delay if it is not, then rank by expected cost including the cost of waiting:

from dataclasses import dataclass

@dataclass
class Option:
    name: str
    usd_per_gpu_hour: float
    p_available: float      # chance the capacity is there at the planned start
    slip_days: float        # expected delay when it is not
    commit_hours: float     # hours you pay for regardless of use (0 for on-demand)

def expected_cost(o, gpu_hours, gpus, delay_cost_per_day):
    pay_hours = max(gpu_hours, o.commit_hours * gpus)
    compute = pay_hours * o.usd_per_gpu_hour
    waiting = (1 - o.p_available) * o.slip_days * delay_cost_per_day
    return compute + waiting

def rank(options, gpu_hours, gpus, delay_cost_per_day):
    return sorted(((expected_cost(o, gpu_hours, gpus, delay_cost_per_day), o.name)
                   for o in options))

The rates are yours to fill in from current quotes; the model deliberately contains none. What it captures is that in scarcity, p_available for on-demand falls toward zero and the waiting term dominates, so expensive committed capacity wins; in a commodity market the waiting term vanishes and the cheapest flexible option wins. The same inputs come from capacity planning.

Worked example: one run in three years

Worked example: a team needs 32 H100s for a fine-tuning run of 20,000 GPU-hours, about 26 days on 32 GPUs, and each day of delay costs them a fixed amount in salaries and a missed launch, call it D.

In mid-2023 on-demand P5 was rarely obtainable, so p_available was near zero with an open-ended slip. A neocloud contract with a one-year minimum meant paying for 32 GPUs times 8,760 hours, about 280,000 GPU-hours, to use 20,000. Unless D was very large, the rational choices were a smaller run on A100s, or sharing a contract with other teams to fill the remaining hours.

From November 2023, Capacity Blocks let the same team buy four instances for a dated window: committed hours matched used hours, and p_available became a question of which start date the offering search returned. By mid-2025, on-demand P5 at the reduced price often had capacity, making a plain on-demand launch with checkpointing the cheapest plan, with a Capacity Block as backup if the start date was fixed. The run and its GPU-hours did not change; the right instrument did. Estimating the GPU-hours is covered in H100-hours per model.

Habits that outlived the shortage

Several engineering habits formed during the shortage are worth keeping in any phase:

  • Portable stacks. Teams whose images, drivers and launchers ran on any provider could take capacity wherever it appeared. Pin CUDA and NCCL versions in a container, not in a host image.
  • Checkpoint as if preempted. Frequent, asynchronous checkpoints to object storage let a run move between providers or survive a reservation ending.
  • Elastic training. Code that tolerates a changing world size can start on what is available and grow.
  • Measure delivered throughput. Scarce GPUs were often in poorly connected clusters. Run a communication benchmark before accepting a cluster, not after the run is slow.
  • Watch the next generation. The pattern repeats; the time to book the new part is before its scarcity peak, and the time to buy the old one is after.

Reading where a part sits in the cycle

Because the cycle repeats, it pays to watch for where a part sits in it. Useful signals, roughly in the order they move:

  • Instrument availability. Whether short reservations and on-demand launches for the part succeed in your regions. Run a weekly offering search and launch test and keep the history; it is the most direct measurement you will get.
  • Contract terms. Minimum commitment length and prepayment demanded by providers. They shorten as scarcity ends.
  • Reported lead times. Server vendors comment on these publicly; treat single reports as noise and trends as signal.
  • List price changes. Hyperscaler price cuts, like the 2025 P5 cut, arrive late in the cycle and confirm a phase change that the other signals showed first.

Record these alongside your own launch success rate. A team that can say how often a 32-GPU request succeeded last month in each region has a planning input no market report provides.

Failure modes

Failure modes seen repeatedly across the three phases:

  • Quota mistaken for capacity, with launch dates set from a quota email.
  • Overcommitting at the peak. Multi-year contracts signed in 2023 at scarcity prices kept running through the 2025 price cuts.
  • Single-zone launchers that fail on the first capacity error instead of searching zones and types.
  • Reservation cliffs. A Capacity Block ends at a fixed time; a run without a recent checkpoint loses the work since the last one.
  • Fabric surprises. Capacity obtained in pieces lands in different placement groups or providers, with far less bandwidth between them than inside one.

Trade-offs

Committing early buys certainty at the price of flexibility and of paying for idle hours. Staying on-demand keeps flexibility and benefits from price declines but carries start-date risk that, at the height of a shortage, is close to certain failure. Short dated reservations sit between them. Newer generations change the arithmetic again: H200 versus H100 economics shows how a faster part at a higher price can still win per unit of work. Choose by which phase the part you need is in, not by what worked last time.

What to do next

  1. Write down your next run's GPU-hours, GPU count, earliest start and cost of each day of delay.
  2. Fill in the procurement model with current quotes for on-demand, short reservations, contracts and spot, and rank them.
  3. Make your launcher retry across zones, regions and instance types on capacity errors, and log every failure.
  4. Run a Capacity Block offering search (or your provider's equivalent) for your window to learn real start dates before you need them.
  5. Containerise drivers and frameworks so capacity from another provider is usable within a day.
  6. Test checkpoint and resume by killing a run on purpose.
  7. Benchmark interconnect bandwidth on any new cluster before starting the real job.
Key takeaway: H100 availability moved from scarcity in 2023 through loosening in 2024 to commodity pricing in 2025, and each new generation is likely to repeat the cycle. Treat quota as permission, match the purchase instrument to the phase and your deadline, and keep your stack portable and checkpointed so you can take capacity wherever it appears.