Most engineers meet the GPU supply chain as a single line in a planning document: the nodes arrive in week 10, or the quota request was approved for half of what we asked. Behind that line is a chain of factories, each with its own capacity, lead time and yield, and the slowest stage decides when your training run starts. You cannot change the chain, but you can stop being surprised by it, and you can write plans and code that keep working when the hardware arrives late, in pieces, or as a different model from the one you designed for.
This article walks the chain from silicon to your scheduler, explains why a few stages keep becoming the bottleneck, and then turns that into engineering practice: a capacity planner that works in tranches, training code that adapts to the device it finds, a burn-in routine for new nodes, and the contract trade-offs. Market numbers change every quarter and most published figures are estimates, so this page deliberately avoids quoting them. Every number below is labelled illustrative, and you should replace it with figures from your own vendors.
The chain, stage by stage
Logic die. The GPU compute die is manufactured at a leading-edge foundry; for the major data-center accelerators today that is TSMC. A wafer takes months to move through hundreds of process steps, and the largest dies sit near the maximum size a lithography scanner can expose in one shot, so only a few dozen fit on a wafer and a defect anywhere in a die can scrap it or force parts of it to be disabled.
HBM. High-bandwidth memory is a stack of DRAM dies connected vertically by through-silicon vias and sitting on a base logic die. Three companies make it: SK hynix, Samsung and Micron. Each stack is itself a small yield problem, because every layer and every vertical connection must work. HBM capacity also competes with ordinary DRAM for the same wafer fabs. The HBM architecture article explains why training needs this memory; here the point is that there are only three sources and each GPU needs several stacks.
Advanced packaging. The compute die and the HBM stacks are placed side by side on a silicon interposer or bridge, which carries thousands of short wires between them, and the assembly is mounted on an organic substrate. TSMC's family of processes for this is called CoWoS (chip on wafer on substrate). Packaging used to be a cheap back-end step. For these accelerators it is a capacity-limited, high-precision process, and it has been one of the most frequently cited constraints on output.
Substrate, module and test. Large package substrates have their own suppliers and lead times. The finished package is mounted on a module or baseboard, tested and burned in. Server and rack. Original design manufacturers integrate GPUs, CPUs, NICs, power supplies and cooling into servers or whole racks. Network. A training cluster also needs switches, NICs, cables and optical transceivers in matching quantities, and optics have been short on their own. Site. Finally the racks need a data hall with enough power and cooling, which the datacenter power article and deployment article cover. A cloud provider then turns all of this into quotas and reservations, which is the only part you see.
Why packaging and memory become the choke points
A GPU package is an assembly, and an assembly is only as good as all of its parts. If you bond a compute die to several HBM stacks and one stack is faulty, you usually cannot rework it: the whole package, including the expensive compute die, is lost. The probability of a good package is roughly the product of every part's yield and the assembly yield, so adding stacks lowers it even if each stack is very reliable. The industry fights this with known-good-die testing, which tests parts before bonding them, and with redundancy inside each die, but the arithmetic still explains why each new generation, with more stacks and larger interposers, ramps slowly.
def package_yield(die_yield, stack_yield, n_stacks, assembly_yield):
"""Probability that an assembled accelerator package is good.
Without known-good-die testing every part is bonded blind, so the
package survives only if every part and the assembly step survive.
"""
return die_yield * (stack_yield ** n_stacks) * assembly_yield
# Illustrative numbers only, not vendor data.
for stacks in (4, 6, 8):
y = package_yield(die_yield=0.80, stack_yield=0.97, n_stacks=stacks, assembly_yield=0.98)
print(stacks, "stacks ->", round(y, 3))
# 4 stacks -> 0.694 6 stacks -> 0.653 8 stacks -> 0.614The second reason is capacity lead time. A new packaging line or HBM fab takes years to plan and build, and demand for AI accelerators rose faster than those plans. When one stage is short, everything downstream idles: finished compute dies wait for packaging, and packaged GPUs wait for network gear or a powered hall. For you the useful lesson is that the delivery date you are quoted is the minimum over several independent stages, so a slip anywhere moves it, and the stages late in the chain (network, power, integration) slip as often as the famous ones.
What reaches you: allocation, not a price list
At the end of the chain, capacity is allocated, not simply sold. Cloud providers ration new-generation accelerators through quotas per region and per project, time-bound reservations and committed-use contracts, and short-term capacity products that let you book a fixed block of nodes for days or weeks. On-demand capacity for the newest parts is often unavailable in the region you want, while the previous generation is easier to get. If you buy hardware, you receive a delivery schedule in tranches, and each tranche still has to be racked, cabled and burned in before it is useful.
Three consequences shape engineering. First, you will usually receive capacity in pieces, so a plan that needs every GPU on day one will wait for the last piece. Second, capacity may land in a region you did not choose, so data, images and credentials must be ready to move; the region selection article treats quotas and capacity as a hard filter for that reason. Third, generations overlap: you may run two GPU models side by side for a year, and code that assumes one memory size or one numeric format will break or waste capacity on the other.
Worked example: planning a run around tranches
Suppose a team needs about 1.5 million GPU-hours for a pre-training run. The vendor promises 128 GPUs in week 4, another 256 in week 10 and 512 more in week 16. These are illustrative numbers. The naive plan waits for all 896 GPUs and starts in week 16. The planner below walks the schedule week by week, assumes 85 percent of wall-clock time turns into useful training after failures and checkpoints, and reports when the job finishes.
from dataclasses import dataclass
@dataclass
class Tranche:
week: int # week the nodes are handed over and pass burn-in
gpus: int
def weeks_to_finish(gpu_hours_needed, tranches, efficiency=0.85,
min_gpus_to_start=256):
"""Walk week by week; training starts once enough GPUs exist.
efficiency covers failures, restarts and checkpoint overhead.
Returns the week the job finishes, or None within the horizon.
"""
done, gpus = 0.0, 0
arrivals = {t.week: t.gpus for t in tranches}
for week in range(0, 104):
gpus += arrivals.get(week, 0)
if gpus >= min_gpus_to_start:
done += gpus * 24 * 7 * efficiency
if done >= gpu_hours_needed:
return week
return None
plan = [Tranche(4, 128), Tranche(10, 256), Tranche(16, 512)]
print(weeks_to_finish(1_500_000, plan)) # 25
print(weeks_to_finish(1_500_000, plan, min_gpus_to_start=128)) # 24With a 256-GPU minimum, training starts in week 10 and finishes in week 25. Starting on the first 128 GPUs in week 4 finishes in week 24. One week is a modest gain, but the six early weeks are also time to find data-loader bugs, validate loss curves and harden checkpointing before the expensive part begins. So design the run to start small and grow. That needs two properties from your code: it must tolerate a changing world size, and its checkpoints must load at a different parallel layout.
Engineering for the hardware you actually get
Elastic world size. PyTorch's torchrun accepts a node range such as --nnodes=4:16, restarting workers with a new world size when membership changes. That only helps if your job can resume from a checkpoint at the new size, and if global batch size stays constant by adjusting gradient accumulation rather than letting it drift with the number of GPUs, which would silently change the optimisation.
Reshardable checkpoints. Save sharded state with a format that records the logical tensor layout, such as torch.distributed.checkpoint, which can load a checkpoint written by one sharding configuration into another. Test this path before you need it: save on 64 GPUs, load on 128, and compare the loss on a fixed batch.
Derive settings from the device. Micro-batch size, activation checkpointing and numeric format should come from the device you find, not from constants tuned on one SKU. The sketch below reads memory and compute capability at start-up; static_state_gb stands for your own estimate of weights, gradients and optimizer state per GPU. FP8 tensor cores first appeared with the Hopper and Ada generations, so an FP8 path must have a BF16 fallback.
import torch
def plan_from_device(model_params_b, seq_len, target_tokens_per_step):
"""Derive micro-batch and accumulation from the GPU actually present,
instead of hard-coding values tuned for one SKU."""
props = torch.cuda.get_device_properties(0)
mem_gb = props.total_memory / 2**30
world = torch.distributed.get_world_size()
# Crude, calibrated per model: measure peak memory at micro-batch 1, then scale.
per_sample_gb = 0.004 * model_params_b * seq_len / 1024 # fit this from a probe run
headroom_gb = mem_gb * 0.80 - static_state_gb(model_params_b, world)
micro = max(1, int(headroom_gb // per_sample_gb))
tokens_per_micro_step = micro * seq_len * world
accum = max(1, round(target_tokens_per_step / tokens_per_micro_step))
fp8 = props.major >= 9 or (props.major, props.minor) == (8, 9) # Hopper, Ada
return dict(micro_batch=micro, grad_accum=accum, fp8_capable=fp8,
device=props.name, mem_gb=round(mem_gb, 1))Serving on mixed fleets. For inference, treat each GPU model as a separate pool with its own measured throughput and memory limit, route requests by model size and context length, and keep a quantized variant ready for pools whose memory is too small for the full-precision weights. The router's capacity table should come from benchmarks you run on each pool, not from the spec sheet.
Burn-in and early failures
New hardware fails most in its first weeks: marginal HBM, bad cables, optics that flap, cooling that was never commissioned under full load. A node that passes a quick health check can still corrupt a large run by producing NaNs or slow collectives. Before handing a tranche to training, run a burn-in: NVIDIA's DCGM diagnostics at a long level (dcgmi diag -r 3, or the longer levels your DCGM version provides), a multi-hour stress load, NCCL all-reduce benchmarks across every node pair and across the full set, and a short real training job with a known loss curve. Record per-node results, drain anything that is an outlier in bandwidth or error counters, and repeat the collective tests after any recabling.
Keep the burn-in suite in version control and run it again whenever a node returns from repair. Collect the failures by component, because the vendor's replacement and warranty process needs this evidence, and because the pattern tells you whether you have a bad batch or a site problem such as cooling.
Failure modes
- Plan assumes all capacity on day one. The run slips to the latest tranche and the team idles. Design for start-small-and-grow.
- Code pinned to one SKU. Hard-coded batch sizes or an FP8-only path cannot use the previous-generation capacity you can actually get.
- Capacity lands in the wrong region. Data sets, container images and secrets are only in the original region, and replicating many terabytes adds weeks.
- Stranded GPUs. Accelerators arrive but switches, optics or power do not, and paid-for hardware sits dark. Track every stage, not only the GPUs.
- Skipping burn-in. A marginal node joins a large job and fails repeatedly, costing far more than the burn-in would have.
- Take-or-pay without a workload. A long commitment signed in a panic is paid for months after the project that needed it ends.
Trade-offs
| Option | Gets you | Costs you |
|---|---|---|
| On-demand cloud | No commitment, fast start when available | Highest unit price; newest parts often unavailable |
| Reservations and committed use | Guaranteed capacity and a lower rate | Payment whether used or not; region and SKU locked |
| Short-term capacity blocks | Fixed dates for a bounded run | Must be ready exactly on time; limited sizes |
| Specialist GPU clouds | Earlier access to some parts | Smaller footprint, fewer managed services, counterparty risk |
| Buying hardware | Lowest cost at high utilisation, full control | Long lead times, site build, staff, obsolescence |
The decision turns on utilisation and certainty. If you can keep GPUs busy for most of their life and your demand is predictable, owning or committing is cheaper. If demand is uncertain, paying more per hour for flexibility is the rational choice, and the code properties above keep you able to use whatever capacity turns up.
What to do next
- Write down, for your next run, which stage of the chain you depend on: cloud quota, reservation, or vendor delivery, and who owns each date.
- Run the tranche planner with your real schedule and set a minimum GPU count that lets you start early.
- Make batch size, accumulation and numeric format derive from the device, and keep a BF16 fallback for every FP8 path.
- Test checkpoint resharding across two world sizes and a resume under torchrun with a node range.
- Put a burn-in suite (DCGM diagnostics, NCCL tests, a known-loss job) in version control and gate every new tranche on it.
- Keep data, images and secrets replicable to a second region so capacity in the wrong place is still usable.