"Is owning GPUs cheaper than renting them?" has a useless true answer: it depends on utilization. This article turns that into a number you can compute. There are four ways to buy a GPU-hour: on-demand cloud instances, committed cloud capacity, rented bare-metal servers, and hardware you own and host in colocation. They differ in price, in how long you are locked in, in what you control at the hardware level, and in who fixes a failed node. Each is cheapest for a different slice of your demand, so the real question is not which one to choose but how much of each to hold.
The method is a demand-duration curve, a commit-level rule that follows from first principles, and a four-year scenario model that prices the risk of committing to capacity you later do not need. Every number below is the output of the model shown in the article, run on illustrative rates. None of them are market prices. GPU pricing moves too quickly and varies too much by contract to quote. Put in your own quotes and the method stays the same. The cost of owned hardware per hour, including capital recovery, power and staff, is derived in GPU infrastructure cost and is used here as a single input.
Four ways to buy a GPU-hour
| Mode | Commitment | You control | Provider handles | Illustrative rate per GPU-hour |
|---|---|---|---|---|
| On-demand cloud VM | None; per second or hour | Guest OS, containers | Hardware, hypervisor, repairs, network | 4.00 |
| Committed cloud | Typically one or three years | Same as on-demand | Same; capacity is held for you | 2.80 |
| Rented bare metal | Monthly to multi-year per server or block | Full OS, drivers, firmware settings, topology | Hardware swap, power, network | 2.20 |
| Owned in colocation | Hardware life, here four years | Everything | Space and power only | 1.90 fully loaded |
The ratio of each rate to the on-demand rate is what drives the analysis: 0.70 for committed cloud, 0.55 for bare metal and about 0.47 for owned. The owned rate is per allocated hour and already includes capital recovery over four years, power, space, support and staff. If your owned rate is computed from the purchase price alone, it is too low and every conclusion below will lean towards owning.
What bare metal changes for training and inference
"Bare metal" means no hypervisor between your operating system and the hardware. For training and inference, the hypervisor's compute overhead with GPU passthrough is usually small. What changes is control and visibility. On bare metal you can read the real topology with nvidia-smi topo -m, choose the driver and CUDA versions, set GPU clocks and persistence mode, partition GPUs with MIG if the hardware supports it, and tune NCCL for the actual InfiniBand or RoCE fabric instead of a virtualised one. For large collective-heavy training, the fabric often matters more than the hypervisor. A rented bare-metal cluster with a non-blocking fabric can beat a cloud instance type with a thinner network, and the reverse is just as possible. Measure step time on your own job before accepting anyone's benchmark.
Control has a price, paid in engineering time. On bare metal you own the image, the driver upgrades, the health checks, and the decision to drain a node whose GPU is throwing ECC errors. Hyperscalers sell all of that as part of the instance. The CoreWeave and Latitude.sh profiles describe two providers' positions on this spectrum. In the cost model, that engineering time belongs in the bare-metal and owned rates, not in a footnote.
The commit-level rule from first principles
Suppose you commit to L GPUs at a rate r_c per GPU-hour, paid every hour, and buy anything above L on demand at r_od. Should you add one more committed GPU, at level L+1? It costs r_c every hour. It saves r_od only in the hours when demand is above L, because only then would you otherwise rent a GPU on demand. If P(D > L) is the share of hours in which demand exceeds L, the extra GPU pays for itself exactly when:
P(D > L) * r_od >= r_c <=> P(D > L) >= r_c / r_odSo the rule is: commit up to the demand level that is exceeded in at least r_c / r_od of all hours. This is the classic newsvendor result, and it gives each buying mode a utilization break-even. Committed cloud at a ratio of 0.70 is worth it for capacity used at least 70 percent of the time. Bare metal at 0.55 for capacity used 55 percent of the time. Owned at 0.47 for capacity used about half the time. Utilization here means the share of hours the capacity is actually needed, not GPU busy-percent inside those hours. That second inefficiency affects every mode equally and is covered in GPU cost optimization.
Worked example: one year of demand
The worked example is a team with two workloads. A training cluster runs 64-GPU jobs and is busy on a random 75 percent of days. Between runs it falls back to 16 GPUs for experiments. Inference follows a daily cycle between 8 and 40 GPUs. Simulated hour by hour for a year, demand averages 75.8 GPUs, ranges from 24 to 104, and totals 664,360 GPU-hours. Sort the hours by demand and you get the demand-duration curve:
| Share of hours demand is at least | 95% | 70% | 55% | 50% | 30% |
|---|---|---|---|---|---|
| GPUs | 27 | 73 | 77 | 80 | 93 |
Applying the rule to the curve: committed cloud alone should cover 73 GPUs, bare metal alone 77, and owned alone 80. A brute-force search over every commit level gives the same three answers, which is a useful check that the model and the rule agree. Year-one costs, each with on-demand burst above the commit level:
| Strategy | Commit level | Year-1 cost | vs all on-demand |
|---|---|---|---|
| All on-demand | 0 | 2,657,440 | baseline |
| Committed cloud + burst | 73 | 2,189,064 | -18% |
| Bare metal + burst | 77 | 1,796,064 | -32% |
| Owned + burst | 80 | 1,588,560 | -40% |
Two things stand out. Committing at all saves 18 percent, and the choice of mode moves the saving from 18 to 40 percent, so both decisions matter. And the floor sits near the median of demand, not near the minimum. Teams that size committed capacity to the level met in 95 percent of hours (27 GPUs here) leave most of the savings unclaimed.
Adding time: four years and three futures
One year of known demand flatters ownership. Owned hardware is a four-year commitment made before you know what the next three years look like, while bare metal and committed cloud can be resized each year. The model therefore runs three four-year scenarios. In the growth scenario, demand scales by 1, 1.5, 2 and 2.5 in successive years. In the flat scenario it stays at 1. In the shrink scenario it goes 1, 0.8, 0.5 and 0.3. The scenarios are weighted 0.4, 0.4 and 0.2. Each year above the owned floor is filled optimally with bare metal plus on-demand burst.
| Owned floor (GPUs) | Grow | Flat | Shrink | Expected (4 years) |
|---|---|---|---|---|
| 0 (bare metal + burst only) | 12.62M | 7.21M | 4.69M | 8.87M |
| 32 | 12.28M | 6.88M | 4.42M | 8.55M |
| 48 | 12.11M | 6.71M | 4.61M | 8.45M |
| 64 | 11.94M | 6.54M | 5.03M | 8.40M |
| 80 | 11.78M | 6.38M | 5.60M | 8.38M |
| 104 | 11.73M | 6.92M | 6.92M | 8.85M |
For comparison, all on-demand has an expected cost of 13.22M and committed cloud with burst 10.81M. The expected cost of the owned-plus-bare-metal mix reaches its minimum at a floor of 80, but the curve is nearly flat from 48 to 80. The difference across that whole range is 0.07M, under 1 percent. The shrink scenario is not flat at all: a floor of 80 costs 5.60M there against 4.61M at 48. Choosing 48 gives up less than 1 percent of expected savings to remove about 1M of downside. That is the general lesson: choose the owned floor from the flat part of the expected-cost curve that has the smallest regret in the bad scenario, and let shorter commitments absorb the rest. Past about 110 GPUs, owning is worse than the bare-metal-only option in expectation.
The model as code
The model fits in a page. It is deliberately simple: hourly demand, flat rates, on-demand burst with no availability limit. Extend it with your own demand logs before trusting its output.
import math, random
HOURS = 8760
RATES = {"on_demand": 4.00, "cloud_reserved": 2.80, "bare_metal": 2.20, "owned": 1.90} # illustrative
def year_demand(scale, seed=7):
rng, d, training_on = random.Random(seed), [], True
for h in range(HOURS):
if h % 24 == 0:
training_on = rng.random() < 0.75 # training busy on ~75% of days
infer = 24 + 16 * math.sin(2 * math.pi * (h % 24) / 24) # daily cycle, 8..40 GPUs
d.append(math.ceil(((64 if training_on else 16) + infer) * scale))
return d
SCALES = {"grow": ([1, 1.5, 2, 2.5], 0.4), "flat": ([1, 1, 1, 1], 0.4), "shrink": ([1, 0.8, 0.5, 0.3], 0.2)}
scenarios = {name: ([year_demand(f, seed=7 + i) for i, f in enumerate(fs)], prob)
for name, (fs, prob) in SCALES.items()}
def burst(d, level):
return sum(max(0, x - level) for x in d) # GPU-hours above the commit level
def best_flat(d, rate):
"""Commit level minimising committed spend plus on-demand overflow (brute force)."""
return min(((c, c * HOURS * rate + burst(d, c) * RATES["on_demand"])
for c in range(max(d) + 1)), key=lambda t: t[1])
def owned_plus_bm(d, n_owned):
"""Owned floor fixed for the hardware's life; bare metal resized yearly above it."""
return min(((b, (n_owned * RATES["owned"] + b * RATES["bare_metal"]) * HOURS
+ burst(d, n_owned + b) * RATES["on_demand"])
for b in range(max(d) + 1)), key=lambda t: t[1])
def expected_cost(n_owned, scenarios):
# scenarios: {name: (list_of_hourly_demand_per_year, probability)}
return sum(prob * sum(owned_plus_bm(d, n_owned)[1] for d in years)
for years, prob in scenarios.values())
for n in range(0, 130, 8):
print(n, round(expected_cost(n, scenarios) / 1e6, 2))Feed it real hourly GPU demand from your scheduler or capacity plan, not averages. The curve's shape is the whole answer, and averages hide it.
Costs the rate card leaves out
The flat rates leave out costs that often decide the outcome in practice. Check each one before you sign.
- On-demand availability. The model assumes burst capacity always exists. For high-end GPUs it often does not, at least not in the region or quantity you need. If burst is unreliable, part of the burst layer must be committed, or covered by preemptible capacity as described in GPU spot instances.
- Lead time. Owned clusters take months to procure, rack and burn in. Rented bare metal takes weeks, and on-demand takes minutes when capacity exists. Cost the months of on-demand spend while you wait.
- Data gravity. Egress charges and storage near the GPUs can outweigh a rate difference when datasets are large or checkpoints move between providers. Put a per-terabyte line in the model.
- Failure handling. GPUs and optics fail. With owned hardware you hold spares and staff. With bare metal, read the contract's replacement window and what it credits. A node down for days in a 64-GPU job idles the whole job.
- Obsolescence. A new GPU generation can cut the cost per unit of work, which shrinks the value of a four-year commitment. Shorter hardware life in the owned rate is the honest way to model that risk.
- Software and people. Bare metal and owned capacity need people who can run drivers, schedulers and fabrics. If you do not have them, their cost belongs in the rate.
Failure modes in the analysis
- Sizing to average demand. The average (75.8 here) says nothing about the curve. Use hourly data.
- Purchase price as the owned rate. Without capital recovery, power and staff, ownership looks far cheaper than it is.
- Ignoring demand uncertainty. A one-year analysis picks the largest owned floor. Run scenarios and look at regret.
- Comparing peak benchmarks. Vendor numbers on different fabrics do not transfer. Measure cost per completed step or per million tokens on your own job.
- Assuming burst is free to obtain. When on-demand capacity is unavailable, the real alternative is waiting, and waiting has a cost.
- Locking every layer at once. Owning the floor and committing the middle for the same four years removes the flexibility that made the mix cheap.
What to do next
- Export a year of hourly GPU demand per workload from your scheduler and build the demand-duration curve.
- Collect quotes for each mode and compute each rate's ratio to on-demand. Include staff and capital recovery in the owned rate.
- Apply the commit-level rule to find each mode's break-even level, then confirm it with the brute-force model.
- Write three demand scenarios with probabilities, run the four-year model, and choose the owned floor with the smallest bad-case regret on the flat part of the curve.
- Add availability, lead time, egress and failure-window terms, and re-run.
- Benchmark one real training step or inference load on bare metal and in the cloud before committing.
- Re-run the model every quarter with new demand data and resize the shorter-commitment layers.