The price of a GPU is the number everyone quotes and the number that matters least on its own. What a business pays for is useful GPU time: hours in which a GPU runs work someone wanted, at a cost that includes the server around it, the network fabric, the storage, the electricity and cooling, the building, support contracts, spares and the people who keep it running, all spread over a life that ends when the hardware stops being worth its power.
This article builds that cost from first principles as a model you can run. It turns capital into an hourly charge, adds operating costs, divides by the fraction of hours that do useful work, then uses the result to test sensitivity, compare against rented capacity and price a unit of work. It deliberately stops before the decisions those numbers feed: choosing hardware and renting versus buying are covered in GPU selection, in depth, sizing in GPU infrastructure planning and the levers for reducing spend in GPU cost optimization.
The cost stack
Every owned GPU-hour carries six kinds of cost. Capital: the servers, the share of network switches, optics and cables, and the storage cluster that feeds them. Cost of capital: money tied up in hardware has a price, whether it is borrowed or would otherwise earn a return. Power and cooling: the energy the IT equipment draws, multiplied by the facility's power usage effectiveness (PUE) to include cooling and distribution losses. Facility: colocation or data hall space, usually priced per kilowatt of IT capacity per month whether you use it or not. Support and spares: hardware maintenance contracts, software licences and replacement parts. People: the engineers who rack, image, monitor and repair the fleet.
Two denominators then turn spend into a unit price. The first is allocated hours: every hour in the year, for every GPU you own. The second, and the one that matters, is useful hours: hours spent on work that completed and was wanted. The gap between them is idle time, scheduling gaps, failed jobs, restarts from checkpoints and capacity held back as headroom.
Turning capital into an hourly charge
Dividing purchase price by years of life ignores the cost of capital and understates the hourly charge. The standard tool is the capital recovery factor: the constant annual payment that repays an amount over n years at interest rate r, CRF = r(1+r)^n / ((1+r)^n - 1). At 8 percent over 4 years it is about 0.302, so each dollar of capex costs about 30 cents a year, not 25.
The life you choose matters more than almost any price. Accelerators lose economic value as newer generations deliver more work per watt and per dollar, so a GPU can be physically healthy but uneconomic to run long before it fails. Finance teams often depreciate servers over five or six years; for GPUs bought for frontier training, three to four years of economic life is a more defensible planning assumption, with any residual value treated as upside. Use the same life in your cost model that you would bet your budget on.
Power, facility and the rest
Power is easier to compute than people expect. An H100 SXM GPU has a maximum board power of 700 W; an 8-GPU server with CPUs, memory, NICs and fans is rated around 10.2 kW at maximum for NVIDIA's DGX H100. Servers rarely average their nameplate rating, so multiply by an average load factor, then by PUE, then by the energy price and 8,760 hours. For facility charges, use the contracted IT kilowatts, because you pay for reserved capacity whether you draw it or not.
Power matters far more as a constraint than as a cost: at a few cents per GPU-hour it rarely dominates the bill, but the megawatts available at a site cap how many GPUs you can install at all, as GPU datacenter power explains. Support contracts are usually quoted as a percentage of hardware cost per year. Staff cost is best expressed per GPU per year, which scales with fleet size and automation maturity.
The model as code
from dataclasses import dataclass
HOURS = 8760
@dataclass
class Node: # one 8-GPU server and its share of the cluster
server_capex: float = 300_000 # illustrative purchase price, USD
fabric_storage_frac: float = 0.25 # network and storage, as a share of server capex
life_years: int = 4
cost_of_capital: float = 0.08
it_kw: float = 10.2 # rated maximum IT power
avg_load: float = 0.75 # average draw as a share of rated
pue: float = 1.3
usd_per_kwh: float = 0.10
facility_usd_per_kw_month: float = 150 # charged on rated kW
support_frac_per_year: float = 0.08 # of capex
staff_usd_per_gpu_year: float = 1_000
gpus: int = 8
useful_frac: float = 0.60
def crf(r, n):
return 1 / n if r == 0 else r * (1 + r) ** n / ((1 + r) ** n - 1)
def gpu_hour_cost(x: Node) -> dict:
capex = x.server_capex * (1 + x.fabric_storage_frac)
yearly = {
"capital": capex * crf(x.cost_of_capital, x.life_years),
"power": x.it_kw * x.avg_load * x.pue * x.usd_per_kwh * HOURS,
"facility": x.it_kw * x.facility_usd_per_kw_month * 12,
"support": capex * x.support_frac_per_year,
"staff": x.staff_usd_per_gpu_year * x.gpus,
}
gpu_hours = x.gpus * HOURS
out = {k: v / gpu_hours for k, v in yearly.items()}
out["per_allocated_hour"] = sum(yearly.values()) / gpu_hours
out["per_useful_hour"] = out["per_allocated_hour"] / x.useful_frac
return outEvery input is a field with a default you are expected to replace with your own quotes and measurements. Keep the model in version control next to the assumptions document, so that when someone asks why the number changed, the diff answers.
Worked example
Run the model with its defaults: an 8-GPU server at an illustrative 300,000 dollars, plus 25 percent for fabric and storage, so 375,000 dollars all in; four years at 8 percent; 10.2 kW rated, 75 percent average load, PUE 1.3, 10 cents per kWh; 150 dollars per kW-month of facility; 8 percent support; 1,000 dollars of staff per GPU-year; 60 percent useful.
| Component | Per allocated GPU-hour (USD) | Share |
|---|---|---|
| Capital recovery | 1.62 | 64 percent |
| Support and spares | 0.43 | 17 percent |
| Facility | 0.26 | 10 percent |
| Power and cooling | 0.12 | 5 percent |
| Staff | 0.11 | 4 percent |
| Total | 2.54 | 100 percent |
| Per useful hour at 60 percent | 4.24 |
Two things stand out. Capital is nearly two thirds of the cost, so the purchase price, the life and the cost of capital are where the money is. And the useful fraction turns 2.54 into 4.24: 1.70 of every useful hour pays for hours that did nothing. These prices are placeholders, not quotes; the proportions are what carry across.
Sensitivity: which inputs move the answer
Change one input at a time and record the per-hour cost. The model gives the following on the worked example.
| Change | Per allocated hour | Per useful hour |
|---|---|---|
| Baseline (4 years, 60 percent useful) | 2.54 | 4.24 |
| 3-year life | 3.00 | 5.01 |
| 5-year life | 2.27 | 3.78 |
| Zero cost of capital | 2.27 | 3.78 |
| Energy price doubled to 0.20 per kWh | 2.67 | 4.45 |
| 40 percent useful | 2.54 | 6.36 |
| 80 percent useful | 2.54 | 3.18 |
| 90 percent useful | 2.54 | 2.83 |
Doubling the energy price adds 5 percent. Moving from 60 to 80 percent useful cuts the cost of useful work by a quarter, and assuming a 5-year life instead of 3 cuts it by about a quarter as well, on paper. That is why utilization engineering, which LLM FinOps turns into team-level accountability, pays back faster than negotiating power contracts, and why the life assumption deserves scrutiny in every budget review.
Measuring the useful fraction
The useful fraction is the input most often guessed and most worth measuring. Break it into factors you can observe. Allocation: the share of GPU-hours assigned to any job, from the scheduler. Busy: the share of allocated time the GPUs were actually executing kernels, from DCGM or the driver's utilization counters, which measure time with work queued, not arithmetic efficiency. Goodput: the share of busy time that produced kept results, after subtracting work lost to failures and replayed from the last checkpoint, evaluation runs nobody read and duplicate experiments.
For inference fleets, the useful fraction is bounded by the headroom needed to meet latency objectives at peak; a service sized for its daily peak runs well below its capacity at night. Report the factors separately: a cluster that is 95 percent allocated and 50 percent busy has a data loading or scheduling problem, not a demand problem.
def useful_fraction(gpu_hours_owned, gpu_hours_allocated, gpu_hours_busy,
gpu_hours_lost_to_failures, gpu_hours_discarded):
allocation = gpu_hours_allocated / gpu_hours_owned # scheduler accounting
busy = gpu_hours_busy / gpu_hours_allocated # DCGM or driver counters
kept = gpu_hours_busy - gpu_hours_lost_to_failures - gpu_hours_discarded
goodput = kept / gpu_hours_busy # job records, checkpoint logs
return {"allocation": allocation, "busy": busy, "goodput": goodput,
"useful": allocation * busy * goodput}
# A month on a 512-GPU cluster: 368,640 owned GPU-hours
print(useful_fraction(368_640, 331_776, 265_421, 13_271, 13_271))
# allocation 0.90, busy 0.80, goodput 0.90 -> useful about 0.65The example month shows how quickly healthy-looking factors compound: 90, 80 and 90 percent multiply to about 65 percent useful. Each factor has a different owner, the scheduler team, the training framework team and the researchers, which is why reporting only the product hides who can fix it.
Break-even and unit cost as outputs
Once you have the cost per allocated hour, comparison with rented capacity is one division. If a provider charges P per GPU-hour and you can rent only the hours you use, owning breaks even when your useful fraction exceeds your allocated cost divided by P. With the example's 2.54, a hypothetical rental price of 3.00 needs 85 percent useful to break even, and a price of 4.00 needs 64 percent. In practice rental capacity is often committed for a term, which makes it an allocated cost too, so compare like with like.
Unit cost follows the same way: divide cost per useful hour by measured throughput. If a GPU sustains, say, 1,500 output tokens per second across its batch, it produces 5.4 million tokens per useful hour, and 4.24 dollars per useful hour becomes about 0.79 dollars per million tokens. For training, multiply cost per useful hour by the GPU-hours the run consumed, including restarts.
Failure modes
- Server price only. Fabric, optics and storage can add a quarter or more to capex and are easy to leave in someone else's budget.
- Accounting life as economic life. A six-year depreciation schedule makes year five look cheap until the hardware cannot compete with what replaced it.
- Allocated hours as the denominator. Pricing products on 2.54 when useful hours cost 4.24 loses money on every unit sold.
- Nameplate power for energy, average power for capacity. It should be the reverse: average draw sets the energy bill, rated draw sets how many servers fit under the site limit.
- Ignoring stranded capacity. Racks you cannot fill because of power, cooling or network ports still cost facility fees.
- Comparing owned allocated cost with on-demand rental prices. Committed contracts, egress and storage charges change the comparison.
What to do next
- Copy the model, replace every default with your own quotes, contracts and tariffs, and commit it alongside a written assumptions file.
- Measure allocation, busy and goodput fractions for the last 30 days instead of assuming a useful fraction.
- Run the sensitivity table for your inputs and present the life and useful-fraction rows to whoever approves hardware budgets.
- Compute break-even utilization against each rental quote you have, on a committed-to-committed basis.
- Publish cost per useful GPU-hour and cost per unit of work monthly, and track the trend, not just the level.
- Revisit the economic life assumption when each new accelerator generation ships.