Every GPU you own runs on two clocks. The first is the accounting clock: the finance team picks a useful life, divides the cost by it and books the same depreciation charge every year until the asset reaches zero. The second is the economic clock: what an hour on that GPU is actually worth, which falls whenever a faster or more efficient generation ships, whenever rental prices drop, and whenever your power and floor space could earn more with something newer in the slot. The two clocks almost never agree, and most bad hardware decisions come from reading one while believing it is the other.
This article is about the timeline. GPU infrastructure cost covers how capital becomes an hourly charge, and LLM TCO covers a three-year serving cost model. Here you will learn what useful-life numbers in filings mean, why economic value is front-loaded, how GPUs cascade from training to inference, and how to decide when to keep or replace a fleet. Every dollar figure is illustrative.
Two clocks: book life and economic life
Start with the vocabulary, because the words are used loosely. The cost basis is what you capitalised: the GPUs, the servers around them and usually installation. The useful life is the period over which accounting spreads that cost. Salvage value is what you expect to recover at the end, often set to zero for servers. Straight-line depreciation charges (cost minus salvage) divided by life each year, so book value falls in a straight line. If the asset's recoverable value falls well below book value, accounting may require an impairment, a one-off write-down.
Economic value is different: it is the present value of the margin the GPU can still earn, meaning revenue or avoided cloud spend minus power, space, support and staff, discounted for time. Economic life ends when that margin turns negative, or when the slot the GPU occupies would earn more holding something else. Nothing forces the two lives to match. A six-year book life can sit on a GPU whose economic value is mostly gone by year three, and an old GPU can keep earning after its book value hits zero.
One rule follows immediately and saves real money: book value is irrelevant to a keep-or-replace decision. The purchase is sunk. What matters is future cash: what the GPU will earn, what it will cost to run, and what someone will pay for it now. Book value matters to the income statement, because selling below it books a loss, and that is a reason to talk to your controller, not a reason to keep a GPU that loses money every hour.
What the useful-life disclosures say
Public filings are where most people meet GPU depreciation, and they are easy to misread. The large cloud providers disclose useful lives for broad pools of server and network equipment, not for GPUs specifically, and they have changed those lives several times:
| Company | Change | Stated effect |
|---|---|---|
| Microsoft | Servers and network equipment from 4 to 6 years, from fiscal 2023 | Announced July 2022 |
| Alphabet | Servers from 4 to 6 years, some network gear from 5 to 6 years, in 2023 | About $3.4 billion lower depreciation for the year |
| Meta | Certain servers and network assets to 5.5 years from January 2025 | About $2.9 billion lower depreciation for 2025 |
| Amazon | A subset of servers shortened from 6 to 5 years in 2025 | Cited the faster pace of AI and machine learning technology |
Read these as statements about averages across very large fleets, most of which are general-purpose CPU servers that genuinely do last six years. They tell you nothing direct about how long an accelerator stays competitive. Amazon's move is the one that explicitly points at AI hardware, and it went the other way from everyone else. There is a live public argument that GPU lives are too long: critics say the economic life of a frontier GPU is closer to two or three years, defenders point to old generations still renting out and running inference. Both sides are describing the economic clock; neither side is describing the accounting rule. Your plan needs a number for each, and it should not borrow either from a filing.
Why economic value decays
Economic value decays for five reasons, and each one is measurable.
- New generations raise throughput per dollar. When a successor delivers several times the useful work for less than several times the price, the market rate for an hour on the older part has to fall to stay competitive. Rental prices for a generation typically drop as its successor ramps.
- New number formats are not backported. Hopper added FP8 tensor cores; Blackwell added FP4. An A100 cannot run an FP8 training recipe at FP8 speed, because it has no FP8 tensor cores. As frameworks and models adopt a format, older GPUs lose more than raw FLOPs suggest.
- Memory capacity caps the models you can serve. A GPU with 40 GB of HBM cannot hold what a 141 GB or 192 GB part holds, so it drops to smaller models or needs more parallelism, which costs efficiency.
- Software support ends. CUDA eventually drops old architectures; CUDA 13 removed Maxwell, Pascal and Volta as compilation targets. Framework wheels follow. A GPU frozen on an old toolkit is a security and maintenance cost.
- Power and space become the scarce input. In a site with a fixed power budget, an old GPU's real cost includes the work a newer GPU could have done with the same kilowatts. This is often the decisive driver; see GPU datacenter power.
The cascade from training to inference
Because value is front-loaded, well-run fleets move GPUs down a cascade rather than holding them in one role. A new generation goes to the most demanding work, frontier pre-training, where its speed and interconnect matter most and the opportunity cost of a slow run is highest. When the next generation arrives, it moves to fine-tuning, research and smaller training runs. Later it serves inference, batch scoring and evaluation, where a model that fits in memory runs fine on older silicon and the latency target, not peak FLOPs, decides the hardware. At the end it is resold or recycled.
The cascade has preconditions. Inference needs the model to fit; a cluster built around a training fabric may be over-provisioned for serving, so check whether the InfiniBand you paid for is still useful. And a site that is out of power cuts the cascade short: the old GPU still works, but its slot is worth more to a newer one.
A model you can run
The following program puts both clocks side by side and answers the replacement question for a power-capped site. It is deliberately small, so every assumption is visible. The economic value function sums discounted yearly margin while the margin is positive; the swap threshold tells you how much a unit of useful work must be worth before replacing an owned fleet pays.
from dataclasses import dataclass
HOURS = 8760
def straight_line(cost, life, salvage=0.0):
step = (cost - salvage) / life
return [max(salvage, cost - step * y) for y in range(life + 2)]
def economic_value(age, p0, decay, util, opex, r=0.10, max_age=8):
"""Present value of remaining margin for a GPU of this age.
p0: price per GPU-hour when new; decay: yearly fractional drop."""
v = 0.0
for t in range(max_age - age):
price = p0 * (1 - decay) ** (age + t)
margin = price * HOURS * util - opex
if margin <= 0:
break
v += margin / (1 + r) ** (t + 1)
return v
@dataclass
class Fleet:
gpus: int
capex_per_gpu: float # 0 for a fleet you already own (sunk)
kw_per_gpu: float # GPU plus its share of host, network, storage
units_per_gpu: float # useful work per GPU-year, normalised
ops_per_gpu: float # support, staff, spares per GPU-year
def annual_cost(f, life, pue=1.3, usd_kwh=0.08):
power = f.gpus * f.kw_per_gpu * pue * HOURS * usd_kwh
return f.gpus * f.capex_per_gpu / life + power + f.gpus * f.ops_per_gpu
def swap_threshold(old, new, life, resale_per_old):
"""Value per unit-year above which replacing old with new pays."""
extra_units = new.gpus * new.units_per_gpu - old.gpus * old.units_per_gpu
extra_cost = annual_cost(new, life) - annual_cost(old, life)
extra_cost -= old.gpus * resale_per_old / life
return extra_cost / extra_unitsTwo modelling choices matter. The old fleet's capex is zero because it is sunk, which is the rule from the first section written as code. And the comparison is per site, not per GPU: in a power-capped building the new fleet has fewer, hungrier GPUs, so the fair question is what the same megawatt produces under each option.
Worked example: one GPU, then one megawatt
Take a GPU bought for $25,000 with a six-year book life and no salvage. Assume, purely for illustration, that it rents for $3.00 an hour when new, that the market rate falls 30% a year, that it is busy 70% of the time, and that power plus support costs $2,911 a year: 1.0 kW all-in at a PUE of 1.3 and $0.08 per kWh is $911, plus $2,000 of support. With a 10% discount rate the program gives:
| Age (years) | Book value | Economic value | Price per hour | Yearly margin |
|---|---|---|---|---|
| 0 | $25,000 | $30,258 | $3.00 | $15,485 |
| 1 | $20,833 | $17,798 | $2.10 | $9,966 |
| 2 | $16,667 | $9,612 | $1.47 | $6,103 |
| 3 | $12,500 | $4,470 | $1.03 | $3,399 |
| 4 | $8,333 | $1,518 | $0.72 | $1,506 |
| 5 | $4,167 | $164 | $0.50 | $181 |
| 6 | $0 | $0 | $0.35 | -$747 |
Under these assumptions margin stays positive through age 5, so the economic life is six operating years, the same as the book life. Yet economic value drops below book value within the first year and is a third of it by year three. That gap is the hidden write-down: an asset worth less than the balance sheet says, even when the two lives match. Change the decay rate to 45% and margin turns negative at age 4, a four-year life; change it to 20% and the GPU still earns at age 8, past its book life.
Now the site question. One megawatt of IT power holds 1,000 of the old GPUs at 1.0 kW each, or 555 new GPUs at 1.8 kW each that do three times the work, so 1,665 units against 1,000. The new GPUs cost $40,000 each, amortised over four years, with the same $2,000 support per GPU. The old fleet costs $2.91 million a year to run; the new one costs $7.57 million including capital. The extra 665 units cost $4.66 million, about $7,006 per unit-year. If the old GPUs resell for $5,000 each, the threshold falls to about $5,126. So: if a unit of useful work is worth more than roughly $5,100 a year to you, replace; if it is worth less, keep the old fleet even though the new one is far more efficient. Notice that per unit the old fleet is cheaper, $2,911 against $4,547, because its capital is sunk. Efficiency alone never justifies a swap; scarce power plus enough demand does.
What the accounting does to the decision
On the accounting side, three mechanics are worth knowing before you talk to finance. A change in useful life is normally treated as a change in estimate and applied prospectively: the remaining book value is spread over the new remaining life, and past years are not restated. That is why extending lives produced immediate drops in reported depreciation. Selling a GPU below its book value books a loss on disposal, which can make an economically sound early retirement look bad in one quarter. And if a class of assets has clearly lost value, for example a cluster stranded by a power shortfall, an impairment review may be required. None of these change the cash decision, but they change how it is reported, so surface them early.
Failure modes
- Planning with the book life. A capacity plan that assumes six years of competitive service from a training GPU will under-budget the next purchase.
- Treating list rental prices as market value. Committed and spot rates can sit far below list. Use transacted prices, or your own avoided cloud spend.
- Ignoring the format gap. Comparing peak dense BF16 FLOPs hides the advantage of FP8 or FP4 on newer parts for workloads that can use them.
- Assuming resale. Resale prices are thin, volatile and fall when a successor ramps; model resale as a range, including zero.
- Keeping GPUs because book value is high. The sunk-cost error.
- Forgetting the slot. An old GPU in a power-capped site is not free just because it is paid for.
Trade-offs
Long book lives smooth reported earnings but hide obsolescence and can force a large impairment later. Short lives are conservative and make early retirement painless, at the cost of heavier charges now. Renting moves the decay risk to the provider, who prices it into the hourly rate; that is worth paying for when demand is uncertain or your generation is about to be superseded. Owning wins when utilisation is high and stable and you have a cascade to absorb older GPUs. H200 versus H100 economics shows the same logic applied to two adjacent generations.
What to do next
- Write down both clocks for every GPU pool: the book life finance uses and your own estimate of economic life.
- Collect transacted rental or avoided-cloud prices for each generation quarterly, and fit a decay rate instead of guessing one.
- Run the economic value function per pool and list where economic value is far below book value; brief finance before it becomes an impairment surprise.
- Map a cascade: which workloads each generation will move to, and whether those workloads fit in its memory and need its fabric.
- For any power-capped site, compute the swap threshold per unit of work and compare it with what a unit is worth to your business.
- Track software support dates for each architecture in the CUDA and framework release notes.
- Re-run the model when a new generation ships or prices move, not on a fixed annual schedule.