Trend Watch 2026-09-22

Trend Watch: The Jevons Paradox Is Eating the AI Efficiency Story

Cheaper AI is supposed to mean less compute burned. In 2026 it means the opposite -- and that's the 160-year-old economic pattern everyone in AI infrastructure is now quoting.

Jevons Paradox is an 1865 observation about coal: making steam engines more fuel-efficient didn't reduce coal consumption, it increased it, because efficiency made coal-powered machinery cheap enough to deploy everywhere. In 2026, it has become the single most-quoted economic argument in AI infrastructure debates, and the numbers back it up. The price of a unit of intelligence fell by roughly two orders of magnitude over the past year -- yet the industry's total spend on inference rose about 75% in the same period. Cheaper tokens didn't shrink the bill. They exploded demand for tokens.

The token-mix data makes the mechanism visible: open-weight models handled 29% of all gateway token volume in June 2026, up from 11% in April, with DeepSeek alone accounting for 22.6% of that volume. Falling per-token cost didn't just make existing workloads cheaper -- it unlocked entirely new categories of workload (longer agent loops, more retries, more speculative tool calls) that weren't economical to run before. The rebound effect crossed the classic Jevons threshold within a single year: energy use per unit of AI output fell, while aggregate energy consumption climbed anyway.

That's showing up in the physical infrastructure numbers too. Global data-centre electricity demand is forecast to nearly double between 2025 and 2030, with the AI-specific slice of that demand expected to roughly triple over the same window. High-bandwidth memory shortages are projected to persist through the end of 2027. And there's a sharper edge to it: several inference providers are reportedly losing money on frontier-model serving even as total demand explodes, because each new unit of demand pays less than it costs to serve at the frontier -- efficiency gains are being competed away into lower prices faster than they're being banked as margin.

The part of this I keep coming back to: every "AI just got X% more efficient" headline this year has been read as good news for the power-grid and GPU-shortage story. The Jevons data says read it the opposite way -- efficiency is the accelerant, not the brake, on compute demand. If you're planning infrastructure, capacity, or even just a training budget around the assumption that efficiency gains buy you headroom, this is the pattern that says they won't; they'll get spent on doing more, not on doing the same thing for less.

Efficiency gains in AI are not reducing total compute demand -- they're the mechanism driving it higher, exactly as Jevons Paradox predicts. Plan infrastructure and budget around rising aggregate demand, not around efficiency buying you slack.