A capacity forecast answers one question: how many GPUs must be serving at the busiest hour of a future week, with what confidence, ordered how far ahead. Most LLM teams answer it with requests per second and a growth guess, and most get it wrong in the same two ways. Requests are the wrong unit, because cost scales with tokens and tokens per request move whenever the product or the model changes. And a single line is the wrong shape, because what you buy is a tail, the peak hour in a bad week.
This article builds a forecast in tokens from first principles: what to measure, how to decompose growth, weekly shape and peaks, how mix shifts change tokens per request, a readable forecaster with quantiles, backtesting with pinball loss and coverage, and the last step, turning a P90 peak into replicas through a measured rating. The code runs on NumPy alone, and the worked example uses its output. Measuring the per-replica rating is covered in LLM capacity analysis; fleet ledgers and buy triggers in GPU capacity planning.
Measure tokens, not requests
Log three numbers per request at the gateway: input tokens, output tokens and model. Prefill work scales with input tokens and is compute-bound; decode scales with output tokens and is memory-bandwidth-bound, with the KV cache growing with context length. The two use the GPU differently, so a replica's capacity is not one number but a pair, and the mix decides which binds. GPU serving capacity estimation explains why.
Aggregate to hourly sums per model and per traffic tier, interactive and batch, because batch traffic can be deferred to troughs and should not be forecast into the peak. Keep at least sixteen weeks; one launch or holiday distorts anything shorter. And log rejected and throttled requests too, because a forecast trained only on served traffic learns your past capacity limit, not your demand.
The forecasting pipeline
The separation matters. The rating, tokens per second per replica at your latency objective, changes with every inference-engine upgrade, quantisation change or new GPU. If the forecast is in GPUs, each of those changes invalidates history. If it is in tokens, the history stays valid and only the last conversion changes.
Growth, weekly shape and the peak hour
LLM traffic decomposes well into three multiplicative parts. Growth is the week-over-week trend in total tokens, best fitted on a log scale so that 3% a week compounds correctly. Weekly shape is the share of the week's tokens in each of the 168 hours: business-hours peaks for enterprise traffic, evening peaks for consumer apps, weekend dips. Noise is the remainder, roughly log-normal hour to hour.
The ratio that matters for buying is peak to mean. In the synthetic series below the mean hour carries 58% of the peak hour's tokens, so a fleet sized for the peak sits about 40% idle on average. That idle share is the budget for batch work, and the reason a single global region with users across time zones is cheaper per token than several regional ones.
Mix shifts and launches
Tokens per request is not constant, and the changes come from decisions you can see coming. A reasoning model that thinks before answering multiplies output tokens per request. Retrieval or agent features that stuff context multiply input tokens. Moving a tier to a smaller model changes the rating, not the tokens. A launch adds a step, not a trend.
Keep these out of the statistical model. Fit growth and shape on history, then apply known changes as explicit overlays from an event calendar: this feature reaches 30% of traffic in week 6 and triples output tokens for that slice. Overlays are reviewable and can be wrong in public; a model that tries to learn them from three past launches cannot.
A forecaster you can explain
The forecaster below is deliberately simple: a log-linear trend on weekly totals, an average weekly shape, and quantiles taken from its own one-week-ahead errors. You can explain every number to a finance partner, and it is a baseline that fancier models must beat in the backtest before they replace it.
import numpy as np
H = 24 * 7 # hours per week
def forecast(y, horizon_weeks, qs=(0.5, 0.9, 0.99), fit_weeks=8):
"""y: hourly tokens, whole weeks. Growth x weekly shape, quantiles from backtest."""
w = y[-fit_weeks * H:].reshape(fit_weeks, H)
totals = w.sum(axis=1)
slope, icpt = np.polyfit(np.arange(fit_weeks), np.log(totals), 1)
shape = (w / totals[:, None]).mean(axis=0) # share of week per hour
errs = [] # one-week-ahead log errors
for k in range(4, fit_weeks):
s, i0 = np.polyfit(np.arange(k), np.log(totals[:k]), 1)
errs.append(np.log(w[k] / (np.exp(i0 + s * k) * shape)))
errs = np.concatenate(errs)
out = {}
for h in range(1, horizon_weeks + 1):
base = np.exp(icpt + slope * (fit_weeks - 1 + h)) * shape
widen = np.sqrt(h) # uncertainty grows with horizon
out[h] = {q: base * np.exp(np.quantile(errs, q) * widen) for q in qs}
return out, np.exp(slope) - 1 # forecasts, weekly growth
def pinball(actual, pred, q):
d = actual - pred
return np.mean(np.maximum(q * d, (q - 1) * d))The square-root widening assumes errors behave like a random walk across weeks. It is a convention, not a law; the backtest tells you whether it is too wide or too narrow for your traffic.
Quantiles, not a single line
Buy to a quantile, not to the median. A P50 forecast is exceeded half the time; sizing the fleet to it means saturating at the peak hour every other week. P90 of the peak hour is a common target for interactive traffic, with a burst plan, such as shedding batch work, degrading to a smaller model or spilling to a cloud pool, for the remaining tail. P99 is what you show when someone asks about the launch week.
The quantile you choose is a cost decision. Each step up buys idle GPUs for most weeks to avoid saturation in a few, and the price of saturation, failed requests or long queues, is a product number, not a statistical one. Write it down beside the forecast.
Backtesting
A forecast you have not scored is a guess. Hold out recent weeks, forecast them from the weeks before, and score two things. Coverage: the fraction of actual hours below the P90 line should be near 0.9; far below means you are underbuying, far above means you are wasting money. Pinball loss: the standard score for a quantile forecast, which penalises under-forecasts by q and over-forecasts by 1 - q, so the P90 line is pulled up toward the tail. Track both per horizon, because a forecast that is calibrated at one week can be badly wrong at twelve, and twelve is often the horizon you buy at.
Score the peak hour separately as well. Coverage over all 168 hours can look fine while the forecast is consistently low at the weekday peak and high overnight, which is the worst combination for buying. A simple check is the ratio of actual to forecast P90 at each week's busiest hour, plotted over time; a drift above 1 is the earliest warning that growth has changed. Finally, compare every new model against the trend-times-shape baseline on the same held-out weeks. If a gradient-boosted or neural forecaster cannot beat it on pinball loss at the horizon you buy at, keep the baseline: it is cheaper to run, easier to explain and fails in ways people can see. Re-run the backtest monthly, and record the scores next to each capacity decision so the next review can see how much the forecast was trusted and whether that trust was earned.
From tokens to replicas
Converting tokens to replicas needs a rating measured on your engine, model and latency objective. Treat it as two numbers, decode-bound output tokens per second and prefill-bound input tokens per second, and combine them linearly, as below. The linear mix is an approximation that holds when one replica interleaves prefill and decode; validate it with one load test at your real mix, and if you disaggregate prefill and decode onto separate pools, forecast each pool separately instead.
def replicas(out_tok_s, in_tok_s, rate_out, rate_in, headroom=0.8):
"""rate_out / rate_in: measured per-replica capacity at SLO, each alone.
Linear mix: fraction of a replica used = out/rate_out + in/rate_in."""
load = out_tok_s / rate_out + in_tok_s / rate_in
return int(np.ceil(load / headroom))Headroom covers what the rating does not: a replica restarting, uneven load balancing, a node draining for maintenance. Spare capacity for failures is a fleet question, covered in the capacity planning article.
Worked example: four weeks ahead
Twenty weeks of synthetic hourly output-token demand with a weekday afternoon peak, weekends at 60% of weekdays, 3.5% weekly growth and 8% log-normal noise. Train on the first sixteen weeks and forecast the last four. The fit recovers 3.46% weekly growth. One week ahead, 89.3% of actual hours fall under the P90 line, which is calibrated. Four weeks ahead, 98.2% do: the square-root widening is too generous for this series, so the P90 line is effectively a P98 and would overbuy. That is the backtest doing its job.
The week-4 P90 peak is about 27,600 output tokens per second. With a hypothetical rating of 9,000 output and 60,000 input tokens per second per replica, an input-to-output ratio of 6 to 1, and 80% headroom, the load is 3.06 plus 2.76 replica-equivalents, which divided by 0.8 rounds up to 8 replicas. Now apply an overlay: a retrieval feature doubles context, moving the ratio to 12 to 1. Output tokens are unchanged, but the answer becomes 11 replicas, nearly 40% more, from a change no request-count forecast would have seen.
Horizons and lead times
Match the horizon to how fast supply moves. Autoscaling inside an existing pool reacts in minutes; adding nodes from a cloud reservation takes days; a new reservation or a hardware purchase takes months. Run the same forecaster at each horizon, and make each decision from the horizon that matches its lead time with the backtest error for that horizon. Long horizons deserve wide bands and commitments you can adjust, which is the subject of LLM reserved capacity.
Failure modes
- Forecasting requests. A reasoning model or longer context changes cost per request; forecast tokens.
- Training on served traffic only. Throttled demand disappears and the forecast learns your old limit.
- Forecasting in GPUs. Every engine upgrade breaks the history; convert at the end.
- Daily totals. Daily sums hide the peak hour, which is what you buy for.
- Launches in the trend. A step change fitted as growth extrapolates forever; use overlays.
- No backtest. Bands that look reasonable can cover 70% or 99%; only scoring tells you.
- Stale rating. The tokens-per-replica number drifts with prompts and engine versions; re-measure.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Simple trend x shape | Explainable, robust baseline | Misses regime changes |
| Gradient boosting or neural models | Can capture interactions | Opaque; needs long history; overfits launches |
| Higher quantile target | Fewer saturated peaks | More idle GPUs most weeks |
| Per-model forecasts | Accurate conversion per rating | Noisy for small models; sum for the fleet |
| Event overlays | Reviewable product knowledge | Depend on honest launch estimates |
What to do next
- Log input tokens, output tokens, model and tier per request, including throttled ones.
- Build hourly series per model and tier, at least sixteen weeks deep.
- Fit trend times weekly shape as a baseline and backtest it at 1, 4 and 12 weeks.
- Choose a quantile target per tier and write down the price of saturation.
- Keep an event calendar of launches and mix changes, applied as overlays.
- Measure a two-number rating per model at your SLO and re-measure on engine changes.
- Convert the quantile peak to replicas last, and feed each horizon to the matching decision.