Capacity planning answers one question ahead of time: will we have enough of every resource, at the moment we need it, at a cost we accept? Autoscaling answers a different question, what to do in the next five minutes, and it only works inside limits someone planned earlier: quotas, reserved instances, database size, network links, licences and the time it takes to buy any of them.
This guide treats capacity planning as a small system with inputs, a model, a policy and a feedback loop. It shows how to turn a business forecast into requests per second for each resource, how to measure what one instance can really carry, how much headroom to keep and why, and how to tell from the data that the plan is drifting. A Python model and a worked example carry the arithmetic so you can adapt it to your own services.
What capacity planning is, and what it is not
A capacity plan is a statement of the form: for the next N months, at the forecast demand plus a stated margin, these resources in these quantities keep the service inside its SLO, and here is the date by which each purchase or change must start. It is a decision document, refreshed on a cadence, with numbers anyone can recompute.
It is not a scaling design. Choosing between caches, replicas and shards is covered in Scaling patterns in depth, which also derives Little's law and queueing behaviour; this guide uses those results rather than repeating them. It is also not autoscaling. Cloud autoscaling moves the instance count between a floor and a ceiling; capacity planning decides where the floor and ceiling sit, and makes sure the ceiling is actually purchasable when it is reached.
The architecture: six stages and a loop
The process has six stages. Demand drivers are the business quantities that cause load: active users, orders, documents ingested, model calls. A forecast projects those drivers forward. A demand model converts each driver into a planning peak for each resource, in units that resource understands: requests per second for an API tier, writes per second and bytes per day for a database, tokens per second for an inference fleet. A unit cost model, measured by load testing, says how much of each unit one instance can carry. A headroom policy adds the margins the business has agreed to pay for. The provisioning plan turns the result into concrete actions with dates.
The seventh element, the review, is what makes it architecture rather than a one-off estimate. Every cycle compares last cycle's forecast with what happened and every unit cost with production telemetry. Errors are expected; unexplained errors that persist are the signal that a model input is wrong.
Demand: from business drivers to a planning peak
Start from drivers rather than from last month's CPU graph. CPU history tells you what happened under the old code and the old traffic mix; drivers let you ask what happens if orders double or a new feature triples calls per order. For each service write down the ratio from driver to work: API calls per order, rows written per order, bytes stored per order, tokens per conversation. Measure those ratios from production logs and re-measure them every cycle, because product changes move them silently.
Daily totals are useless for sizing; peaks matter. The peak-to-average ratio, peak-minute rate divided by daily average rate, captures the daily shape. Consumer services often sit between 2 and 4; batch-heavy internal systems can be far higher. Measure yours at the granularity at which saturation hurts: per minute for an API, per second for a bursty queue.
Forecast the driver with an explicit uncertainty band rather than a single line. A trend plus weekly and yearly seasonality is enough for most services; the important habit is to plan against a high percentile of the forecast, such as the P90, rather than the median, and to keep a calendar of known events, such as sales, launches and tax deadlines, as separate multipliers. Events break statistical forecasts because they are rare in the history.
Supply: measuring what one instance can carry
The unit cost model comes from a load test of one instance, or one shard, with production-like data and request mix. Step the offered load up and record throughput and latency percentiles at each step. Throughput rises roughly linearly, then flattens; latency stays flat, then climbs steeply. The point where the latency SLO, for example p99 under 300 ms, is first violated is the knee. Load testing architecture explains how to build a harness whose numbers you can trust, including open-loop load generation, which matters here because closed-loop tools hide the knee.
Never plan to run at the knee. Queueing theory says latency grows non-linearly with utilisation, so a small traffic surprise near the knee becomes a large latency surprise. Choose a safe operating point, a fixed fraction of the knee throughput that you allow at the planning peak; it doubles as your burst margin. Seventy percent is a common starting value; the right value depends on how steep your latency curve is, which the load test shows directly. Read percentiles from properly merged histograms, as described in Latency histograms and quantile estimation, not from averaged per-host percentiles.
Repeat this for every resource that can saturate: CPU and memory per pod, connections and IOPS per database, partition throughput per broker, GPU memory per model replica, egress bandwidth per NAT gateway. The resource with the smallest runway is the binding constraint, and it is frequently not the one on the dashboard everyone watches.
Headroom policy: margins you decide on purpose
Headroom is the gap between the planning peak and the capacity you provision. Each margin should be named, justified and owned, because unnamed margins get stacked twice or cut entirely in a cost review.
- Failure headroom. If the service must survive the loss of one of z zones, the remaining z - 1 zones must carry the whole peak, so total capacity must be the required capacity times z / (z - 1). With three zones that is 1.5 times; with two zones it is 2 times, which is why two-zone designs are expensive.
- Growth during lead time. If adding capacity takes L weeks, today's plan must cover demand L weeks from now plus a review interval. Lead times differ wildly: seconds for autoscaling within quota, days for quota increases, weeks for reserved capacity or a database migration, months for hardware or large accelerator reservations.
- Deploy and maintenance headroom. Rolling deploys and node drains temporarily remove instances. A surge setting of 25 percent means a quarter more pods exist during a deploy, which must fit inside quota.
Runway is the time until a resource crosses its alert threshold at forecast growth. Express every finding as runway compared with lead time: a resource whose runway is shorter than its lead time needs action now, even if it is only 60 percent full.
A capacity model in Python
The arithmetic is simple enough to live in a short, reviewed script checked into the repository next to the service. Keeping it in code rather than a spreadsheet makes the assumptions diffable and lets the review compare the inputs from cycle to cycle.
import math
from dataclasses import dataclass
@dataclass
class Resource:
name: str
knee_per_instance: float # units/s one instance sustains at the latency SLO
safe_fraction: float # share of the knee allowed at the planning peak
zones: int # instances are spread evenly across zones
survive_zone_loss: bool = True
def planning_peak(daily_units, peak_to_avg, forecast_uplift, event_mult):
# daily driver volume -> peak units per second for one resource
return daily_units / 86_400 * peak_to_avg * forecast_uplift * event_mult
def instances_needed(peak, r):
n = math.ceil(peak / (r.knee_per_instance * r.safe_fraction))
if r.survive_zone_loss and r.zones > 1:
n = math.ceil(n * r.zones / (r.zones - 1)) # z-1 zones carry the peak
return math.ceil(n / r.zones) * r.zones # whole instances per zone
def runway_days(capacity, used, growth_per_day, alert_fraction=0.8):
free = capacity * alert_fraction - used
return math.inf if growth_per_day <= 0 else free / growth_per_day
api_peak = planning_peak(1_200_000 * 25, peak_to_avg=3.0,
forecast_uplift=1.15, event_mult=1.5)
api = Resource("checkout-api pod", knee_per_instance=220,
safe_fraction=0.7, zones=3)
print(round(api_peak), instances_needed(api_peak, api)) # 1797 18
print(round(runway_days(2000, 1100, 7.2))) # 69 (GB, GB, GB/day)
Worked example: a checkout service for the next two quarters
The business forecasts 1.2 million orders per day by the end of the planning window. Production logs show 25 API calls, 4 database writes and about 6 KB of stored rows and indexes per order. The peak minute runs at 3 times the daily average, the P90 forecast sits 15 percent above the median, and a planned sale is expected to multiply peak traffic by 1.5.
API tier: 30 million calls per day is 347 per second on average, 1,042 at the daily peak, 1,198 at the P90 forecast and about 1,797 during the sale. A load test puts one two-vCPU pod's knee at 220 requests per second where p99 crosses 300 ms; at a safe fraction of 0.7 each pod is planned at 154. That needs 12 pods, and surviving the loss of one of three zones raises it to 18, six per zone. The quota check that follows is the real deliverable: 18 pods plus a 25 percent deploy surge, plus every other service in the namespace, must fit the regional vCPU quota.
Database writes: 4.8 million writes per day becomes about 288 writes per second at the planned sale peak, comfortably inside the primary's measured knee. Storage is different. At 7.2 GB per day, a 2 TB volume holding 1.1 TB reaches its 80 percent alert line in about 69 days. Moving to a larger volume class or archiving old orders is a project the team estimates at twelve weeks, so storage, not CPU, is the binding constraint, and work must start this cycle even though every dashboard is green.
Provisioning: matching the instrument to the lead time
| Instrument | Lead time | Good for | Watch out for |
|---|---|---|---|
| Autoscaling within quota | Seconds to minutes | Daily shape, unforecast bursts | Quota and cold start set the ceiling |
| Quota increase | Hours to days | Raising the autoscaling ceiling | Not guaranteed; request before you need it |
| Reserved or committed capacity | Days to weeks | The steady base load | Commitment outlives the forecast if demand falls |
| Data migration or resharding | Weeks to months | Storage and write limits | Must start when runway equals lead time |
| Hardware or large accelerator reservations | Months | GPU fleets, on-premises | Forecast error is expensive in both directions |
A sound pattern is to commit to the base load you are confident of, autoscale the daily shape on demand pricing, and keep failure headroom as real, running capacity rather than as something you hope to buy during an incident, because a zone outage is precisely when everyone else is buying too. Cloud cost analysis shows how to attribute the resulting bill to rate, volume and mix so the cost review argues about the right lever.
The review loop: keeping the plan honest
Run a review on a fixed cadence, monthly for fast-growing services and quarterly for stable ones, plus before every major event. Bring four things. First, forecast error: compare last cycle's predicted driver and peak with actuals, using mean absolute percentage error, and look at its sign; a forecast that is always low is a bias, not noise. Second, unit cost drift: production throughput per instance at observed latency, compared with the load-test knee; a release that costs 20 percent more CPU per request eats a fifth of your headroom without any traffic change. Third, utilisation at the actual peak for each resource, not the daily average. Fourth, runway versus lead time for every resource, sorted by the difference.
Tie the review to the SLO process described in SLO program rollout: latency SLO burn during peaks is the most direct evidence that the safe fraction is set too high.
Failure modes
- Planning on averages. Daily averages hide the peak minute, and averaged percentiles hide the tail. Size on peaks from merged histograms.
- One resource only. CPU gets planned while connections, IOPS, file descriptors, IP addresses in a subnet or a third-party rate limit run out first.
- Stale unit costs. The load test was run a year ago on different code. Re-run it every cycle or after any major release.
- Failure headroom spent on growth. Utilisation creeps to 80 percent and nobody notices that a zone loss now overloads the survivors. Alert on utilisation against the post-failure ceiling.
- Quota surprises. The plan says 18 pods; the region allows 14 more vCPUs than you use. Check quotas as part of the plan, not during the incident.
- Shared dependencies. Your service scales, the shared database or identity service behind it does not. Include downstream owners in the demand model.
Trade-offs
Every margin is money. Planning to P90 rather than the median, surviving zone loss, and keeping a 0.7 safe fraction together more than double the instance count relative to a naive estimate. The answer is not to remove margins but to make each one an explicit, priced decision: a team can choose to accept degraded latency during a zone loss and shed load instead, as long as the decision is written down and the load shedding is tested.
What to do next
- List the demand drivers for one service and measure the driver-to-work ratios from production logs.
- Measure the peak-to-average ratio at one-minute granularity for the last 90 days.
- Run a stepped load test on one instance and record the knee for each saturating resource.
- Write the headroom policy down: zone-loss rule, safe fraction, event multipliers, deploy surge.
- Put the capacity model in a script in the repository and compute runway for every resource.
- Compare runway with lead time and open work for any resource where runway is shorter.
- Check regional quotas against the planned ceiling plus deploy surge.
- Schedule the review and record forecast error so the next cycle can correct its bias.