Someone will ask what a new system will cost to run long before it exists. The usual answers are a guess, a number from a pricing calculator that covers only the virtual machines, or a refusal until there is a bill. All three lead to the same outcome: a surprise in month two, usually from a line item nobody estimated, such as data transfer, log ingestion or a managed gateway charged per gigabyte.
This guide gives a method that produces a defensible estimate in a day: start from the drivers the business can predict, translate them into the resources the design consumes, price every metered line item, express uncertainty honestly as a range, and write down each assumption with an owner and a date to check it. It is about estimating before you build. Explaining a bill that has already arrived is covered in cloud cost analysis, and the organisational practice around cost in cloud FinOps.
Why estimates go wrong
Cloud bills are the product of usage and rate, summed over hundreds of meters. Estimates fail in predictable ways. They price the obvious meters, compute and storage, and miss the ones that scale with traffic, such as bytes out, requests to object storage, load balancer processing, NAT gateway processing and observability ingestion. They size for average load when capacity is set by peak. They forget redundancy: a database with a standby costs twice the instance, and an N+1 tier always has an idle instance. And they present one number, which hides the fact that the inputs are uncertain by a factor of two.
The method below addresses each of these. It forces every meter onto the page, derives capacity from peak, adds redundancy explicitly, and carries uncertainty through to the result. The point is not precision. The point is that when the bill arrives, you can say which assumption was wrong.
The method in one picture
There are two passes through the same model. The first uses each driver's likely value and produces a line-by-line bill you can read and challenge. The second samples each driver from its range many times and produces a distribution of totals, from which you take the median, P50, and a pessimistic figure, P90. Both passes use identical code, so the range and the line items can never disagree.
Step 1: drivers, not servers
Start with quantities the business or product can estimate, not with instance types. For a request-serving system the drivers are usually monthly requests, the ratio of peak to average traffic, the size of responses, the volume of stored data and how fast it grows, and the volume of logs and metrics per request. For a batch or ML system they are rows or tokens processed, jobs per day and accelerator-hours per job; LLM cost analysis works through that case from a GPU-hour to cost per million tokens.
Each driver should come with a source. Requests per month may come from the product forecast, response size from measuring the existing API, and the peak factor from the traffic shape of a similar service. If there is no source, that is an assumption, and it goes into the assumption log with a wide range.
Step 2: the bill of materials
Translate drivers into resources by walking the architecture diagram hop by hop and asking, for each component, which meters it runs. A typical request path touches a load balancer (hours and processed bytes), compute (instance-hours, sized for peak), a database (instance-hours, a standby, storage, backups, I/O on some engines), object storage (GB-months and requests), data transfer out to the internet and between zones or regions, and the observability stack (log ingestion, retention, metrics series, traces). Private networking adds NAT gateway hours and per-GB processing, which can be large if workloads pull container images or call external APIs through NAT.
Capacity comes from peak, not average. Measure how many requests one instance serves at your target utilisation with a load test, divide peak requests per second by that, round up, and add redundancy for the failure you design for. The capacity planning guide covers how to measure per-instance throughput and choose headroom; the estimate simply consumes its numbers.
Step 3: unit rates
Get every rate from the provider's own price list for the region, operating system and tier you will use, not from a blog post. AWS, Google Cloud and Azure each publish a pricing calculator and a machine-readable price list API. Record for each rate the region, the SKU or meter name and the date you read it, because prices and free tiers change. Note tiered pricing, where the per-GB rate falls as volume rises, and free allowances, which matter for small systems and disappear in large ones.
Decide explicitly whether the estimate uses on-demand rates or committed-use discounts. Estimate on demand first, because commitments are a financing decision you make once the usage is real. Then show the committed figure as a separate line, with the term and the coverage you assumed. Every rate in the code below is a labelled placeholder; replace each one before using the result.
Step 4: a model in code
A spreadsheet works, but a short program has two advantages: the same function computes both the likely bill and the simulated range, and the model can be reviewed and versioned like any other code. This estimator covers the request-serving example used in the rest of the guide.
import math, random
HOURS = 730 # hours in an average month
# Placeholder unit rates. Replace every one from your provider's price list.
RATE = {
"vm_hour": 0.17, # per instance-hour, 4 vCPU general purpose
"db_hour": 0.50, # per managed-database instance-hour
"db_gb_month": 0.115, # per GB-month of database storage
"obj_gb_month": 0.023, # per GB-month of object storage
"egress_gb": 0.08, # per GB to the internet, after any free tier
"lb_hour": 0.025, # per load-balancer hour
"lb_gb": 0.008, # per GB processed by the load balancer
"log_gb": 0.50, # per GB of log ingestion
}
# Drivers as (low, likely, high). These are the numbers you argue about.
DRIVERS = {
"requests_month": (400e6, 600e6, 900e6),
"resp_kb": (25, 40, 70),
"peak_factor": (2.5, 3.0, 4.5),
"rps_per_vm": (120, 150, 170), # measured at target utilisation
"db_gb": (300, 500, 800),
"obj_gb": (1500, 2000, 3500),
"log_bytes_req": (400, 600, 1200),
}
def bill(d):
avg_rps = d["requests_month"] / (HOURS * 3600)
peak_rps = avg_rps * d["peak_factor"]
vms = math.ceil(peak_rps / d["rps_per_vm"]) + 1 # N+1 for a failed instance
egress_gb = d["requests_month"] * d["resp_kb"] * 1e3 / 1e9
log_gb = d["requests_month"] * d["log_bytes_req"] / 1e9
return {
"compute": vms * HOURS * RATE["vm_hour"],
"database": 2 * HOURS * RATE["db_hour"] + d["db_gb"] * RATE["db_gb_month"],
"storage": d["obj_gb"] * RATE["obj_gb_month"],
"egress": egress_gb * RATE["egress_gb"],
"lb": HOURS * RATE["lb_hour"] + egress_gb * RATE["lb_gb"],
"logs": log_gb * RATE["log_gb"],
}
def likely():
return bill({k: v[1] for k, v in DRIVERS.items()})
def simulate(n=20000, seed=7):
rng = random.Random(seed)
totals = []
for _ in range(n):
d = {k: rng.triangular(lo, hi, mode) for k, (lo, mode, hi) in DRIVERS.items()}
totals.append(sum(bill(d).values()))
totals.sort()
return totals[n // 2], totals[int(n * 0.9)]The structure matters more than the numbers. Drivers are separate from rates. The bill function is the resource model: it turns drivers into quantities and multiplies by rates, one line per meter group. Ranges are triangular distributions defined by low, likely and high values, which are easy for people to give and good enough for planning. The simulation reuses bill unchanged.
Step 5: ranges, P50 and P90
A single number invites false confidence. Ask whoever owns each driver for three values: the lowest plausible, the most likely, and the highest plausible without assuming disaster. Wide ranges are honest, not weak. Then sample all drivers together many times and read off percentiles. Use P50 as the planning figure and P90 as the budget alarm, the level at which you expect to investigate rather than be surprised.
Two effects make the simulated median higher than the bill computed from likely values. Ranges are usually skewed upward, because traffic and payload sizes are bounded below but not really above. And step functions such as whole instances round up. Both are real, and an estimate built from likely values alone will be low for exactly these reasons.
Worked example: a public catalogue API
A team is launching a public product catalogue API. The product forecast is 600 million requests a month, the existing v1 API returns 40 KB on average with compression, and similar services peak at three times average. A load test shows one 4-vCPU instance serves 150 requests per second at 60 percent CPU. Running the likely pass of the estimator gives:
| Line item | Quantity at likely values | Monthly (placeholder rates) |
|---|---|---|
| Compute | 228 rps average, 685 peak, 5 instances plus 1 spare | 745 |
| Database | primary and standby, 500 GB storage | 788 |
| Object storage | 2,000 GB-months | 46 |
| Egress to internet | 24,000 GB | 1,920 |
| Load balancer | 730 hours plus 24,000 GB processed | 210 |
| Log ingestion | 360 GB at 600 bytes per request | 180 |
| Total | 3,888 |
Three things stand out. Data transfer out is half the bill and the compute that everyone discussed is a fifth, which is typical for APIs with sizeable payloads and is why egress cost architecture deserves a design review of its own. The database standby doubles that line, which is the price of the availability target. And logs cost a quarter as much as the servers, at only one structured line per request.
The simulation with the ranges shown in the code gives a P50 of about 4,388 and a P90 of about 5,517, well above the 3,888 from likely values, because response size and request volume are both skewed upward. The team budgets 4,400 a month, sets a billing alert at 5,500, and immediately sees which lever matters most: cutting average response size from 40 KB to 25 KB through field selection and better caching headers would save more than the whole compute line.
The assumption log
Every driver and every non-obvious modelling choice gets an entry: the value, who supplied it, the evidence, the main risk, and when and how it will be checked against reality. The log turns the post-launch review from an argument about the total into a check of specific claims.
assumption: resp_kb likely 40
owner: api team (J. Rao)
evidence: median response of the v1 API in staging, gzip on, 2026-09-20
risk: mobile clients request the full catalogue on cold start
check_by: first week after launch, from load balancer bytes-outAfter launch, compare the first full month line by line with the estimate. Where a line is off, find which driver or rate was wrong and fix that input, then re-run the model. Over a few months the estimator becomes a unit-cost model, cost per thousand requests or per active customer, which is far more useful than the original estimate because it predicts the cost of the next feature.
Failure modes
- Missing meters. NAT gateway processing, inter-zone transfer, object storage requests, snapshot storage and log retention are the usual omissions. Walk every hop of the diagram.
- Average sizing. Capacity priced at average load is too small by the peak factor; the real system either costs more or fails at peak.
- Forgotten environments. Staging, load-test and developer environments often add a significant fraction of production. Estimate them as separate, smaller copies.
- Free-tier optimism. Free allowances make a small prototype look nearly free and vanish at production volume.
- Commitment double counting. Applying a committed-use discount to the estimate and then again in the financial plan.
- Unit confusion. Mixing GB and GiB, or KB as 1,000 and 1,024 bytes, shifts transfer lines by several percent; pick one and state it.
- Stale rates. Rates copied once and reused for a year. Record the date read and refresh before each review.
Trade-offs
| Approach | Good for | Weakness |
|---|---|---|
| Provider pricing calculator | Quick check of a known configuration | Only covers what you remember to add; no ranges |
| Spreadsheet model | Shared review with finance | Ranges and simulation are awkward; hard to version |
| Code model with simulation | Ranges, sensitivity, reuse as a unit-cost model | Needs an engineer to maintain |
| Analogy with an existing service | Sanity check on the total | Hides differences in payload, traffic shape and design |
What to do next
- List the drivers for your system and give each an owner and a source.
- Walk your architecture diagram hop by hop and list every meter each component runs, including transfer, NAT, requests and observability.
- Load-test one instance at target utilisation and size compute from peak, not average, with explicit redundancy.
- Read every unit rate from your provider's price list for your region, and record the meter name and the date.
- Adapt the estimator above, keeping drivers, rates and the bill function separate.
- Collect low, likely and high values for each driver, simulate, and report P50 and P90 rather than one number.
- Write the assumption log and set a billing alert at P90.
- Reconcile the first full month line by line, fix the wrong inputs, and turn the model into a cost per unit of business volume.