Someone will ask what a new system will cost to run long before it exists. The usual answers are a guess, a number from a pricing calculator that covers only the virtual machines, or a refusal until there is a bill. All three lead to the same outcome: a surprise in month two, usually from a line item nobody estimated, such as data transfer, log ingestion or a managed gateway charged per gigabyte.

This guide gives a method that produces a defensible estimate in a day: start from the drivers the business can predict, translate them into the resources the design consumes, price every metered line item, express uncertainty honestly as a range, and write down each assumption with an owner and a date to check it. It is about estimating before you build. Explaining a bill that has already arrived is covered in cloud cost analysis, and the organisational practice around cost in cloud FinOps.

Advertisement

Why estimates go wrong

Cloud bills are the product of usage and rate, summed over hundreds of meters. Estimates fail in predictable ways. They price the obvious meters, compute and storage, and miss the ones that scale with traffic, such as bytes out, requests to object storage, load balancer processing, NAT gateway processing and observability ingestion. They size for average load when capacity is set by peak. They forget redundancy: a database with a standby costs twice the instance, and an N+1 tier always has an idle instance. And they present one number, which hides the fact that the inputs are uncertain by a factor of two.

The method below addresses each of these. It forces every meter onto the page, derives capacity from peak, adds redundancy explicitly, and carries uncertainty through to the result. The point is not precision. The point is that when the bill arrives, you can say which assumption was wrong.

The method in one picture

From business drivers to a cost rangeDriversrequests, bytes, GB, peakResource modelVMs, GB, GB movedUnit ratesprice list APILine itemsone per meterThree-point rangeslow, likely, highSimulationsample, recomputeP50 and P90budget and alarmAssumption logwho, why, check byAfter launch: compare the bill line by line with the estimate, fix the driver that was wrong, not the totalthe estimate becomes the first version of a unit-cost model
Drivers flow through a resource model and unit rates into line items; ranges on the drivers flow through the same model into a P50 and P90.

There are two passes through the same model. The first uses each driver's likely value and produces a line-by-line bill you can read and challenge. The second samples each driver from its range many times and produces a distribution of totals, from which you take the median, P50, and a pessimistic figure, P90. Both passes use identical code, so the range and the line items can never disagree.

Advertisement

Step 1: drivers, not servers

Start with quantities the business or product can estimate, not with instance types. For a request-serving system the drivers are usually monthly requests, the ratio of peak to average traffic, the size of responses, the volume of stored data and how fast it grows, and the volume of logs and metrics per request. For a batch or ML system they are rows or tokens processed, jobs per day and accelerator-hours per job; LLM cost analysis works through that case from a GPU-hour to cost per million tokens.

Each driver should come with a source. Requests per month may come from the product forecast, response size from measuring the existing API, and the peak factor from the traffic shape of a similar service. If there is no source, that is an assumption, and it goes into the assumption log with a wide range.

Step 2: the bill of materials

Translate drivers into resources by walking the architecture diagram hop by hop and asking, for each component, which meters it runs. A typical request path touches a load balancer (hours and processed bytes), compute (instance-hours, sized for peak), a database (instance-hours, a standby, storage, backups, I/O on some engines), object storage (GB-months and requests), data transfer out to the internet and between zones or regions, and the observability stack (log ingestion, retention, metrics series, traces). Private networking adds NAT gateway hours and per-GB processing, which can be large if workloads pull container images or call external APIs through NAT.

Capacity comes from peak, not average. Measure how many requests one instance serves at your target utilisation with a load test, divide peak requests per second by that, round up, and add redundancy for the failure you design for. The capacity planning guide covers how to measure per-instance throughput and choose headroom; the estimate simply consumes its numbers.

Step 3: unit rates

Get every rate from the provider's own price list for the region, operating system and tier you will use, not from a blog post. AWS, Google Cloud and Azure each publish a pricing calculator and a machine-readable price list API. Record for each rate the region, the SKU or meter name and the date you read it, because prices and free tiers change. Note tiered pricing, where the per-GB rate falls as volume rises, and free allowances, which matter for small systems and disappear in large ones.

Decide explicitly whether the estimate uses on-demand rates or committed-use discounts. Estimate on demand first, because commitments are a financing decision you make once the usage is real. Then show the committed figure as a separate line, with the term and the coverage you assumed. Every rate in the code below is a labelled placeholder; replace each one before using the result.

Step 4: a model in code

A spreadsheet works, but a short program has two advantages: the same function computes both the likely bill and the simulated range, and the model can be reviewed and versioned like any other code. This estimator covers the request-serving example used in the rest of the guide.

import math, random

HOURS = 730  # hours in an average month

# Placeholder unit rates. Replace every one from your provider's price list.
RATE = {
    "vm_hour": 0.17,          # per instance-hour, 4 vCPU general purpose
    "db_hour": 0.50,          # per managed-database instance-hour
    "db_gb_month": 0.115,     # per GB-month of database storage
    "obj_gb_month": 0.023,    # per GB-month of object storage
    "egress_gb": 0.08,        # per GB to the internet, after any free tier
    "lb_hour": 0.025,         # per load-balancer hour
    "lb_gb": 0.008,           # per GB processed by the load balancer
    "log_gb": 0.50,           # per GB of log ingestion
}

# Drivers as (low, likely, high). These are the numbers you argue about.
DRIVERS = {
    "requests_month": (400e6, 600e6, 900e6),
    "resp_kb":        (25, 40, 70),
    "peak_factor":    (2.5, 3.0, 4.5),
    "rps_per_vm":     (120, 150, 170),   # measured at target utilisation
    "db_gb":          (300, 500, 800),
    "obj_gb":         (1500, 2000, 3500),
    "log_bytes_req":  (400, 600, 1200),
}

def bill(d):
    avg_rps = d["requests_month"] / (HOURS * 3600)
    peak_rps = avg_rps * d["peak_factor"]
    vms = math.ceil(peak_rps / d["rps_per_vm"]) + 1          # N+1 for a failed instance
    egress_gb = d["requests_month"] * d["resp_kb"] * 1e3 / 1e9
    log_gb = d["requests_month"] * d["log_bytes_req"] / 1e9
    return {
        "compute":  vms * HOURS * RATE["vm_hour"],
        "database": 2 * HOURS * RATE["db_hour"] + d["db_gb"] * RATE["db_gb_month"],
        "storage":  d["obj_gb"] * RATE["obj_gb_month"],
        "egress":   egress_gb * RATE["egress_gb"],
        "lb":       HOURS * RATE["lb_hour"] + egress_gb * RATE["lb_gb"],
        "logs":     log_gb * RATE["log_gb"],
    }

def likely():
    return bill({k: v[1] for k, v in DRIVERS.items()})

def simulate(n=20000, seed=7):
    rng = random.Random(seed)
    totals = []
    for _ in range(n):
        d = {k: rng.triangular(lo, hi, mode) for k, (lo, mode, hi) in DRIVERS.items()}
        totals.append(sum(bill(d).values()))
    totals.sort()
    return totals[n // 2], totals[int(n * 0.9)]

The structure matters more than the numbers. Drivers are separate from rates. The bill function is the resource model: it turns drivers into quantities and multiplies by rates, one line per meter group. Ranges are triangular distributions defined by low, likely and high values, which are easy for people to give and good enough for planning. The simulation reuses bill unchanged.

Step 5: ranges, P50 and P90

A single number invites false confidence. Ask whoever owns each driver for three values: the lowest plausible, the most likely, and the highest plausible without assuming disaster. Wide ranges are honest, not weak. Then sample all drivers together many times and read off percentiles. Use P50 as the planning figure and P90 as the budget alarm, the level at which you expect to investigate rather than be surprised.

Two effects make the simulated median higher than the bill computed from likely values. Ranges are usually skewed upward, because traffic and payload sizes are bounded below but not really above. And step functions such as whole instances round up. Both are real, and an estimate built from likely values alone will be low for exactly these reasons.

Worked example: a public catalogue API

A team is launching a public product catalogue API. The product forecast is 600 million requests a month, the existing v1 API returns 40 KB on average with compression, and similar services peak at three times average. A load test shows one 4-vCPU instance serves 150 requests per second at 60 percent CPU. Running the likely pass of the estimator gives:

Line itemQuantity at likely valuesMonthly (placeholder rates)
Compute228 rps average, 685 peak, 5 instances plus 1 spare745
Databaseprimary and standby, 500 GB storage788
Object storage2,000 GB-months46
Egress to internet24,000 GB1,920
Load balancer730 hours plus 24,000 GB processed210
Log ingestion360 GB at 600 bytes per request180
Total3,888

Three things stand out. Data transfer out is half the bill and the compute that everyone discussed is a fifth, which is typical for APIs with sizeable payloads and is why egress cost architecture deserves a design review of its own. The database standby doubles that line, which is the price of the availability target. And logs cost a quarter as much as the servers, at only one structured line per request.

The simulation with the ranges shown in the code gives a P50 of about 4,388 and a P90 of about 5,517, well above the 3,888 from likely values, because response size and request volume are both skewed upward. The team budgets 4,400 a month, sets a billing alert at 5,500, and immediately sees which lever matters most: cutting average response size from 40 KB to 25 KB through field selection and better caching headers would save more than the whole compute line.

The assumption log

Every driver and every non-obvious modelling choice gets an entry: the value, who supplied it, the evidence, the main risk, and when and how it will be checked against reality. The log turns the post-launch review from an argument about the total into a check of specific claims.

assumption: resp_kb likely 40
owner: api team (J. Rao)
evidence: median response of the v1 API in staging, gzip on, 2026-09-20
risk: mobile clients request the full catalogue on cold start
check_by: first week after launch, from load balancer bytes-out

After launch, compare the first full month line by line with the estimate. Where a line is off, find which driver or rate was wrong and fix that input, then re-run the model. Over a few months the estimator becomes a unit-cost model, cost per thousand requests or per active customer, which is far more useful than the original estimate because it predicts the cost of the next feature.

Failure modes

  • Missing meters. NAT gateway processing, inter-zone transfer, object storage requests, snapshot storage and log retention are the usual omissions. Walk every hop of the diagram.
  • Average sizing. Capacity priced at average load is too small by the peak factor; the real system either costs more or fails at peak.
  • Forgotten environments. Staging, load-test and developer environments often add a significant fraction of production. Estimate them as separate, smaller copies.
  • Free-tier optimism. Free allowances make a small prototype look nearly free and vanish at production volume.
  • Commitment double counting. Applying a committed-use discount to the estimate and then again in the financial plan.
  • Unit confusion. Mixing GB and GiB, or KB as 1,000 and 1,024 bytes, shifts transfer lines by several percent; pick one and state it.
  • Stale rates. Rates copied once and reused for a year. Record the date read and refresh before each review.

Trade-offs

ApproachGood forWeakness
Provider pricing calculatorQuick check of a known configurationOnly covers what you remember to add; no ranges
Spreadsheet modelShared review with financeRanges and simulation are awkward; hard to version
Code model with simulationRanges, sensitivity, reuse as a unit-cost modelNeeds an engineer to maintain
Analogy with an existing serviceSanity check on the totalHides differences in payload, traffic shape and design

What to do next

  1. List the drivers for your system and give each an owner and a source.
  2. Walk your architecture diagram hop by hop and list every meter each component runs, including transfer, NAT, requests and observability.
  3. Load-test one instance at target utilisation and size compute from peak, not average, with explicit redundancy.
  4. Read every unit rate from your provider's price list for your region, and record the meter name and the date.
  5. Adapt the estimator above, keeping drivers, rates and the bill function separate.
  6. Collect low, likely and high values for each driver, simulate, and report P50 and P90 rather than one number.
  7. Write the assumption log and set a billing alert at P90.
  8. Reconcile the first full month line by line, fix the wrong inputs, and turn the model into a cost per unit of business volume.
Key takeaway: Estimate cloud cost from drivers, not servers. Turn predictable business quantities into the resources your design consumes, walk every hop to list each metered line item, size capacity from peak with explicit redundancy, and price everything from the provider's dated price list. Carry uncertainty as low, likely and high values through the same model to a P50 for planning and a P90 for alerting. Log each assumption with an owner and a check date, then reconcile the first bill line by line and fix the input that was wrong, not the total.