Engineers are asked "how long will this take?" constantly and are rarely taught how to answer. The usual result is a single number produced by intuition, quietly padded, then turned into a deadline that the team misses, which teaches everyone that estimates are useless. They are not useless; single-number estimates are. Software work has a skewed distribution of outcomes: many tasks finish near the likely case, a few run far over, and almost none finish dramatically early. Any method that ignores that shape will be optimistic on average.

This guide gives a method you can run on a real project this week: separate the estimate from the target and the commitment, decompose until you can describe each task, give ranges, add them up properly with a short simulation, correct with your own history, convert effort to calendar time, and communicate the result so that the people relying on it can make good decisions. A worked example carries the numbers through, and all code is runnable.

Advertisement

Estimate, target, commitment: three different things

Most estimation arguments are really arguments between three different things that share one word. An estimate is a prediction: given what we know, the work will probably take between this and that. A target is a business wish: we would like it by the conference. A commitment is a promise: we will deliver this scope by this date, and we accept the risk of the chosen confidence level.

Keeping them apart is the first skill. When a manager says the estimate is too long, they usually mean the target is earlier than the estimate; the honest responses are to change scope, change resources, or accept a lower probability of hitting the date, not to edit the prediction. An estimate that changes because someone disliked it has stopped being an estimate.

The method at a glance

An estimate is a pipeline that ends in a distribution, then loopsScopegoal, non-goals, doneDecomposetasks you can describeRangeslow / likely / highSimulatesum distributionsCalibratepast actual / estimateEffort P50 / P80person-daysCalendarcapacity, dependenciesCommunicaterange, risks, assumptionsCommita chosen percentileTrack and re-estimateat each milestonenew informationSpikes turn the widest ranges into narrow ones before you commit.
Scope is decomposed into describable tasks; each gets a range; a simulation sums the ranges into a distribution; history calibrates it; capacity turns effort into dates; the commitment is a chosen percentile; and new information loops back into the estimate.
Advertisement

Scope and decomposition

You cannot estimate what you cannot describe. Start with a one-paragraph goal, an explicit list of non-goals, and a definition of done that includes release, not just merge: deployed, behind a flag or rolled out, monitored, documented. Then break the work down until every task is something an engineer on the team could explain step by step. A useful test is that the likely case for each task is a few days at most; a task whose likely case is three weeks hides decisions nobody has made yet.

The tasks people forget are the ones that are not writing feature code, and they are often a large share of the total:

  • Design and review time, including waiting for reviewers.
  • Data migrations and backfills, with their dry runs.
  • Test infrastructure, fixtures and the tests themselves.
  • Integration with other teams' services, and the waiting that comes with it.
  • Security, privacy or compliance review.
  • Rollout: feature flags, staged release, monitoring and dashboards, rollback plan.
  • Documentation, support runbooks and handover.

When a task cannot be described because nobody knows how it will work, do not guess: schedule a spike, a time-boxed investigation of one or two days whose output is knowledge, and estimate the real task afterwards. A spike converts the widest range in the plan into a narrow one, which is the cheapest risk reduction available.

Ranges, not points

For each task, give three numbers in ideal engineer-days: an optimistic case (things go well, but not miraculously), the most likely case, and a pessimistic case (the things that plausibly go wrong do go wrong, short of disaster). A practical calibration target is that the real outcome should fall inside your low-to-high range about nine times in ten. Most people's first ranges are far too narrow; you learn to widen them by checking them later.

Two arithmetic facts matter. First, the expected value of a skewed task is above its most likely value; the classic PERT approximation of the mean is (optimistic + 4 * likely + pessimistic) / 6. So the sum of the likely cases is systematically below the expected total, and plans built from likely cases are late on average even when every individual number was honest. Second, you cannot add pessimistic cases to get a pessimistic total; it is very unlikely that every task hits its worst case at once, so that sum is an overstatement. Summing distributions properly needs either statistics or a simulation, and the simulation is easier.

Adding ranges with a Monte Carlo simulation

A Monte Carlo simulation draws one possible duration for every task, adds them, records the total, and repeats many times. The sorted totals form a distribution from which you read percentiles. The script below uses a triangular distribution per task, which needs exactly the three numbers you already have, plus one shared multiplier that models risks affecting every task together: an unfamiliar codebase, a slow review culture, a key person on leave. Without that shared factor, independent draws cancel each other out and the simulation is overconfident.

import random

# (task, optimistic, likely, pessimistic) in ideal engineer-days
TASKS = [
    ("spike: IdP requirements",        1, 2, 4),
    ("data model + migration",         2, 3, 6),
    ("OIDC login flow",                3, 5, 9),
    ("SAML login flow",                4, 6, 12),
    ("JIT provisioning + role mapping",2, 4, 8),
    ("admin settings UI",              3, 5, 8),
    ("session, logout, enforce SSO",   2, 3, 6),
    ("test against three IdPs",        2, 4, 9),
    ("security review + fixes",        1, 3, 7),
    ("docs, runbook, rollout flag",    1, 2, 4),
]

def simulate(n=100_000, seed=7, shared_sigma=0.15):
    rng = random.Random(seed)
    totals = []
    for _ in range(n):
        base = sum(rng.triangular(o, p, m) for _, o, m, p in TASKS)
        totals.append(base * rng.lognormvariate(0.0, shared_sigma))   # correlated risk
    totals.sort()
    return totals

t = simulate()
pct = lambda q: t[int(q * len(t))]
print(f"P50 {pct(0.50):.1f}  P80 {pct(0.80):.1f}  P95 {pct(0.95):.1f} engineer-days")

Note the argument order: Python's random.triangular takes (low, high, mode), not low, mode, high, an easy bug to make. The shared sigma of 0.15 is a judgement, not a constant of nature; raise it when the team or technology is new to you, and calibrate it from history as described below.

Worked example: adding single sign-on to a B2B product

The task list above is for adding SAML and OIDC single sign-on to an existing B2B web application with just-in-time user provisioning and an admin settings page. Running the script gives these figures:

MethodTotal (engineer-days)
Sum of likely cases37
Sum of PERT means40.3
Simulation, independent tasks: P50 / P80 / P9543.6 / 46.7 / 49.7
Simulation with shared risk (sigma 0.15): P50 / P80 / P9543.5 / 50.3 / 57.7
Sum of pessimistic cases73

Read the table from the top. The number most teams would give, 37 days, is below even the median outcome. The sum of pessimistic cases, 73, is so unlikely that quoting it would be sandbagging. The shared-risk simulation barely moves the median but widens the tail: the P80 is 50 days and the P95 almost 58, because correlated risks are exactly the ones that do not cancel out. The simulation's median sits above the PERT sum because a triangular distribution's mean is a third of the sum of its three points, which weights the pessimistic case more heavily than PERT's formula does: 43.7 against 40.3 here. Neither is the truth; both say the likely-case sum is too low.

Calibrate with your own history

Every team has a bias, and the cheapest correction is your own record. For recent projects, collect the original estimate and the actual effort and compute the ratio. This is reference-class forecasting, the outside view: instead of asking only how this project looks from inside, ask how projects like it have actually gone.

history = [  # (project, estimated days, actual days)
    ("billing export", 20, 27), ("audit log", 35, 41), ("search revamp", 60, 95),
    ("2FA", 15, 18), ("tenant sharding", 80, 118), ("webhooks v2", 25, 30),
]
ratios = sorted(actual / est for _, est, actual in history)
median = ratios[len(ratios) // 2]
print("ratios", [round(r, 2) for r in ratios], "median", round(median, 2))

These six projects are illustrative; their median ratio is 1.35, every ratio is above 1, and the two largest projects overran most. With a median like that, either multiply new estimates by it or, better, find which tasks you keep missing and add them to your decomposition template. Re-check the ratio each quarter. Six data points are noisy; even so, they beat no data.

From effort to calendar time

Engineer-days are not calendar days. Convert using the team's real capacity: people assigned, the fraction of their time actually available for project work after meetings, support, on-call, reviews of others' work and interruptions, and holidays. A focus factor between 0.5 and 0.7 is common; measure yours from past sprints rather than assuming it.

For the SSO example, two engineers at a focus factor of 0.6 provide 1.2 engineer-days per working day. The P50 of 43.5 engineer-days becomes about 36 working days, roughly seven and a quarter weeks; the P80 of 50.3 becomes about 42 working days, roughly eight and a half weeks. Then check the dependency structure: SAML testing cannot start before the SAML flow exists, and the security review has a queue. Adding a third engineer helps only where tasks can run in parallel, and adding people late to a project adds onboarding and coordination cost, the effect Fred Brooks described in The Mythical Man-Month.

Communicating the estimate

Present a range with a confidence level, the assumptions behind it, and what would change it. For example: "Fifty-fifty within seven and a quarter weeks of starting, eighty percent within eight and a half, assuming two engineers and that our three pilot customers' identity providers behave like the documentation says. The biggest risk is SAML edge cases; the spike in week one will narrow this range, and we will re-estimate then."

Then let the business choose the confidence level for the commitment. A launch tied to a public event may need P90; an internal tool may be fine at P50. Re-estimate at every milestone with the remaining work only, and publish the new range even when it moves the wrong way; an estimate that is never updated becomes a fiction everyone politely ignores.

Failure modes

  • Anchoring. Someone states a number first and every later estimate drifts towards it. Collect estimates independently before discussing them.
  • The planning fallacy. Imagining the project going to plan instead of consulting how similar projects went. Use the reference class.
  • Happy-path scope. Feature code estimated, rollout, migration and review not. Use a decomposition template.
  • Hidden padding. Each layer of management adds a buffer silently, so nobody knows the real estimate. Make contingency explicit, as a percentile, in one place.
  • Estimates turned into deadlines. A P50 quoted as a promise fails half the time by definition.
  • Parkinson's law. Work expands to fill a generous single-number deadline. Commit to a percentile, and track against the median.
  • Precision theatre. Estimating in hours or arguing about story points for small tasks. Spend the effort on the riskiest items instead.

Trade-offs: how much estimating is enough

Estimation has a cost and should be proportional to the decision it informs. A two-week change needs a sentence. A quarter-long project that decides hiring or a contract needs the full method. An alternative for teams with steady flows of similar-sized work is to forecast from throughput: count how many items the team finishes per week, split the project into items of similar size, and simulate completion dates from past weekly throughput. It avoids per-task guessing but needs stable history and comparable items.

Estimating well also depends on good design. A written design with explicit non-goals makes decomposition possible; see how to write a design doc and how to review one. For the money side of the same plan, read how to estimate cloud cost, and for rollout tasks that belong in every estimate, the launch guide.

What to do next

  1. Write the goal, non-goals and definition of done for your next project, including rollout and documentation.
  2. Decompose until each task's likely case is a few days, and schedule spikes for anything nobody can describe yet.
  3. Have two or three engineers give independent low, likely and high figures, then discuss only the large disagreements.
  4. Run the simulation script with a shared-risk factor and read P50 and P80, not the sum of likely cases.
  5. Collect estimate-versus-actual ratios for your last five or more projects and apply what they tell you.
  6. Convert effort to calendar time with a measured focus factor and the real dependency order.
  7. Communicate a range, a confidence level, assumptions and the date of the next re-estimate, and keep that date.
Key takeaway: Treat an estimate as a distribution, not a number: decompose until tasks are describable, give each a low, likely and high figure, sum them with a short simulation that includes shared risk, correct the result with your own history, and convert effort to dates with real capacity. Then let the business choose the percentile it commits to, and re-estimate at every milestone as spikes and progress narrow the range.