Most load tests fail before the first request is sent. Someone points a tool at staging, turns a dial until something breaks, screenshots a graph, and the team learns nothing it can act on. The graph cannot say whether production will survive next month, because nobody wrote down what next month looks like, the environment was not comparable, and the test never checked that the load generator itself kept up.

This guide is the procedure, end to end: write the question, build a workload model from real traffic, prepare the environment, run a fixed sequence of tests, read the results, and write a report with a decision in it. It uses k6 for the examples, but every step applies to any tool. For how to build a durable load testing platform, see Load testing architecture; for WebSocket and gRPC streams, see Load Testing Bidi Servers with k6.

Advertisement

Step 1: write the question as a hypothesis

A load test is an experiment, so it needs a falsifiable hypothesis with numbers in it. Fill in this sentence before anything else: at a sustained arrival rate of X requests per second, with traffic mix M, the system meets latency target L and error target E, for duration D, in environment V.

Our running example is a checkout API before a seasonal sale. The hypothesis reads: at 2,400 requests per second, with the production mix from last year's peak day, checkout keeps p99 latency under 400 ms and errors under 0.5 percent for two hours, in the performance environment at full production size. Each number has a source. The 400 ms and 0.5 percent come from the service level objective (see running an SLO programme). The 2,400 comes from step 2. Also write down what result would change a decision: if the hypothesis fails, the team adds capacity or delays a feature launch. A test whose outcome changes nothing is not worth running.

Step 2: build the workload model from production

The workload model says how much load, of what kind, arriving how. Derive it from access logs or metrics, never from guesses. Last year's peak day served 2.9 million checkout API calls in its busiest hour: an average of about 806 per second. Within that hour, the busiest minute ran 1.5 times the hourly average, about 1,200 per second. Marketing expects double the traffic, and the team wants a margin, so the target is 2 x 1,200 = 2,400 per second sustained.

EndpointShare of callsMedian payloadNotes
GET /cart52%3 KBRead-heavy, cache-friendly
POST /cart/items27%1 KBWrites to the cart store
POST /checkout12%2 KBPayment provider call, the slow path
GET /orders/{id}9%4 KBOrder status polling

Two modelling choices matter more than any other. First, model arrivals as an open system: real users arrive whether or not your server is slow, so the test must keep sending at the target rate even when responses lag. A closed model, where a fixed number of virtual users each wait for a response before sending the next, quietly reduces load as latency rises and hides the very collapse you are looking for. Second, use realistic data: thousands of distinct carts and products, not one user ID hammering one hot row and one warm cache line.

Little's law turns the model into generator settings. Mean requests in flight equal arrival rate times mean time in system. Sizing conservatively with the p99 of 0.4 s instead of the mean gives 2,400 x 0.4 = 960 in flight, so the generator needs roughly 1,000 virtual users ready, with headroom to 2,000 if latency degrades.

Advertisement

Step 3: prepare the environment

Results only transfer to production if the environment differs in known, small ways. Walk this list and write the answer to each in the test plan:

  • Size and shape. Same instance types, same replica counts or a documented fraction, same autoscaling rules. A half-size environment does not give half the capacity: shared databases and caches rarely scale linearly.
  • Data volume. Production-sized tables and indexes. A query that is instant on 10,000 rows can take seconds on 400 million.
  • Dependencies. Decide per dependency whether to hit the real thing, a sandbox, or a stub with recorded latency. A payment sandbox with a rate limit of 50 calls per second will end your test early for the wrong reason.
  • Observability. Server-side metrics (CPU, memory, connection pools, queue depths, garbage collection) on the same dashboard as client-side results, with synchronized clocks.
  • Load generators. Sized and placed so they are never the bottleneck: keep them under about 70 percent CPU, and put them in the network position real traffic comes from.
  • Permission. Tell the owners of every shared component and every third party. An unannounced test looks exactly like an attack.

Step 4: run a fixed sequence of tests

One test cannot answer every question, so run several, in order, each building on the last.

TestShapeDurationQuestion it answers
Smoke1 to 5 percent of target5 minDoes the script work, and are errors zero at trivial load?
BaselineCurrent production peak15 minDo results match production metrics at known load?
StepIncrease rate in steps past target7 min per stepWhere is the knee, and what saturates first?
SoakTarget rate, held2 to 8 hoursDo leaks, log growth or pool exhaustion appear over time?
SpikeJump from baseline to 2x target in seconds10 minDoes autoscaling react before queues overflow?

The baseline is the calibration step that most teams skip. If the environment shows 120 ms p99 at 1,200 per second while production shows 180 ms at the same rate, the environment is flattering you and every later number needs that caveat. Fix the gap or record it.

Step 5: script the step test

k6 has arrival-rate executors for the open model. ramping-arrival-rate starts iterations at the rate you give it, regardless of response time, using a pool of preAllocatedVUs and growing to maxVUs if needed. When it runs out of virtual users it does not slow down silently: it counts each iteration it could not start in the dropped_iterations metric. That metric is your proof the generator kept up, so put a threshold on it.

import http from "k6/http";
import { check } from "k6";

const steps = [];
for (let rate = 600; rate <= 3600; rate += 300) {
  steps.push({ target: rate, duration: "1m" });   // ramp to the next step
  steps.push({ target: rate, duration: "6m" });   // hold it
}

export const options = {
  scenarios: {
    step: {
      executor: "ramping-arrival-rate",
      startRate: 300, timeUnit: "1s",
      preAllocatedVUs: 1000, maxVUs: 2000,
      stages: steps,
    },
  },
  thresholds: {
    http_req_failed: ["rate<0.005"],
    http_req_duration: [{ threshold: "p(99)<400", abortOnFail: true, delayAbortEval: "2m" }],
    dropped_iterations: ["count<1"],
  },
};

const carts = JSON.parse(open("./carts.json"));   // 50,000 real-shaped cart IDs

export default function () {
  const cart = carts[Math.floor(Math.random() * carts.length)];
  const r = Math.random();
  let res;
  if (r < 0.52)      res = http.get(`${__ENV.BASE}/cart/${cart}`, { tags: { ep: "get_cart" } });
  else if (r < 0.79) res = http.post(`${__ENV.BASE}/cart/${cart}/items`, item(), { tags: { ep: "add_item" } });
  else if (r < 0.91) res = http.post(`${__ENV.BASE}/checkout`, pay(cart), { tags: { ep: "checkout" } });
  else               res = http.get(`${__ENV.BASE}/orders/${cart}`, { tags: { ep: "order" } });
  check(res, { "status 2xx": (x) => x.status >= 200 && x.status < 300 });
}

The abortOnFail threshold stops the test once p99 breaks the objective, after a delay so a warm-up blip does not end the run. The ep tags let you split latency per endpoint, which matters: an aggregate p99 can look fine while the 12 percent of checkout calls are failing.

Step 6: run day

  1. Announce the window in the team and dependency channels, with a named person who can stop the test.
  2. Confirm dashboards show both server and client metrics, and note the start time in UTC.
  3. Run the smoke test. Any error at trivial load is a script or environment bug; fix it first.
  4. Run the baseline and compare it to production. Record the gap.
  5. Run the step test. At each step, write down the first resource that moves toward its limit.
  6. Stop immediately if a shared dependency degrades or anyone outside the test is affected.
  7. Run the soak at the target rate, then the spike.
  8. Export raw results, not just the summary, and store them with the script version and environment build.

Step 7: read the results and find the knee

A step test: offered load against p99 latency and achieved throughputoffered load (requests per second), one step every 7 minutesp99SLO: p99 400 msknee: about 2,700 req/sLeft of the kneelatency flat, throughput tracks offered loadqueues are short; extra load is absorbedRight of the kneequeues grow without boundp99 explodes, errors follow
Latency stays flat while queues are short, then bends sharply once a resource saturates. The knee, not the crash point, is your usable capacity.

Plot each step's p99 and achieved throughput against offered load. Left of the knee, latency is flat and throughput equals offered load. At the knee, some resource reaches saturation, queues begin to grow, and latency curves upward. Beyond it, achieved throughput stops rising and errors appear. Usable capacity is the last step left of the knee that still meets the objective, not the rate at which the system falls over.

def find_knee(steps, slo_p99_ms=400, max_err=0.005):
    # steps: [{"rate": offered, "achieved": rps, "p99": ms, "err": fraction}], in order
    best = None
    for prev, cur in zip([None] + steps, steps):
        kept_up = cur["achieved"] >= 0.98 * cur["rate"]
        ok = cur["p99"] <= slo_p99_ms and cur["err"] <= max_err and kept_up
        bending = prev is not None and cur["p99"] > 1.5 * prev["p99"]
        if not ok or bending:
            return best, cur      # last good step, first bad step
        best = cur
    return best, None             # never found the knee: test more load

In the example, steps up to 2,700 per second hold p99 near 180 ms; at 3,000, p99 jumps to 520 ms and the database connection pool shows 100 percent use. The hypothesis passes at 2,400 with 12 percent headroom to the knee, and the report names the pool as the first bottleneck. Cross-check with the utilization law, utilization equals throughput times service time: if each request holds a connection for 40 ms and the pool has 120 connections, it saturates at 120 / 0.04 = 3,000 per second, which matches what the test showed. When theory and measurement agree, trust the number. When they disagree, find out why before you publish anything. Capacity planning shows how to carry the number forward into a plan.

Step 8: check the test before you trust it

Failure of the testHow to detect itFix
Generator saturatedGenerator CPU over 70 percent, dropped_iterations above zeroAdd generators, distribute the run
Closed model hid the collapseThroughput falls while latency rises and errors stay lowUse an arrival-rate executor
Cache too warmHit rate far above productionUse production-sized key space and realistic randomness
Stubbed dependency too fastStub latency far below production p99Replay recorded latency distributions
Averages hid the tailOnly means in the reportReport p50, p95, p99 per endpoint
Wrong build testedVersion in results differs from the releaseRecord build IDs with every run

Step 9: write the report

Keep it to one page. State the hypothesis and whether it passed. Give the knee, the first saturated resource and the evidence, the environment differences and their likely effect, and the decision: ship, add capacity, or fix a bottleneck and retest. Link the raw results and the script commit. A report without a decision is a status update, not a test result.

What to do next

  1. Write the hypothesis sentence for your next test, with every number sourced.
  2. Pull the busiest hour and busiest minute from production logs and derive your target rate and mix.
  3. Switch any closed-model scripts to an arrival-rate executor with a dropped_iterations threshold.
  4. Run a baseline at current production peak and record the gap from production.
  5. Run a step test past the target and identify the first saturated resource.
  6. Compute that resource's limit with the utilization law and compare it to the measured knee.
  7. Schedule a soak of at least two hours before every major release.
  8. Publish a one-page report with a decision and archive the raw results.
Key takeaway: A load test is an experiment. Write a hypothesis with sourced numbers, derive an open workload model from production traffic, and make the environment comparable or document how it differs. Then run smoke, baseline, step, soak and spike tests with an arrival-rate generator that proves it kept up. Usable capacity is the knee, not the crash point. Name the first saturated resource, check it against the utilization law, and end the report with a decision.