A retry is a bet that the next attempt will meet a different world than the last one did. When the failure was a dropped packet, a leader election or a pod being replaced, the bet is good: a second attempt a few milliseconds later usually succeeds and the user never notices. When the failure was an overloaded dependency, the bet is bad, and every client making it at once adds load to the thing that is already failing. The same three lines of code that hide a blip also turn a brownout into an outage.
Exponential backoff and jitter are the two controls that make the bet safe. Backoff spaces attempts further apart each time, so a struggling server sees less traffic the longer it struggles. Jitter randomises the wait, so thousands of clients that failed at the same instant do not all come back at the same instant. Neither is enough alone. This page builds a retry policy from first principles: what to retry, how long to wait, how to cap retry traffic with a budget, and where in a call chain retries belong, then works through a service that multiplied one failure by 27.
Why retries help and why they hurt
Each attempt consumes a connection, a client thread, a server request slot and a share of whatever is scarce downstream. When failures are rare and independent, the extra attempts are noise. When they are correlated, because the dependency is down or overloaded, every client retries at once and offered load rises exactly when capacity has fallen.
The multiplier compounds with depth. If a front end calls a service that calls a database, and each layer makes up to three attempts, one user request can become 3 x 3 x 3 = 27 database queries during an outage. The database was failing because it had too much work; now it has up to 27 times as much. This is retry amplification, and it is the usual reason a short dependency hiccup turns into a long incident: the dependency recovers, the queued retries hit it, and it falls over again.
What is safe to retry
Two questions decide whether an attempt may be repeated. First, is the error transient? Connection resets, timeouts, HTTP 502, 503 and 504, gRPC UNAVAILABLE and throttling responses such as HTTP 429 usually are. Validation, authentication, not-found and conflict errors are not; retrying them only adds latency and load. Classify by error code before status code, because some services put a throttling code inside a 5xx.
The second question is whether the operation is safe to perform twice. A timeout tells you nothing about whether the server acted. A GET is safe. A PUT that writes a full value is safe. A POST that charges a card or appends to a ledger is not, unless the server deduplicates. The standard fix is an idempotency key: the client generates a unique key per logical operation, sends it with every attempt, and the server records the key with the result so a repeat returns the stored result instead of acting again. The key must be created once, before the first attempt, not inside the retry loop; generating it per attempt defeats the point.
Exponential backoff
Exponential backoff sets the wait before retry n to base * 2^n, capped at some maximum. With a 100 ms base the waits are 100, 200, 400, 800 ms and so on. Doubling matters because it makes the total retry rate from a population of clients fall geometrically during a long failure: a client that has failed five times is sending one request every few seconds, not ten per second. The cap matters because without it the twentieth wait is more than a day.
Backoff alone has a flaw at scale. Clients that failed together compute the same schedule and retry together at 100, 300 and 700 ms. The server sees synchronised spikes separated by silence: idle between them, overloaded during them. Jitter spreads each wave across its window.
Jitter variants and a retry helper
The AWS Architecture Blog post "Exponential Backoff And Jitter" compared three randomised variants by simulating many clients contending for one resource. Full jitter draws the wait uniformly from zero to the exponential value. Equal jitter keeps half the exponential value and randomises the other half. Decorrelated jitter grows from the previous sleep rather than from the attempt number. In that simulation, full and decorrelated jitter did the least total work and finished soonest, equal jitter was somewhat worse, and plain backoff without jitter was far worse than all three.
| Variant | Wait before retry n | Character |
|---|---|---|
| No jitter | min(cap, base * 2^n) | Synchronised waves; avoid |
| Full jitter | random(0, min(cap, base * 2^n)) | Widest spread; some retries happen almost immediately |
| Equal jitter | t/2 + random(0, t/2), t = min(cap, base * 2^n) | Guarantees a minimum wait, less spread |
| Decorrelated | min(cap, random(base, prev * 3)) | Depends on the previous sleep, not the attempt count |
Full jitter is the sensible default, and it is what the AWS SDKs document for their updated, opt-in behavior. Use equal jitter when an immediate retry is known to be pointless, for example after a throttling response. Here is a complete retry helper with full jitter, a deadline, error classification and a shared budget:
import random, time
class RetryBudget:
# Token bucket shared by every call from this client. Failures spend, successes refill.
def __init__(self, capacity=100.0, cost=5.0, refill=1.0):
self.capacity, self.cost, self.refill = capacity, cost, refill
self.tokens = capacity
def try_spend(self):
if self.tokens < self.cost:
return False
self.tokens -= self.cost
return True
def on_success(self):
self.tokens = min(self.capacity, self.tokens + self.refill)
class Retryable(Exception):
pass # raised by call() for timeouts, resets, 503, 429
def call_with_retry(call, budget, deadline_s=2.0, base=0.05, cap=1.0, max_attempts=3):
deadline = time.monotonic() + deadline_s
attempt = 0
while True:
remaining = deadline - time.monotonic()
try:
result = call(timeout=remaining) # per-try timeout never exceeds the deadline
budget.on_success()
return result
except Retryable:
attempt += 1
if attempt >= max_attempts or not budget.try_spend():
raise
sleep = random.uniform(0, min(cap, base * 2 ** (attempt - 1)))
if time.monotonic() + sleep >= deadline:
raise # no time for another real attempt
time.sleep(sleep)The budget lives outside the call so concurrent requests share it, the per-try timeout shrinks as the deadline is used, and non-retryable exceptions are never caught, so they fail on the first attempt.
Retry budgets and retrying at one layer
A maximum attempt count bounds the retries for one request, but not for the system. If every request is failing, three attempts each still means three times the load. A retry budget bounds the ratio: retries are allowed only while they are a small fraction of traffic. Once failures are widespread the budget is empty, clients fail fast, and the dependency sees roughly its normal request rate instead of a multiple of it.
gRPC builds this in. Its retry design (proposal A6) gives the client a token count that starts at maxTokens; every failed RPC subtracts 1, every successful RPC adds tokenRatio, and retries stop while the count is at or below half of maxTokens. The AWS SDK documentation describes the same shape for its updated retry behavior, which as of 2026-10-03 is opt-in through the AWS_NEW_RETRIES_2026 environment variable: a 500-token bucket per client, 14 tokens per retry after a transient error, 5 per retry after throttling, with a successful retry refunding its own cost and a first-try success adding 1. Those numbers are the documented values for that opt-in mode; check the current page before relying on them.
The second structural rule is to retry at one layer only. In a chain of services, pick the layer closest to the failure that can see a meaningful error, usually the client of the flaky dependency, and make every other layer pass errors through. Combine this with a circuit breaker, which stops calls entirely when a dependency is clearly down, and with bulkheads, which stop one failing dependency from consuming every thread.
Deadlines and per-try timeouts
Every retry policy needs a deadline that the user experience can afford. If a page must render in two seconds, the whole retry loop gets less than two seconds, and each attempt gets what remains. Without one, three attempts with ten-second socket timeouts hold a request for thirty seconds. Propagate the deadline downstream, as gRPC does with its deadline metadata, so a service three hops away does not keep working on a request whose caller has already timed out.
Set per-try timeouts from measured latency: a little above the dependency's 99th percentile cuts off the slow tail, while a timeout below the median just manufactures failures.
Server-directed delays
Sometimes the server knows better than the client when to come back. HTTP defines the Retry-After response header, which carries either a number of seconds or a date and is commonly sent with 429 and 503 responses. gRPC uses a trailer named grpc-retry-pushback-ms: a non-negative integer tells the client exactly how long to wait, and a negative or unparseable value tells it not to retry at all. The AWS SDK documentation describes an x-amz-retry-after header in milliseconds, which it clamps between the computed backoff and that backoff plus 5 seconds, and does not jitter because the service is expected to have jittered it already.
Honour these signals, but clamp them: wait at least the computed backoff and at most the remaining deadline, so a buggy one-hour value cannot freeze a client and a zero cannot disable backoff. Note also that gRPC's own backoff multiplies each wait by random(0.8, 1.2) rather than using full jitter, and treats any maxAttempts above 5 as 5; it relies on its throttling budget for the rest.
Worked example: a failover that became an outage
Consider an order page. The browser calls an API gateway, the gateway calls an order service, and the order service calls an inventory database. Each layer was written by a different team and each uses a client library with three attempts and a 100 ms backoff without jitter. Normal traffic is 2,000 requests per second, and the database comfortably serves 4,000 queries per second.
A failover makes the database reject connections for 20 seconds. With three layers of three attempts, offered load at the database peaks at up to 2,000 x 27 = 54,000 queries per second, in synchronised waves 100 and 300 ms after each failure. The new primary comes up into that backlog, its latency rises past the client timeouts, the timeouts count as failures and trigger more retries, and the 20-second failover becomes a 15-minute outage.
The fix has four parts. Retries move to the order service, which can see the database's error codes; the gateway and browser stop retrying server errors. The order service uses full jitter with a 50 ms base and a 1 s cap, under a 1.5 s deadline. And a retry budget caps retries near 10 percent of successful calls. In the same failover, database load stays near 2,000 x 1.1 because the budget empties within the first second, and the new primary comes up to a load it can serve. Users see 20 seconds of errors instead of 15 minutes.
Failure modes
- Retry storm after recovery. Retries queued during an outage arrive together when the dependency returns. Jitter, budgets and a circuit breaker that half-opens gradually prevent the second collapse.
- Duplicate side effects. A timed-out payment request is retried and the customer is charged twice. Use idempotency keys created once per logical operation, and make the server store the result against the key.
- Retrying the unretryable. A 400 or 403 retried three times wastes time and budget and hides the real error. Classify explicitly; never retry by default on every exception.
- Timeouts longer than the deadline. The per-try timeout is the library default of 30 s, so one attempt outlives the user. Derive per-try timeouts from the remaining deadline.
- Hidden nested retries. An HTTP client, a database driver and a service mesh sidecar each retry silently. Inventory every layer that retries before tuning any of them.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| More attempts | Survives longer blips | More amplification during outages; longer tail latency |
| Longer cap | Gentler on a recovering dependency | Requests wait longer before failing |
| Full jitter | Best spread, least total work | Some retries almost immediate |
| Retry budget | Load stays near normal during outages | Some recoverable requests fail fast |
| Retrying at one layer | No multiplication | Needs agreement across teams |
| Hedged requests instead | Cuts tail latency on healthy systems | Extra load all the time, not only on failure |
Hedging sends a second copy before the first fails, so it adds load continuously and suits only idempotent reads under a budget. Servers still need their own protection; see rate limit system architecture, and gray failure for the degraded-but-not-down case that retries quietly hide.
What to do next
- List every layer in one request path that retries: application code, HTTP or gRPC client, database driver, mesh sidecar, load balancer. Write down attempts and backoff for each.
- Multiply the attempt counts along the path. If the product exceeds about 3, remove retries from all but one layer.
- Classify errors explicitly in that layer: retry timeouts, resets, 503, 504, 429 and throttling codes; never retry 4xx validation or auth errors.
- Add idempotency keys to every non-idempotent operation that is retried, generated once per logical operation.
- Switch to full jitter with a small base and a cap of about one second for interactive paths.
- Set an end-to-end deadline from the user-facing latency target, and derive per-try timeouts from what remains.
- Add a retry budget shared per client instance, and alert when it empties.
- Honour Retry-After and gRPC pushback, clamped to the remaining deadline.
- Export first-attempt failure rate, retries per request and budget level; test the policy by injecting a 30-second dependency outage in staging.