Exponential Backoff
The retry discipline that prevents failed calls from becoming outages: wait longer after each failure, add randomness, and give up on purpose.
Last reviewed: July 25, 2026
What is exponential backoff?
Exponential backoff answers “when should I retry?” with “later each time”: wait 1s, then 2s, 4s, 8s… doubling until a cap. The doubling matters because naive immediate retries are an accelerant — a service hiccups, every client retries at once, the retry wave is larger than the original load, and a blip becomes an outage. Backoff spreads the recovery pressure across time.
Jitter is not optional
Pure exponential backoff synchronizes clients: everyone who failed at T retries at T+1, T+3, T+7 — coordinated stampedes at predictable instants (the thundering herd). Adding randomness breaks the synchronization; AWS’s analysis landed on full jitter — sleep a uniform random amount between 0 and the exponential cap — as the best-behaved variant, and it’s the one to copy. One line of randomness is the difference between recovery and rhythmic self-DDoS.
The judgment calls around the loop
Retry only what’s retryable: 429s and 5xx/timeouts yes; 400s and auth failures never — retrying a malformed request 8 times just multiplies noise. Idempotency is a prerequisite: a timed-out call may have succeeded; retrying non-idempotent operations without idempotency keys double-charges customers (see message queues for the same law). Bound the total: caps and retry budgets, because a request retrying for two minutes holds resources everywhere. And in microservices, beware retry amplification — three layers each retrying 3× turns one user request into 27 backend calls; retry at one layer, use circuit breakers above.
What people get wrong
- Backoff without jitter — the textbook half-solution.
- Retrying everything, including the un-retryable and the non-idempotent.
- Every layer retrying — the multiplication that melts databases during partial outages.
Primary source: AWS Architecture Blog — Exponential Backoff and Jitter
Combining Backoff With Jitter
Pure exponential backoff, where every failed client retries after exactly the same escalating delay, can itself create a problem: if many clients fail simultaneously (during a brief service outage, for example), they’ll all retry at exactly the same calculated intervals, creating synchronized bursts of retry traffic that can overwhelm a recovering service right as it’s trying to come back online — a phenomenon sometimes called the “thundering herd” problem. Adding jitter (a small random variation to each client’s calculated backoff delay) spreads out these retries over a range of times instead of a single synchronized moment, which is why production-grade retry implementations almost always combine exponential backoff with jitter rather than using pure exponential backoff alone.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.