Put a limit on every retry before you tune anything, because a loop that retries forever turns one failed call into an outage. Then the answer to what is exponential backoff fits in a line: each retry waits about twice as long as the last, 1, 2, 4, then 8 seconds, plus a little randomness called jitter, so a thousand clients do not return in the same second.
What is exponential backoff: a retry that waits longer each time
Exponential backoff is a retry strategy in which each failed attempt waits about twice as long as the one before, up to a cap, and stops at a fixed limit. With a 1 second base the waits run 1, 2, 4 and 8 seconds. A random offset called jitter keeps many clients from retrying together.
This control is one of the checks in hardening SaaS applications for resilience, and it applies to every call your app makes to a service it does not run: the AI provider, the payment provider, the email service. In the Production Hardening Sprint it is deliverable 6.5, which adds bounded retries with backoff to transient failures where repeating the operation is safe.
The control has four parts, and dropping any one of them breaks it:
- A base delay: the first wait, such as 1 second.
- A multiplier: how much each wait grows. A multiplier of 2 is the convention, and it is where the doubling comes from.
- A cap on any single wait, so the tenth retry does not sleep for 512 seconds, about 8.5 minutes.
- A limit on the number of attempts or on the total time spent, so the retry ends.
As one line of pseudo-code, the wait before retry number n, counting from 0, is delay = min(cap, base * 2 ** n). With full jitter the wait is a random number between 0 and that value, the variant AWS’s post on exponential backoff and jitter calls “a small change to the sleep function”. Worked through with a 1 second base, a 30 second cap and 5 retries, that exponential backoff formula gives this schedule:
| Retry | Wait without jitter | Wait with full jitter | Total waited so far, no jitter |
|---|---|---|---|
| 1 | 1 s | 0 to 1 s | 1 s |
| 2 | 2 s | 0 to 2 s | 3 s |
| 3 | 4 s | 0 to 4 s | 7 s |
| 4 | 8 s | 0 to 8 s | 15 s |
| 5 | 16 s | 0 to 16 s | 31 s |
A sixth retry would compute 32 seconds, and the cap would cut it to 30. With jitter the total is shorter on average and never longer, because each wait is drawn from below its no-jitter value.
The exponential backoff algorithm is older than web APIs: networks that detect or avoid collisions, Ethernet among them, use it when they retransmit frames. AWS’s retry with backoff pattern states the purpose of an exponential backoff retry for applications: it “improves application stability by transparently retrying operations that fail due to transient errors.” Google Cloud’s Memorystore documentation adds that the maximum wait “is typically 32 or 64 seconds.”
You may already run a bounded schedule like this. Claude Code reconnects a dropped remote MCP server with exponential backoff: up to five attempts, starting at a one-second delay and doubling it each time. When an HTTP or SSE server’s first connection fails with a transient error, such as a 5xx response, a refused connection or a timeout, it retries up to three times; an authentication or not-found error is not retried, because it needs a configuration change, though an authentication error is retried when a headersHelper is the server’s only source of the Authorization header. Stdio servers are local processes, and Claude Code does not reconnect them automatically.
Why add jitter to exponential backoff?
Jitter is a random amount added to, or drawn from, each backoff delay. Without it, 1,000 clients that failed in the same second all wait 1, 2, 4 and 8 seconds together, and the recovering service meets the same spike after every wait.
Plain backoff spaces the retries out in time, but it does not spread the clients. AWS’s simulation found “There are still clusters of calls”, with quiet gaps between them. To retry with exponential backoff and jitter, you pick one of the three variants the post compares:
- Full jitter: the whole wait is random, between 0 and the capped exponential value.
- Equal jitter: in the post’s words, “we always keep some of the backoff and jitter by a smaller amount”.
- Decorrelated jitter: “similar to ‘Full Jitter’, but we also increase the maximum jitter based on the last random value”.
In AWS’s results, “The no-jitter exponential backoff approach is the clear loser.” Equal jitter “does slightly more work than ‘Full Jitter’, and takes much longer,” and between the other two, “The ‘Full Jitter’ approach uses less work, but slightly more time.” Those are AWS’s results from its own simulation, not mine. My working rule from them: use full jitter unless you have a reason not to.
What goes wrong without it
Each row is a way retry code goes wrong in a live app, with what you or the provider would notice and which part of the control stops it.
| What the code does | What the user or provider sees | The part of the control that prevents it |
|---|---|---|
| Never retries | One dropped connection to the AI provider shows the user an error that a second attempt could have cleared | A small attempt limit on transient errors |
| Retries instantly in a loop | Calls go out as fast as the app can send them to a provider that is already failing; a 429 arrives on top of the 500, and the shared rate limit runs out for every other user | The base delay and the multiplier |
| Backs off on the same schedule as every other client | The recovering service is knocked over again by a synchronized wave of retries | Jitter |
| Retries at several layers | The browser, the API route and the SDK each make 3 attempts, and one click becomes 27 calls | One layer owns the retries |
| Retries work that is not safe to repeat | A charge, an email or an insert runs twice, because the first attempt succeeded and only the response was lost | Retry only safe work, or make it safe first |
When a retry storm hits, the server is overloaded by the very clients waiting for it to recover. Microsoft’s retry storm antipattern puts it this way: “When a service becomes unavailable or busy, frequent client retries can prevent the service from recovering and worsen the problem.” Its fixes are the parts of the control above: limit the attempts and their duration, increase the wait between them, and honor the retry-after header when the server sends one.
The layer row looks small at 27. The Amazon Builders’ Library on timeouts, retries and backoff runs the same arithmetic on a five-deep stack with three retries at each layer and concludes that “the load on the database will increase 243x, making it unlikely to ever recover.”
My June and July 2026 audits scored 21 third-party apps on 12 pillars, and Reliability and Correctness averaged 31.4 out of 100, ranked the weakest of the 12. Those 21 are a set I chose and audited, not a random sample, so the average describes those apps, not AI-built apps in general.
How to retry failed API calls safely
A safe retry policy has four rules: retry only transient errors, retry only work that gives one result when run twice, wait with exponential backoff and jitter, and stop after a fixed number of attempts; my working rule is about 3 for a request a user is waiting on.
Should we retry on 500? The errors worth retrying and the ones to fail at once
Retryable errors are the ones that can clear on their own: network failures, 408, 429, 502, 503, 504, and 500 limited to a few attempts. Under my working rule, client errors such as 400, 401, 403, 404 and 422 are not retried, because the same request fails the same way.
A transient fault, in Microsoft’s guidance on transient faults, is a “Temporary failure that’s self correcting and likely to succeed if retried after a suitable delay.” The meaning column below follows RFC 9110’s status definitions (RFC 6585 for 429); the retry column is my rule.
| Status or error | Meaning | Retry? (my rule) | Note |
|---|---|---|---|
| Network error, connection refused or reset | No HTTP response arrived | Yes | The server may already have received the request |
| 408 Request Timeout | The server did not receive a complete request in the time it was prepared to wait | Yes | |
| 429 Too Many Requests | The user has sent too many requests in a given amount of time | Yes, after the wait the server names | Use Retry-After when present |
| 500 Internal Server Error | The server hit an unexpected condition that prevented it from fulfilling the request | Yes, a few attempts only | A passing fault or a bug |
| 502 Bad Gateway | A gateway or proxy got an invalid response from the server behind it | Yes | |
| 503 Service Unavailable | A temporary overload or scheduled maintenance | Yes | Use Retry-After when present |
| 504 Gateway Timeout | A gateway or proxy got no timely response from the server behind it | Yes | |
| 400 Bad Request | The server cannot or will not process the request because of a perceived client error | No | Fix the request |
| 401 Unauthorized | The request lacks valid authentication credentials | No | Fix the key or the session |
| 403 Forbidden | The server understood the request but refuses to fulfill it | No | |
| 404 Not Found | The server found no current representation of the resource | No | |
| 409 Conflict | The request conflicts with the current state of the resource | No | Read the state again, then decide |
| 422 Unprocessable Content | The syntax is correct, but the server could not process the instructions | No | Fix the data |
When a 429 or a 503 carries a Retry-After header, wait that long instead of the computed delay. Servers send RFC 9110’s Retry-After header “to indicate how long the user agent ought to wait before making a follow-up request”, and its value is either a number of seconds, as in Retry-After: 120, or an HTTP date. RFC 6585’s 429 status says the response “MAY include a Retry-After header indicating how long to wait before making a new request”, so plan for both cases: header present and header missing.
My working rule gives a 500 a small number of attempts, never the full budget: it can be a passing fault or a bug in the provider, and the limit is what tells them apart, because a bug fails on every attempt. The 4xx rows are the opposite case. Microsoft’s retry storm page says of a 400 that “retrying the same request likely won’t help because the server has already informed you that your request isn’t valid.”
Only retry work that is safe to repeat
Work that is safe to repeat leaves one result however many times it runs. Reads qualify as they are. A charge, an email or an insert qualifies only with an idempotency key, a sent marker or a unique constraint standing between the retry and a duplicate.
The test for a write is whether running it twice still leaves one result, and the common calls sort like this:
| Call | Safe to repeat as is? | What makes it safe |
|---|---|---|
| Fetch a record | Yes | Nothing extra |
| AI completion | Yes, though a repeat can cost twice | A low attempt limit |
| Send an email | No | The email provider’s idempotency key where it offers one; a sent marker checked before sending stops a repeat only after the send was recorded |
| Create a charge | No | The payment provider’s idempotency key, the same one on every attempt |
| Insert a row | No | A unique constraint, or an upsert on one |
The insert row has a real case behind it. I audited an AI coding workspace, in June and July 2026, that saved every chat message through a database function, one copy of which was an upsert keyed on project and sequence number, but the table had no unique constraint on those columns. Depending on which schema copy was live, every save errored, or retries and concurrent saves could write duplicate messages. The lesson I take from it: a retried insert is only as safe as the unique constraint behind it, and a chat save that assumed a constraint the live schema lacked is the full story.
For the charge row, Stripe’s idempotent requests are the model. You send an Idempotency-Key header, and Stripe saves “the resulting status code and body of the first request made for any given idempotency key, regardless of whether it succeeds or fails”; later requests with that key get the same result back. Keys can be removed automatically once they are at least 24 hours old, and a key reused after its original is removed starts a new request.
A payment call that timed out is unknown, not failed: the charge may exist. Stripe’s advanced error handling guide says a client in that state should “retry such requests with the same idempotency keys and the same parameters until they’re able to receive a result from the server”, and it advises against a new key “because the original key may have produced side effects.”
Background jobs need the key designed properly, which is its own topic: an idempotency key for safe retries. For events arriving from a provider, the matching question is how to make a webhook handler idempotent.
Retry with exponential backoff and jitter in JavaScript and Python
Check the SDK before you write a loop. Stripe’s guide says “Most client libraries can generate idempotency keys and retry requests automatically, but need to be configured to do so”, and its Node example is stripe.setMaxNetworkRetries(2). The stripe-node README adds that since v13 the library makes one retry by default for failed requests that are safe to retry. OpenAI’s client libraries retry some failures on their own, with an option to change the count: maxRetries in Node, max_retries in Python. A hand-written loop on top of either is the multiplying row from the failure table, so keep one and set the other to zero.
If no SDK covers the call, this JavaScript helper uses full jitter, the retry column from the status table, and Retry-After when the server sends it in seconds:
// fn throws an error with the response's status and headers (no status = network error)
const RETRYABLE = new Set([408, 429, 500, 502, 503, 504]);
const shouldRetry = (err) => !err.status || RETRYABLE.has(err.status);
const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
export async function retry(fn, { base = 1000, cap = 10000, attempts = 3 } = {}) {
for (let attempt = 0; ; attempt++) {
try { return await fn(); } catch (err) {
if (attempt + 1 >= attempts || !shouldRetry(err)) throw err;
const retryAfter = Number(err.headers?.get?.('retry-after')); // seconds form only
await sleep(retryAfter > 0 ? retryAfter * 1000 : Math.random() * Math.min(cap, base * 2 ** attempt));
}
}
}
In Python, tenacity’s documentation gives the pieces: wait_random_exponential is a “Random wait with exponentially widening window”, stop_after_attempt stops after a set number of attempts, and retry_if_exception takes a predicate.
from tenacity import retry, retry_if_exception, stop_after_attempt, wait_random_exponential
RETRYABLE = {408, 429, 500, 502, 503, 504}
class CallFailed(Exception):
def __init__(self, status=None): # status None means a network error, no response
self.status = status
def transient(exc):
return isinstance(exc, CallFailed) and (exc.status is None or exc.status in RETRYABLE)
@retry(retry=retry_if_exception(transient), wait=wait_random_exponential(multiplier=1, max=10),
stop=stop_after_attempt(3), reraise=True)
def call_provider():
... # make the request; raise CallFailed(status) when it fails
With 3 attempts, the limit on 500s and the overall limit are the same; on a background job’s longer budget, my working rule is to count 500s separately and stop them sooner. The Python block leaves Retry-After out to stay short, so add it before you point it at a provider that sends the header. Both blocks are written from the libraries’ documentation and the formula above; neither has been run against a provider.
Starting numbers, as my working rule: for a request a user is waiting on, a base of about half a second to a second, a cap around 8 to 10 seconds and about 3 attempts, because the user’s patience is the real limit. For a background job, my working rule is a base around 2 seconds, a cap of a minute or more and about 5 to 8 attempts. Treat these as starting points to tune against each provider’s limits, not as a standard. Microsoft’s guidance points the same way for interactive work: “the interval should be short and only a few retries should be attempted.”
OpenAI 429: a rate limit you wait out, or a billing limit you cannot
An OpenAI 429 comes in two kinds, and the error tells you which. A rate limit 429 clears when you slow down, so back off and follow the Retry-After header. A billing 429, such as an exhausted credit balance or a reached spend limit, no retry fixes, so fail at once and alert the account owner.
OpenAI’s error codes list six 429 entries. The table gives each with its cause in OpenAI’s words; the last column is my reading of the fix OpenAI gives.
| OpenAI error and code | Cause, in OpenAI’s words | Wait or stop (my reading) |
|---|---|---|
| Rate limit reached for requests (no code printed) | “You are sending requests too quickly.” | Wait: pace requests and follow Retry-After when present |
Slow down, slow_down (type rate_limit_error) | “Your request rate increased too quickly.” | Wait, then raise the rate gradually |
Credit balance exhausted, credit_balance_exhausted | ”Your organization has no prepaid credits remaining.” | Stop: credits must be added |
Organization spend limit reached, organization_spend_limit_exceeded | ”Your organization reached its enforced spend limit.” | Stop: the limit must be raised or removed |
Project spend limit reached, project_spend_limit_exceeded | ”Your project reached its enforced spend limit.” | Stop: the project limit must be raised or removed |
Organization usage limit reached, organization_usage_limit_exceeded | ”Your organization reached its OpenAI-assigned usage limit.” | Stop: a higher limit must be requested |
OpenAI is plain about the stop rows: “Retrying billing, spend, or quota errors won’t restore API access.” For the two spend limits, access otherwise “resumes after the monthly limit resets.” If your logs show insufficient_quota, the page explains why that word is not enough: inspect error.code for billing errors, because “The broader error.type can still be insufficient_quota.” A forum thread titled “Error Code 429, but there is money” is one more reason to read the code before blaming the balance. If the very first request returns a 429, read the code before tuning any backoff, since a billing 429 fails the first call as readily as the thousandth. In Python, OpenAI’s library “raises RateLimitError for 429 responses”, so catch that class and branch on the code.
For the rate-limit kind, OpenAI’s rate-limit guide describes Retry-After as “The minimum number of seconds to wait before retrying a temporary rate-limit error, when present,” and warns that “unsuccessful requests contribute to your per-minute limit, so continuously resending a request won’t work.”
Three neighbors look similar and are not this. A call that never answers at all is a timeout, not a 429, and why an OpenAI API call hangs, then times out has a different fix. Supabase’s built-in email service is another limit that no backoff schedule waits out in useful time; the answer to the two Supabase email rate limits is custom SMTP. The 429 your own API sends to its callers is the other side of the same header: rate limiting in API endpoints.
One layer retries, and it stops: budgets, breakers and the rest of the stack
My working rule: pick one layer to own retries for each call, usually the server code closest to the provider, and set every other layer to zero. The Builders’ Library states it as a rule for cheap calls: “our best practice is to retry at a single point in the stack.” Microsoft’s guidance shows the cost at just two layers: “If you implement retry with a count of three on both calls, there are nine retry attempts in total against the service.”
Every attempt also needs its own deadline, so knowing how to set a timeout on fetch comes before any retry loop; otherwise one hung attempt eats the whole budget. Work slower than the user will wait belongs in a queue, and once you know how to run long tasks in the background, the queue’s own retry settings replace the loop.
A retry budget caps retries across all calls rather than per call. The Builders’ Library limits retries locally with a token bucket, which “allows all calls to retry as long as there are tokens, and then retry at a fixed rate when the tokens are exhausted.” A circuit breaker goes further and stops calling a provider that keeps failing for a cool-off period. Its states, and the fallback the user sees while it is open, belong with what is graceful degradation.
Report the final failure once, with the attempt count, to your error tracker, which is little help if its stack traces are minified and unreadable. The screen a user sees after the last attempt fails is a separate job that starts with how to add a frontend error boundary.
Providers run the same pattern toward you. In live mode Stripe retries a failed webhook delivery for up to 3 days with exponential backoff; in a sandbox it retries three times over a few hours. Your handler must therefore expect repeats, and a Stripe webhook that returns 200 and updates nothing is the failure waiting on that side of the delivery.
How to verify it
Retries are verified with 6 checks in staging: a failure that clears succeeds within the limit, a failure that persists stops at the limit, a 400 is never retried, Retry-After is honored, a lost response creates no duplicate, and one click never multiplies across layers.
To test retry limits on transient failures, run the checks against your own app in staging, with the dependency replaced by a stub or a local mock server you control, so the stub’s request log becomes your evidence. Microsoft’s guidance suggests the same kind of stand-in: “Create a mockup of the resource or service that returns a range of errors that the real service might return.”
- 01 Transient case: make the stub fail twice and then succeed, with an attempt limit of 3 or more. Expect one successful result and exactly 3 calls in the stub's log, each gap no longer than that retry's capped delay plus the time the failed call took. With full jitter a later gap can be shorter than an earlier one, so run it once with jitter off to watch the gaps grow. Evidence: the log lines with timestamps.
- 02 Persistent case: make the stub fail every time. Expect exactly the configured number of attempts, then a clean failure state for the user and one error report. Evidence: the attempt count and the error event.
- 03 Permanent error: return a 400 or a 401. Expect exactly one call. Evidence: one line in the stub's log.
- 04 Retry-After: return a 429 with the header Retry-After: 5. Expect the next call no sooner than 5 seconds later. Evidence: the two timestamps.
- 05 Side effects: for each write path that retries, make the stub record the write and then drop the reply, so the first attempt succeeded and only the response was lost. Expect one charge or one row afterwards, and one email where the email provider takes an idempotency key; a payment retry must carry the same idempotency key, and a database insert run twice with the same key values must leave one row. Evidence: the stub's record or the row count, before and after.
- 06 Layers: during the persistent case, count calls at the stub for one user click, leaving every SDK's built-in retries at their defaults unless your code sets them. The count must equal the owning layer's limit, not a product of several layers. Evidence: the call count.
Keep the settings for each call (base, cap, attempts), the six results and the date together as the record. A full provider-outage rehearsal, with the messages customers would see, goes a step further, and that is what chaos testing is for. The sprint’s check for deliverable 6.5 follows the same idea: we simulate transient and persistent failures and verify retry limits and side effects.
Where the sprint does this
Deliverable 6.5, Retries with backoff, adds bounded retries with backoff to transient failures where repeating the operation is safe. It sits in area 06, Error handling & reliability, which holds 10 of the 123 deliverables in the published sprint scope. Two neighbors matter here: deliverable 6.4 sets deliberate timeouts for AI, email, payment and other external calls, and deliverable 6.7 makes jobs idempotent and provides a dead-letter or failed-job path with a recovery procedure. The production readiness report, deliverable 13.1, accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. Hosting, paid tools, and API usage remain in your accounts.
Common questions about retries and backoff
What is the typical retry time for exponential backoff?
There is no standard one. My starting point is a base of about a second, doubling up to a cap near 10 seconds when a user is waiting and about a minute for a background job, with about 3 attempts for the first case and 5 to 8 for the second.
How long does a 429 error last?
A rate-limit 429 lasts until the limit’s window resets, and the server can say how long in a Retry-After header, either in seconds or as a date. On OpenAI, the x-ratelimit-reset-requests header gives “The time until the rate limit (based on requests) resets to its initial state.” A billing 429 on OpenAI lasts until credits are added or the limit is raised, or, for a spend limit, until the monthly limit resets.
Is error 429 my fault?
Yes, in the sense the standard defines: a 429 means your client “has sent too many requests in a given amount of time” for the limit that applies to it. OpenAI’s list of causes includes a loop or script making frequent or concurrent requests, an API key shared with other users or applications, a free plan with a low rate limit, and a project limit. It is never fixed by retrying faster.
What is a retry pattern?
A retry pattern is a design in which a caller repeats a failed operation a limited number of times, with a growing wait between attempts, and only for errors that are likely to clear. “Retry with backoff” names the same idea, with exponential backoff and jitter as its usual form, and AWS’s pattern page adds that “Operations should be idempotent when you use the retry with backoff pattern.”
How do you fix jitter?
That question usually means network jitter, the variation in delay that makes calls and video stutter, and it is fixed on the network, not in your code. In retries, jitter means something else: a random delay you add on purpose, which is a feature and not a fault.
If you have a working app built with these tools and need it ready for real customers, this is what we do.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase