Reliability and correctness came last of the 12 areas I rate, averaging 31.4 out of 100 over 21 audited apps. What is graceful degradation for an app built on other companies’ APIs? My answer is one rule: an outage at one provider must not take down the pages that never needed it.
What is graceful degradation in an app that depends on other companies’ APIs
Graceful degradation is a design rule for apps that depend on outside services: when one provider fails, every feature that does not need it keeps working, the feature that does shows a fallback or a plain message, and nothing waits forever. The app gets smaller for a while instead of going dark.
Those 21 are third-party apps I audited in June and July 2026, a selected set I chose, not a random sample, so the average is not a rate for AI-built apps in general. Degradation is one part of hardening SaaS applications for resilience, the part that decides what a customer sees on the day a provider has a bad hour.
The graceful degradation pattern has three parts, in my reading. Isolate the dependency, with a timeout and a circuit breaker, so a slow provider cannot hold every request. Substitute something for the missing answer: a stored value, a cached answer, a queue, or a second provider. Tell the user, with a message on that one feature and nowhere else.
| Service that is down | Without degradation | With degradation |
|---|---|---|
| The model provider | The AI button spins, requests pile up, and pages that never call the model slow down with it | The AI panel says drafting is paused; the dashboard, history and billing work as normal |
| The payment provider | Every signed-in page checks billing live, so customers who owe nothing are locked out | Access comes from the plan stored in your own database; only checkout and the billing portal show a pause message |
| The email API | Signup sends the welcome email inside the request, fails, and no account is created | The account is created and the email waits in a queue until the provider answers again |
A degraded mode is a state with a name in the code, such as aiPaused, that someone chose and tested. Nobody discovers it on the day of the outage.
Graceful degradation vs fault tolerance, and vs progressive enhancement
Fault tolerance keeps full service through a failure by having a spare: a replica, a second region, a standby. Graceful degradation accepts less service on purpose and keeps the most important parts running. That is my reading of the two terms. A small app usually gets its fault tolerance from its hosting platform and its database provider, and has to build its degradation itself, because only the app knows which features matter.
Web designers use the same phrase for something else. MDN’s glossary entry defines it as a design philosophy for a site that “will work in the newest browsers, but falls back to an experience that while not as good still delivers essential content and functionality in older browsers”, and says progressive enhancement is related but is “often seen as going in the opposite direction to graceful degradation”. That browser sense is not the subject here.
What goes wrong without it
An entire app down from one API outage is the result of a design choice nobody made on purpose. The four patterns below are my reading of how it happens, not findings about your app. Check your own code for each one.
| What the user saw | The cause in the code | The fix on this page |
|---|---|---|
| Every page errors while one outside API is down | A root layout or middleware awaits that provider on every request: a subscription check, a feature-flag fetch, an analytics call | Move the call off the shared path |
| A payment provider outage blocked every page, and signed-in customers who owe nothing could not open the dashboard | The billing status is fetched live from the payment provider on each load instead of read from the app’s own database | Read access from your own database, kept current by webhooks |
| The AI feature hangs, then the whole server slows | The model call has no timeout, so a slow provider holds open connections until slow becomes down | A timeout, then a circuit breaker |
| Signup fails whenever the email API fails, and the account is never created | The email send sits inside the signup request, before the account is saved | An outbox table and a background job |
Reliability was the weakest area in those same audits. The third row is easy to miss in testing, because a slow provider looks fine on the day you test it. Every outside call needs a deadline first, and how to set a timeout on fetch covers that part.
These fixes are design work for a calm week. If a feature is broken right now and customers are paying, the questions of whether to switch the broken part off and what to tell them belong to my app is broken and I have paying customers. A traffic spike that trips a provider’s limits is a different first day again, and what to do when your app goes viral starts there.
How to add a fallback for a dependency
A fallback for a dependency is chosen from 5 options: a stored value, a cached answer, a queue that runs the work later, a second provider, or a paused feature with a message. First move the call off the shared path so only the journeys that need it can fail.
My working method has four steps, in this order.
- 01 List every outside service the server or the browser calls. Search the code for SDK imports and for fetch( calls to other domains.
- 02 For each service, mark the customer journeys that truly cannot run without it, and the ones that must not care.
- 03 Move the call off the shared path: out of the root layout, the middleware and the auth callback, so only the journeys that need it ever touch it.
- 04 Choose a fallback for each service from the menu: a stored value, a cached answer, a queue, a second provider, or a paused feature with a message.
The result is a dependency map. Here is a filled example for a small AI writing tool, with its assumptions stated: a Next.js app on a serverless host, Supabase for sign-in and the database, Stripe subscriptions, a transactional email API, one model API and an analytics script. Your services and journeys will differ; the shape of the table will not.
| Service | Journeys that need it | Journeys that must not care | Fallback | What the user sees |
|---|---|---|---|---|
| Model API | Generate a draft | Sign in, dashboard, saved drafts, billing, settings | Pause the feature, or queue the request | A panel that says drafting is paused and saved drafts are safe |
| Stripe | Checkout, the billing portal | Everything a paying customer already has | Read the plan from your own database | A pause message on the billing page only |
| Email API | Sending the verification email and receipts | Creating the account, everything else | Write the message to an outbox table and send it from a job | Nothing, or a note that the email may take a few minutes |
| Supabase Auth | New sign-ins, password resets | Pages for people already signed in, where the token is verified locally | None for new sign-ins | A plain “sign-in is unavailable right now” on the login page |
| Database | Every journey | None | None: a static maintenance page served by the host | The maintenance page and a status update |
| Analytics script | None | All of them | Load it async and never await it | Nothing |
The database row is there on purpose. Some dependencies cannot degrade, and the map should say so.
A fallback when a model provider is down
A model provider fallback has 3 levels: pause the AI feature and keep the rest of the app working, queue the request and deliver the result later, or fail over to a second model. Start with the first; it needs no second contract.
That is my working order, cheapest first. The paused feature costs a message and a flag. Queuing costs a job runner and a way to tell the user the result is ready, which is the same machinery as how to run long tasks in the background. A failover to a second model or provider behind one interface costs the most: the output will differ, there is a second bill, and my working rule is that the second provider gets the same spending cap as the first. A second provider does not make the feature immune either, because both can fail on the same day.
Decide what counts as “down” from the provider’s own error list, not from any error at all. In OpenAI’s error codes, a 500 means “Issue on our servers” and a 503 means “The requested model is temporarily overloaded”. In Anthropic’s API errors, a 500 is api_error and a 529 is overloaded_error: “The API is temporarily overloaded.” Those, plus timeouts and connection errors, point at the provider.
Other errors need a closer read. OpenAI uses 429 for rate limits and for spend limits, and its docs say “Retrying billing, spend, or quota errors won’t restore API access.” Anthropic returns a 400 for a malformed request and also “when usage reaches an organization or workspace spend limit you set” (a Claude Code workspace limit can return a 429 instead). A spend limit is a good reason to show the paused message too: the feature cannot run until the limit changes, and retrying only burns time.
In one app I audited, a spam classifier hard-coded a dated model identifier, so the day the provider retired that model the product would stop working, with one generic error and no fallback. It sits with the other failure cases in hardening SaaS applications for resilience. The lesson I take from it: a model the product cannot live without is a dependency like any other, and the plain-message fallback is the one you can ship this week.
The model call’s own deadline, streaming the answer, and answering straight away while the result arrives later are covered in why an OpenAI API timeout hangs your app’s AI feature. This section only decides what the rest of the app does when that call fails.
When the payment provider is down
If a payment provider outage blocked every page, the app was asking the provider a question its own database could answer. Existing customers keep their access when the plan they paid for is stored in the app’s own database, kept current by the payment provider’s webhooks, and every page reads it from there. Only checkout and the billing portal need the provider live, so only those two pause, with a message that says billing changes are paused and that the customer’s choice is saved for later.
Never grant paid access on a payment that has not been confirmed and call it a fallback. A stored plan is a fallback; a guessed one is a free account.
Webhooks have their own recovery path. In live mode Stripe retries a failed webhook delivery for up to 3 days with exponential backoff; in a sandbox it retries three times over a few hours. That covers a delivery your endpoint failed to accept, for example during your own outage. Stripe’s docs also say an endpoint “might occasionally receive the same event more than once”, so the handler has to be safe to run twice on the same event. How to make a webhook handler idempotent is that job. Moving payments from test mode to live mode is a separate checklist, payment go-live, and stays out of this section.
Auth, the database and email: what can and cannot degrade
Some services cannot degrade much, and saying so is part of the design. If the auth provider is down, new sign-ins fail; there is no honest fallback for proving who someone is.
People already signed in can keep working for a while if the server checks their token without calling the auth provider. Supabase’s server-side auth guide says getClaims() verifies “locally against a cached copy of the project’s public keys” on projects with asymmetric signing keys, “the default for new projects”, while projects “still using a symmetric secret” call the Auth server instead. The same page says that when the access token is close to expiring, getClaims refreshes the session before it verifies. In my reading, a session survives an auth outage only until its token needs that refresh.
If the database is down, the app is down. The degraded mode there is a maintenance page your host can serve without the database, plus an update on your status page, not a fallback in the code.
Email can degrade almost completely. Write the message to an outbox table inside the same transaction that creates the account, and send it from a job, so signup never depends on the email API being up. The job then retries safely with an idempotency key for safe retries. Analytics, chat widgets and other third-party scripts load async and are never awaited by anything a customer is waiting on.
Circuit breaker pattern example for third-party APIs
A circuit breaker wraps calls to one provider and has 3 states: closed lets calls through and counts failures, open fails them at once and runs the fallback, half-open tries a call to see if the provider is back. It turns a slow outage into a fast, handled one.
In Martin Fowler’s description of the circuit breaker, once failures reach a threshold the breaker trips, and “all further calls to the circuit breaker return with an error, without the protected call being made at all”. In the half open state, “the circuit is ready to make a real call as trial to see if the problem is fixed”. Without a breaker, every request still waits the full timeout on a dead provider; with one, the fallback is immediate. That is the whole gain.
| State | What happens to calls | What moves it on |
|---|---|---|
| Closed | Calls reach the provider; failures are counted | The error percentage passes errorThresholdPercentage (default 50), and it opens |
| Open | Calls fail at once without reaching the provider; the fallback runs if you set one | resetTimeout passes (default 30000 ms), and it goes half open |
| Half open | The next call is a trial | Success closes it; a failure or a timeout opens it again |
Here is a circuit breaker pattern for third-party APIs in Node, using opossum and the option names its README documents. The values are the README’s example values, not recommendations; choose your own per provider.
const CircuitBreaker = require('opossum');
// Example values from opossum's README. Set your own per provider.
const breaker = new CircuitBreaker(callModel, {
timeout: 3000,
errorThresholdPercentage: 50,
resetTimeout: 30000,
});
breaker.fallback(() => ({ paused: true }));
breaker.on('open', () => console.warn('model provider circuit open'));
breaker.fire(prompt).then(showAnswer).catch(showError);
callModel is your own function around the SDK call. When the fallback runs, fire() resolves with the fallback’s return value, so showAnswer receives { paused: true } and draws the paused panel. opossum’s errorFilter option keeps an error you choose, such as a malformed-request 400, out of the failure count; in the current source, a filtered error skips the fallback and goes to catch.
On a serverless host, breaker state kept in memory belongs to one short-lived instance, in my reading, so each new instance starts closed. opossum’s README shows toJSON() for reading a breaker’s state and a state option for starting a new breaker from it, naming serverless platforms such as AWS Lambda as the use case; where you store that state between invocations is up to your code. The cruder alternative is a manual switch that pauses the feature, the job of open source feature flags.
Where the breaker goes relative to retries depends on who does the retrying. Azure’s circuit breaker pattern combines the two by “using the Retry pattern to invoke an operation through a circuit breaker”, with retry logic that stops “if the circuit breaker indicates that a fault isn’t transient”. A retry loop you wrote therefore calls breaker.fire() on each attempt and stops once breaker.opened is true; with a fallback set, an open circuit answers with the fallback at once, so there is nothing left to retry.
An SDK’s built-in retries run inside the wrapped call: Anthropic’s SDK retries transient failures “twice by default”, and OpenAI’s Node SDK retries certain errors “2 times by default”. With the breaker’s timeout at 3 seconds, it fails the call and runs the fallback even if the SDK is still retrying underneath; opossum’s README shows an AbortController option for cancelling that request when its timeout is reached. How retries wait between attempts is a separate question: what is exponential backoff. AWS’s circuit breaker pattern states the intent in one line: the pattern “can prevent a caller service from retrying a call to another service (callee) when the call has previously caused repeated timeouts or failures”.
What the user sees, and what the health check says
The paused feature shows a message in its own panel, and my working rule is that the message says four things: what is unavailable, that the rest of the app works, whether the user’s work is saved or queued, and when to try again, with the button disabled rather than spinning. This is the designed paused state that ships with the feature. Telling customers about a live incident is a different job, done during the incident itself.
A crashed component is a different case again. An exception thrown while rendering belongs to how to add a frontend error boundary; a paused feature is a normal state the app chose.
The health check needs the same thinking. If an optional provider’s outage makes the health check fail, a platform that restarts or removes failing instances can take a working app away from everyone. Let the endpoint report “degraded” and still count as healthy. The status design belongs to what a health check endpoint is, and container checks to a Docker Compose health check.
Log the breaker opening as an event, so an alert can fire. Fowler’s text says “Any change in breaker state should be logged”. A handled fallback reaches your error tracker only if your code sends it there.
How to verify it
Graceful degradation is verified in staging in 5 steps: disable one integration the way it really fails, check the fallback appears within the timeout, walk the journeys that should not care, check the health status and error volume, then re-enable and confirm recovery without a redeploy.
To test the app with an integration disabled, run it in staging, never against live customers, and keep the evidence each step names.
- 01 Switch off one integration so it fails like a real outage: point its base URL at an address that refuses connections, or at a stub that accepts the connection and never answers. Do not delete the feature or its code. Evidence: which method you used, and the date.
- 02 Use the feature that needs it. The fallback or the paused message appears within the wait you set, counting any retries the SDK makes inside that wait, and nothing spins forever. Evidence: a screenshot of the degraded feature.
- 03 Walk every journey the dependency map marks as must not care: sign in, open the dashboard, read existing data, change a setting. All of them pass. Evidence: the map with pass or fail in each cell.
- 04 If the app has a health endpoint with a degraded state, it reports degraded; if it has none, record that as a gap. If the code sends the fallback or the breaker opening to the error tracker, the tracker shows one grouped event, not a flood; if it sends nothing, record not reported as a finding for the logging work. Nothing was written twice. Evidence: row counts of the tables the journeys wrote to, before and after.
- 05 Re-enable the integration and confirm recovery without a redeploy: the breaker closes after its trial call, and queued work drains. Evidence: the empty queue and the working feature, with the time.
Repeat it for each integration in the map. A revoked API key is a different test: OpenAI documents a bad key as a 401, “The requesting API key is not correct”, which tests your configuration-error path, not an outage. The full drill, with several failure scenarios, customer messages and recovery timing, is what chaos testing is.
In the Production Hardening Sprint, deliverable 6.8 is verified this way: disable an integration and check its fallback plus unaffected customer journeys.
Where the sprint does this
Deliverable 6.8 of the Production Hardening Sprint keeps unaffected functions usable when an external service is unavailable, checked as the last paragraph of the verify section says. The result goes into the production readiness report, deliverable 13.1, which accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. New features that change the product’s core capabilities are separate work; the sprint includes only the supporting interfaces the listed controls need, such as session management, account deletion and billing self-service. Hosting, paid tools, and API usage remain in your accounts. Every deliverable is listed in the sprint’s published scope.
Common questions about degraded mode and fallbacks
What is an example of graceful degradation?
A model provider goes down, and the AI panel says drafting is paused while sign-in, the dashboard, saved drafts and billing keep working. The same pattern keeps paying customers inside the app during a payment provider outage, because their plan is read from the app’s own database and only checkout pauses.
What does graceful shutdown mean?
Graceful shutdown means a server stops taking new requests and finishes the ones in progress before the process exits, after it receives a stop signal such as SIGTERM. In Node on non-Windows platforms, SIGTERM and SIGINT have default handlers that exit the process, and installing a listener removes that default, so your code can call server.close(), which “Stops the server from accepting new connections” and closes the connections that are not sending a request or waiting for a response. Node lists these under Node’s signal events. Shutdown is about stopping cleanly; degradation is about staying up with less.
What is a fallback model?
A fallback model is the second model an AI feature calls when the first one is unavailable, from the same provider or a different one. Its answers will not match the first model’s, so test the feature’s output on it before you need it, and give it a spending cap like the one described in the model provider section above.
What are fallback values?
Fallback values are stored defaults a feature returns when it cannot fetch the live value, such as the last known plan from your own database or a cached copy of a list. A fallback value is never a default that grants access, a discount or a paid feature the user has not been confirmed for.
What happens if the API gateway is down?
If the API gateway is down, every route behind it is unreachable, so the app is down for every journey that passes through it. In my reading, a gateway belongs with the database: a dependency that cannot degrade. The answers there are the redundancy your platform provides and a static maintenance page served from somewhere else, plus a status update.
If you have a working app built with these tools and need it ready for real customers, this is what we do.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase