At two in the morning the AI provider your app calls starts timing out, and every request waits with it. In 17 of the 21 third-party apps I audited in June and July 2026, nothing would have recorded those errors: no error tracking, no alerting. Hardening your SaaS applications for resilience takes ten checks that make a failure visible, bounded and recoverable.

What hardening SaaS applications resilience covers: the ten checks

Hardening SaaS applications for resilience comes down to ten controls: honest exception handling, a global backend handler with frontend error boundaries, error tracking that reaches a person, a timeout on every external call, bounded retries with backoff, background jobs for long work, jobs that run once, graceful degradation, a truthful health endpoint and a staging outage drill.

The 21 apps behind the opener’s number were a selected set that I audited, not a random sample, so 17 of 21 is not a rate for AI-built apps in general. Error handling and reliability is one of the 13 areas of production hardening, one of those that answer how the app behaves when a provider fails or the load climbs, and it holds 10 of the 123 checks. For a small team, SaaS reliability is mostly decided in the code paths those ten cover.

At enterprise scale, application resiliency starts with redundancy and load balancing. My reading for an app on managed hosting is that this layer is the host’s to run, and how much of it a given plan includes, a second instance for example, is in that host’s own docs. This page is about the app’s layer, the software resilience work that stays with you whichever host you pick.

ControlWhat it isWhy it mattersNext topic
Honest exception handlingEvery failure is caught on purpose and reportedSilent failures can make an unsuccessful operation look completeAPI error handling best practices
Backend handler and frontend boundariesOne handler for server errors, a boundary around each risky part of the screenA localized failure should not leave the customer with an unexplained blank pagehow to add a frontend error boundary
Error trackingErrors land in a tracker, readable and labeled, with an alertCustomer complaints should not be the only way you discover application faultsstack traces are minified and unreadable
External-call timeoutsA deadline on each call to an AI, email or payment providerSlow providers can leave requests and workers waiting indefinitelyhow to set a timeout on fetch
Retries with backoffA failed call repeated a limited number of times, with growing gapsTemporary failures need recovery without creating duplicate actions or retry stormswhat is exponential backoff
Background jobsSlow work leaves the request and runs as a job with a statusLengthy work can exceed request limits or be interrupted when a browser disconnectshow to run long tasks in the background
Jobs that run onceA job gives one result however often it runs, and failed jobs are keptRetries can duplicate actions, while discarded failures leave work unfinishedidempotency key for safe retries
Graceful degradationThe rest of the app keeps working while one provider is downOne provider outage should not unnecessarily block the entire applicationwhat is graceful degradation
Health endpointOne URL that says whether the app can serve requestsMonitoring needs a reliable signal for whether the application can serve requestswhat is a health check endpoint
Staging outage drillA provider failure rehearsed in staging, with the results written downFallback code needs to be exercised before customers depend on itwhat is chaos testing

What is a resilient application? The four failures a five-person SaaS has to survive

A resilient application keeps serving the customers it can while one part is failing, and tells you it is failing. For a five-person SaaS that means surviving four failures: a dependency down, a job stuck, a bad deploy and wrong data. Error handling contains the first two; deploys and data have their own checks.

Resilience in software is that property written into the code, call by call, rather than bought from the host. The mapping below is mine: what the customer sees in each failure, and which part of the work holds it.

FailureWhat the customer seesWhere it is contained
A dependency is down (an AI, email or payment provider)A spinner that never ends, or an errorTimeouts, retries and graceful degradation, rehearsed by the staging outage drill
A job is stuckAn export that never arrivesBackground jobs and jobs that run once
A bad deployThe app broke after a releaseReleases and rollback, outside error handling: DevOps for startups
Wrong dataA half-written orderData checks, outside error handling: data consistency checklist for SaaS
Any of the four, unrecordedNothing, until a customer writes inException handling, error boundaries and error tracking, in every case, or nobody learns it happened

A website bug, to a customer, is whichever of these four reaches their screen; what it means when your app has a bug is a separate question with its own answer. The wider frame, readiness as a whole, is what makes code production ready; this page stays with failure.

What goes wrong without it

Each symptom below is the one a founder notices first, followed by the control that was missing and the topic to read next.

Something broke and nothing recorded it

A customer writes to say checkout failed yesterday, and there is no trace of it: no event, no alert, no log line you can find. That was where most apps stood in the same audits behind the opener. With no tracker, a complaint is the only error report you get, and it arrives after the moment has passed. If a tracker does catch the error but the stack traces are minified and unreadable, it is only half set up. If it is happening right now, the first-hour steps are in my app is down, what do I do first.

The AI provider hangs and every request waits with it

Pages that call the model spin, then the whole app slows, because every request that touched the provider is still waiting and nothing told it to stop. A slow provider can hold requests and workers open with no end point unless the code sets one. The fix is a deadline on every outside call, and for plain HTTP requests that starts with how to set a timeout on fetch. The model call has clocks of its own, which is the subject of why an OpenAI call hangs, then fails.

The AI model the app was built on is retired

A feature that has always worked starts returning the same error to everyone, and nobody changed the code. In one app I audited, a spam classifier hard-coded a dated model identifier, so the day the provider retired that model the product would stop working, with one generic error and no fallback. The lesson I take from it: a model id is a dependency with an end date, so where it is set, and what the feature shows when the call fails, are decisions to make before the provider makes them. One provider’s change should not stop more of the product than it has to, and a fallback nobody has run is a guess. Two questions follow for a small app: what is graceful degradation for a single model call, and what is chaos testing when staging is the only place to try it.

A retry charges the customer twice

A payment call times out, the code tries again, and the customer can end up with two charges for one order. Recovery from a brief failure has to happen without duplicate actions and without a flood of repeated calls. It also has to deal with the failures retries give up on: a job dropped after its last attempt leaves the work half done. For the spacing between attempts, the question is what is exponential backoff; for making the second attempt harmless, the answer is an idempotency key for safe retries.

The export dies at the request timeout

A customer clicks export, watches a spinner, and the file never arrives. Long work that runs inside the request can run past the limit on how long a request may live, or stop when the customer closes the tab. Each host sets that limit its own way, and the Vercel function timeout is one of them. The fix is the background-jobs control below: the work moves into a job, and the customer sees its status while it runs; work that belongs on a schedule instead of a click can run as a cron every few minutes.

A blank white screen with no explanation

One widget throws, and the customer gets a blank white page where the app used to be, with no message and no way back. A failure in one part of the screen should not take the whole screen with it and leave the customer guessing. The frontend half of the fix is knowing how to add a frontend error boundary, so the broken panel shows a message while the rest stays usable. The backend half is a global handler that answers with a safe error instead of a stack trace.

The app is live, something broke after an AI change, and customers are on it

When the app is live and something broke an hour ago, the first job is recovery, and two articles already cover it: my app is broken and I have paying customers for the customers, and what actually changed when an AI edit broke the app for the code. Worry about AI-generated code causing outages is fair, and the controls that contain a bad change are the same whoever or whatever wrote it. What this page adds is the setup for next time, and three topics cover the aftermath: an incident response plan template written before the next one, what Sev1 means so everyone rates an outage the same way, and what RCA means in software for the write-up.

The ten controls, one by one

Each control below gives what to build, the test that proves it, and the page with the detail.

1. Every failure is caught, answered honestly and recorded

Remove swallowed failures and empty catch blocks, replacing them with deliberate recovery or useful error reporting. The pattern to hunt is a catch that logs nothing and lets the caller carry on as if the write worked; how to spot that and its cousin, a route that answers success with no data, is when your app fails silently. A Python backend applies the same rule with its own idioms, a separate subject: Python error handling best practices. The test: inject failures and verify accurate responses and diagnostic records. The rules per failure type fall under API error handling best practices, the first row of the table.

2. The backend has a global handler and the frontend has boundaries

Implement global backend error handling and frontend error boundaries with clear recovery states. The recovery state is the part people skip: a short message and a button that retries or goes back, where there would otherwise be a frozen screen. On the server, the global handler is the last net, so an error nobody planned for still returns a safe answer and still gets recorded. The test: trigger server and component errors and confirm safe responses and contained UI failures. For the frontend steps, the topic is how to add a frontend error boundary.

3. Errors are tracked, readable and routed to a person

Install error tracking such as Sentry with protected source maps, environment labels, and alerts. Sentry is an example, not a requirement: any tracker that reads your source maps, tells production from staging and sends an alert does the job. The alert has to reach a channel someone reads, or the tracker is a log nobody opens. The test: send a test error and verify symbolication, environment attribution, and alert delivery. Symbolication means the trace shows your own file and line, and when it does not, the problem is the one the first table names: stack traces are minified and unreadable.

4. Every external call has a timeout

Set deliberate timeouts for AI, email, payment, and other external calls. Deliberate means chosen per call: a model call writing a long answer needs a different deadline from a payment lookup, and neither should inherit a default nobody picked. When the deadline passes, the code has to decide what the customer sees, which is where the other controls take over. The test: simulate a slow dependency and confirm bounded waits and a useful failure state. For plain requests, the mechanics sit under how to set a timeout on fetch.

5. Retries back off, with a limit

Add bounded retries with backoff to transient failures where repeating the operation is safe. Bounded means a maximum number of attempts; transient means a timeout or a rate limit, where a declined card is a different kind of failure; safe means repeating the call cannot charge or send twice. The test: simulate transient and persistent failures and verify retry limits and side effects. The timing between attempts is the question of what is exponential backoff.

6. Long work runs in the background with visible status

Move long-running AI generation, exports, and bulk communication into background jobs. The request starts the job and returns straight away; the page then shows the job as queued, running, done or failed, so a customer who closes the tab can come back to the result. The job also needs somewhere to report its own failure, which is the next control. The test: run a long task beyond the normal request window and verify completion and user-visible status. Queue choices and job patterns belong to running long tasks in the background, the sixth row of the first table.

7. Jobs run once and failures land somewhere recoverable

Make jobs idempotent and provide a dead-letter or failed-job path with a recovery procedure. Idempotent means a job that runs twice still leaves one result. The failed-job path means a job that runs out of retries lands where a person can see it and replay it, with written steps for doing that. The test: replay jobs, exhaust retries, and verify one intended result plus a recoverable failure record. The key that makes a repeat harmless is the idempotency key for safe retries in the first table.

8. One provider down does not take the app down

Keep unaffected functions usable when an external service is unavailable. In practice, if the AI provider is down, login, billing and saved work still load, and the AI feature says it is unavailable instead of spinning forever. The same goes for email: a signup should still succeed while the welcome message waits in a queue. The test: disable an integration and check its fallback plus unaffected customer journeys. Patterns for each kind of dependency come under what is graceful degradation.

9. A health endpoint tells the truth without telling secrets

Provide a safe health endpoint reporting the readiness of required services without exposing secrets. Readiness means the endpoint checks what the app needs to serve a request, such as the database, and reports degraded when one of those is down. Safe means no connection strings or keys in a public answer. What the endpoint should return is its own question, what is a health check endpoint; for apps that run in containers, the matching piece is a Docker Compose health check. The test: check healthy and degraded responses and confirm sensitive details are not public.

10. An outage has been simulated in staging

Simulate AI or payment-provider failure in staging and verify the expected recovery behavior. Staging, so no customer pays for the lesson, and written down, so the result outlives the afternoon: which provider was cut off, what queued, what the customer saw, and how the app came back. Run it again after any change of provider or queue. The test: record outage scenarios, queue behavior, customer messages, and recovery results. How far to take it at this size is the subject of what is chaos testing.

Who runs this when there is no SRE

A small team without a site reliability engineer keeps the question and drops the job title: how does the app fail, and who finds out? My working rule for a five-person SaaS is ten controls, one alert channel with a named person on it, and a written first step for each alert.

For the role itself at a larger company, the first question is what is a site reliability engineer; at five people you keep a few parts of it, below.

Self-healing code, as I read it for a small app, is mostly controls 4, 5, 7, 8 and 9 doing their work: a call that gives up on time, a retry that stops, a job that can be replayed, a feature that steps aside and an endpoint that reports the truth. Self-healing software adds the host restarting a crashed process, where the host does that and the plan allows it: restart settings and their limits differ by host and plan, some entry plans cap them, and the host’s own docs say which. Neither one replaces a person reading the alert.

The other half of my working rule is that the named person changes each week, and the rest of the team knows who it is without asking. The written first step can be short, such as open the provider’s status page, then post in the customer channel; what matters is that nobody woken by an alert has to invent it. The severity words that decide who else gets woken come from what Sev1 means, above.

How to verify the whole area in an afternoon

Verifying this area takes ten tests in staging and, as my working estimate, about an afternoon: a test error that reaches a person, a thrown route, a thrown component, a slow stub, a forced retry, a task past the request window, a replayed job, one integration switched off, a health read and one simulated outage.

The order is mine, for a one-person team: tracking goes first because every later test needs somewhere to land. Each item points back to its control’s test above and says what to keep.

  1. 01 Error tracking (control 3): send a test error from staging. It counts when a person gets the alert and the trace shows your own file and line. Keep the link to the error event.
  2. 02 Exception handling (control 1): make one route fail on purpose. It counts when the response is honest and a diagnostic record exists. Keep the response body and the event link.
  3. 03 Handler and boundaries (control 2): throw inside one component and one route. It counts when the rest of the page stays usable and the server answers safely. Keep a screenshot and the response.
  4. 04 Timeouts (control 4): point one external call at a slow stub. It counts when the wait ends at the deadline you chose, with any retries the provider's client adds on its own inside that wait. Keep the time waited.
  5. 05 Retries (control 5): fail a call briefly, then for good. It counts when retries stop at the limit and nothing happened twice. Keep the attempt count.
  6. 06 Background jobs (control 6): run a task longer than the request window. It counts when it finishes and the customer can see its status throughout. Keep the job id.
  7. 07 Jobs that run once (control 7): replay one job and exhaust another. It counts when there is one result and one failure record you can recover from. Keep both job ids.
  8. 08 Graceful degradation (control 8): switch one integration off and click through everything else. It counts when the other journeys still work. Keep the list of journeys you tried.
  9. 09 Health endpoint (control 9): read it while healthy, then with one required service down. It counts when both answers are right and neither shows a secret. Keep both responses.
  10. 10 Outage drill (control 10): cut one provider off in staging. It counts when the scenario, the queue, the customer message and the recovery are written down. Keep that record.

The timeout item times the whole wait because OpenAI’s Python client retries a request that times out by default, so one slow call can become several attempts; the setting and its default are in the OpenAI timeout article above. These ten tests cover one area, and the full production readiness checklist runs the same kind of test across the rest of the app.

Where the sprint stops

After handover from the Production Hardening Sprint, your team operates the application and receives its alerts. The included cover is 14 calendar days of fixes for defects in the delivered sprint work, plus 30 calendar days of async questions about the handover and architecture. Hosting, paid tools, and API usage remain in your accounts, and we explain any required third-party costs before enabling them. Your app’s current framework and hosting setup are the starting point, and we refactor or replace components where the production work requires it.

Where the sprint does this

In the sprint, area 06, Error handling & reliability, covers these ten controls, 10 of its 123 deliverables. Each one is built as its section above describes and checked with the test written under it. Every result goes into the production readiness report with its verification evidence, and the report accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. Each control’s wording and test sit under area 6 of the published scope.

Common questions about keeping a small app up

What does resilience mean in software?

In software, resilience means the code keeps the parts that still work running while one part fails, then recovers once the failing part is back, without anyone restarting things. I’d define it from the code’s side because that is where a small team can change it: each call, job and screen decides what happens when something under it breaks.

What is resilience vs reliability?

Reliability is how often the app works; resilience is what happens when it does not. That is my framing, and it gives each word its own test: a small SaaS measures reliability with uptime and error rates, and tests resilience with a staging outage drill, the tenth control above.

How do I know if an app has a bug?

You know from the app’s own signals, before a customer tells you: error tracking that alerts a person, and a health endpoint that a monitor reads. Complaints will still come; the point is that they are not your only signal. Some failures throw nothing, such as a request that answers success with no data behind it, and which signal catches each of those is the subject of when your app fails silently, above.

Did Amazon have an outage due to AI code?

Amazon says no. After a Financial Times report, Amazon published a statement saying the interruption was “the result of user error”, specifically “misconfigured access controls”, and “not AI as the story claims”. It described “an extremely limited event” affecting a single service, AWS Cost Explorer, in one of its “39 Geographic Regions”. My reading: whatever wrote the change, access limits and a review before production are what contain it, and the safeguards Amazon lists include “mandatory peer review for production access”.