What is chaos testing? It is breaking one dependency on purpose to check that the app does what you designed it to do when that dependency fails. Netflix’s version runs in production, against live servers. A small SaaS needs 3 outage drills in staging instead: the AI provider, the payment provider, and one integration switched off.

What is chaos testing, and what an outage drill is

Chaos testing is injecting one failure on purpose and comparing what the system does with what you expected. It comes from chaos engineering, which Netflix practiced by terminating production instances at random. For a small SaaS the useful form is an outage drill: one named dependency failed in staging, with the expected behavior written first.

The principles of chaos engineering define the discipline as “experimenting on a system in order to build confidence in the system’s capability to withstand turbulent conditions in production.” Netflix’s Chaos Monkey follows those principles, and its repository says it “randomly terminates virtual machine instances and containers that run inside of your production environment.”

Chaos testing takes the same idea and uses it as a test: inject one failure, then compare what happens with what you wrote down beforehand. The purpose, as I read it, is coverage of the code that runs least: fallback code, retry code and error messages almost never execute in normal use, and the only way to know they work is to make them run. An outage drill is chaos testing at the size a small SaaS can afford, a planned failure of one named dependency in staging, and it is one of the checks in hardening SaaS applications for resilience.

The neighboring terms get mixed up. This is how I separate them:

TermWhat it meansThe scale it assumesWhere it lives on this site
Chaos engineeringOngoing experiments on a running system, usually in productionRedundant servers and an on-call teamNot covered: production experiments are out of scope here
Chaos testing, or an outage drillOne dependency failed on purpose, with the result compared against a written expectationA staging copy of the appThis page
Fault injectionForcing a chosen error at a chosen point in the code or the networkAny app with a test suiteThe four methods below
Failover testingProving that a standby takes over, most often a server or databaseA cluster with a standbyHere, only the app’s switch to its fallback, in the AI drill
Disaster recovery drillRehearsing the loss of your own database or hostA backup and a fresh environmentThe disaster recovery checklist for SaaS
Restore drillProving that a backup loads and the data is wholeA backup fileA restore drill for your backups

Those last two rehearse losing your own data. This page rehearses losing someone else’s service. And it stays in staging for a plain reason, as I see it: a small app has no spare servers to absorb a real failure, so an experiment in production is an outage your customers pay for. Every drill below runs in staging with sandbox keys.

What goes wrong without it

An AI coding assistant writes a catch block, a retry and a please-try-again message because the prompt asked for error handling. None of it runs until the provider has a bad afternoon, and then it runs for every user at the same moment. That is fallback code never tested until a real outage, and in my reading it is the default state of any app nobody has drilled.

My audit numbers point the same way. Across the 21 third-party apps I audited in June and July 2026, the Reliability & Correctness pillar averages 31.4 out of 100, ranked worst of the 12 pillars I score. Those 21 apps are a selected set, not a random sample, so the figure is no rate for AI-built apps in general. In the same audits, 17 of the 21 apps had no error tracking or alerting: when a user hits an error, nothing records it. So the first outage is also the first time anyone sees the fallback run.

Five fallbacks often fail on first use. They are patterns, my reading of how this code gets written, not findings about your app:

The fallback as writtenWhat happens the first time it runsWhat a drill shows
A retry with no limit and no backoffEach failed call repeats at once, adding load to a provider that is already strugglingHow many calls one user request turns into
A provider call with no timeoutEvery request waits on the slow provider, and the workers fill upHow long a request hangs before anything gives
A payment retry with a fresh keyThe first attempt had in fact succeeded, so the retry charges the card againWhether the retry reuses the first attempt’s idempotency key
A queue that accepts work during the outageJobs are taken in, then dropped when the worker restartsWhether every job submitted during the fault still exists afterward
An error message that says something went wrongThe user gets no next step, and nothing records the errorThe exact text the customer sees, and whether an error event exists

Each fix is a topic of its own: what exponential backoff is, how to set a timeout on fetch, an idempotency key for safe retries and API error handling best practices.

Provider outages happen on a schedule nobody chooses. On April 28, 2026, Anthropic’s status page first reported an issue preventing users from reaching Claude.ai, then identified elevated errors on the Anthropic API, later described as elevated authentication errors for requests to the API and Claude Code. An update at 18:59 UTC put the impact at 17:34 to 18:52 UTC, and the incident was marked resolved at 19:15. The lesson I take from it: in a window like that, an app built on the API runs its fallback for every user at once, and a staging drill is where that fallback should have run first.

None of this says your app has these gaps. The drill is how you find out. If an outage is happening to you right now, the drill can wait: start with what to do first when your app is down.

How to do it: three outage drills in staging

Every drill follows the same loop: pick a way to break the dependency, write the expected behavior, inject, watch, remove the fault, record. The first section below covers the ways to break things, the next two are the payment drill and the AI drill, and the integration drill runs on the kill switch.

Test error handling by injecting failures: four ways to break a dependency

Error handling is tested by injecting failures in 4 ways: breaking a credential or base URL in staging, mocking the HTTP call in automated tests, putting a fault proxy between the app and the provider, or flipping a kill switch in the app. Slow and half-finished responses matter more than a clean outage.

Fault injection means forcing a chosen error at a chosen point. I order the four methods cheapest first, which is my working order, not a ranking of the tools:

MethodHow it is doneWhat it simulates wellCaution
Break the credential or base URLIn staging’s environment, set the provider key to a revoked key, or the base URL to an address nothing answers onA revoked key gets a 401 from OpenAI and from Anthropic; a dead address gives connection errorsTests only the provider-unreachable path, and is easy to leave broken
HTTP mock in automated testsA Mock Service Worker handler returns an error status, a 429 with Retry-After, a body that does not parse, or a network errorEvery status code and malformed reply, the same way every run, in CINever touches the deployed stack
Fault proxyToxiproxy sits on the app’s connection and adds latency, timeouts or resets; Dev Proxy reads a devproxyrc.json file that lists the URLs to watch and points to a file of the errors to returnSlow, hanging and reset connections (Toxiproxy); API errors, rate limits and slow APIs (Dev Proxy)Toxiproxy is a TCP proxy, and its README describes no setup for an HTTPS API (checked 2026-10-04)
Kill switch in the appA feature flag makes the provider client throw before it calls outThe provider gone, or one integration switched offNeeds a small code change

Mock Service Worker is an API mocking library for browser and Node.js, and its HttpResponse.error() stands in for DNS errors, connection timeouts or a client going offline. A test fixture for the rate-limited case and the network error looks like this, with your provider’s endpoints in place of the placeholders:

import { http, HttpResponse } from 'msw'

export const providerDown = [
  http.post('https://api.your-provider.example/v1/generate', () =>
    HttpResponse.json({ error: 'rate limited' }, { status: 429, headers: { 'Retry-After': '2' } })),
  http.post('https://api.your-provider.example/v1/embed', () => HttpResponse.error()),
]

Toxiproxy calls itself “a framework for simulating network conditions,” made for testing, CI and development environments: you route the app’s connection through the proxy and change its health over HTTP. Its toxics include latency, timeout and reset_peer, and disabling a proxy takes the service down. Given the caution in the table, I’d point it at a database, a cache or a local mock server rather than a hosted model API. Microsoft Dev Proxy is the HTTP-level option: an API simulator that intercepts network requests and returns API errors, rate limits and slow responses, without changes to the app’s code.

The kill switch is the method I’d add first to an app that has none, because it doubles as the off switch a degraded mode needs; what is graceful degradation is the design question behind it. It also runs the third drill on its own: switch one integration off (email sending, a CRM sync, a search index) and walk every other journey. The expected result is that only that feature says it is unavailable, and nothing else slows down or errors.

A clean outage is the rare case, in my reading, and five kinds of failure are worth injecting: slow (the response arrives after your timeout), erroring (500s), refusing (429s), lying (a 200 whose body does not parse), and partial (the call succeeded and the confirmation never arrived).

The enterprise end exists and is not needed here. AWS Fault Injection Service is “a managed service that enables you to perform fault injection experiments on your AWS workloads,” and Azure Chaos Studio is “a managed service for chaos engineering and Azure resilience testing.” Kubernetes chaos tooling sits at the same scale, and Chaos Monkey itself requires apps managed with Spinnaker. For a Python backend, the code side of all this is Python error handling best practices.

How to simulate a payment provider outage, and mock Stripe downtime in staging

A payment provider outage is simulated in a Stripe sandbox with 4 scenarios: the API call times out, the call succeeds but the reply is lost, webhooks stop arriving, and the hosted billing page is unreachable. Each has a written expectation for the customer and the database. stripe-mock is a fast test server, but it returns success instead of errors.

stripe-mock is a mock HTTP server “based on the real Stripe API,” and it looks like the obvious way to mock Stripe downtime in staging. Its README rules that out: the server is stateless, and “Testing for specific responses and errors is currently not supported. It will return a success response instead of the desired error response.” Stripe’s automated testing guide describes the route that does work: “You can generate Stripe API responses manually for various errors and mock the response returned in back-end automated testing.” The same guide says Checkout and the Payment Element have security measures that prevent automated testing, so the drill drives your app up to the hand-off and fakes Stripe’s answers behind it.

The four scenarios follow. The Stripe behavior in each row comes from Stripe’s docs; what the customer and the database should show is my working rule.

ScenarioHow to inject itWhat the customer should seeWhat the database should show afterwards
Creating the Checkout session times outThe mocked client or the kill switch fails the callPayments are temporarily unavailable, and you have not been chargedNo order marked paid; a logged error with the request id
The call succeeds at Stripe and the app never sees the replyAn app-side test flag discards the successful response and throwsOne confirmation, after the retryOne payment: the retry sends the same idempotency key and parameters, and Stripe returns the saved result of the first request
Webhooks stop arrivingDisable the sandbox webhook endpoint, complete a test payment, re-enable the endpoint, then resend the eventsPayment received, your account is being activatedAccess granted once the resent events are processed, each event handled once
The customer portal or hosted page is unreachableThe kill switch fails the call that creates the sessionThe billing screen says billing is unavailable and offers a support contactNo change to the subscription

The partial case is the one that costs money. Stripe’s idempotent requests work by “saving the resulting status code and body of the first request made for any given idempotency key, regardless of whether it succeeds or fails,” and “Subsequent requests with the same key return the same result, including 500 errors.” The key only protects a retry that reuses it with the same parameters, because Stripe’s idempotency layer errors when the parameters differ.

Run the webhook scenario against a registered webhook endpoint in the sandbox, pointed at staging. The retry and resend rules in Stripe’s webhook docs are written for registered destinations; they say nothing about events sent while stripe listen is stopped (checked 2026-10-04). Disabling the endpoint is the harsh version: Stripe “won’t attempt to resend any events generated while the destination is disabled,” so recovery means resending them. Dashboard Resend works up to 15 days after event creation; the CLI stripe events resend <event_id> --webhook-endpoint=<endpoint_id> works up to 30 days. The gentler version leaves the endpoint on and makes the staging handler return an error. In live mode Stripe retries a failed webhook delivery for up to 3 days with exponential backoff; in a sandbox it retries three times over a few hours. A drill that outlasts the sandbox window ends with a resend too. Stripe does not guarantee event order, so the backlog has to be handled once per event ID, in whatever order it arrives.

Everything here runs in a Stripe sandbox with test keys, never in live mode. A real webhook failure in production is a different job, covered under Stripe webhooks failing.

Test failover behavior in staging: the AI provider is down

Failover behavior is tested by injecting 4 model-provider failures in staging: a 429, a server error, a response slower than the timeout, and a 200 with unusable output. The app should back off, fall back, keep queued work, and return to normal by itself when the fault is removed.

Failover here means the app switching to its fallback, not a database standby taking over. The status codes below come from OpenAI’s error codes page and Anthropic’s API errors; the expected behavior is my working rule.

Injected failureExpected app behaviorExpected queue behaviorRecovery check
429 rate limitWaits at least as long as Retry-After says, retries a bounded number of times, then shows a deferred stateThe request is queued or deferred, never retried in a tight loopNew requests succeed without a redeploy
500, 503 or 529 server errorRetries a bounded number of times with backoff, then falls backWork stays queued, with a retry capThe circuit closes and the queue drains
Response slower than the timeoutThe call is cut at the deadline, and the user sees the degraded stateThe job is retried later or expires, never left hangingResponse times are back to normal
200 with unusable outputValidation rejects the output, and the same fallback runsNothing downstream consumes the bad outputValid output flows again

OpenAI lists 429 for rate limits, 500 for a server error and 503 with the code server_is_overloaded, and tells you to follow the Retry-After header when it is present. Anthropic lists 500 api_error, with the advice to retry with exponential backoff, and 529 overloaded_error for when “The API is temporarily overloaded.” Not every 429 clears with time: Anthropic’s tier spend-cap 429 “has no retry-after header and keeps failing until access resumes,” and OpenAI warns that retrying billing, spend, or quota errors “won’t restore API access.” So inject one 429 with Retry-After and one without, and expect the second to reach the fallback instead of a long wait.

Count attempts at the network, not only in your code. Anthropic’s official SDK retries transient failures, such as connection errors, rate limits and 5xx errors, twice by default, honoring the retry-after header when present, and accepts max_retries to change or disable that. OpenAI’s defaults, including what happens to a timed-out call, are under OpenAI API timeout. Any retries in your own code multiply on top of the SDK’s.

The fallback itself is chosen when you design the degraded mode; the drill only exercises it. The options are a second provider or a smaller model, a cached or template answer, or the feature switched off with a message while the rest of the app keeps working. For the queue, I hold four working rules: work submitted during the outage is kept, not dropped; it is retried with a cap; it drains after recovery without repeating side effects; and stale jobs expire instead of firing hours late.

Then comes the half that is easy to skip. Remove the fault and check that the app returns to normal without a redeploy: the circuit closes, the queue drains, and any alert that fired clears.

In the Production Hardening Sprint, deliverable 3.11, AI feature hardening, is verified this way: run injection test cases against each feature, a spend simulation that hits the limit, and a provider-outage simulation in staging.

Outage drill checklist for small teams

An outage drill for a small team has 12 steps in three groups. Before: name one dependency, one failure and the expected behavior. During: inject, walk the main journeys, screenshot the customer message, watch the alert. After: remove the fault, check recovery, file the gaps, record the drill.

Before the drill:

  • Name the one dependency and the one failure you will inject.
  • Write the expected behavior for the user, the data and the queue, in three sentences.
  • Confirm staging uses sandbox keys and cannot email real customers.
  • Tell anyone else who uses staging that it will misbehave.
  • Set a time box. I’d allow about an hour.

During the drill:

  • Inject the failure.
  • Walk the three most important journeys the way a customer would.
  • Read the exact customer-facing message, and screenshot it.
  • Watch the logs, the error tracker and the alert channel, and note whether an alert fired and when.

After the drill:

  • Remove the fault and watch the app recover.
  • Compare observed with expected, and file each gap as a ticket with its screenshot.
  • Record the drill in the table under the next heading.

If no alert fires because none is set up, that is the first finding. The customer message belongs to the drill as much as the code does. What to publish, and on which channel, during a real outage is a separate decision, starting with whether a small SaaS needs a status page, and who does what is the job of an incident response plan template. My working rule for cadence: once before launch, and again whenever a provider or the fallback changes.

How to verify it

An outage drill is verified by its record: one row per scenario with 8 columns, from how the failure was injected to what the customer saw and how the app recovered. Expected behavior is written before the drill, and no row ends on an unhandled error without a fix and a re-run.

ScenarioInjected howExpectedObservedQueue behaviorCustomer messageRecoveryDate
AI provider rate-limitsMocked 429 with Retry-AfterDeferred state, bounded retries, then fallback
Stripe webhooks stopSandbox endpoint disabled, then events resentActivating message, access after the resend, one grant per event
Email integration offKill switchOnly email reports unavailable; signup still works

Four checks decide whether the drill passed, and each one can fail:

  1. 01 Every dependency whose failure would stop a paying customer has at least one row. Evidence: the dependency list kept beside the table.
  2. 02 Every row's expected behavior was written before the drill. Evidence: the file's history shows the expected column saved before the observed one.
  3. 03 No observed cell says unhandled exception, blank screen, double charge or job lost without a linked fix and a re-run date. Evidence: the ticket link in the row.
  4. 04 The recovery column is filled in for every row. Evidence: no empty recovery cell.

Keep the evidence beside the table: the screenshot of each customer message, the log excerpt with its request id, and the alert with its timestamp if one fired. Timeouts, retries with backoff and graceful degradation each have checks of their own on the pages named earlier, and they are not repeated here.

For the Production Hardening Sprint, deliverable 6.10, staging outage simulation, is verified this way: record outage scenarios, queue behavior, customer messages, and recovery results.

Where the sprint does this

In the Production Hardening Sprint, deliverable 6.10 simulates AI or payment-provider failure in staging and verifies the expected recovery behavior, with the record described in the section above. The loss of your own data is a separate deliverable, 4.12, which restores the application and data into a fresh environment, times the recovery, and documents the procedure. The results go into the production readiness report, deliverable 13.1, which accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. Hosting, paid tools, and API usage remain in your accounts. Every deliverable is listed in the published scope.

Common questions about chaos testing

What is the difference between a chaos test and a stress test?

A stress test raises the load until something gives, and answers how much the app can take; a chaos test removes one dependency at normal load, and answers what happens when it fails. That split is my reading, and the two catch different bugs. Load and stress testing is a topic of its own, with its own tools.

What is the difference between chaos testing and fuzz testing?

Fuzz testing feeds a program bad input, while chaos testing takes away something the program depends on. OWASP on fuzzing defines it as “a software testing technique aimed at identifying bugs, vulnerabilities, or unexpected behavior by automatically providing a program with unexpected, malformed, or semi-malformed inputs.” Chaos testing leaves the input alone and fails a dependency, to check how the app recovers.

What is an example of chaos testing?

Disabling a Stripe sandbox webhook endpoint during a test payment, then re-enabling it and resending the events, is one. The expectation written beforehand is an “activating” message for the customer and access granted once per event after the resend; at production scale, Netflix’s Chaos Monkey ends instances at random.

How did Netflix use chaos engineering?

Netflix built Chaos Monkey, a tool that ends virtual machines and containers in its production environment at random, on the reasoning that “Exposing engineers to failures more frequently incentivizes them to build resilient services.” I think it works because Netflix runs many redundant instances, so losing one is survivable. A small SaaS with one server and one model provider does not have that cushion, which is why the drills here stay in staging.

Can you provide an example of a system outage message?

An in-app outage message has four parts: what is not working, what still works, whether the user lost anything or was charged, and when to look again. Here is one written for an app with an AI summary feature:

Smart summaries are unavailable because our AI provider is having problems. Everything else in your workspace works as normal. Nothing you wrote was lost, and you were not charged for this request. We will retry your queued summaries when the provider recovers, and this banner disappears when they are back.

That covers the screen inside the app; a public status update is a separate piece of writing.