On the day a region goes dark or the hosting account gets locked, a plan written that morning is too late. A disaster recovery checklist for SaaS is 6 lines: what you run and where, the RTO and RPO you accept, who does what, the restore procedure, one timed drill, and the evidence it left behind.

The disaster recovery checklist for SaaS: the plan, the drill, the evidence

A disaster recovery checklist for a SaaS has 6 lines: the inventory of services and the accounts they live in, the RTO and RPO the business accepts, named roles, a numbered restore procedure, one timed drill into a fresh environment, and a folder holding the evidence from that drill.

Recovery is one control in the data consistency checklist for SaaS, and it is the one that asks whether the whole service comes back when someone follows written steps, not just whether a backup exists. Each line below records one thing and leaves one piece of proof.

Checklist lineWhat it recordsWho owns itThe evidence
InventoryEvery service, the account it lives in, the region, and who can log inWhoever holds the owner role on those accountsA dated list, one row per service
TargetsThe RTO and RPO the business acceptsThe founder, because the business sets themThe targets line, with the last tested time beside each
RolesWho declares, who restores, who talks to customers; on a two-person team, two names and a fallbackThe person named to declareA contact list kept outside production
Restore procedureNumbered steps in the order the services come backThe person named to restoreThe steps, dated at their last change
Timed drillThe whole service rebuilt in a fresh environment from the plan aloneThe person named to restoreStart and stop times, services recovered, the result
Evidence folderThe drill record, the data checks, the gaps and their fixesThe person named to declareA folder that lives outside the production account

Everything that should be included in a DRP for a small SaaS fits in those six rows. If you think of disaster recovery as 5 steps (audit and assess, set targets, back up, assign roles and communication, test and update), they map onto the six lines: assessing is the inventory, setting targets is the targets line, backups feed the restore procedure, roles and communication are the roles line, and testing and updating are the drill and the evidence folder.

For a startup with no servers of its own, disaster recovery starts from a short list of what counts as a disaster:

  • A provider region outage: the platform is fine elsewhere, but your project lives in the region that is down.
  • A deleted or locked account: a billing lapse, a mistaken click or a compromised login takes the project, and anything stored inside it, out of reach.
  • A destructive migration: a schema change or a script removes or corrupts production data, and a rollback of the code does not bring it back.
  • A compromised key: someone else holds a production secret, and you have to rebuild on fresh credentials from a copy you trust.

In the Production Hardening Sprint, deliverable 4.12 is the disaster recovery drill: we restore the application and data into a fresh environment, time the recovery, and document the procedure.

IT disaster recovery plan for small business: the one-page template

An IT disaster recovery plan for a small business with no IT department fits on one page: what runs where and in which account, the two recovery numbers, who declares and who restores, the restore order, the customer message, and the date and result of the last drill.

IT disaster recovery planning for a company like this is mostly about accounts. DR plan templates that start from servers and laptops are written for a different company. Ready.gov’s IT disaster recovery plan page says the plan “begins by compiling an inventory of hardware (e.g. servers, desktops, laptops and wireless devices), software applications and data.” My reading: for an app with no hardware of its own, that inventory is a list of accounts and managed services, each with a region and a set of people who can log in.

What follows is an example disaster recovery plan for one common stack: treat it as a sample DRP and change every line that does not match yours. The assumed stack: a Next.js app on Vercel, Postgres on a managed host, Stripe for payments, a transactional email provider, and a domain at a registrar. The disaster recovery plan below is the template: copy it into the place your team keeps documents that have to survive an outage, and fill the brackets. The RTO, RPO and drill lines stay blank until you set them and test them.

DISASTER RECOVERY PLAN: [product]
Version date: [ ]    Next review: [ ]

SCOPE
  Production only. Stack: Next.js on Vercel, Postgres on a managed host,
  Stripe, a transactional email provider, a domain at a registrar.

INVENTORY (service / owning account / region / who can log in)
  Registrar and DNS host ... [account] / - / [names]
  Vercel project ........... [team] / [region] / [names]
  Postgres host ............ [org, project] / [region] / [names]
  Copy outside the host .... [where] / [region] / [names]
  Stripe ................... [account] / - / [names]
  Email provider ........... [account] / - / [names]
  Secrets vault ............ [password manager] / - / [names]

TARGETS
  RTO: [ ]    last tested whole-service time: [ ]
  RPO: [ ]    copy the last test restored from, and its timestamp: [ ]

ROLES (kept on paper and in the vault, not only in production tools)
  Declares: [name, phone]    fallback: [name, phone]
  Restores: [name, phone]    fallback: [name, phone]
  Talks to customers: [name]

DECLARE RULE
  [Declarer] declares when production cannot be recovered in place,
  or when the outage is expected to last longer than the RTO.
  Any other outage: roll back first. This plan is not in use.

RESTORE ORDER (each step on a new project, never over production)
  a. Domain and DNS: log in to the registrar and DNS host; records kept at [where]
  b. Database: restore [which copy] into a new Postgres project; run the data checks
  c. Secrets: load them from the vault into the new Vercel project
  d. Application: deploy [branch or commit] to the new Vercel project
  e. Webhooks: point Stripe and the email provider at the new deployment;
     if a provider issues a new signing secret, add it and redeploy
  f. Smoke test: [sign up, log in, one paid action, one email received]
  g. Switch DNS to the new deployment only after f passes

CUSTOMER MESSAGE
  Where: [status page or channel outside the app]    Posted by: [name]

LAST DRILL
  Date: [ ]   Scenario: [ ]   Whole-service time: [ ]   Result: [ ]
  Gaps found, and the date each was fixed in this plan: [ ]

The template is free to copy and there is no disaster recovery plan PDF or Word file to download: the block is the whole IT disaster recovery plan. Keep the secret values in the vault, not only in the hosting dashboard, because the dashboard’s account is one of the things that can be lost. I put secrets before the application so the first deploy already has them, and the declare rule is mine. The customer message line needs a place to publish that does not go down with the app, which is the job of a status page for a small SaaS. The commands for each restore step belong in a runbook template, with the plan pointing to them rather than repeating them.

What goes wrong without it

Your team needs evidence of recovery time before a real incident. With no plan, if the server behind your app dies (on managed services that means a region, an account or a bad migration), recovery turns into improvising under pressure. Here are four ways the day goes, as patterns (my list, not events from any one company):

The dayWhat the team discoversThe checklist line that prevents it
A provider region goes downThe backups sit in the same region, and nobody wrote down how the service comes back anywhere elseInventory (regions) and restore procedure
The account is locked or the project deletedThe backups lived inside the thing that is goneInventory (where each copy lives, which login can delete it)
A destructive change reaches productionThe data restores, but the environment variables, webhook endpoints and DNS records that made the app run were never written downRestore procedure, built from the inventory
A reviewer asks for the planAsked for the plan and the date of its last test, the team has neitherTimed drill and evidence folder

On Supabase, deleting a project removes its backups too, which Supabase backups walks through.

A public case shows how much sits outside the data. Google Cloud’s account of the incident, from May 2024, says that in early 2023 Google operators used an internal tool to deploy one of the Google Cloud VMware Engine Private Clouds of its customer UniSuper, and left one input parameter blank. The system assigned a then unknown default fixed one-year term. At the end of that period, that Private Cloud, one of the customer’s multiple Private Clouds across two zones, was deleted, with no customer notification because the deletion was not a customer request. The customer and Google teams worked 24x7 over several days to recover the Private Cloud, restore the network and security configurations and the applications, and recover data. The post says the recovery was assisted by the customer’s architectural approach, and that data backups stored in Google Cloud Storage in the same region were not impacted and, along with third-party backup software, were instrumental in aiding the rapid restoration. Google says it deprecated the internal tool and corrected the system behavior that set such Private Clouds for deletion.

The lesson I take from it is mine, not Google’s: a recovery plan lists what lives outside the thing that can be deleted, and it covers more than the data, because the days of work went into network and security configurations and applications as well as data.

Two of the pillars I score in audits sit closest to recovery. In the third-party apps I audited in June and July 2026, the Reliability & Correctness pillar averages 31.4 out of 100 across the 21 scored on it, and Data Integrity & Safety averages 51.6 out of 100 across the 20 scored on it. Those 21 are the 11 public third-party apps and the 10 held-out third-party apps I audited blind, a selected set of audited apps and not a random sample or a rate for AI-built apps in general.

How fast you find out is a separate control, covered in uptime monitoring for founders. If the day is a deleted database, the first hour belongs to the first-hour database recovery runbook, and this plan takes over once the data is safe from further damage.

How to do it: the targets, the fit with backups, the test

The definitions, activation criteria and test vocabulary below come from NIST SP 800-34, AWS and the providers’ own pages, not from a drill I ran. The four parts come in the order a team does them: name the targets, match the plan to the backups, decide what the provider owns, then test.

RTO and RPO: the two numbers the plan has to name

RTO and RPO are the two targets a disaster recovery plan names. In NIST’s glossary, RPO is the point in time data must be recovered to after an outage, and RTO is how long a system’s components can stay in recovery before the organization’s mission or business processes are harmed. The plan writes each target beside the last test’s result.

Those definitions come from the glossary of NIST SP 800-34. With RTO and RPO explained that way, the table shows where each one lives in the plan and what the test writes beside it.

NumberNIST’s definitionWhere the plan writes itWhat the test records beside it
RPOAs in the capsule aboveThe TARGETS line of the plan blockThe timestamp of the copy the test restored from
RTOAs in the capsule aboveThe TARGETS line of the plan blockThe whole-service time, from the declare to a passing smoke test

The last column is the disaster recovery time to restore the whole service, which is a different number from how long the database restore took. NIST lists “Expected duration of the outage lasting longer than the RTO” among the criteria a plan may be activated on, which is why the plan block’s declare rule names the RTO. For a security incident such as a compromised key, the RPO question becomes which copy predates the compromise, and that copy may be older than the most recent backup.

AWS’s disaster recovery options group disaster recovery strategies into four approaches, backup and restore, pilot light, warm standby, and multi-site active/active, “ranging from the low cost and low complexity of making backups to more complex strategies using multiple active Regions”, and AWS gives the rule for choosing among them: “Use your RTO and RPO needs to help you choose between these approaches.”

What sets each number, the arithmetic behind a nightly backup, and how a small app picks its targets are all in test your backups with a restore drill.

Backup and disaster recovery: how the two fit together

Backup and disaster recovery are two jobs. A backup is a copy of the data. Disaster recovery is the rehearsed way back to a running service: the data, the application, the secrets, the DNS, the third-party settings, the people and the clock.

The difference between disaster recovery and backup shows up in everything the recovery plan adds to a backup, and the table puts the two side by side.

BackupsThe disaster recovery plan
What it isCopies of the database, and sometimes of files, taken on a scheduleA page that says how the whole service comes back, in what order, and by whom
What it answersDoes a copy of the data from a given point exist?Can a named person rebuild the running service from nothing but the page, and how long did it take last time?
Who checks itWhoever watches the backup job and restores a copyThe person who declares, at every drill and review

The restore order in the plan block is one example of a backup and recovery procedure for a managed stack; with the commands for each step written out, it becomes the SOP. The cloud backup and recovery best practices I hold a plan to on managed services come down to three working rules, whether the plan starts from the block above or from a backup and recovery plan template you already keep:

  • The plan names where each copy lives and which login can delete it.
  • The plan lists what the backups do not hold, such as secrets, DNS records and webhook settings, and where each one is written down.
  • The plan carries the date of the last test and the copy that test restored from.

Whether a copy is really separate from production (a copy outside the platform, separate credentials, retention locks) is its own section of the restore drill article. What the provider keeps, how long it keeps it, and how to add the off-site copy are in the database backup checklist for startups. I treat backup-and-recovery software and disaster-recovery services as a different product from this plan: they make and move copies, and the plan still has to say who uses them and in what order.

Disaster recovery in cloud computing when everything you run is managed

Disaster recovery in cloud computing is getting the service and its data back after an outage; on managed platforms it is shared: the provider recovers its platform, you recover your project. A deleted project, locked account or destructive migration is yours, and a second region counts only once you have proven you can restore into it.

The provider’s pages draw that line themselves. Supabase’s shared responsibility model says that “you are always responsible for” your Supabase account, access management, data, and applying security controls. What Supabase’s own backups hold is covered in the Supabase backups article linked above. AWS’s shared responsibility model for resiliency puts it in two phrases: AWS is responsible for “resiliency of the infrastructure that runs all of the services offered in the AWS Cloud”, and for managed services, “You are responsible for managing resiliency of your data including backup, versioning, and replication strategies.”

My reading, for a small app: disaster recovery in the cloud does not have to mean paying for a standby copy in a second region. A second region is a restore target you have proven you can create, with the project, the database copy and the secrets ready to go there, and the drill below is how you prove it. Cloud disaster recovery services, as I read them, are built for fleets of virtual machines, which a SaaS on managed platforms does not run.

Two terms from cyber security sit one layer above this plan. BCP in cyber security is the business continuity plan, which keeps the company running while systems are down: people, decisions, customers. BIA in cyber security is the business impact analysis, which decides which systems matter most and how long each can be down; both, and how they relate to recovery, are in a business continuity plan template.

How to run a disaster recovery test

A disaster recovery test for a small SaaS tests the written plan, not the backup. Someone declares by the plan’s rule, rebuilds each service in the plan’s order using nothing but the plan, runs the smoke test and stops the clock. Every step they had to improvise or ask about becomes a fix to the plan.

The database restore itself is not taught here. The isolated target, side effects switched off, the records to check and what counts as a pass are the restore drill article’s procedure; run it as written there. What this test adds is the plan around the restore:

  1. 01 Pick one scenario from the four disasters listed near the top of this page and apply the plan's declare rule exactly as written.
  2. 02 Hand the written plan to the person it names to restore, and allow nothing else: no memory, no chat history, no one's laptop settings. Every question they have to ask is a gap in the plan.
  3. 03 Start the clock at the declare.
  4. 04 Bring the services back in the plan's order, each on a new project or deployment that is not production: domain and DNS access (log in to the registrar and the DNS host and check the records are written down, changing nothing), then the database, the secrets, the application, and the third-party webhooks in test mode.
  5. 05 Run the smoke test the plan names against the new deployment.
  6. 06 Stop the clock when the smoke test passes, and write the whole-service time beside the RTO in the plan.
  7. 07 Write down every step that was improvised, missing or wrong, fix the plan the same day, and date the change.

In a real outage, the plan comes into use only when its declare rule is met: an ordinary outage starts with a rollback, as in what to do first when the app is down. The order in step 4 is mine; NIST asks for the same shape, recovery procedures “written in a stepwise, sequential format so system components may be restored in a logical manner.” The restore drill article keeps the database time and the whole-flow time apart; the clock in this test records the second one. NIST’s rule for each test or exercise matches step 7: results go into an after-action report, and corrective actions are captured for updating the plan.

Whether you count five testing types for a disaster recovery plan or fewer, NIST SP 800-34 is where to check the terms. It separates tests, training and exercises, and names two exercise types: tabletop exercises, which are discussion-based only, and functional exercises, where people perform their duties in a simulated operational environment. It lists the areas a contingency plan test should address, as applicable: notification procedures, system recovery on an alternate platform from backup media, internal and external connectivity, system performance using alternate equipment, restoration of normal operations, and other plan testing where coordination is identified.

My reading: on managed services, a functional test that actually rebuilds the service is within reach of a two-person team, so a tabletop is where the plan starts, not where it ends. Nothing in this test touches production.

How to verify it: auditing disaster recovery

Auditing disaster recovery comes down to 7 checks: a dated plan with owners, written targets, a test record with declare and stop times, integrity checks, the observed limits, a named copy outside the production account, and a plan that changed after the test.

Run them as a backup and disaster recovery audit checklist against the plan, and keep the list as your audit template: the evidence for each check, filed together, is the record a reviewer reads.

  1. 01 The plan exists, names its owners, keeps contacts reachable outside production, and carries a review date that has not passed.
  2. 02 The RTO and RPO are written down, with the tested whole-service time beside the RTO.
  3. 03 The test record shows a date, the declare and stop times, and the services recovered, in order.
  4. 04 The integrity checks from the restore drill article ran on the restored data, and their results are filed.
  5. 05 The observed limits are written down: what could not be recovered, and why.
  6. 06 The plan names the copy the test restored from and where it lives, outside the production account.
  7. 07 The plan changed after the test: the dated fix list from the last step of the test is in it.

This is a self-check you can run before a DR audit, not an audit opinion. A check that has no evidence behind it is a failed check, even when everyone remembers doing the work.

Sprint deliverable 4.12, the disaster recovery drill, is verified this way: record the timed drill, recovered services, integrity checks, and observed recovery limits.

Where the sprint does this

In the Production Hardening Sprint, the disaster recovery drill is deliverable 4.12, delivered and verified as described above. Deliverable 4.7, verified backups and restore, is verified this way: restore a backup into an isolated environment and check representative data integrity. The production readiness report, deliverable 13.1, delivers the result for every scope item, the work completed, and its verification evidence. Formal third-party certifications and independent audit opinions are separate from these engineering deliverables. Hosting, paid tools, and API usage remain in the client’s accounts. Every deliverable, with its verify line, is in the published scope.

Common questions about disaster recovery for a SaaS

How often should a DRP be updated?

Update it after every test, after any change of provider, architecture or the people it names, and at least once a year: that is my working rule. NIST SP 800-34 says the plan should be reviewed for accuracy and completeness at an organization-defined frequency or whenever significant changes occur to any element of the plan.

What are the NIST guidelines for disaster recovery plans?

NIST’s guidance is Special Publication 800-34, Revision 1, the Contingency Planning Guide for Federal Information Systems, published in May 2010. It describes itself as “recommended guidelines for federal organizations” that set out technology practices for information system contingency planning, and says it “has been prepared for use by federal agencies” and “may be used by nongovernmental organizations on a voluntary basis.” Its planning principles are written for three platform types, client/server, telecommunications and mainframe systems, so the inventory and restore order above are my translation of them for managed services.

Can RPO be higher than RTO?

Yes. As I read them, the two measure different things, a point in time for the data and a length of time for the recovery, so neither limits the other. NIST keeps them apart too: “Unlike RTO, RPO is not considered as part of MTD”, the maximum tolerable downtime.

When would a DRP be activated?

A DRP is activated when one or more of its written activation criteria are met. NIST SP 800-34 says those criteria may be based on the extent of damage to the system, how critical the system is to the organization’s mission, and an outage expected to last longer than the RTO. The rule in the plan block above is mine: declare when production cannot be recovered in place, or when the outage is expected to outlast the RTO.