Put the date on the calendar: the customer rollout, the newsletter feature or the paid campaign that brings more users than the app has ever seen. This scaling readiness checklist for startups is what has to be true before that date. 13 of the 21 third-party apps I audited had no rate limiting on their most expensive endpoint.
The scaling readiness checklist for startups: what has to be true by the surge, by area
A scaling readiness checklist for a startup’s app has eight gates: a measured capacity number, the first limit found and moved, the database pool sized, rate limits and a spend ceiling, timeouts and fallbacks, alerts that reach a person, a restored backup, and a tested rollback. Each gate has an artifact to keep.
The 21 apps behind the opening number are two groups I audited in June and July 2026: 11 third-party public vibe-coded apps, audited exhaustively across all 12 pillars, and 10 disjoint third-party apps held out and audited blind. They are a selected set, not a random sample, so the numbers describe those apps and are not a rate for AI-built apps in general. The same audits also counted the spend side: 12 of the 14 AI apps had a confirmed denial-of-wallet path, where a stranger or free account can burn the owner’s paid AI or compute bill without limit.
This is the milestone after the go-live checklist, for an app that already has users. If this is the app’s first launch, the seven launch gates and the plan from seven days out to the first week live in the launch calendar, and the steps outside engineering are in how to launch an app; this page keeps no timeline. It is the growth checklist for SaaS engineering, not for the company, and the business lists get their own section at the end.
| Area | The gate | The artifact to keep | The page that owns the detail |
|---|---|---|---|
| 1. Capacity | A load test at the expected peak with margin (my working rule: about twice the peak you expect) | A report of the workload, duration, environment, concurrency, latency, and error rate before and after changes, and the number written as a capacity statement that links capacity claims to load-test evidence and identifies untested projections as projections | load testing a web application, what a capacity plan is |
| 2. The first limit | The test’s first failure fixed, then the same workload run again | The two runs side by side, the one that failed and the one after the fix | scaling web applications |
| 3. The database | The pool size compared with the connection limit for the plan’s compute size | Both numbers written down together; on Supabase, the compute size’s “Database Max Connections” and “Connection Pooler Max Clients” beside the pool size set in the “Connection pooling” section of the Database Settings (Supabase’s connection management guide) | connection pooling in Postgres |
| 4. Rate limits and spend | A rate limit and a spend ceiling on every expensive endpoint and on the AI feature | The limit and the ceiling written per endpoint, and a record of a request past each one being refused | rate limiting in an API |
| 5. Third parties | A timeout and a fallback for each outside service | A record of disabling an integration and checking its fallback plus unaffected customer journeys | hardening SaaS applications for resilience |
| 6. Alerts and status | Alerts that reach a person, and a public status page | A record of triggering each configured condition, with its alert threshold and behavior; a test incident published, with the page verified as reachable separately from the app | logging and monitoring |
| 7. Backups | A backup restored, not only taken | The timed drill, recovered services, integrity checks, and observed recovery limits, recorded | a disaster recovery checklist for SaaS |
| 8. Releases | A rollback tested | A test release rolled back, with application behavior and retained data verified | release readiness checklist |
Every compute add-on on Supabase has a pre-configured direct connection count and Supavisor pool size, so row 3 starts from numbers the platform already set. The compute page lists the connection limits for each compute size, and on the Teams and Enterprise plans the Database client connections chart also shows a reference line for the compute size’s maximum connection limit.
On Supabase, row 7 depends on the plan. Supabase says “We automatically back up all Pro, Team, and Enterprise Plan projects on a daily basis,” recommends that free tier projects regularly export their data with the CLI’s db dump command and keep off-site backups, and says restore to a new project “is only accessible to paid plan users” with physical backups enabled. On the Free plan, then, the drill restores that dump into a separate database.
What kind of growth is coming: a spike, a step or a ramp
Growth arrives in one of three shapes. A spike comes fast and falls away, as a newsletter feature can; a step jumps and stays, as a customer’s rollout to its staff can; a ramp climbs over weeks, as a paid campaign can. Each shape tests different gates first.
The table is my reading of how each shape tends to behave, not a vendor fact, and the headroom column is my working rule. The row numbers point at the gates table above.
| The shape | How the users arrive | An example | The gates it tests first | How long the headroom has to hold |
|---|---|---|---|---|
| Spike | Fast, then falling away | A newsletter feature or a post that travels can arrive this way | Rows 4, 6 and 1: the spend ceiling, the alerts and status page, the capacity at the peak | Until the measured request rate has fallen back |
| Step | A jump that stays | A customer rolling the app out to its staff, or a new large account, can arrive this way | Rows 1, 3 and 7: capacity, the database pool, a restore with the larger data | For good: the capacity statement is rewritten on the new workload |
| Ramp | A climb over weeks | A paid campaign or a seasonal peak can arrive this way | Rows 1, 6 and 4: the capacity statement’s projections, the alert thresholds, the spend per user | Until the next measured test |
My rule: which shape is coming is read from the measured request rate once users arrive, never from where they come from; the name of the source does not set how long a spike lasts.
A spike tends to find the limit nobody measured. A founder schedules a post about an app that has never had many users at once, and nobody has run a load test. When the post goes out, the app’s short-lived serverless code opens direct database connections instead of going through the pooler, and the database reaches the connection limit for its compute size. A load test at the expected peak, reporting concurrency and error rate, would have found the same limit on a quiet day, when the move to the pooler or a larger plan could be made calmly.
A spike in payments meets the payment provider’s limits too. The documentation on Stripe’s rate limits puts it this way: “A sudden increase in charge volume, such as a flash sale, might result in rate limiting.” Stripe says it tries to set its rates high enough that legitimate payment traffic never exceeds the limits, and asks you to contact Stripe Support if you suspect an upcoming event might push you over them; the global limit in live mode is 100 requests per second per account. Stripe also generally discourages load testing against a sandbox, because API limits are lower there, and recommends a configurable way to mock out the Stripe calls during the test instead.
Whether the stack can handle a tenfold increase in volume is a question only a measured capacity statement answers, and any figure past the tested workload is a projection, labeled as one in row 1.
Turning a user count into requests, and which platform ceiling gives first, is covered in why an AI app stalls at 100 concurrent users. For a Product Hunt day, the schedule from the week before to the week after is in the Product Hunt launch technical checklist. The load test variant that matches each shape, a spike test or a soak test, belongs with the load test in row 1, and what to change first, in what order, with the scaling work in row 2.
What to do if something fails
The undo for each limit a surge can find should already be written down: a bigger pool, a pooler or a larger compute size for connections, a tighter limit or a queue for an expensive endpoint, the fallback for a third party, and the rollback for a bad build. Customers hear about it on the status page.
Connections have three undos, and each has a condition. Supabase’s general rule for pool size: if the app heavily uses the PostgREST database API, be conscientious about raising it past 40% of the Database Max Connections; otherwise, you can commit 80% to the pool. Serverless and edge functions open many short-lived connections, so Supabase’s guide to connecting to Postgres points them at the shared pooler in transaction mode, which does not support prepared statements: turn them off in the connection library. A larger compute size is the undo to schedule before the date rather than make mid-surge, because on Supabase changing it “will incur downtime”. What each connection error means, and how to stop a pool outage, is in connection pool exhausted.
For an expensive endpoint, my undo is a tighter limit, or a queue in front so work waits instead of piling onto the database. A third party that fails falls back the way row 5 proved it would. A bad build goes back through the rollback from row 8, and the steps on Vercel or Netlify are in roll back a deployment. What to write on the page while it is happening is in a status page for a small SaaS, and if a surge has already broken the app, the first hours are in my app went viral, now what.
A quiet dashboard proves little. In those same June and July audits, 17 of the 21 third-party apps had no error tracking or alerting: when a user hits an error, nothing records it. Like the opening number, it describes the apps I chose to audit, not every app.
Where the sprint fits
In the Production Hardening Sprint, this work is a set of named deliverables. In deliverable 9.2 we simulate concurrent users, identify the first bottlenecks, fix them and rerun the workload. Deliverable 9.6 documents measured concurrent capacity on the current infrastructure and the changes needed to plan for five times that workload. Two more cover the alerts and the restore: 8.3 alerts on error spikes, latency, connection pressure, and queue backlog, and 4.12 restores the application and data into a fresh environment, times the recovery, and documents the procedure. Deliverable 13.1, the production readiness report, delivers the result for every scope item, the work completed and its verification evidence; it is verified by accounting for all 123 IDs, keeping failures visible until resolved and explaining genuine non-applicable items. How each one is checked is listed in the published scope.
What the business scaling checklists cover instead
The pages that rank for this phrase are about the company: raising capital, the board, the team and hiring, cash flow, market validation. Sorted, that material falls into product and market, operations and infrastructure, finance and funding, and team and leadership. This page is the operations and infrastructure part, turned into gates you can check. If the list is more than the team you have can carry, who to bring in is the question in fractional CTO vs agency, and the rest of what production hardening covers starts from its own page.
If you have a working app built with these tools and need it ready for real customers, this is what we do.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase