“How much traffic can it take?” comes up when an investor asks about scaling limits or a customer’s review asks about load. So what is a capacity plan? One page: the concurrent load you measured, the conditions you measured it under, and the changes needed at 5 times that load, with every untested number marked as a projection.

What is a capacity plan for one web app?

A capacity plan for a web app is a one-page statement with five fields: the workload it was measured on, the environment, the measured concurrent capacity at a stated latency and error rate, the limit that bound first, and the changes needed for five times that load, with their assumptions.

An investor asking about scaling limits wants this page with the five-times field filled in, and a customer’s security or procurement questionnaire wants the same page attached to its load question. A team that says “we have no idea how much traffic we support” has not run its first load test yet: analytics shows today’s peak, not what the app can take, and the steps below get from one to the other. Which ceiling a given platform tends to hit first is a separate question, worked through in why your AI app stalls at 100 concurrent users; that stalls article comes up again in the steps below.

FieldWhat it recordsWhy it is on the page
WorkloadWhat users did and how many did it at once: the journeys, their mix, the length of each test stepA capacity figure means nothing until you know which traffic produced it
EnvironmentThe hosting plan, the region and the database size the test ran againstA test on a near-empty staging database answers a different question from production
Measured concurrent capacityThe concurrent users the app held, at a stated latency and error rate”Held” needs a line drawn in advance, or any number passes
Limit that bound firstThe resource that slowed or failed before anything elseIt is the next thing to change, and the first thing a reviewer asks about
Changes for five times the loadWhat has to change to carry five times the measured workload, each change with its assumptionIt turns a measurement into a plan and shows which parts are projected

As a capacity plan template, those five rows are enough: copy them, fill each field, and link every measured value to its test record. The capacity statement is one control in the wider checklist for web performance optimization.

Capacity planning in business and operations: the other meaning

In business and operations, capacity planning answers a different question: how an organization plans people, machines and materials against demand. IBM’s guide defines it as a process that “examines the production capacity and resources an organization needs to meet current and future demand” and names three strategies, lead, lag and match. Asana’s guide frames it around a project’s resources, 6sigma’s around an organization’s resources and projected demand, and Atlassian’s is titled as a project management guide. The words overlap (demand, capacity, growth) and the method does not. This page uses the software meaning, for one web app.

What goes wrong without it

Three failures follow when the number exists but the page does not.

FailureWhat you seeWhat a statement would have said
The number was a guessA concurrent-user figure repeated in a deck with no test behind it, and the first busy launch finds the real ceilingThe measured figure, and the limit that bound first
The number had no conditionsA figure measured on a small or empty staging database, quoted as if it held for productionThe environment, including the database size, beside the number
A projection was presented as a measurementA deck line a reviewer asks to see the evidence for, and there is noneThe word projected and the assumption behind it

The second row has a mechanism behind it: a page that is fast on a few rows can slow down on many, which the stalls article walks through in its own section. The third row is where a reviewer starts, working from a wider list of what investors look for in code.

Picture a founder answering a customer’s questionnaire with the concurrent-user figure from a load test run once on staging. The customer’s reviewer writes back asking what the test ran, on what environment and with what data, and the founder has the figure but none of its conditions. A number without its conditions is a poor basis for growth decisions, and here it is also a poor answer to a buyer.

In my June and July 2026 audits, the Performance & Scale pillar averages 53.3 out of 100, scored on all 21 third-party apps. Those 21 are the third-party apps I audited in that period, a selected set rather than a random sample, so the average describes them and is no score for AI-built apps in general.

How to produce the number, and write the statement

Producing the number takes six steps: turn users into a request rate, pick the journeys real users run, load test staging at rising multiples of today’s peak, record the workload, duration, environment, concurrency, latency and error rate at each step, note what bound first, and write the page.

  1. 01 Turn users into a request rate. A user count says little until it becomes requests per second and the time each request holds a connection; the stalls article does that arithmetic in its section "Turn users into requests before you size anything".
  2. 02 Pick the workload: the three or so journeys that matter most, run in roughly the proportion real users run them.
  3. 03 Load test staging, seeded with production-shaped data, at rising multiples of today's peak, up to five times. The ramp and the pass mark are in the stalls article's section "Run a controlled load test".
  4. 04 At each step, record the workload, the duration, the environment, the concurrency, the latency and the error rate.
  5. 05 Note the limit that bound first: the connection pool, a slow query, function concurrency, a third-party API, whichever gave out before the rest.
  6. 06 Write the statement: the five fields, each measured value linked to its test record, every untested number marked projected.

Step 2 is my working rule: about three journeys keeps the result readable, and the mix should follow what real users do. Step 4 is my working list of what a load-test record needs. The test itself, from script to first fix, belongs to load testing a web application, with a k6 load testing example as a script to start from. Finding which resource gives out first, and what to do about it, is part of scaling web applications.

Here are the five fields filled for an illustration, a two-founder SaaS on a Supabase database and Vercel functions, with one added row for the platforms’ published limits. The measured values are bracketed placeholders, because they only exist once you run the test. The platform limits are real values, read from Supabase’s connection limits and Vercel’s function limits on 4 October 2026. The host’s limits are its Pro plan’s, because Vercel says “the Hobby plan restricts users to non-commercial, personal use only”.

FieldFilled in (illustration)Recorded, measured, published or projected
Workload[journey A], [journey B] and [journey C] in [their share of sessions], [length of each step] per stepRecorded
EnvironmentStaging on Vercel Pro and a Supabase Micro compute (the smallest size a paid Supabase plan can launch), database seeded to [row count], region [region]Recorded
Platform limits in that environmentSupabase Micro: Database Max Connections 60, Connection Pooler Max Clients 200. Vercel Pro functions on Fluid compute: 300s default and 800s maximum duration; concurrency auto-scales up to 30,000Published, read 4 October 2026
Measured concurrent capacity[measured concurrent users] at [p95] and [error rate]Measured
Limit that bound first[the resource that gave out first] at [concurrency]Measured
Changes for five times the load[change 1], then [next limit] and [change 2], each with [its assumption]Projected until rerun

Supabase calls its connection figure a recommended value that “can be customized via max_connections”, so the statement records the value the project actually runs with.

To estimate server capacity growth, build the last row of that table: take the measured capacity and the limit that bound, name the change that moves that limit (a bigger pool, a cache, a queue), then name the next limit after it, and mark each one measured or projected. Today’s peak, the base the test multiplies, comes from analytics. GA4’s Realtime report shows “Active users in last 5 minutes” and “Active users in last 30 minutes”. The five-minute card is the nearest thing it shows to concurrent users; that is my reading, not Google’s definition. Page weight in the browser is a separate budget, handled by an image optimization checklist for websites rather than by this test.

A pre-launch capacity testing checklist is the same six steps run once before the go-live date, as one line of a go-live checklist. Google’s SRE launch coordination checklist, which the book calls Google’s original checklist from “circa 2005”, asks for the same things: under “Volume estimates, capacity, and performance” it lists traffic and bandwidth estimates, the launch spike and the traffic mix six months out, a load test and “capacity per datacenter at max latency”; under “Growth issues” it lists spare capacity, “10x growth” and growth alerts. What has to be true area by area before a surge is a longer list, the scaling readiness checklist for startups.

The projection rule: what “five times” means, and how to label it

A projection is a number nobody measured. It goes on the statement only with the assumption that produced it and the word projected beside it. The usual assumptions are that the next limit is the one the test showed second, that the fix scales in a straight line, and that the workload mix stays the same.

Each of those can fail, and this is my reading of how. The second limit in the test may not be the second limit after the fix, because a fix can push load onto a resource the test never stressed. A bigger pool or a cache helps until something behind it, such as the database’s own CPU or memory, becomes the ceiling, so the gain flattens. Growth also need not arrive evenly: one new feature or one large customer can change which journeys dominate.

Five times is the planning horizon this page uses. Google’s checklist plans further, with the “10x growth” line quoted above.

How to verify it

A capacity statement is verified by following its numbers back: every capacity claim links to a dated load-test record, every untested projection is marked as a projection, and the environment matches production or the difference is written down. Pick one number and follow it to its test.

To verify capacity numbers with load evidence, run five checks, each of which leaves something you can see:

  1. 01 Every number on the statement links to a load-test record with a date. Evidence: click the number and a record opens.
  2. 02 The record holds everything step four of the method lists, for each load step. Evidence: no blank columns in the record.
  3. 03 The test environment matches production, or the statement writes the difference down. Evidence: plan, region and database size side by side for both.
  4. 04 Every untested number is labeled projected and names its assumption. Evidence: no bare figure without either a test link or that label.
  5. 05 A stranger can rerun the test from the statement alone. Evidence: the script path, the data set and the command are on the page.

Checking a statement someone else wrote works the same way with one number. Pick it, follow its link, and check the record’s date and environment against what the statement claims. If the link opens nothing, or the record describes a smaller database or a different plan, treat the number as a projection whatever the page calls it.

In the Production Hardening Sprint, deliverable 9.6 is verified this way: link capacity claims to load-test evidence and identify untested projections as projections.

Where the sprint does this

Deliverable 9.6 documents measured concurrent capacity on the current infrastructure and the changes needed to plan for five times that workload. The numbers come from deliverable 9.2, where we simulate concurrent users, identify the first bottlenecks, fix them, and rerun the workload. The capacity statement then goes into the technical due diligence pack, bundled with the readiness report, architecture diagram, data model and security checklist into one PDF. The production readiness report accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. The app’s current framework and hosting setup are the starting point, and components are refactored or replaced where the production work requires it. The full entry, with its why-it-matters line, is the capacity statement in the published scope.

Common questions about capacity plans

What are the three types of capacity planning?

Workforce, product and tool capacity planning are the three types IBM’s guide gives as examples, and it names three main strategies: lead, lag and match. For digital services, it says tool capacity planning “includes procuring cloud resources, physical servers, virtual machines (VM) and the IT infrastructure needed for service delivery.” For one web app, the three that matter are measured capacity, projected capacity and the change plan between them; that split is my framing, not a standard.

What is server capacity planning?

Server capacity planning is working out how much load the servers and database behind an app can take before latency or errors cross a line you set, measured by a load test and written down with its conditions. For one app, the written version records the workload, the environment, the measured capacity, the first limit and the changes for five times the load.

Can you provide an example of a capacity plan?

Yes: the illustration table under the method fills all five fields for a two-founder SaaS on Supabase and Vercel Pro, with every measured value left as a bracketed placeholder and only the platforms’ published limits, read on 4 October 2026, as real figures.

How do you perform capacity planning?

You convert expected users into requests per second, choose the few journeys that carry most of the traffic, ramp a load test on staging past today’s peak, log the conditions at every step, find the first resource to give out, and write the page.

What is the best tool for capacity planning?

For one web app, the best tool is a load-testing tool plus a written page: k6 is one such tool, and the page is the five-field statement above. Planning software built for staff and projects answers the business meaning of capacity planning, not this one.