Make the cheap changes before you buy servers. Scaling web applications for one app follows an order by cost: an index, the query shape, the connection pool, a cache, background jobs, a bigger box, and only then a second box behind HAProxy, NGINX or the host’s balancer. Skip any step your measurements rule out, and measure again after each change.
What scaling web applications means when you have one app
Scaling a web application means keeping the same response time as load grows. For one app that is an order of changes, cheapest first, with a measurement after each one. A second server is the last step, not the first.
This is one part of web performance optimization: the part where the app is quick for one visitor and slow for many at once. Finding which limit your app is actually hitting, from the symptom, the platform it runs on and a controlled load test, is the job of why your AI app stalls at 100 concurrent users. What follows starts once you know, and covers what to change and in what order.
For a single app, scalable web applications come from cheap changes made in order, long before they come from a second server. Most of learning how to scale an application with one codebase and one database is learning that order. Scaling a SaaS application here means the app and its database, not the company’s growth; web-scale applications, spread over many regions and many teams, are a later problem, and the architecture section below draws the one-app version instead.
A scaling checklist for a web application does not need a template download or a GitHub repository: the order table and the step log below are the whole checklist, and both copy into a README in your own repo.
Scaling a no-code app keeps the same idea, fix the query before buying capacity, but the platform decides which steps you can take and adds a ceiling of its own. The builder platforms’ ceilings are listed in that same concurrent-users article, and Bubble’s capacity and workload meters are covered in why a Bubble app slows down under traffic. On Supabase, ask is Supabase slow, or is it your queries before anything else here. The front-end half of speed, page weight and Core Web Vitals, is a different measurement; learn how to run a Lighthouse audit for that half.
Why it matters before the traffic arrives
Across the 21 third-party apps I audited in June and July 2026, the Performance & Scale pillar averages 53.3 out of 100, scored on all 21. Those 21 are a selected set, 11 public vibe-coded apps plus a held-out set of 10 others audited blind, not a random sample and not a rate for AI-built apps in general.
In the same audits, 13 of the 21 third-party apps had no rate limiting on their most expensive endpoint. When a spike arrives, the most expensive endpoint with no limit on it is spending nobody planned, and no step in the order below caps it: a bigger box only lets it spend faster. The guard for it sits beside the order, in the next section.
None of this says your app is slow. It says the scaling work is worth checking before a launch rather than during one.
What to change, in what order
The order of changes for one app runs cheapest first: the index, then the query shape, then the connection pool, then a cache, then long work moved off the request, then a bigger box, and only then a second box behind a load balancer. Each step waits for a measurement.
The order is my working rule: cheapest first, a measurement between each step, and you start at the step your measurement points to, which the concurrent-users article helps you find. The “what it costs you” column is my reading of what each change asks of you as the owner, not a price.
| Step | The change | What it costs you | Where it is taught |
|---|---|---|---|
| 1 | An index on the column the slow query filters or joins on | A migration, no new bill | analyze a Postgres query |
| 2 | The query shape: fetch only the rows and columns the page shows, one query instead of one per row | A code change, no new bill | the N+1 query problem solution |
| 3 | A connection pool in front of the database | A setting or a pooler, usually no new service to run | connection pooling for Postgres |
| 4 | A cache for repeated reads | One more thing to invalidate, and to keep separate per customer | the caching strategies web applications need |
| 5 | Long work moved off the request into a background job | A queue and a worker to run | how to run long tasks in the background |
| 6 | A bigger box | A higher tier or more metered CPU and memory, usually a bigger bill, same code | horizontal scaling vs vertical scaling |
| 7 | A second box behind a load balancer | A second bill, and the app must keep no state of its own on the instance | The horizontal and vertical scaling question from step 6, then the load balancer section below |
The rows-and-columns half of step 2 is the “Oversized reads” section of the concurrent-users article. At step 2, one finding shows why the order starts in the code: in one of the apps I audited in June and July 2026, an AI coding workspace fetched the full text of every file in a workspace just to draw the file tree, which needs only names and folders. That is a query-shape fix, in my reading, and no second server makes a read of every file cheaper.
Which lever fits which kind of saturation during a spike, a CDN, autoscaling, a read replica or a bigger tier, is set out in what to do in the first 48 hours after an app goes viral. When one endpoint is slow rather than the whole app, start with where an API loses its milliseconds.
Three guards sit beside the order rather than in it, because they cap cost or failure instead of buying capacity. The first is rate limiting in an API, applied to the most expensive endpoint. The second is a timeout on every outside call; in browser and Node code that means knowing how to set a timeout on fetch. The third applies where the app runs on functions: the platform’s concurrency ceiling, which on AWS means Lambda concurrency.
What a scalable web app architecture actually looks like at your size
A scalable architecture for one app is one relational database with a pool, one cache, one queue for background work, a CDN for static assets, and the host’s own load balancing. That diagram holds until the second database or the second team.
That is where the seven steps above leave a small app. Steps 1 to 3 live inside the database box, step 4 is the cache, step 5 is the queue and its worker, and steps 6 and 7 are the size and the count of the app box. Drawn as a list, from the browser inward:
- The CDN, serving images, scripts and styles so most of those requests never reach the app.
- The host’s balancer, which on a managed platform you do not run yourself.
- One app instance, or a few identical ones once step 7 is done.
- The cache beside the app, and the queue with its worker.
- The pool, then the one relational database.
A high-performance web application architecture, drawn as an edge layer, a balancer, a cache, compute and a data tier, is the same set of layers at a bigger size. Large-scale web application architecture, the web-scale kind with many services, many databases and a team for each, is for the tenth engineer, not the first customer.
Microsoft’s architecture guide says “A traditional three-tier application has a presentation tier, an optional middle tier, and a database tier”, and the one-app list above is that shape. Whether you need two tiers or three is a three-tier web application architecture question, not a scaling one.
CPU and memory limits: when the container ceiling is the bottleneck, not the code
CPU throttling means work gets less CPU than it asks for. In a container, it happens once the kernel’s CFS bandwidth control has assigned all of the group’s quota for the period: threads asking for more are throttled until the next period. The kernel counts each time in nr_throttled. Memory is a separate limit.
This sits before step 6 because a bigger box only helps if the box is the ceiling. CFS throttling is how Linux enforces a CPU limit on a group of processes, and a container’s CPU limit is such a group quota, in my reading. In the kernel’s CFS bandwidth control docs, a group is allocated up to a quota of CPU time within each period, and once the quota has been assigned, further requests for it leave those threads throttled: “Throttled threads will not be able to run again until the next period when the quota is replenished.”
The kernel exports a group’s bandwidth statistics as five fields in its cpu.stat file; the three that show throttling, in its own words:
nr_periods: “Number of enforcement intervals that have elapsed.”nr_throttled: “Number of times the group has been throttled/limited.”throttled_time: the total time, in nanoseconds, for which the group’s entities have been throttled.
Those are the cgroup v1 names; the docs note that “The cgroupfs files described in this section are only applicable to cgroup v1.” On cgroup v2, the same cpu.stat file reports nr_periods, nr_throttled and throttled_usec when the CPU controller is enabled. For an app that is CPU-bound against its limit, the docs expect nr_periods to roughly equal nr_throttled.
The symptom, in my reading: response times climb while the host’s CPU graph looks well under the limit, because a CPU graph shows an average over its sampling window while the throttling happens inside each short period (the cgroup v1 default period is 100ms). Reading the counters needs a shell inside the container, on a VPS or a container you can exec into. On a managed host without one, the host’s CPU graph and the plan’s CPU limit are what you have; Railway, for example, publishes the RAM and CPU a service may use, “Depending on the plan you are on”. Platform-by-platform concurrency ceilings are in the concurrent-users article from the first section. The laptop meaning of CPU throttling, a processor slowing itself down when it runs hot, is a different subject.
A memory limit is a separate ceiling, and hitting it usually shows up as an OOM kill rather than as slow requests.
On Microsoft Azure the ceiling is written into the plan. Microsoft’s App Service scaling guide says scale up gets “more CPU, memory, or disk space” by changing the pricing tier of the App Service plan, and scale out increases the number of VM instances that run the app; “Basic, Standard, and Premium service plans scale out to as many as 3, 10, and 30 instances, respectively.” In my reading, that makes Azure App Service scaling a tier decision first. To configure scaling for a web app on an Azure App Service plan, pick the tier before the instance count, then decide on auto scaling inside it. The difference between autoscale rules and automatic scaling is not a throttling question; it belongs with the scaling question at step 6.
Load balancers: when one earns its place, and how to prove traffic is actually spread (HAProxy vs NGINX)
A load balancer earns its place when there is a second instance to balance. HAProxy and NGINX both do the job for one app. The proof that traffic is spread is an instance id in a response header, counted across a batch of requests, sent from several clients when the balancer pins each client to one instance.
This is step 7 of the order. On a managed host the balancer is part of the platform, in my reading. Railway runs extra instances as replicas, and “If you are using a single region with multiple replicas, Railway will randomly distribute public traffic to the replicas of that region.” Its Free plan row lists 1 replica per service. A VPS with a second instance needs its own balancer in front, and what the app must stop keeping on one instance before that helps is the stateless half of the step 6 question.
HAProxy vs NGINX, in each project’s own words: HAProxy’s own description calls it “a free, very fast and reliable reverse-proxy offering high availability, load balancing, and proxying for TCP and HTTP-based applications.” NGINX’s load-balancing guide covers load balancing across multiple application instances, lists three methods, round-robin, least-connected and ip-hash, and describes a fourth, least-time, in a section of its own. NGINX describes itself as a good deal more than a balancer. For one app either does the job, and in my reading the choice is what else you need the process to do: balancing and proxying alone, or serving files and caching as well.
Testing a load balancer takes one response header and a batch of requests. To test the load balancer’s spread on a system you own, add the instance id to a response header (on Railway, the RAILWAY_REPLICA_ID variable each replica gets), send a batch of requests (my working rule is about 20), and count the ids per instance. The count per instance is the evidence. On Railway the header matters more, because its metrics tab sums all replicas into one graph.
Before reading the count, check the balancing method, because it decides what a passing load balancing test looks like. With round-robin or least-connected, one instance taking every request means the balancer is not balancing; NGINX’s guide says round-robin gives a more or less equal spread, provided there are enough requests and they are processed uniformly and completed fast enough. With ip-hash, “the requests from the same client will always be directed to the same server except when this server is unavailable”. Any sticky-session setting does the same, so there a batch from one client lands on one instance by design, and my method is to send the batch from more than one client address. Railway’s docs say it “does not support sticky sessions” and spreads traffic at random, so on Railway an uneven count still passes as long as every replica shows up.
For HAProxy performance metrics, three fields in its Management Guide show a queue building, a server going down and errors coming back: qcur, “current queued requests”; status, which reads UP, DOWN or another state; and hrsp_5xx, “http responses with 5xx code”. NGINX performance data from its stub_status module is thinner: active connections, accepts, handled, requests, reading, writing and waiting, with no per-server status or response codes, and the module “is not built by default”. For one app, the header count and the 5xx counter answer most of what a dashboard would.
How to check your own app
Checking your own app against the order means one line per step: done, with the measurement it moved, or skipped, with the measurement that ruled it out. Seven dated lines are the evidence that the app was scaled on purpose, and they feed the capacity statement.
The load test that produces those measurements is the “Run a controlled load test” section of the concurrent-users article from the first section, and the method for load testing a web application, with a k6 load testing example to run, is its own subject.
The step log is my working rule: one dated line for each row of the order table, and every line names the number that justified the call. A line reads like “done: the dashboard query stopped scanning the whole table” or “skipped: the pool never ran short during the test”.
| Step | Done or skipped | The measurement | Date |
|---|---|---|---|
| 1. Index on the slow query’s filter or join column | |||
| 2. Query shape: rows, columns, one query per page | |||
| 3. Connection pool | |||
| 4. Cache for repeated reads | |||
| 5. Long work in a background job | |||
| 6. Bigger box | |||
| 7. Second box behind a load balancer |
Once all seven lines are filled, the next document is a capacity statement: it names the workload it was measured on and marks anything beyond it as a projection, which is what a capacity plan is for a small app. Keep the filled step log in the repository, next to the code it describes.
Where the sprint does this
In the Production Hardening Sprint, we do this as deliverable 9.2: simulate concurrent users, identify the first bottlenecks, fix them, and rerun the workload, then report the workload, duration, environment, concurrency, latency, and error rate before and after changes. Deliverable 9.1, appropriate caching, caches repeated reads where it is safe, with explicit invalidation and customer-data boundaries, and is verified by comparing cache hits and misses, testing invalidation, and verifying tenant separation. Deliverable 9.6, the written capacity statement, documents measured concurrent capacity on the current infrastructure and the changes needed to plan for five times that workload, with capacity claims linked to load-test evidence and untested projections identified as projections. Your app’s current framework and hosting setup are the starting point, and components are refactored or replaced where the production work requires it. Each of the three, with its verify line, is in the published scope.
Common questions about scaling a web app
How to scale from 0 to million users?
One step at a time, through the same order: index and query first, then the pool, a cache and background jobs, then a bigger box, then more boxes, with a measurement between each. The one-app layout of a database with a pool, a cache, a queue, a CDN and the host’s balancer holds until a second database is needed, in my reading. The arithmetic that turns users into requests belongs to the concurrent-users article.
What are scalable web applications?
Scalable web applications are apps whose response time stays level as traffic and data grow. That property is measured, not declared: a load test at the workload you expect, repeated after each change, is what shows it.
Is NGINX just a load balancer?
No. NGINX describes itself as “an HTTP web server, reverse proxy, content cache, load balancer, TCP/UDP proxy server, and mail proxy server”, so balancing is one of several jobs it does.
How do I tell if my CPU is being throttled?
Read nr_throttled and the throttled time in the container’s cpu.stat file, throttled_time on cgroup v1 or throttled_usec on cgroup v2, and watch whether they rise under load. That needs a shell inside the container; on a managed host without one, compare the host’s CPU graph with the plan’s CPU limit.
What are the four types of load balancers?
On AWS, the four are the Application, Network and Gateway Load Balancers, which Elastic Load Balancing lists as its current generation, plus the Classic Load Balancer, which AWS calls “the previous generation” and recommends migrating away from. That is AWS’s own list. For one app, the balancer that matters is the one your host already runs.
If you have a working app built with these tools and need it ready for real customers, this is what we do.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase