The average performance and scale score in my June and July 2026 audits was 53.3 out of 100, across 21 third-party apps. Web performance optimization is making the pages load fast, the API answer fast, and the app hold up when many people use it at once. For one app it comes down to seven controls, and each has a test.

The web performance optimization techniques that matter for one app

Web performance optimization for one app comes down to seven controls: deliberate caching, a load test with the first bottleneck fixed, a split bundle that loads routes when needed, sized and compressed images, a Lighthouse run against written targets, a capacity statement backed by the load test, and core flows that meet WCAG AA.

The 21 apps behind that opening average are a selected set of third-party codebases I audited, each scored on this pillar, not a random sample, so 53.3 is not a rate for AI-built apps in general. This area holds 7 of the 123 checks in a full production hardening pass, and it answers the question every such pass asks about load: how the app behaves when many people use it at once.

The familiar front-end work has a course of its own: web.dev’s Learn Performance has modules on image performance, code-splitting JavaScript and lazy loading. In web application performance optimization, that front-end work maps to controls 3 and 4 below, with caching as part of control 1. Web performance solutions for a live app go further, to the load test and the capacity statement built on it: neither MDN’s web performance index nor the module list of web.dev’s course mentions load testing. Putting accessibility in the same pass is my own addition.

The table below and the afternoon list at the end together make a web performance audit checklist for the whole app; rows 3, 4 and 5 on their own are the frontend performance optimization checklist. Web performance management means going through the same list again after every large change; that is my working rule. The “number it moves” column is my own mapping to the five numbers in the next section.

ControlWhat it isWhy it mattersThe number it movesPage to open
1. CachingReads that repeat are served from a cache that expires on purposeRepeated expensive work wastes resources; unsafe caches can expose stale or private dataAPI response timecaching strategies web applications
2. Load testMany simulated users against staging, the first slowdown fixedLaunch capacity should be measured before real traffic becomes the testThroughput under loadload testing a web application
3. BundleThe first page ships only the code it usesLarge initial downloads slow first use, especially on constrained devicesPage experiencehow to reduce JavaScript bundle size
4. ImagesPictures sent at the size they are shownOversized assets waste bandwidth and slow customer-facing pagesPage experienceimage optimization checklist for websites
5. LighthouseA lab run on chosen pages against written targetsSlow or difficult-to-use pages create avoidable frictionPage experience and the accessibility resultLighthouse audit
6. Capacity statementA written limit with the test behind itA number without its conditions is a poor basis for growth decisionsCapacitywhat is a capacity plan
7. AccessibilityCore flows usable by keyboard and screen readerEnterprise and public-sector buyers require it, and generated interfaces fail it by defaultThe accessibility resultaccessibility audit

What performance means in software: the five numbers, and who each one is for

Performance in software is how fast and how far an app holds up, measured as five numbers: page experience as the three Core Web Vitals, API response time, throughput under load, capacity as concurrent users before the first failure, and the accessibility result. Customers feel the first two; investors and buyers ask for the rest.

Performance in software engineering can reach down to an algorithm’s running time or a program’s memory use. For a live web app, I narrow software performance to five web application performance metrics, because each one answers a question someone outside the codebase asks. The first comes from web.dev’s Core Web Vitals page: Largest Contentful Paint (LCP) “measures loading performance”, Interaction to Next Paint (INP) “measures interactivity”, and Cumulative Layout Shift (CLS) “measures visual stability”. The “who it is for” column is my reading of who asks.

NumberWho it is forWhere it is measuredThe page
Page experience (LCP, INP, CLS)A customer on the pageIn the browser, on the pages you choseCore Web Vitals for AI-built apps
API response timeA customer waiting on a clickOn the server, per route, with percentilesAPI response time and tail latencies
Throughput under loadAn investor or a buyer’s technical reviewerIn a load test: requests per second at a stated error rateControl 2 below
CapacityAn investor planning growthIn the same load test: concurrent users before the first failureControl 6 below
Accessibility resultA buyer’s procurement checkAn automated audit plus a keyboard-only walkControl 7 below

What goes wrong without it

The demo was fast; the launch crowd made it slow

Your app feels quick in every demo with a handful of test accounts, and then the first busy day does not. Nobody measured how much it could take, so real users do the measuring. HealthCare.gov’s launch is a public case, set out in the HHS inspector general’s 2016 case study. The report says one CMS technical official “characterized the launch itself as a test of the system”, and that limited performance testing on September 26, 2013 found the site “could support far fewer concurrent (simultaneous) users than planned.” On October 1, 2013 the site “experienced 250,000 concurrent users, much greater than the planned capacity,” outages began within 2 hours, and the report adds the problems were “not caused solely by a higher number of visitors” but also by core problems in website performance. The lesson I take from it, my opinion rather than the report’s: the first real load test should not be the launch. The fix is load testing on staging before launch day; the controlled load test section of why an AI app stalls at 100 concurrent users sets out a first run against staging, and when one server is not enough, scaling web applications is the next decision.

The first page is heavy before anyone clicks

On a phone, the landing page shows a blank screen or a spinner before anything can be tapped. The browser has to fetch and run everything the first page asks for, including code for screens the visitor never opens and photos far larger than the boxes they sit in, and slower phones on slower networks pay for it most. The fixes split in two: a smaller JavaScript bundle is control 3, and properly sized images are control 4.

Every page asks the database for the same rows

The database dashboard shows the same query running on every page view, and response times climb as users arrive. Nothing keeps a copy of reads that rarely change, so each visitor pays for work the app has already done. A cache added in a hurry brings the opposite risk, the one the table’s first row names: stale or private data served to the wrong person. Caching strategies differ in where the copy lives and how it expires, and control 1 below sets the bar; the caching section of the concurrent-users article above covers when a cache raises the ceiling and when it hides a slow query.

The process was killed for memory

The app restarts under load, and the host’s log shows the process was stopped for using too much memory. On a Linux host, containers included, an OOM kill is the kernel killing processes to free up memory when it detects there is not enough, and the first job is finding what grew. On AWS Lambda, each concurrent request gets its own copy of the function’s environment, so the question becomes Lambda concurrency rather than one machine’s memory. Choosing between a bigger machine and more machines is horizontal scaling vs vertical scaling.

Nobody can say how many users it can take

An investor or a buyer asks how many users the app can take, and the honest answer is a guess. A figure with no workload, environment or test behind it cannot be checked, so nobody can plan hiring, hosting or a launch around it. For a small app, a capacity plan starts as a short written statement built from the load test’s numbers, control 6 below.

A keyboard user cannot check out

Someone who moves through the page with the Tab key reaches the checkout, and the focus vanishes or a button has no name a screen reader can announce. Generated screens fall short of WCAG AA by default unless someone checks them, and the table’s last row names the buyers who require it. An accessibility audit of signup, checkout and the main task finds these before a buyer’s reviewer does.

The seven controls, one by one

1. Caching is deliberate, invalidated, and tenant-safe

Cache repeated reads where it is safe, with explicit invalidation and customer-data boundaries. In practice that means deciding which reads are worth a copy (a pricing page, a settings lookup, a slow report), what event clears each copy, and whether the cache key includes the customer or user, so one account’s cached answer never reaches another. The test: compare cache hits and misses, test invalidation, and verify tenant separation. Choosing among caching strategies for web applications starts with the layer: the browser, a CDN or the server.

2. A load test has been run and the bottleneck fixed

Simulate concurrent users, identify the first bottlenecks, fix them, and rerun the workload. Run it on staging, with payment and email providers stubbed or in sandbox mode, and check your host’s rules on load tests before you start. For an API with no front end, API load testing hits the routes directly. The tool matters less than running one: a k6 load testing example, Locust load testing and Artillery load testing all do the job, and picking between JMeter, k6, Locust or Gatling is a separate decision. The test: report the workload, duration, environment, concurrency, latency, and error rate before and after changes.

3. The bundle is split and routes load when needed

Reduce unnecessary frontend code and load routes or features when needed. In practice that can mean the admin area, the charting library and the rich text editor stop shipping with the landing page and arrive when someone opens them. The same build should minify that code, and the server or CDN should compress text responses such as JavaScript, CSS and HTML. The test: compare built bundle sizes and test route loading with representative network conditions. Reducing JavaScript bundle size starts with reading the build output.

4. Images are sized, compressed and served right

Optimize image dimensions, formats, compression, and delivery while preserving useful quality. One case: a photo uploaded straight from a phone camera and shown in a small card, sent at full size to every visitor. Resize on upload or through an image service, pick a modern format, and serve from a CDN with sensible caching. The test: compare transferred sizes and inspect rendered images at intended display dimensions. The second half of that test matters: a smaller file that looks blurry on the product page is not a win. An image optimization checklist for websites turns this into a page-by-page pass.

5. Lighthouse has been run with targets, on a named device profile

Run Lighthouse on representative pages, implement performance and accessibility improvements, and deliver a passing result against recorded acceptance targets. Representative means the pages a customer actually lands on: the home page, signup, and the busiest screen after login. Write the targets down before the first run, so the result is a pass or a fail rather than a number you argue about afterward. The test: report the chosen pages, device profile, numerical targets, and final results, and check the key accessibility interactions yourself. A Lighthouse audit set up this way repeats cleanly: same pages, same profile, same targets.

6. There is a written capacity statement backed by the load test

Document measured concurrent capacity on the current infrastructure and the changes needed to plan for five times that workload. The statement names the workload, the environment and the number of concurrent users the app held before the first failure, then lists what would need to change to plan for five times that load. A figure the test did not reach is never quoted as a limit. The test: link capacity claims to load-test evidence and identify untested projections as projections. A capacity plan starts from this statement.

7. The core flows meet WCAG AA and pass a keyboard walk

Bring the core flows to WCAG AA: keyboard navigation, contrast, labels, and focus order. Core flows are the few paths that make or lose money: signup, checkout and the main task the app exists for. Contrast and labels show up in an automated audit; focus order and keyboard traps usually only show up when a person tries it. The test: score 100 on an automated accessibility audit of the core pages and complete a keyboard-only walk through signup, checkout, and the main task. An accessibility audit of those three flows is the place to start.

How to verify the whole area in an afternoon

Verifying this area takes seven tests and, as my working estimate, about an afternoon: Lighthouse on chosen pages, an accessibility audit with a keyboard walk, bundle sizes and route loading on a representative network, image transfer sizes, a cache hit and an invalidation, a load test before and after the fix, and a capacity statement written from it.

The order below is my working order for a one-person team, cheapest first, with the load test near the end because it needs staging ready. Each item points back to the test written under its control; this list says what to record.

  1. 01 Control 5, Lighthouse: write the page list, the device profile and the targets first, then run it; record the scores beside the targets.
  2. 02 Control 7, accessibility: run the automated audit on the core pages, then do the keyboard-only walk; record the score and any step where focus was lost.
  3. 03 Control 3, bundle: build, note the bundle sizes, and load one route on a throttled network; record the sizes and load time before and after.
  4. 04 Control 4, images: note the transferred image sizes on your heaviest page, before and after; check the images still look right at their display size.
  5. 05 Control 1, caching: record hits and misses, change one cached record to prove it clears, and try a read across two accounts that must fail.
  6. 06 Control 2, load test: run it on staging with third parties stubbed or in sandbox mode, before and after the fix; keep the full report.
  7. 07 Control 6, capacity: write the statement from those load-test numbers, with every untested figure labeled as a projection.

These seven sit inside the full production readiness checklist, next to the security, data and release checks that a launch also depends on.

Where the sprint stops

In the Production Hardening Sprint, your app’s current framework and hosting setup are the starting point, and we refactor or replace components where the production work requires it. Formal third-party certifications and independent audit opinions are separate from these engineering deliverables. Hosting, paid tools and API usage remain in your accounts, and we explain any required third-party costs before enabling them.

Where the sprint does this

Area 9 of the sprint is these seven controls, 7 of its 123 deliverables, each one delivered and checked by the test written under its control above. Every result goes into the production readiness report with its verification evidence; that report accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. Each control and its check are listed in area 9 of the published scope.

Common questions about a faster app

What are the top 3 website performance metrics to monitor?

The top three are the Core Web Vitals, which web.dev says “should be measured by all site owners”: Largest Contentful Paint for loading, Interaction to Next Paint for interactivity, and Cumulative Layout Shift for visual stability. For an app with an API, I’d watch the API’s response time as well, the second number in the five-numbers table.

What does performance optimization do?

Performance optimization removes the waiting a customer notices and moves the point where the app starts to fail further out. On the front end, pages appear and respond sooner; on the back end, the app keeps answering as more people use it at the same moment. Each gain is judged against a number recorded before the change, which is why every control above comes with a test.

What is the first rule of optimization?

Measure before you change anything: that is my working rule, not a quotation. Run Lighthouse on your chosen pages and the load test on staging (controls 5 and 2) first, so each fix is judged against a before number rather than a feeling.

Can you give me an example of performance monitoring?

One example is a scheduled check from outside the app that requests one chosen page and one API route, records how long each takes, and alerts a person when either goes over the target you wrote down. Running the full list again after each large change is the web performance management described near the top of this page. Who gets the alert, and how, belongs to the app’s alerting setup.