JMeter’s and Locust’s home pages explain how to generate traffic. What they leave out is the rest of load testing a web application: how much load to send, which number counts as a fail, and which bottleneck to fix first when the run fails. Those three decisions are what this page is about.

Load testing a web application: what it is, what it measures, and what it does not

Load testing a web application means sending simulated, signed-in users through the journeys real customers take, at the numbers you expect and a step beyond, and recording latency, error rate and throughput for each step. It measures per-user, database-bound work, which a speed test of one page never reaches.

A website and a web application are different jobs here. Load testing a website made of brochure pages mostly measures static pages that a CDN can cache, while web app load testing measures logged-in, per-user work that ends in the database, and that is where load lands. It is one of the performance checks in web performance optimization, next to caching, bundle size and query speed.

For SaaS applications, performance testing usually starts because someone outside asks for a number: a customer with a seat count in the contract, or an investor asking how many users the app can take. Load testing is how a SaaS answers with a measured figure instead of a guess.

What you want to knowDoes a load test measure it?Where it is measured instead
Latency, error rate and throughput per journey step as users growYesThis test
Front-end speed for one visitorNoA speed test of one page load
Whether each feature works correctlyNoFunctional and end-to-end tests
Whether the app is secureNoA security review and a targeted security test
Capacity above the load you actually ranNoNowhere: an untested figure is a projection

The second row is the easiest to mix up. PageSpeed Insights provides “both lab and field data about a page”, which describes how the page loads for a visitor, not what the server does with many visitors at once. Choosing among speed tools belongs to Core Web Vitals for AI-built apps. Testing only the endpoints, with no pages and no browser session, is the narrower job of API load testing.

What is load testing, in software testing and for a web app

Load testing is a type of performance testing that checks how a system behaves under the load you expect: simulated users make real requests while latency, error rate and throughput are recorded. It matters because without it, the first test of that load is real customers on the busiest day.

In software testing, load testing is the kind of performance testing that answers one question: does the app hold at the load you expect? Load testing is important for a plain reason of cost. A test costs a staging copy and some machine time, and a failed launch costs the customers who tried to sign up that day.

A run records three numbers at each step as the user count climbs: latency, read as percentiles rather than an average; error rate; and throughput, meaning the requests per second the app actually completed. The test types below change the shape and length of the load. The three numbers stay the same.

Load, stress, soak and spike: examples of performance testing, and the question each one answers

Performance testing comes in four common shapes. A load test holds average production load. A stress test goes above average to see how the system manages. A soak test holds average load for hours. A spike test adds a sudden, short, very high burst. Names vary by tool, and k6 adds smoke and breakpoint tests.

These four are the types of software performance testing this page uses, but the names are one tool’s words, not a standard. k6’s guide to load test types says “no consensus even exists about the names of these test types”. The table takes its load and duration columns from k6’s cheat sheet; the question and the sample are mine, as examples of performance testing on a small SaaS.

TestLoad (k6’s words)Duration (k6’s words)The question it answersA sample for a small SaaS
LoadAverage productionMid (5-60 minutes)Can we take a normal busy hour?The two main journeys at the busiest hour’s arrival rate
StressHigh (above average)Mid (5-60 minutes)Where is the ceiling, and how does it fail?The same journeys stepped up until the fail line breaks, recording what broke
SoakAverageLong (hours)Does anything leak: memory, connections, disk, a queue?The load test’s rate held for hours while memory and open connections are watched
SpikeVery highShort (a few minutes)What happens when arrivals jump in seconds?Arrivals jump from normal to the highest rate the request logs show, then fall back

A soak test is the easiest to skip, and it is the one that finds a connection that is never returned or a queue that never drains. Stress testing web applications answers a different question from a load test: not whether the app holds, but where the ceiling is and what the failure looks like when it arrives. Breakpoint tests, in k6’s words, “gradually increase load to identify the capacity limits of the system”.

A performance testing sample only means something if its numbers come from the app’s own request logs, so size a spike from the highest arrival rate those logs show, never from the name of whatever sent the traffic. The article on what to do when an app goes viral says the same: treat the source label as context, not as a duration or a workload model.

On a small SaaS I’d run them in this order: load first, spike second, stress third, soak last. Load tells you whether an ordinary busy hour works, the spike tells you whether launch day works, and stress and soak earn their time once those two pass. Deciding which kind of growth is coming (a spike, a step or a ramp) is part of the scaling readiness checklist for startups.

What goes wrong without it

Each failure below looks like bad luck on the day, and each has a test that would have shown it earlier.

Without a load testWhat the users or the business seeWhat the test would have shown
The launch is the first load testCustomers find the ceiling, on the busiest dayThe same ceiling, on a staging copy, before launch
The upgrade that did not helpA bigger plan, and the same slowdownWhich resource saturated first, before money went on more of the wrong one
The number nobody can backA customer or investor asks how many users the app takes, and the answer is a guessA measured figure with its conditions written next to it
The test that measured the wrong thingA clean benchmark, then a slow launchThat the signed-in journeys were never exercised

The second row is why the advice in the viral-app article is to find the bottleneck before buying capacity: the resource that saturated first is the one to fix.

The last row, as a case. Before a launch, a founder points a command-line HTTP benchmark at the site’s home page, a tool its own docs describe as designed to give “an impression of how your current Apache installation performs”, and gets a fast, clean result. The home page is cached at the edge, so the responses carry Cloudflare’s cache status HIT, “The resource was found in Cloudflare’s cache”, and the app’s servers barely see the test; the signed-in dashboard, which no request in the test touched, is the part launch traffic reaches. Launch capacity should be measured before real traffic becomes the test.

I audited 21 third-party apps in June and July 2026: 11 public vibe-coded apps audited across all 12 pillars, then 10 held-out apps the audit method had never seen, audited blind. Across those 21, the Performance & Scale pillar averages 53.3 out of 100, ninth of the 12 pillars counting from the weakest. Those 21 are a selected set of apps I audited, not a random sample, so the figure is not a rate for AI-built apps in general.

If your app is already slowing down, start with why an app slows down under concurrent users, which maps each symptom to the limit behind it.

How to do load testing: six steps from one user journey to a fixed bottleneck

Load testing is done in six steps: model real user journeys with think time, size the load and write the fail line first, spike signup and the CDN, record error rate and percentiles per step, find the first bottleneck, then fix that one thing and rerun the identical test.

The article on why an app slows down under concurrent users has a short controlled-load checklist for the same job. This is the full method, and it does not repeat that checklist. The six steps are the performance testing strategy I’d use for a web app with one team and one database, and they apply whether you perform load testing on a website with a login or on an internal tool. Load testing best practices, such as randomized think time and a fail line written before the run, appear inside the step they belong to, and performance testing web applications this way gives you the same report whichever tool runs it.

Model one real user journey, not one URL

To load test a web application, start with what one person does, not with a URL. A virtual user is a script of a visit: sign in, land on the dashboard, open a list, open an item, do the core action, and maybe pay. Two or three journeys cover most apps (the reader who browses, the writer who creates, the new sign-up), and each gets a share of the virtual users that matches your analytics.

StepRequest or requestsThink timeData it needsShare of users
Sign inThe login request, keeping the session cookie or tokenA few seconds, randomizedA test account from a pool, one per virtual userEvery journey
Land on the dashboardThe dashboard page or its data callsA few secondsAn account shaped like a real customer’sEvery journey
Open a listThe list endpoint at its default page sizeSeveral seconds, as a person scansLists as long as your biggest customer’sReaders and writers
Open an itemThe item’s page or API callSeveral secondsItems with real attachments and historyReaders and writers
Do the core actionThe write: create, update or sendLonger, as a person typesThe CSRF token and form fields carried from the pageWriters
Pay, if the journey paysThe checkout hand-off in the payment provider’s test modeAs long as entering a card takesTest-mode card detailsA small share of writers

Here is the journey as a tool-neutral outline. It is pseudocode, not any tool’s syntax:

journey "writer" (share: the writer share from your analytics)
  setup: take a test account from the pool; sign in; keep cookie and CSRF token
  GET  /dashboard              -> expect 200 and the account name in the body
  wait random(think time)
  GET  /projects               -> expect 200 and a non-empty list
  wait random(think time)
  GET  /projects/{random id}   -> expect 200
  wait random(longer think time)
  POST /projects/{id}/items    (with CSRF token) -> expect 201 and no error field
  record per step: latency, status, body check passed or failed

Think time is what makes the users human. Without it every virtual user fires requests back to back, and a handful of them send the traffic of a crowd, so the test measures a robot. Randomize it, or the virtual users fall into step and hit the server in waves.

Logged-in state takes the most setup. Give each virtual user its own test account, or draw from a pool, sign in during setup, and carry cookies and CSRF tokens the way the browser would. Size the test data like production, because a list of a dozen rows proves nothing about the list your biggest customer opens. Leave static assets to the CDN and out of the script, unless the CDN is what you are testing.

Protocol-level scripts, plain HTTP requests, are the default. Browser-level load is for the page where client rendering is the question, and k6’s hybrid approach to browser load describes an option that is “much less resource-intensive”: “combining a small number of virtual users for a browser test with a large number of virtual users for a protocol-level test”. For a runnable script in one tool, a k6 load testing example turns this outline into code.

Size the load, set the fail line, pick the place

Turn users into requests before choosing a number. The article on why an app slows down under concurrent users has that arithmetic in its section on turning users into requests, so I won’t repeat it. My starting rule for the target is about twice the busiest hour you have seen, or the seat count in the customer’s contract, whichever is larger.

Write the fail line before the run, per journey step: a p95 latency, an error rate, and no server-side resource above about 80 percent, which is my working rule, not a standard. A line drawn after the run always passes, because it gets drawn around whatever happened.

Run against staging with production-sized data, with third parties stubbed or in their sandbox mode, and put the generator on a cloud machine near the app, not on a laptop over home broadband. Two web server load test best practices I hold to belong here too: the generator never runs on the server it is testing, and the test never runs against production. Only point a load test at systems you own or have written permission to test, and read your host’s policy on load tests before the first run; the arrival-rate settings and host rules for endpoint-only tests belong with API load testing.

Test burst traffic against signup, and test a load burst at the CDN

Signup is the journey a launch hits hardest, and it carries the most limits that are not yours: the auth provider’s rate limits, the email provider’s sending limit, a captcha, and a password hash that is slow on purpose. Test burst traffic against signup with a spike of new sign-ups using generated addresses on a test domain, in the email provider’s sandbox or test mode where it has one, and record where it stops: which limit, and what the new user saw.

On Amazon SES, a new account starts in the sandbox, where “You can only send mail to verified email addresses and domains, or to the Amazon SES mailbox simulator”, with a maximum of 200 messages per 24-hour period and 1 message per second. There, the generated addresses sit on a verified test domain or go to the mailbox simulator, whose mail does “not count toward your sending quota or your bounce and complaint rates” but is still “limited by your account’s maximum sending rate”. So a signup burst in the sandbox that stalls at 1 message per second has found the sandbox, not your app. When the limit really is yours, the fix I’d try first is a friendlier failure and a queue for the email, not more capacity.

To test a load burst at the CDN, send a spike at the public pages and the app shell and read the cache-status header on every response. On Cloudflare that is CF-Cache-Status, which “indicates whether a resource is cached or not”; Cloudflare’s cache status values include HIT, MISS (eligible, but served from the origin) and DYNAMIC (sent to the origin without a cache lookup). Expect DYNAMIC on HTML and JSON unless a rule caches them, since Cloudflare does not count those among its default cached file extensions.

Count the share of HIT responses among the assets you expect to be cached, and that is the hit ratio for the burst. A low ratio means the origin took the burst. Deciding what to cache, and keeping one customer’s page out of another customer’s response, is part of caching strategies for web applications.

Measure error rate during load test, and write the web server load test report

To measure error rate during a load test, divide errors by requests for each journey step, split them by status code, and read them over time rather than as one total. Time matters because of how errors cluster. In my illustration, 0.5 percent spread evenly across a run is noise, while 0.5 percent that is every request for 9 seconds of a 30-minute run is an outage.

Two kinds of error never reach the server’s logs. Timeouts and refused connections happen at the generator, so count them there. A success status with an error message in the body is an error too, so check the response content, not only the status code.

The web server load test report has six fields, each recorded before and after the fix, plus the fail line and whether it held. Record latency as p50, p95 and p99, since p95, p99 and tail latencies show the queue that an average hides. Keep it to one dated page, with the tool’s raw summary attached.

FieldWhat to writeBefore the fixAfter the fix
WorkloadJourneys, their shares, the target arrival rate or user countFirst runIdentical to the first run
DurationRamp-up, hold and ramp-down timesFirst runIdentical to the first run
EnvironmentWhere the app and the generator ran, data size, app versionFirst runIdentical except the app version
ConcurrencyThe users or arrival rate actually reachedFirst runSecond run
Latencyp50, p95 and p99 per journey stepFirst runSecond run
Error ratePer step, by status code, over timeFirst runSecond run

A capacity statement cites this report as its evidence. For the wider question of what is a capacity plan, the report supplies the measured part and nothing else.

Find database bottleneck under load, and load test database connection limits

To find a database bottleneck under load, record two things during the run so you can read them afterwards. The first is connections in use against the connection limit, on the same time axis as the tool’s latency graph. The second is the top statements for the test window.

pg_stat_statements handles the second, “tracking planning and execution statistics of all SQL statements executed by a server”. Reset it just before the run, since pg_stat_statements_reset “discards statistics gathered so far”; but “By default, this function can only be executed by superusers. Access may be granted to others using GRANT.” On a managed database whose role cannot reset, take a snapshot of the view just before and just after the run and compare the two.

The view exists only where the module is loaded: “The module must be loaded by adding pg_stat_statements to shared_preload_libraries”, which means a server restart, and it is then enabled per database with CREATE EXTENSION pg_stat_statements. Supabase’s docs say all projects “come with the pg_stat_statements extension installed” and show how to enable it from the dashboard (Database, then Extensions, then pg_stat_statements), but do not state whether the default role can reset it. Confirm it is on before the test day, not during it.

To rank those statements by total time, find the slow query with pg_stat_statements. Matching the pattern to a limit is the which-limit table in the article on why an app slows down under concurrent users. To load test database connection limits on purpose, the ramp, the stop below the hard limit and the values to record at each step are in every connection pool exhausted error and its fix.

The two usual query causes have their own pages: one query per row is the N+1 query problem, and a policy evaluated per row is Supabase RLS performance. The one thing I’d add from this page is to run the test with your biggest tenant’s data, not the average tenant’s, because the biggest account is where a per-row cost shows first.

Thread blocking and a blocked event loop: the bottleneck a load test surfaces first

Thread blocking is a thread waiting on I/O or a lock, or busy on CPU, while other work queues behind it. In Node.js one Event Loop thread runs your JavaScript callbacks, so a slow synchronous call holds up every other client. Under load it shows as every route slowing together while the database stays idle.

There are two server models to know. In the thread-per-request model, common on the JVM and in many Python and PHP setups, each blocked thread is one fewer worker, and when the whole pool is blocked, requests queue and time out. In Node.js, Node’s guide to not blocking the event loop describes “one Event Loop” and “a pool of k Workers”, and says: “While a thread is blocked working on behalf of one client, it cannot handle requests from any other clients.”

The guide names the usual offenders: the synchronous encryption, compression, file system and child process APIs (a list it calls “reasonably complete as of Node.js v9”), vulnerable regular expressions (REDOS), and JSON.parse and JSON.stringify on large input. Java questions about threads blocking describe the same idea on the JVM, where code that blocks a thread holds a worker while it waits.

In results it looks like this: every route slows together, including trivial ones such as a health check, while the database sits idle and one CPU core stays busy. To confirm it in Node.js, monitorEventLoopDelay from perf_hooks “samples and reports the event loop delay over time”. On the JVM, take a thread dump during the run: many request threads waiting in the same stack is the sign.

The fixes are the asynchronous version of the call, a worker thread or a background job for CPU-heavy work, which is how to run long tasks in the background, and a timeout on every outbound call, so a slow provider cannot hold your threads.

Fix the first bottleneck, then rerun the identical test

Stop when the fail line holds at the target, not when the graphs look perfect. The second bottleneck only appears once the first is gone, so fix one thing, then rerun with the same script, the same data and the same place, as the short checklist’s one-change rule says. The result is a tested load capacity for the website and the app behind it, with its conditions attached. What to change once the code fixes run out, and in what order, is scaling web applications.

Load testing software tools: what to run the test with

Load testing software comes in five types: scripted open-source generators such as k6, Locust, Gatling and Artillery, the GUI veteran JMeter, managed cloud services that run those engines, commercial hosted platforms, and browser-level tools. For a small app, an open-source generator on one cloud machine covers load and spike tests.

This section sorts web application load testing tools by type and fit; it does not rank them. In software testing, performance testing tools split into the same five types whether the label says load, performance or stress. Which tool is used for load testing matters less than whether the test lives in your repository and models a real journey.

Tool typeExamplesThe test is written inRuns fromFits when
Scripted open-source generatorsk6, Locust, Gatling, ArtilleryJavaScript (k6); Python (Locust); Java, JavaScript, TypeScript, Scala or Kotlin (Gatling); YAML, TypeScript or JavaScript (Artillery)Your own machine or a cloud machine you rentThe team writes code and wants the test in the repository
The GUI veteranJMeterA test plan built in its GUIYour own machine or a cloud machine, run in CLI modeThe team prefers building tests visually, or already has JMeter plans
Managed cloud servicesAzure Load TestingJMeter or Locust scripts, or a single URLThe provider’s cloudYou want the generator machines managed, and the app runs on that cloud
Commercial hosted platformsBlazeMeter, LoadViewCheck the vendor’s docsThe vendor’s cloud, or on-premises where offeredYou want dashboards and vendor support and will pay for them
Browser-level toolsk6 browser, Artillery with PlaywrightBrowser scriptsReal browsers on the generator’s machinesClient-side rendering under load is the question

The commercial row sells performance testing solutions rather than a tool: the generator, the machines and the dashboards in one paid account.

My choice rule has four parts: the language your team already writes, a test that lives in the repository beside the code, support for an arrival rate (new users per second) rather than only a fixed user count, and a free tier or open-source license that covers a small app’s test. The best performance and load testing tools for web applications are the ones your team will keep running, which is why this section gives types and fit rather than a ranking. For a head-to-head choice between JMeter, k6, Locust or Gatling, weigh those four points against your own team.

For SaaS performance testing tools, the first question is whether the tool can sign in and carry a session, because a SaaS journey starts behind a login. Check that with a short trial run of the sign-in step before writing the whole journey in any of the web app load testing tools above. Browser-level tools are the UI load testing tools: they drive real browsers, cost more machine per virtual user, and belong to the page where rendering is the question. Website performance testing tools usually means speed tests of one visitor, a different measurement; the FAQ below draws the line.

Open source load testing tools, and the free ones

Five open source load testing tools cover most small apps. The license and script language come from each project’s own repository or docs.

ToolLicenseScript languageIn its own words
k6AGPL-3.0JavaScript, run by a Go engine”A modern load testing tool, using Go and JavaScript”
LocustMITPython”define your tests in regular Python code”
GatlingApache-2.0Java, JavaScript, TypeScript, Scala or Kotlin”Test scenarios are defined as code”
ArtilleryMPL-2.0 (some Azure modules BSL)YAML, TypeScript or JavaScript”Load test HTTP APIs, GraphQL, WebSocket, and more”
JMeterApache-2.0A test plan built in its GUI”a 100% pure Java application designed to load test functional behavior and measure performance”

All five are open source performance and load testing tools for web applications and API endpoints, and you pay nothing for the software itself, with one exception: Artillery’s repository says some Azure-specific modules are under the BSL license, and “commercial and/or production usage requires a commercial license” on Azure. Locust and Apache JMeter show their licenses on the repository page, as do the other three. Which of these open source load testing tools is best for your team comes down to the script language column.

Free means three different things when you look for free load testing tools for a web application or an API. Open source you run yourself is free software on a machine you pay for. A vendor’s free tier comes with a cap, and the cap is whatever the vendor’s own page says on the day you read it. A free trial is a paid product with a clock on it. For a small web application, free performance and load testing tools cover the load and spike tests: an open source generator on one rented cloud machine, where the machine is the only cost. The same free load testing tools work for a website’s public pages; the logged-in journeys are what need the scripting.

Locust load testing and Artillery load testing both take the journey outline above as a script in each tool’s own format.

JMeter for a web application

To perform load testing on a web application using JMeter, build the test in the GUI and run it from the command line. The same plan does performance testing for a web application using JMeter at any load level; only the thread count changes. These seven steps use the element names as JMeter’s own manual prints them.

  1. 01 Add a Thread Group and set the number of threads (users), the ramp-up period, and a loop count or, with the scheduler, a duration.
  2. 02 Add HTTP Request Defaults with the Server Name or IP filled in, so every request inherits the host.
  3. 03 Add an HTTP Cookie Manager and an HTTP Header Manager, so each thread keeps its own session and sends the headers your app expects.
  4. 04 Add the journey as HTTP Request samplers, recorded with the HTTP(S) Test Script Recorder or written one at a time.
  5. 05 Add timers for think time, such as a Uniform Random Timer between steps.
  6. 06 Add one Response Assertion per step that checks the body, so a 200 with an error message counts as a failure.
  7. 07 Save the plan, run it in CLI mode, and have JMeter write its HTML report at the end of the run.

The manual’s best-practices page says “Use as few Assertions as possible”, which is why step 6 stops at one per step. JMeter’s getting-started manual sets the rule behind step 7: “GUI mode should only be used for creating the test script, CLI mode (NON GUI) must be used for load testing”, and in CLI mode JMeter can “generate an HTML report at end of Load Test”. Each JMeter thread keeps its own cookie store, which is what lets one thread stand in for one signed-in user.

Running JMeter against an API alone, with no pages and no sessions, is API load testing with JMeter. Whether JMeter is the right tool for your team at all is the head-to-head question from the tools section above.

Load testing Azure web apps

For an app on Azure App Service, Microsoft’s own web application load testing software is Azure Load Testing, which the MS docs call “a fully managed load-testing service”. Azure Load Testing “supports running Apache JMeter-based tests or Locust-based tests”, and its quick test creates a load test “by using a URL”, for “a single URL-based HTTP endpoint”. If your application is hosted on Azure, it “collects detailed resource metrics” to help find bottlenecks.

The quick test has the same weakness as the cached home page in the case above: one URL, no session. Use an uploaded JMeter or Locust script for the real journeys. On the App Service side, watch instance CPU and memory, connection limits, and whether your autoscale rules fired or not during the run. The methodology stays the six steps above: MS tooling changes who runs the generators, not how a web application load test is designed.

Spring Boot load testing and microservices performance testing

For Spring Boot load testing, start from the server model: a Spring Boot service handles each request on its own thread unless it is built reactive, so the blocked-thread section above applies directly. Watch the server’s thread pool and the database connection pool together.

Spring Boot’s Actuator metrics cover both, with one catch: Tomcat metrics are instrumented “only when an MBean Registry is enabled”. The registry is disabled by default; setting server.tomcat.mbeanregistry.enabled to true turns it on, and the metrics are then published under the tomcat. meter name. Data source metrics appear under jdbc.connections (active, idle, maximum and minimum connections), and Hikari’s under hikaricp.

Any HTTP tool can drive the test. For a JVM team, Gatling’s documentation describes a tool that fits the language: “a high-performance load testing tool built for efficiency, automation, and code-driven testing workflows”, with SDKs for Java and Kotlin among others.

Microservices performance testing takes two passes. First test one service alone, with its dependencies stubbed, to find its own ceiling. Then run the whole journey end to end. Give every service a timeout on the calls it makes, or one slow service holds the threads of every caller upstream.

Online website load testing and the HTTP load testing tool on your laptop

An online website load testing service usually works like this: paste a URL, prove you own the domain, get a graph. That is fine for an anonymous GET against a public page and weak for logged-in journeys. A free online load testing tool that never asks you to prove you own the domain should not be pointed at anyone’s site, including yours.

A command-line HTTP load testing tool hammers one URL. ab “is a tool for benchmarking your Apache Hypertext Transfer Protocol (HTTP) server”, wrk is “a modern HTTP benchmarking tool”, and hey is “a tiny program that sends some load to a web application”. They are good for comparing one endpoint before and after a fix. They are not a load test of an application, because no real user requests one URL over and over with no session.

How to verify it

A load test is verified by its report: six fields, workload, duration, environment, concurrency, latency and error rate, each recorded before and after the fix, against a fail line written before the run. A run where the generator itself ran out of CPU or virtual users does not count.

Each check below is one a reader can fail, and each names the evidence to keep.

  1. 01 The script is in the repository, and a second person ran it from the README. Evidence: the README section and their run's output.
  2. 02 The report has all six fields, before and after the fix. Evidence: the dated report.
  3. 03 The generator was not the limit: it reached the arrival rate or user count it was asked for, and its own CPU and network stayed below saturation. A run where the generator ran out of CPU or virtual users is void; a throughput that fell because the app slowed is a result. Evidence: the tool's summary and the generator machine's CPU graph.
  4. 04 The fail line was written before the run. Evidence: its date against the run's date.
  5. 05 The first bottleneck is named with its evidence, and the rerun shows it gone. Evidence: the statement list, the connection graph or the event loop delay, from both runs.
  6. 06 Any capacity figure the test did not reach is labeled a projection. Evidence: the report's capacity line says which numbers were measured and which were not.

Keep the tool’s raw output, the host’s graphs for the same time window and the dated report together, in the repository if you can. Load testing is one of several checks that end in a measured result, alongside an accessibility audit.

In the Production Hardening Sprint, deliverable 9.2 is verified this way: “Report the workload, duration, environment, concurrency, latency, and error rate before and after changes”, and deliverable 9.6 is verified this way: “Link capacity claims to load-test evidence and identify untested projections as projections”.

Where the sprint does this

Deliverable 9.2, load test and bottleneck fixes, is where we “Simulate concurrent users, identify the first bottlenecks, fix them, and rerun the workload”, for the reason in the launch case above. The written capacity statement, deliverable 9.6, is to “Document measured concurrent capacity on the current infrastructure and the changes needed to plan for five times that workload”. Both land in the production readiness report, deliverable 13.1, which is verified this way: “Account for all 123 IDs; keep failures visible until resolved and explain genuine non-applicable items”. Your app’s current framework and hosting setup are our starting point, and we refactor or replace components where the production work requires it. Hosting, paid tools, and API usage remain in your accounts. Every deliverable is listed in the published scope.

Common questions about load tests

How can I test the load speed of my website?

Use a speed test, which measures one visit rather than many visitors at once. PageSpeed Insights “reports on the user experience of a page on both mobile and desktop devices” and gives lab and field data for that page. A load test is a different measurement: many visitors at the same time, and what the server does under them. Which speed tool to read is a Core Web Vitals question, and raising a Lighthouse score is the job of a Lighthouse audit.

Can you give me an example of load testing?

Yes. Here is my illustration with round numbers, not a real app: a SaaS whose busiest hour had about 400 signed-in users tests its two main journeys at about twice that, for about 30 minutes, with the fail line (a p95 and an error rate per step) written first. It records the six report fields, finds the first bottleneck, fixes that one thing, and reruns the same test to show the fail line now holds.

Is JMeter a free tool?

Yes. JMeter is open source under the Apache-2.0 license. What it costs is the machine it runs on and the time to build a test plan that signs in and follows a real journey.

Is Gatling free to use?

Yes, the open-source Gatling is under the Apache-2.0 license. Gatling’s docs also describe a separate Enterprise Edition that “extends the Community Edition capabilities” and invite you to “Try Gatling Enterprise Edition free for 14 days”.