The Performance & Scale pillar averaged 53.3 of a possible 100 on 21 third-party apps I scored in June and July 2026. API load testing measures your app under traffic: by my working rule, pick 3 to 5 routes users hit, ramp to about twice your busiest hour, hold about 10 minutes, and record p95 latency, error rate and throughput.
API load testing: what the test is, and the five things to decide first
An API load test sends a planned, realistic amount of traffic at the API and records how latency, errors and throughput change as the traffic grows. Five decisions come first: the question, the routes and their mix, the target load and ramp, where it runs, and the numbers that count as a fail.
That 53.3 is an average pillar score, not a pass mark, and the pillar was scored on all 21 third-party apps in my June and July 2026 audits. Those apps are 11 public apps audited across all 12 pillars and 10 held-out apps my review engine had never seen, audited blind. They are a selected set of audited apps, not a random sample, so the average describes them and is no rate for AI-built apps in general.
Load testing for an API is the server-side measurement in web performance optimization, and it stays a measurement, not a pass or fail, until you draw the line. These five decisions turn it into an answer:
| The decision | What to write down | The mistake it prevents |
|---|---|---|
| The question | One of: can we take the launch, where is the ceiling, did the fix work | Stopping at the first green run because nobody said what it had to prove |
| The routes and their mix | 3 to 5 named routes and the share of traffic each gets | Testing the health route, which touches nothing that breaks |
| The target load and ramp | Requests per second, how fast it climbs, how long it holds, and the data behind it | Testing with one user’s data and ten rows |
| Where it runs | The target environment, where the generator sits, what is stubbed | Testing from a laptop on wifi |
| The fail line | p95 and error rate per route at the target rate | Deciding after the run what counted as slow |
An API load test does not see the browser half of speed: rendering, scripts and images belong to a Lighthouse audit and the image optimization checklist for websites. The whole-app version, with pages and logged-in journeys, is load testing a web application, which also names the load, stress, soak and spike tests.
Why it matters for a small app: what breaks first under load
In the same audits, 13 of the 21 third-party apps had no rate limiting on their most expensive endpoint. That count describes the apps I picked to audit, not every app, but the risk is easy to picture: the route that costs the most per request has nothing limiting how often it runs, and a load test is the cheap way to learn what it can take before real traffic does.
In my reading, four things give way first in an AI-built API. None of them shows in a demo, because a demo has one user and ten rows. The middle column is the part to learn: it is how each one looks in a load test’s results.
| What breaks first | How it shows in the results | Why a demo never showed it |
|---|---|---|
| The database connection limit | Errors jump while the latency of successful requests stays flat | One user holds one connection at a time |
| A query that grows with the table | p95 climbs as the data grows, not as the users grow | The demo table has ten rows |
| A third-party call made inline | Route latency follows the provider’s, and its rate limit becomes yours | One user never reaches the provider’s limit |
| The serverless concurrency or timeout ceiling | Flat, then a cliff: timeouts or throttling at one rate | One user never reaches the ceiling |
The connection errors in the first row each have their own error string, and every connection pool exhausted error and its fix matches each string to its layer. Which of these ceilings your hosting platform reaches first is its own question, answered by the article on why an AI app stalls under concurrent users.
How it works: the workload, the place, the numbers, the tools and the fix order
Five steps, in the order you do them: design the workload, choose where it runs, decide what to record, pick a tool, and read the results.
The workload to send: routes, mix, data and ramp
The workload for a small API, as my working rule, is 3 to 5 real routes in their real mix, sent at an arrival rate of about twice the busiest hour, ramped up over a couple of minutes and held for about 10, against production-sized data with one token per virtual user.
Pick the routes from real traffic, not from the route file: the busiest read, the busiest write, the most expensive route (search, an export, the AI call) and login. Take the mix from a day of logs or your host’s request panel. Then state the target as requests per second rather than as a number of users. k6’s open and closed models page explains why: “in a closed model, the start or arrival rate of new VU iterations is tightly coupled with the iteration duration.” A fixed group of virtual users waits for each response, so a slowing server quietly lowers the load the test sends, just when you wanted it held. k6’s API load testing guide adds that API load “is generally reported by request rate”, per second or per minute.
| Line | How to get the number | Worked example (assumptions stated) |
|---|---|---|
| Routes | The busiest read, the busiest write, the most expensive route, login | List projects, create a task, search, login |
| Mix | Each route’s share of a day of logs or the host’s request panel | 60, 25, 10 and 5 percent (assumed) |
| Busiest-hour rate | Requests in the busiest hour divided by 3,600 | 5 requests per second, read from the logs (assumed) |
| Target rate | About twice the busiest-hour rate, as my working rule | 10 requests per second |
| Ramp, hold, ramp down | A couple of minutes up, about 10 held, about 1 down, as a starting shape | 2 minutes up to 10 per second, 10 minutes held, 1 minute down |
| Test data | More than one account and realistic row counts, seeded in the test environment | As many test accounts and rows as production holds today |
| Authentication | One token per virtual user, fetched in setup | Tokens for the test accounts fetched before the first request, so login gets only its own share of the load |
The third column is a worked API load testing example, not a recommendation: your logs set your rates. If you start from a user count instead of logs, converting users into requests is a step of its own, and the method to run a controlled load test covers that step and the replay around it.
The held part of that worked example, written as one k6 scenario:
export const options = {
scenarios: {
hold_at_target: {
executor: 'constant-arrival-rate',
rate: 10, // iterations started per timeUnit
timeUnit: '1s', // so 10 per second
duration: '10m',
preAllocatedVUs: 20, // VUs ready before the start
maxVUs: 50, // k6 may add VUs up to this to keep the rate
},
},
};
With one request per iteration, the rate is requests per second; leave out any sleep at the end of the iteration, because the arrival-rate executors already pace it. This executor keeps the load constant; k6’s guide says to “use the ramping-arrival-rate executor instead” to ramp the rate up or down, which is how the two-minute climb and the one-minute descent are written.
Where to run it, and what to point it at
A load test runs only against systems you own or have written permission to test. Staging with production-sized data is the default target, with every third party stubbed or in its sandbox mode: payments, email and model calls. Read the host’s load testing terms before the first run.
| Target | Safe when | What to stub | Who to tell first |
|---|---|---|---|
| Staging with production-sized data | The default for every run | The payment provider (test mode), the email provider (sandbox), the model provider (a fake) | Your team |
| Production in a quiet hour | Only when staging cannot match the database size, at a capped rate, with a kill switch | The same three, wherever the code allows it | Your team, whoever watches the alerts, and the host if its terms ask |
| Your laptop | Writing and debugging the script, never the run of record | Everything | Nobody |
| A system you do not own | Never, unless the owner gave written permission | Not your call | The owner, in writing, before anything runs |
Stubbing matters because the third parties’ rate limits and bills are real even when your traffic is not. The hosts’ terms matter for the same reason. Vercel’s fair use guidelines list “Load Testing without authorization” under “Never fair use”. The Amazon EC2 Testing Policy covers high-volume tests sent from EC2 instances, “sometimes called stress tests, load tests, or gameday tests”; it asks that your endpoints sit in the local AWS Region if they are hosted within AWS, and says “Volumetric network-based DDoS simulations are explicitly prohibited from the Amazon EC2 platform.”
Run the generator from a cloud machine, not a laptop on home wifi, and write down where it ran; k6’s guide asks you to “ensure that the location of the load generator is constant across test runs, and avoid running the tests from locations that are too close to the SUT”, the system under test. Then decide what your own rate limiter does during the test: allowlist the generator’s address, or test through the limiter on purpose. Either works. Not knowing which one you chose makes every rate-limit error in the results ambiguous.
Under load: what to record, and where the fail line comes from
Response time under load is recorded as 3 percentiles per route, p50, p95 and p99, next to the throughput actually achieved, the error rate by status code and one server-side saturation number. The fail line comes from the product’s own target for each route, written down before the run, not from a universal number.
| Number | Where it comes from | Why it is there |
|---|---|---|
| The workload: routes, mix, rate, ramp, hold | The worksheet above | So a rerun is the same test |
| Duration and the clock times of the run | The tool’s summary | To line up with the host’s graphs for the same window |
| Environment | Where the generator ran, the target, the data size, what was stubbed | A number without its conditions cannot be compared |
| Rate requested and rate achieved | The test script and the tool’s summary | A gap means the generator or the app fell behind |
| p50, p95 and p99 latency, per route | The tool’s summary, split by route | One average hides the slow tail, and one slow route hides behind fast ones |
| Error rate by status code, per route | The tool’s summary or the host’s logs | A 429, a 500 and a 504 point to different causes |
| One server-side saturation number | The host’s or database’s dashboard: connections in use, function concurrency or CPU | It gives the client-side numbers a cause |
Read the status codes before the latencies. In my reading, a 429 is a rate limiter, yours or the host’s, a 500 is your code, and a 504 is a timeout somewhere between the edge and your server. The tail carries more weight than the median for a reason, and p95, p99 and tail latencies are worth understanding before you set a line.
The fail line, the latency and errors each route may show at the target rate, gets written down before the run, from what that route has to do for the product. A login and an export should not share one threshold. The article on why an AI app stalls under concurrent users answers what a good response time under load is, and why fixed numbers and multiples make a poor general baseline, so this page sets none. A p95 at the target rate far above the p95 at a trickle is a sign of queuing to investigate, not a pass or fail line on its own (my reading). What a good number is for a single request, and why the demo felt fast, is answered in what a good API response time is for a solo-built app.
API load testing tools: what to run it with
API load testing tools come in 4 types: scripted open-source generators such as k6, Locust and Artillery, the GUI veteran JMeter, an API client’s built-in runner such as Postman’s, and hosted online runners. The right one is written in a language the team already uses and can send an arrival rate.
| Tool type | Examples | The test is written in | Where the load comes from | Fits when |
|---|---|---|---|---|
| Scripted open-source generator | k6, Locust, Artillery | JavaScript (k6’s guide), “regular Python code” (Locust), “YAML, TypeScript or JavaScript” (Artillery) | Wherever you run the generator | The test should live in the repo and run in CI |
| The GUI veteran | Apache JMeter | A test plan built in the GUI, run from the command line | Wherever you run JMeter, in CLI mode for the load | The team already keeps JMeter plans |
| An API client’s built-in runner | Postman | The collection’s own requests, no separate script | Your machine or CI, or Postman’s managed infrastructure (plans in the FAQ below) | The collection already exists and the test is small |
| Hosted online runner | Web services that run the test for you | Varies by vendor | The vendor’s machines | A quick look at a public route |
k6’s API load testing guide writes its scripts in JavaScript, and Locust’s documentation calls it “an open source performance/load testing tool for HTTP and other protocols” whose tests are “regular Python code”. Artillery’s test script reference says its scripts “can be written using YAML, TypeScript or JavaScript”, with load set as an arrivalRate per phase. Worked scripts exist for each: a k6 load testing example, Locust load testing and Artillery load testing.
JMeter API load testing splits the work in two: build the test plan in the GUI, then run it without the GUI. JMeter’s best practices put the run plainly, “Use CLI mode”, with jmeter -n -t test.jmx -l test.jtl, and advise using “as few Listeners as possible” during the load. The steps for load testing of an API using JMeter, element by element, are in the JMeter answer at the end of this page.
Postman’s performance testing docs describe each virtual user running the collection’s requests “in the specified order in a repeating loop”. API load testing online, through a hosted runner, means the load leaves the vendor’s machines rather than yours, which suits a public route and gets awkward once every request needs a test account’s token (my reading).
When you compare load testing tools for an API, three questions settle it, by my working rule: the language the team already writes, whether the test can live in the repo and run in CI, and whether the tool can send an arrival rate rather than only a fixed number of users. Choosing between JMeter, k6, Locust or Gatling is a comparison of its own; any of the API performance testing tools that pass those three questions will do the job on a small app.
What to fix first when the numbers are bad
The first fix after a bad load test comes from the shape of the results, and there are 4 shapes: errors jump while latency stays flat, p95 climbs steadily with load, p95 is bad even at a trickle, or latency tracks a third party. Fix one thing, then rerun the identical test.
| Shape of the results | Likely bottleneck | First check | First fix |
|---|---|---|---|
| Errors jump while latency stays flat | A hard limit: database connections, function concurrency, a rate limiter | The database’s connection panel and the status codes by route | Put a pooler in front of the database, or queue the work above the limit you found |
| p95 climbs steadily with load | A shared resource queuing: a slow query, a lock, a single worker | The slowest statements during the run, from pg_stat_statements | Index or rewrite the top statement, then move work that need not block the response |
| p95 is bad even at a trickle | Not a load problem: a missing index, an N+1 loop, a fat payload | One request traced end to end | Fix the query shape and trim the payload |
| Latency tracks a third party | A provider call inline on the request path | The route’s timing against the provider’s own timing | Move the call off the request path, add a timeout, cache what is safe |
For the first shape, the fix is in the pool, and each connection error string has its own fix. For the second, pg_stat_statements shows where the database spent its time, but “the module must be loaded by adding pg_stat_statements to shared_preload_libraries”, so on self-managed Postgres turning it on means a server restart. It keeps the statistics gathered so far until a reset, so call pg_stat_statements_reset() just before the run, or the slowest statements it shows may be all-time ones. “By default, this function can only be executed by superusers. Access may be granted to others using GRANT.” Without that grant, snapshot the view before and after the run and compare total_exec_time per statement, which gives the same per-run answer (my reading). On Supabase, its pg_stat_statements page says to turn the extension on from the dashboard: open the Database page, click Extensions, search for “pg_stat_statements” and enable it. Which role may call the reset is not on that page (checked October 2026), so the snapshot method is the safer default, and Supabase’s inspect guide also says to “Compare saved snapshots with the same reset interval when measuring changes.” Any role can read the timings, but Postgres shows “the SQL text and queryid of queries executed by other users” only to superusers and roles with pg_read_all_stats; Supabase’s troubleshooting page grants it with grant pg_read_all_stats to postgres;.
The third shape is not about load at all, and its single-request causes are the subject of the API response time article. My reading: when a route is slow at a trickle, the shape of the work is the problem, and load only makes it louder. One app I audited in June and July 2026, an AI coding workspace, fetched the full text of every file in a workspace just to draw the file tree, which needs only names and folders. The fuller story, and the order of changes once you know where the ceiling is, belong to scaling web applications.
For the fourth shape, keep the provider’s latency out of your request where you can; what is safe to cache, and for how long, depends on the caching strategies web applications use. Then change one thing, rerun the identical test against the identical data, and record the before and after on the same sheet, the replay step of a controlled load test. Stop when the target rate holds with the error line unbroken, and write the result down as tested capacity, never as a projection. That tested number is the starting point once you know what a capacity plan is.
How to check your own app: an API performance review checklist
An API performance review checklist has 7 checks: the test is in the repo, the workload is written down, the run hit its rate, percentiles and errors are recorded per route, one saturation number is kept, every fix has a before and after, and the environment is stated.
- 01 The test file is in the repo, and a second person can run it from one line in the README. Evidence: the README line and the second person's run.
- 02 The workload sheet names the routes, the mix, the rate, the ramp, the hold and the data size. Evidence: the sheet, dated.
- 03 The run hit the rate it asked for, with achieved throughput within about 5 percent of the requested rate, and where it did not, the record says whether the generator fell behind (its CPU, or too few virtual users allocated for the arrival rate) or the app did. Evidence: the tool's summary.
- 04 p50, p95 and p99 per route, and the error rate by status code, are recorded for the baseline and for the target load. Evidence: two summaries on the same sheet.
- 05 One server-side saturation number is on the same sheet as the client numbers. Evidence: the host's graph for the same window.
- 06 Every fix has a before and an after from the identical test. Evidence: both summaries, with the commit between them.
- 07 The environment line says where the test ran, against what data size, with which third parties stubbed. Evidence: the line itself.
The 5 percent is my working rule, not a standard. A saturated app also lowers achieved throughput, so a missed rate can be a finding about the app, not only about the generator (my reading); k6 notes that too low a preAllocatedVUs setting “will reduce the test duration at the desired rate”. Keep the evidence together: the tool’s summary output, the host’s graphs for the same window, and the sheet with its dates. In the Production Hardening Sprint, deliverable 9.2 is verified this way: “Report the workload, duration, environment, concurrency, latency, and error rate before and after changes.”
Where the sprint does this
Load test and bottleneck fixes is deliverable 9.2 in the published scope: “Simulate concurrent users, identify the first bottlenecks, fix them, and rerun the workload,” because “Launch capacity should be measured before real traffic becomes the test.” Deliverable 9.6, the written capacity statement, documents “measured concurrent capacity on the current infrastructure and the changes needed to plan for five times that workload”, and it is verified this way: “Link capacity claims to load-test evidence and identify untested projections as projections.” The result for every item, the work completed and its verification evidence go into the production readiness report, deliverable 13.1, which accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. Hosting, paid tools, and API usage remain in your accounts.
Common questions about API load tests
Can I do a load test using Postman?
Yes. Postman runs a collection’s requests as virtual users that operate in parallel, each looping through them in order, with load generated from your own machine or CI environment on all plans, or from Postman’s managed infrastructure on Solo, Team and Enterprise plans. The Postman CLI’s --pass-if option exits with a non-zero code, failing the build, when the run misses a condition on a percentile such as p95, the error rate or requests per second.
For the run you keep as a record, send the load from a cloud machine in the same place each run rather than a laptop, as my working rule.
How can I use JMeter to load test an API?
Build the test in JMeter’s GUI and run it from the command line. A Thread Group sets the number of users to simulate, HTTP Request Defaults holds the values every request shares, such as the server, and one HTTP Request sampler per route carries the path, with a JSON body in its Body Data tab. JMeter’s getting-started manual is blunt about the split: “GUI mode should only be used for creating the test script, CLI mode (NON GUI) must be used for load testing”. Run it with jmeter -n -t test.jmx -l test.jtl, keep listeners to a minimum, and add -e and -o to have JMeter “generate an HTML report at end of Load Test”.
Whether JMeter is the right tool for your team at all is a separate choice from how to run it.
Is there any online API testing tool?
Yes, for load: hosted runners send the traffic from the vendor’s machines, and Postman’s cloud runs, which generate load from Postman’s managed infrastructure, are one documented example. They suit a quick test of a public route; an authenticated mix with test accounts and tokens is easier in a scripted tool you control (my reading).
Checking that each endpoint returns the right data, rather than how it holds up under traffic, is functional API testing, a different job.
Which tool is used for load testing?
No single one. For an API, the choice is between a scripted generator (k6 in JavaScript, Locust in Python, Artillery in YAML, TypeScript or JavaScript), Apache JMeter, the runner built into an API client like Postman, and a hosted service, and the deciding questions are what your team already writes and whether the tool sends an arrival rate.
If you have a working app built with these tools and need it ready for real customers, this is what we do.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase