The Performance & Scale pillar averaged 53.3 of a possible 100 on 21 third-party apps I scored in June and July 2026. API load testing measures your app under traffic: by my working rule, pick 3 to 5 routes users hit, ramp to about twice your busiest hour, hold about 10 minutes, and record p95 latency, error rate and throughput.

API load testing: what the test is, and the five things to decide first

An API load test sends a planned, realistic amount of traffic at the API and records how latency, errors and throughput change as the traffic grows. Five decisions come first: the question, the routes and their mix, the target load and ramp, where it runs, and the numbers that count as a fail.

That 53.3 is an average pillar score, not a pass mark, and the pillar was scored on all 21 third-party apps in my June and July 2026 audits. Those apps are 11 public apps audited across all 12 pillars and 10 held-out apps my review engine had never seen, audited blind. They are a selected set of audited apps, not a random sample, so the average describes them and is no rate for AI-built apps in general.

Load testing for an API is the server-side measurement in web performance optimization, and it stays a measurement, not a pass or fail, until you draw the line. These five decisions turn it into an answer:

The decisionWhat to write downThe mistake it prevents
The questionOne of: can we take the launch, where is the ceiling, did the fix workStopping at the first green run because nobody said what it had to prove
The routes and their mix3 to 5 named routes and the share of traffic each getsTesting the health route, which touches nothing that breaks
The target load and rampRequests per second, how fast it climbs, how long it holds, and the data behind itTesting with one user’s data and ten rows
Where it runsThe target environment, where the generator sits, what is stubbedTesting from a laptop on wifi
The fail linep95 and error rate per route at the target rateDeciding after the run what counted as slow

An API load test does not see the browser half of speed: rendering, scripts and images belong to a Lighthouse audit and the image optimization checklist for websites. The whole-app version, with pages and logged-in journeys, is load testing a web application, which also names the load, stress, soak and spike tests.

Why it matters for a small app: what breaks first under load

In the same audits, 13 of the 21 third-party apps had no rate limiting on their most expensive endpoint. That count describes the apps I picked to audit, not every app, but the risk is easy to picture: the route that costs the most per request has nothing limiting how often it runs, and a load test is the cheap way to learn what it can take before real traffic does.

In my reading, four things give way first in an AI-built API. None of them shows in a demo, because a demo has one user and ten rows. The middle column is the part to learn: it is how each one looks in a load test’s results.

What breaks firstHow it shows in the resultsWhy a demo never showed it
The database connection limitErrors jump while the latency of successful requests stays flatOne user holds one connection at a time
A query that grows with the tablep95 climbs as the data grows, not as the users growThe demo table has ten rows
A third-party call made inlineRoute latency follows the provider’s, and its rate limit becomes yoursOne user never reaches the provider’s limit
The serverless concurrency or timeout ceilingFlat, then a cliff: timeouts or throttling at one rateOne user never reaches the ceiling

The connection errors in the first row each have their own error string, and every connection pool exhausted error and its fix matches each string to its layer. Which of these ceilings your hosting platform reaches first is its own question, answered by the article on why an AI app stalls under concurrent users.

How it works: the workload, the place, the numbers, the tools and the fix order

Five steps, in the order you do them: design the workload, choose where it runs, decide what to record, pick a tool, and read the results.

The workload to send: routes, mix, data and ramp

The workload for a small API, as my working rule, is 3 to 5 real routes in their real mix, sent at an arrival rate of about twice the busiest hour, ramped up over a couple of minutes and held for about 10, against production-sized data with one token per virtual user.

Pick the routes from real traffic, not from the route file: the busiest read, the busiest write, the most expensive route (search, an export, the AI call) and login. Take the mix from a day of logs or your host’s request panel. Then state the target as requests per second rather than as a number of users. k6’s open and closed models page explains why: “in a closed model, the start or arrival rate of new VU iterations is tightly coupled with the iteration duration.” A fixed group of virtual users waits for each response, so a slowing server quietly lowers the load the test sends, just when you wanted it held. k6’s API load testing guide adds that API load “is generally reported by request rate”, per second or per minute.

LineHow to get the numberWorked example (assumptions stated)
RoutesThe busiest read, the busiest write, the most expensive route, loginList projects, create a task, search, login
MixEach route’s share of a day of logs or the host’s request panel60, 25, 10 and 5 percent (assumed)
Busiest-hour rateRequests in the busiest hour divided by 3,6005 requests per second, read from the logs (assumed)
Target rateAbout twice the busiest-hour rate, as my working rule10 requests per second
Ramp, hold, ramp downA couple of minutes up, about 10 held, about 1 down, as a starting shape2 minutes up to 10 per second, 10 minutes held, 1 minute down
Test dataMore than one account and realistic row counts, seeded in the test environmentAs many test accounts and rows as production holds today
AuthenticationOne token per virtual user, fetched in setupTokens for the test accounts fetched before the first request, so login gets only its own share of the load

The third column is a worked API load testing example, not a recommendation: your logs set your rates. If you start from a user count instead of logs, converting users into requests is a step of its own, and the method to run a controlled load test covers that step and the replay around it.

The held part of that worked example, written as one k6 scenario:

export const options = {
  scenarios: {
    hold_at_target: {
      executor: 'constant-arrival-rate',
      rate: 10,            // iterations started per timeUnit
      timeUnit: '1s',      // so 10 per second
      duration: '10m',
      preAllocatedVUs: 20, // VUs ready before the start
      maxVUs: 50,          // k6 may add VUs up to this to keep the rate
    },
  },
};

With one request per iteration, the rate is requests per second; leave out any sleep at the end of the iteration, because the arrival-rate executors already pace it. This executor keeps the load constant; k6’s guide says to “use the ramping-arrival-rate executor instead” to ramp the rate up or down, which is how the two-minute climb and the one-minute descent are written.

Where to run it, and what to point it at

A load test runs only against systems you own or have written permission to test. Staging with production-sized data is the default target, with every third party stubbed or in its sandbox mode: payments, email and model calls. Read the host’s load testing terms before the first run.

TargetSafe whenWhat to stubWho to tell first
Staging with production-sized dataThe default for every runThe payment provider (test mode), the email provider (sandbox), the model provider (a fake)Your team
Production in a quiet hourOnly when staging cannot match the database size, at a capped rate, with a kill switchThe same three, wherever the code allows itYour team, whoever watches the alerts, and the host if its terms ask
Your laptopWriting and debugging the script, never the run of recordEverythingNobody
A system you do not ownNever, unless the owner gave written permissionNot your callThe owner, in writing, before anything runs

Stubbing matters because the third parties’ rate limits and bills are real even when your traffic is not. The hosts’ terms matter for the same reason. Vercel’s fair use guidelines list “Load Testing without authorization” under “Never fair use”. The Amazon EC2 Testing Policy covers high-volume tests sent from EC2 instances, “sometimes called stress tests, load tests, or gameday tests”; it asks that your endpoints sit in the local AWS Region if they are hosted within AWS, and says “Volumetric network-based DDoS simulations are explicitly prohibited from the Amazon EC2 platform.”

Run the generator from a cloud machine, not a laptop on home wifi, and write down where it ran; k6’s guide asks you to “ensure that the location of the load generator is constant across test runs, and avoid running the tests from locations that are too close to the SUT”, the system under test. Then decide what your own rate limiter does during the test: allowlist the generator’s address, or test through the limiter on purpose. Either works. Not knowing which one you chose makes every rate-limit error in the results ambiguous.

Under load: what to record, and where the fail line comes from

Response time under load is recorded as 3 percentiles per route, p50, p95 and p99, next to the throughput actually achieved, the error rate by status code and one server-side saturation number. The fail line comes from the product’s own target for each route, written down before the run, not from a universal number.

NumberWhere it comes fromWhy it is there
The workload: routes, mix, rate, ramp, holdThe worksheet aboveSo a rerun is the same test
Duration and the clock times of the runThe tool’s summaryTo line up with the host’s graphs for the same window
EnvironmentWhere the generator ran, the target, the data size, what was stubbedA number without its conditions cannot be compared
Rate requested and rate achievedThe test script and the tool’s summaryA gap means the generator or the app fell behind
p50, p95 and p99 latency, per routeThe tool’s summary, split by routeOne average hides the slow tail, and one slow route hides behind fast ones
Error rate by status code, per routeThe tool’s summary or the host’s logsA 429, a 500 and a 504 point to different causes
One server-side saturation numberThe host’s or database’s dashboard: connections in use, function concurrency or CPUIt gives the client-side numbers a cause

Read the status codes before the latencies. In my reading, a 429 is a rate limiter, yours or the host’s, a 500 is your code, and a 504 is a timeout somewhere between the edge and your server. The tail carries more weight than the median for a reason, and p95, p99 and tail latencies are worth understanding before you set a line.

The fail line, the latency and errors each route may show at the target rate, gets written down before the run, from what that route has to do for the product. A login and an export should not share one threshold. The article on why an AI app stalls under concurrent users answers what a good response time under load is, and why fixed numbers and multiples make a poor general baseline, so this page sets none. A p95 at the target rate far above the p95 at a trickle is a sign of queuing to investigate, not a pass or fail line on its own (my reading). What a good number is for a single request, and why the demo felt fast, is answered in what a good API response time is for a solo-built app.

API load testing tools: what to run it with

API load testing tools come in 4 types: scripted open-source generators such as k6, Locust and Artillery, the GUI veteran JMeter, an API client’s built-in runner such as Postman’s, and hosted online runners. The right one is written in a language the team already uses and can send an arrival rate.

Tool typeExamplesThe test is written inWhere the load comes fromFits when
Scripted open-source generatork6, Locust, ArtilleryJavaScript (k6’s guide), “regular Python code” (Locust), “YAML, TypeScript or JavaScript” (Artillery)Wherever you run the generatorThe test should live in the repo and run in CI
The GUI veteranApache JMeterA test plan built in the GUI, run from the command lineWherever you run JMeter, in CLI mode for the loadThe team already keeps JMeter plans
An API client’s built-in runnerPostmanThe collection’s own requests, no separate scriptYour machine or CI, or Postman’s managed infrastructure (plans in the FAQ below)The collection already exists and the test is small
Hosted online runnerWeb services that run the test for youVaries by vendorThe vendor’s machinesA quick look at a public route

k6’s API load testing guide writes its scripts in JavaScript, and Locust’s documentation calls it “an open source performance/load testing tool for HTTP and other protocols” whose tests are “regular Python code”. Artillery’s test script reference says its scripts “can be written using YAML, TypeScript or JavaScript”, with load set as an arrivalRate per phase. Worked scripts exist for each: a k6 load testing example, Locust load testing and Artillery load testing.

JMeter API load testing splits the work in two: build the test plan in the GUI, then run it without the GUI. JMeter’s best practices put the run plainly, “Use CLI mode”, with jmeter -n -t test.jmx -l test.jtl, and advise using “as few Listeners as possible” during the load. The steps for load testing of an API using JMeter, element by element, are in the JMeter answer at the end of this page.

Postman’s performance testing docs describe each virtual user running the collection’s requests “in the specified order in a repeating loop”. API load testing online, through a hosted runner, means the load leaves the vendor’s machines rather than yours, which suits a public route and gets awkward once every request needs a test account’s token (my reading).

When you compare load testing tools for an API, three questions settle it, by my working rule: the language the team already writes, whether the test can live in the repo and run in CI, and whether the tool can send an arrival rate rather than only a fixed number of users. Choosing between JMeter, k6, Locust or Gatling is a comparison of its own; any of the API performance testing tools that pass those three questions will do the job on a small app.

What to fix first when the numbers are bad

The first fix after a bad load test comes from the shape of the results, and there are 4 shapes: errors jump while latency stays flat, p95 climbs steadily with load, p95 is bad even at a trickle, or latency tracks a third party. Fix one thing, then rerun the identical test.

Shape of the resultsLikely bottleneckFirst checkFirst fix
Errors jump while latency stays flatA hard limit: database connections, function concurrency, a rate limiterThe database’s connection panel and the status codes by routePut a pooler in front of the database, or queue the work above the limit you found
p95 climbs steadily with loadA shared resource queuing: a slow query, a lock, a single workerThe slowest statements during the run, from pg_stat_statementsIndex or rewrite the top statement, then move work that need not block the response
p95 is bad even at a trickleNot a load problem: a missing index, an N+1 loop, a fat payloadOne request traced end to endFix the query shape and trim the payload
Latency tracks a third partyA provider call inline on the request pathThe route’s timing against the provider’s own timingMove the call off the request path, add a timeout, cache what is safe

For the first shape, the fix is in the pool, and each connection error string has its own fix. For the second, pg_stat_statements shows where the database spent its time, but “the module must be loaded by adding pg_stat_statements to shared_preload_libraries”, so on self-managed Postgres turning it on means a server restart. It keeps the statistics gathered so far until a reset, so call pg_stat_statements_reset() just before the run, or the slowest statements it shows may be all-time ones. “By default, this function can only be executed by superusers. Access may be granted to others using GRANT.” Without that grant, snapshot the view before and after the run and compare total_exec_time per statement, which gives the same per-run answer (my reading). On Supabase, its pg_stat_statements page says to turn the extension on from the dashboard: open the Database page, click Extensions, search for “pg_stat_statements” and enable it. Which role may call the reset is not on that page (checked October 2026), so the snapshot method is the safer default, and Supabase’s inspect guide also says to “Compare saved snapshots with the same reset interval when measuring changes.” Any role can read the timings, but Postgres shows “the SQL text and queryid of queries executed by other users” only to superusers and roles with pg_read_all_stats; Supabase’s troubleshooting page grants it with grant pg_read_all_stats to postgres;.

The third shape is not about load at all, and its single-request causes are the subject of the API response time article. My reading: when a route is slow at a trickle, the shape of the work is the problem, and load only makes it louder. One app I audited in June and July 2026, an AI coding workspace, fetched the full text of every file in a workspace just to draw the file tree, which needs only names and folders. The fuller story, and the order of changes once you know where the ceiling is, belong to scaling web applications.

For the fourth shape, keep the provider’s latency out of your request where you can; what is safe to cache, and for how long, depends on the caching strategies web applications use. Then change one thing, rerun the identical test against the identical data, and record the before and after on the same sheet, the replay step of a controlled load test. Stop when the target rate holds with the error line unbroken, and write the result down as tested capacity, never as a projection. That tested number is the starting point once you know what a capacity plan is.

How to check your own app: an API performance review checklist

An API performance review checklist has 7 checks: the test is in the repo, the workload is written down, the run hit its rate, percentiles and errors are recorded per route, one saturation number is kept, every fix has a before and after, and the environment is stated.

  1. 01 The test file is in the repo, and a second person can run it from one line in the README. Evidence: the README line and the second person's run.
  2. 02 The workload sheet names the routes, the mix, the rate, the ramp, the hold and the data size. Evidence: the sheet, dated.
  3. 03 The run hit the rate it asked for, with achieved throughput within about 5 percent of the requested rate, and where it did not, the record says whether the generator fell behind (its CPU, or too few virtual users allocated for the arrival rate) or the app did. Evidence: the tool's summary.
  4. 04 p50, p95 and p99 per route, and the error rate by status code, are recorded for the baseline and for the target load. Evidence: two summaries on the same sheet.
  5. 05 One server-side saturation number is on the same sheet as the client numbers. Evidence: the host's graph for the same window.
  6. 06 Every fix has a before and an after from the identical test. Evidence: both summaries, with the commit between them.
  7. 07 The environment line says where the test ran, against what data size, with which third parties stubbed. Evidence: the line itself.

The 5 percent is my working rule, not a standard. A saturated app also lowers achieved throughput, so a missed rate can be a finding about the app, not only about the generator (my reading); k6 notes that too low a preAllocatedVUs setting “will reduce the test duration at the desired rate”. Keep the evidence together: the tool’s summary output, the host’s graphs for the same window, and the sheet with its dates. In the Production Hardening Sprint, deliverable 9.2 is verified this way: “Report the workload, duration, environment, concurrency, latency, and error rate before and after changes.”

Where the sprint does this

Load test and bottleneck fixes is deliverable 9.2 in the published scope: “Simulate concurrent users, identify the first bottlenecks, fix them, and rerun the workload,” because “Launch capacity should be measured before real traffic becomes the test.” Deliverable 9.6, the written capacity statement, documents “measured concurrent capacity on the current infrastructure and the changes needed to plan for five times that workload”, and it is verified this way: “Link capacity claims to load-test evidence and identify untested projections as projections.” The result for every item, the work completed and its verification evidence go into the production readiness report, deliverable 13.1, which accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. Hosting, paid tools, and API usage remain in your accounts.

Common questions about API load tests

Can I do a load test using Postman?

Yes. Postman runs a collection’s requests as virtual users that operate in parallel, each looping through them in order, with load generated from your own machine or CI environment on all plans, or from Postman’s managed infrastructure on Solo, Team and Enterprise plans. The Postman CLI’s --pass-if option exits with a non-zero code, failing the build, when the run misses a condition on a percentile such as p95, the error rate or requests per second.

For the run you keep as a record, send the load from a cloud machine in the same place each run rather than a laptop, as my working rule.

How can I use JMeter to load test an API?

Build the test in JMeter’s GUI and run it from the command line. A Thread Group sets the number of users to simulate, HTTP Request Defaults holds the values every request shares, such as the server, and one HTTP Request sampler per route carries the path, with a JSON body in its Body Data tab. JMeter’s getting-started manual is blunt about the split: “GUI mode should only be used for creating the test script, CLI mode (NON GUI) must be used for load testing”. Run it with jmeter -n -t test.jmx -l test.jtl, keep listeners to a minimum, and add -e and -o to have JMeter “generate an HTML report at end of Load Test”.

Whether JMeter is the right tool for your team at all is a separate choice from how to run it.

Is there any online API testing tool?

Yes, for load: hosted runners send the traffic from the vendor’s machines, and Postman’s cloud runs, which generate load from Postman’s managed infrastructure, are one documented example. They suit a quick test of a public route; an authenticated mix with test accounts and tokens is easier in a scripted tool you control (my reading).

Checking that each endpoint returns the right data, rather than how it holds up under traffic, is functional API testing, a different job.

Which tool is used for load testing?

No single one. For an API, the choice is between a scripted generator (k6 in JavaScript, Locust in Python, Artillery in YAML, TypeScript or JavaScript), Apache JMeter, the runner built into an API client like Postman, and a hosted service, and the deciding questions are what your team already writes and whether the tool sends an arrival rate.