Performance and scale, as I scored it on 21 third-party apps over June and July 2026, came to a mean of 53.3 out of 100, and a load test is how you find out where your own app sits. Locust load testing writes that test as a plain Python file: each simulated user is a class, each action a method, and one command runs the load against staging.

What Locust load testing is

Locust load testing is load testing written in plain Python: a class for each kind of user, a method for each action, and one command that starts as many users as you ask for against a host you name. It is open source, runs from a laptop, and reports requests per second, response times and failures as the test runs.

That average comes from a selected set of apps I audited, not a random sample, so it describes those 21 apps and not AI-built apps in general; each of the 21 was scored on the pillar. Load testing is one of the measured controls in web performance optimization: you set a number, run the test and keep the result.

Locust’s own description calls it “an open source performance/load testing tool for HTTP and other protocols”, and says its “developer-friendly approach lets you define your tests in regular Python code.” It is MIT-licensed and lives on GitHub under locustio/locust. On October 4, 2026 the latest release was 2.46.6, published September 17, 2026, and the project’s packaging file for that release requires Python 3.11 or newer.

Each simulated user runs in its own greenlet, a lightweight coroutine, and the docs say the event-based design makes it possible for a single process to handle many thousands of concurrent users. HTTP is built in; for another protocol you write a client for it. A web UI shows the run in real time and lets you change the load while it runs, and a headless mode runs the same file from a script or a CI job.

The file and the commands on this page are written from Locust’s documentation for version 2.46.6, not from a benchmark of mine.

Why it matters for a small app: when a Python load generator is the right choice

Locust is the right load testing tool in 3 situations: the team already writes Python, the journey needs real logic such as signing in and reusing a token, or you want to change the load live in a browser. A team that works only in TypeScript is better served by a JavaScript tool.

The “why” column is my reading of where Locust fits for a small team. It ranks no tool.

Your situationPickWhy
The team writes Python every dayLocustThe test is a Python file, so it sits in the repository and goes through review like the rest of the code
The journey needs logic: sign in, read a token from the response, branch on what came backLocustThat logic is ordinary Python inside a task method, not a plugin or a config format
You want to watch a run and change the load while it runsLocustThe web UI shows live numbers and lets you change the user count mid-run
The team works only in TypeScript and wants pass-or-fail thresholds in the scriptk6k6 runs JavaScript test scripts, and TypeScript by default since v0.57; thresholds are the pass/fail criteria you define for the test’s metrics
The test should be a config file with no codeArtilleryArtillery’s first-test guide writes the script in a YAML-based format, and scripts can also be written in JavaScript or TypeScript
The organization already has JMeter test plans and a GUI workflowJMeterJMeter builds test plans in its GUI mode and runs the load from the command line
There is no staging environment, or no target number yetNone yetNo tool can choose the target for you; the workload and the environment come first

For the TypeScript row, start from a k6 load testing example instead of this file. For the config-file row, Artillery load testing is the closer fit. How JMeter and the other load-testing tools compare with each other is a separate question, and this page does not rank them. For the last row, the method of load testing a web application comes before any tool: the journey, the environment and the number you need to hit. In my reading, the tool matters less than the workload you model and the record you keep.

How it works: an annotated Locust example in Python

A Locust file for a signed-in journey has 5 parts: a user class, a wait time, a sign-in step that runs once per user, weighted tasks for what users actually do, and a response check that fails a 200 carrying an error body. It fits in under 40 lines.

This Locust load testing example follows Locust’s guide to writing a locustfile, and every class, decorator and argument name in it is checked there. The paths and JSON fields are placeholders for your own app’s routes.

# loadtest/locustfile.py: one signed-in journey, written from Locust's docs.
# Paths and JSON fields are placeholders: change them to your app's.
import os

from locust import HttpUser, between, task


class SignedInUser(HttpUser):
    wait_time = between(1, 5)  # each user pauses 1 to 5 seconds after every task

    def on_start(self):
        # Runs once for each simulated user: sign in with a staging test account.
        response = self.client.post("/api/login", json={
            "email": os.environ["LOADTEST_EMAIL"],
            "password": os.environ["LOADTEST_PASSWORD"],
        })
        self.auth = {"Authorization": f"Bearer {response.json()['token']}"}

    @task(6)  # weights: the dashboard is picked 6 times as often as creating a record
    def open_dashboard(self):
        self.client.get("/api/dashboard", headers=self.auth)

    @task(3)
    def list_records(self):
        with self.client.get("/api/records", headers=self.auth, catch_response=True) as response:
            # Locust counts any status below 400 as a success, so check the body too.
            if response.status_code == 200 and "error" in response.text:
                response.failure("200 with an error body")

    @task(1)
    def create_record(self):
        response = self.client.post("/api/records", headers=self.auth, json={"title": "load test"})
        record_id = response.json()["id"]
        # name= reports /api/records/123 and /api/records/456 as one line.
        self.client.get(f"/api/records/{record_id}", headers=self.auth, name="/api/records/[id]")

Because a Locust script is plain Python, load testing with Locust can do whatever the journey needs: read a token from the sign-in response, store it on the user, and send it with every later request. The @task weights decide the mix. Locust picks tasks at random in proportion to their weights, so set them from how often real users do each thing, not from what is easiest to test.

The catch_response block is there because of how Locust scores a request. Locust’s docs state the rule: “Requests are considered successful if the HTTP response code is OK (<400)”. An endpoint that answers 200 with an error message in the body would count as a success without the check, so the file marks it with response.failure(). The name argument does the other job: the statistics group requests by URL, so without it every record ID would get its own line.

Use dedicated test accounts in staging, never real customers’ accounts, and keep the credentials in environment variables as the file does. My working rule for the expensive endpoint, an AI call, an export or a payment, is to stub it in staging or give it its own low-weight task with a spending cap on the provider’s side. Run load only against systems you own or have written permission to test, in staging, with third-party services stubbed or in sandbox mode.

Install and first run: a Locust load testing tutorial in five commands

The first run needs a terminal for the install, then nothing but the browser. Create a virtual environment, install Locust, check the version, and start it with your file and your staging host:

python3 -m venv .venv
source .venv/bin/activate
pip install locust
locust -V
locust -f loadtest/locustfile.py -H https://staging.example.com

Then open http://localhost:8089 in a browser, the address Locust’s quickstart gives. The start form asks for “Number of users (peak concurrency)”, “Ramp up (users started/second)” and “Host”. Start with one user. If every request in the Statistics tab shows zero failures, the journey signs in and works, and you can add load from there.

Users, tasks and wait time: the three ideas in a locustfile

A user class is one kind of visitor, and the user count says how many of them run at once. Tasks are what that visitor does, picked by weight. Wait time is the pause after each task, and the docs spell out what happens without one: “the next task will be executed as soon as one finishes.”

In my reading of the model, wait time decides the request rate as much as the user count does. Each user completes one task, then waits, then picks the next, so the rate is roughly the user count times the requests in a task, divided by how long one task-and-wait cycle takes. That makes the target a request rate on named endpoints, and users and wait time the two settings you adjust to reach it.

Locust has a wait time built for that: constant_throughput gives “an adaptive time that ensures the task runs (at most) X times per second.” The docs work through the arithmetic in a note: “if you want Locust to run 500 task iterations per second at peak load, you could use wait_time = constant_throughput(0.1) and a user count of 5000.” They also warn that “Wait time can only constrain the throughput, not launch new Users to reach the target”, and that wait times apply to tasks, not requests, so a task that sends two requests doubles the request rate per user.

Getting this backwards is an easy mistake. A founder’s target is a request rate on the dashboard endpoints. The locustfile copies the docs’ example, wait_time = between(1, 5), and the run is started with as many users as the target’s requests per second, reading one user as one request a second. Each user pauses between tasks, so the run sends far fewer requests than the target and passes at a load the app never had to carry; the requests-per-second column in Locust’s statistics shows it. The lesson I take from it: set the rate first, then let constant_throughput fix a rate per user and work out the user count from it, the way the docs’ note does.

Which endpoints to include, and in what mix, follows the same reasoning as API load testing: start from the requests real users send most. Spawn rate is how fast users arrive; the docs note that it “does not change peak load, it only changes how fast you get there.” A gradual ramp is kinder to a cold cache than starting every user at once (my reading). For stepped or spiking profiles, a LoadTestShape class with a tick() method sets the user count and spawn rate at every moment of the run.

Headless runs, CSV output and CI: the Locust file in GitHub

Locust runs without its web UI in headless mode: one command sets users, spawn rate and run time and writes CSV files and an HTML report. A pass threshold on latency or error rate is set in code, through a documented hook that sets the process exit code from the run’s results.

locust -f loadtest/locustfile.py --headless -u 50 -r 5 -t 10m \
  -H https://staging.example.com \
  --csv results/run --html results/run.html

Every flag here is in Locust’s configuration reference: -u is the peak user count, -r the users started per second, -t the run time, --csv the prefix for the stats files and --html the report path. The user count, spawn rate and run time in the command are placeholders for your own target. The --csv prefix writes run_stats.csv, run_failures.csv, run_exceptions.csv and run_stats_history.csv into results/.

By default the process exits with code 1 if any request failed, and --exit-code-on-error changes that code. A pass threshold on response time or error ratio is not among the options in Locust’s configuration reference (checked October 4, 2026). The docs’ answer is code: a listener on events.quitting reads environment.stats.total.fail_ratio and environment.stats.total.get_response_time_percentile(0.95) and sets environment.process_exit_code to 1 when either crosses your limit. Put your own numbers in that listener, the ones you wrote down before the run.

Keep the Locust load testing file in your GitHub repository under a loadtest/ folder, with a short README that names the target, the staging host and the command. My working rule is to run it in CI on a manual trigger or a schedule against staging, not on every pull request. Locust’s headless page includes a GitHub Actions job that installs Locust and runs it headless; the rest of the pipeline is ordinary CI CD best practices. The project itself lives in the Locust repository, the source for the code, the releases and the MIT license.

Distributed mode, and when one machine is enough

Distributed mode splits a run between one master and several workers: start one process with --master and the others with --worker and --master-host, or use --processes to launch a master and workers on one machine. The --processes flag relies on fork(), so Locust’s distributed mode docs say it does not work on Windows. They also say to run one worker per CPU core, because a Python process cannot fully use more than one.

The sign you need more than one process is the load generator’s own CPU. When it runs short, Locust logs a warning that begins “CPU usage above 90%”. Locust’s docs also suggest FastHttpUser, which “uses significantly less CPU time, sometimes increasing the maximum number of requests per second on a given hardware by as much as 5x-6x”, for very high throughput on limited hardware; it “does not make individual requests faster.” For a simple test plan with small payloads, the docs put one process at more than a thousand requests per second.

For a small SaaS app, one machine is usually enough in my reading, because the app tends to reach its limits well before the generator does. Run the generator from the same region as your real users, so the latency numbers mean something for them.

How to check your own app: the run record, and what counts as a pass

A Locust run becomes evidence when 6 things are written down: the workload, the duration, the environment, the concurrency, the latency per named request at the median and 95th percentile, and the error rate. Set the pass criteria before the run, then record before and after each fix.

Write the pass criteria down first: the request rate each named endpoint must carry, the latency it must stay under, and the error rate you accept. Then record six fields for every run:

  1. 01 Workload: the journey, the task weights and the wait time, plus the locustfile commit.
  2. 02 Duration: the run time, and how long the ramp took inside it.
  3. 03 Environment: the staging size, the data volume, and what was stubbed.
  4. 04 Concurrency: the users and spawn rate, plus the requests per second the run actually reached.
  5. 05 Latency: the median and the 95th percentile for each named request.
  6. 06 Error rate: failures per request, with the failure messages.

Each number has a place in Locust’s output, in the web UI and in the CSV files the headless run writes:

NumberWhere Locust shows itWrite down
Requests per second”Current RPS” in the Statistics tab; “Requests/s” in _stats.csvThe rate reached per named request and in total
Median latency”Median (ms)”; “Median Response Time” and “50%” in _stats.csvPer named request
95th percentile latency”95%ile (ms)”; “95%” in _stats.csvPer named request
Error rate”# Fails” against ”# Requests”; “Failure Count” against “Request Count” in _stats.csvFailures divided by requests
Failure messagesThe Failures tab; “Error” and “Occurrences” in _failures.csvThe top messages and their counts
Users over time”User Count” in _stats_history.csvWhen the ramp finished

The CSV section of Locust’s configuration reference also points to the Download Data tab, where the web UI offers CSV files. Why the tail of the distribution matters more than the average is the subject of tail latencies.

Then the loop: find the slowest named request, fix one thing, run the same file again, and record before and after. My working rule is that a run with no target, no environment note or no error count is not evidence. What passes is your own number, set from what your customers do; this page gives no threshold, because any number I printed would be invented. The full method and the common bottlenecks sit with load testing a web application.

In the Production Hardening Sprint, deliverable 9.2, Load test and bottleneck fixes, is verified this way: we report the workload, duration, environment, concurrency, latency, and error rate before and after changes.

Where the sprint fits

In the Production Hardening Sprint, deliverable 9.2 is where we simulate concurrent users, identify the first bottlenecks, fix them, and rerun the workload. Deliverable 13.1, the production readiness report, delivers the result for every scope item, the work completed, and its verification evidence. Your app’s current framework and hosting setup are our starting point; we refactor or replace components where the production work requires it. Hosting, paid tools, and API usage remain in your accounts. The exact wording is under deliverable 9.2 in the published scope.

Common questions about Locust

Is Locust free to use?

Yes. Locust is open source under the MIT license, so the tool itself has no fee; the load runs on machines you provide. The project’s site, locust.io, says the maintainers are working on a hosted cloud version, Locust Cloud, which also provides dedicated commercial support; the open-source tool does not need it.

Is locust good for load testing?

Yes, for HTTP journeys a team can write in Python, especially ones that sign in and carry state between requests. The limit to know is that one Locust process cannot fully use more than one CPU core, so a single process caps the load it can generate; distributed mode and FastHttpUser exist for that. Whether it suits you better than another tool depends on your team’s language and workflow more than on the tool.

What is the Python equivalent of JMeter?

Locust is the Python tool that fills JMeter’s role: it does the same job, generating load against a system and reporting the results, with the test written as Python code instead of a test plan built in JMeter’s GUI mode. A feature-by-feature look at the two belongs to a comparison of load-testing tools, not to this answer.

What are the four main types of load testing?

There is no fixed four: k6’s guide to test types lists six types (smoke, average-load, stress, soak, spike and breakpoint) and notes that no consensus exists about the names. In Locust, most of them are a run setting rather than a different tool. A smoke test is one user; an average-load test is your expected users with a ramp; a soak test is the same load with a long --run-time; stress, spike and breakpoint tests fit a LoadTestShape that steps or jumps the user count.