JMeter is the Apache Software Foundation’s free, open-source load testing tool: a Java application that simulates heavy load on a server and measures how the server performs. It is one of 4 open-source load generators a small team should weigh, beside k6, Locust and Gatling, and the right pick is mostly the language your team already writes.
What we are comparing, and for whom: four open-source load generators and one small app
This is for a team of one to five people with a web app on a managed host such as Vercel, Render, Railway or Supabase, and no load test yet. Running one, from a user journey to a fixed bottleneck, is its own subject: load testing a web application. Here the only job is picking the tool.
The set is four open-source load generators, each free to run on a laptop or a CI runner under the license listed in the table below. Artillery load testing is the fifth tool a Node.js team tends to meet, and it stays out of this comparison. Commercial suites such as LoadRunner get one answer in the FAQ and nothing more. Every version, license and price below carries the date I checked it.
This is documented analysis from each project’s own documentation, repository and pricing page, not a benchmark I ran. It quotes no requests-per-second or users-per-machine figure from anyone’s blog. The tool matters less than whether the script and the dataset are reproducible, so the question is which tool your team will keep using, not which one is fastest.
Across the 21 third-party apps I audited in June and July 2026, the Performance and Scale pillar averages 53.3 out of 100, scored on all 21. Those apps are a selected set I chose to audit, not a random sample, so the average is no rate for AI-built apps in general. Load testing is one control among several in web performance optimization.
JMeter, k6, Locust and Gatling: the comparison, row by row
JMeter, k6, Locust and Gatling are 4 free, open-source load generators that differ mainly in how a test is written: JMeter builds an XML test plan in a GUI, k6 runs a JavaScript or TypeScript script, Locust runs Python, and Gatling runs code in Java, Kotlin, Scala, JavaScript or TypeScript. All four run headless from the command line.
The table sets k6 vs JMeter vs Gatling side by side, with Locust added. Row 5 and the last column are my reading, not a vendor fact: a test plan saved as XML reviews as a diff of markup, while the other three review as code, and for a team that reviews every change that row matters most.
| JMeter | k6 | Locust | Gatling | What a small team should care about | |
|---|---|---|---|---|---|
| 1. What it is, in its maker’s words | A Java application that can “simulate a heavy load on a server, group of servers, network or object" | "an open source load and performance testing tool" | "an open source performance/load testing tool for HTTP and other protocols" | "a high-performance load testing tool built for efficiency, automation, and code-driven testing workflows” | Same job; the rows below are where they part |
| 2. Who maintains it | The Apache Software Foundation | Grafana Labs, which acquired k6 in 2021 | Jonatan Heyman, Lars Holmberg and Andrew Baldwin, as its docs list its authors | Gatling Corp, which also sells Gatling Enterprise | Where you file the bug you hit |
| 3. License | Apache-2.0 | AGPL-3.0 | MIT | Apache-2.0 in the repository; the docs add “other specific licenses” | Free to run on your own machines |
| 4. How a test is written | A test plan built or recorded in the GUI, saved as a .jmx XML file | A JavaScript or TypeScript file | Plain Python code | Java, JavaScript, TypeScript, Scala or Kotlin code | Is it a language the team already writes? |
| 5. How a change reads in a pull request | A diff of XML markup | A diff of code | A diff of code | A diff of code | Can the person who reviews your code read the test? |
| 6. Runtime on the machine | Java 8+ for release 5.6.3 | One standalone binary | Python 3.11 or later (Locust 2.46.6) | A 64-bit OpenJDK LTS, 11 to 25; or Node.js v24+ and npm v11+ for JavaScript and TypeScript | One more runtime on the CI image, or none |
| 7. Workload model out of the box | Thread groups, each thread running the whole plan; throughput timers; an experimental Open Model Thread Group | Closed executors such as constant-vus, and arrival-rate executors that hold a rate | A peak user count and a spawn rate (-u, -r) | injectOpen and injectClosed profiles | Can it hold a request rate? See the next section |
| 8. Pass or fail in CI | CI “through 3rd party Open Source libraries for Maven, Gradle and Jenkins”; an exit code for failed samples is not stated in Apache JMeter’s docs | Thresholds; a failed one makes k6 exit with a non-zero code | Exit code 1 on any failed sample by default; process_exit_code set in the script for a latency or failure-ratio line | Assertions: “If at least one assertion fails, the simulation fails” | Does a missed target turn the build red? |
| 9. Report out of the box | An HTML dashboard it can generate from CLI-mode results | An end-of-test summary printed to stdout | A live web UI, --csv files and an --html report file | A static HTML report per run | Is there a file to keep beside the pull request? |
| 10. Protocols beyond HTTP | SOAP and REST, FTP, JDBC, LDAP, JMS, SMTP, POP3 and IMAP, TCP, shell commands, Java objects | WebSockets, gRPC, browser “and more”; extensions for others | Any system you write a client for; third-party extensions | JMS ships with it; its docs have guides for WebSocket and server-sent events; gRPC and MQTT are add-on components, limited to 5 users and 5-minute tests on the free Community Edition | Matters only if the app speaks something other than HTTP |
Checked 2026-09-30 against each project’s documentation, repository and download page.
Apache’s home page says it plainly under the heading “JMeter is not a browser”: “JMeter is not a browser, it works at protocol level”, and it does not execute the JavaScript found in HTML pages. k6 documents a browser API for browser-based performance tests. How fast a page renders in a real browser is a different test from the load tests compared here.
The four tools, one by one
Each tool below ends on a documented property that could make a small team pass on it.
What JMeter is used for, and what it asks of you
JMeter is used to load test web applications, APIs and other protocols, such as databases over JDBC, by simulating concurrent connections and timing the answers. It has two modes: a GUI for building and debugging the test plan, and a command-line mode that Apache’s own manual says to use for the real load test.
In Apache JMeter’s home page words, it is “a 100% pure Java application designed to load test functional behavior and measure performance”, first built for web applications and since “expanded to other test functions”. The getting-started guide is blunt about the modes: “Don’t run load test using GUI mode !” JMeter’s best practices put CLI mode first in their list of ways to cut resource use, and advise switching from BeanShell to JSR223 elements, with Groovy as the language they point to. The HTTP(S) Test Script Recorder captures a browser session into a test plan.
The current release is 5.6.3. Its download page lists it as requiring Java 8+, and the project’s GitHub releases page dates it 9 January 2024 (checked 2026-09-30).
What it asks of you: a Java install, a GUI-built XML file in your repository, and a thread per virtual user, since each thread “will execute the test plan in its entirety”. When one machine is not enough, the manual’s answer is to replicate the test across more computers. Data analysis and visualization plugins extend it.
The reason a code-first team might pass: the test lives in XML, so a reviewer reads each change as markup. The steps to a first JMeter run belong to load testing a web application, not to this comparison.
k6 load testing vs JMeter: a script in JavaScript against a test plan in a GUI
k6 and JMeter differ where a small team feels it most: the test file and the fail line. A k6 test is a JavaScript file, and a failed threshold makes k6 exit with a non-zero code for CI. A JMeter test is an XML plan built in a GUI, and its home page points to third-party libraries for CI.
Grafana Labs documents k6 as an open-source performance testing tool whose tests are JavaScript or TypeScript, with TypeScript on by default since k6 v0.57, run by a standalone binary (the k6 documentation). TypeScript support “strips the type information but doesn’t provide type safety”. Put k6 vs JMeter side by side and three documented differences stand out:
- The test is a text script, not an XML plan, so it diffs like the rest of your code.
- Thresholds are the pass or fail line, and a failed one ends the run with a non-zero exit code; JMeter’s home page sends you to “3rd party Open Source libraries for Maven, Gradle and Jenkins” instead.
- Arrival-rate executors start iterations at a set rate, independent of the response time, as long as VUs are available. In JMeter, each thread in a thread group runs the whole plan, and holding a rate there takes a throughput timer or its experimental Open Model Thread Group.
JMeter still gives people who do not write code more: a full GUI, the recorder, and the protocol list in the table. k6 has a no-code path too. Grafana’s k6 Studio is a desktop app that records a user flow from browser interactions into a HAR file and generates a k6 script from it.
The reason a small team might pass is in how k6 loads code. Its require() “only handles loading of built-in k6 modules, scripts on the local filesystem, and remote scripts over HTTP(S), but it does not support the Node.js module resolution algorithm.” App code that depends on that algorithm will not load as it is, and the docs show Webpack and Rollup templates for bundling it first. Grafana’s own k6-against-JMeter post dates from January 2021 and is the maker’s view. The next step after picking k6 is a k6 load testing example with thresholds in the script.
Locust load testing vs JMeter: Python users against thread groups
Locust and JMeter split on language and on machine use. A Locust test is ordinary Python that can import the app’s own client code, and since Python cannot fully use more than one core per process, you run one worker per core to reach all of a machine’s computing power. JMeter needs no code for a basic plan.
Locust’s own description says it lets you “define your tests in regular Python code”. Every user runs in its own greenlet, the design is event-based so one process holds many users, and a web UI “shows the progress of your test in real-time”, with load you can change mid-run. Tests can also be spread over several machines.
The documented limit, in the docs’ words: “Because Python cannot fully utilize more than one core per process (see GIL ), you need to run one worker instance per processor core in order to have access to all your computing power.”
Locust vs JMeter on tooling: JMeter’s recorder and protocol list are built in. Locust’s docs point to a third-party extension that turns a browser recording (a HAR file) into a locustfile, and to writing your own client for other systems. The report is there too: --html stores an HTML report to a file.
What Locust has that JMeter does not is the plain-Python test. For a Python app, the test can call the same client code, build seed data and sign tokens the way the app does. The reason a small team might pass: with no Python in the stack, the test adds a runtime, Python 3.11 or later for the current release, just for itself. The next step after picking it is a runnable locustfile, the core of Locust load testing.
Gatling open source: what the free edition includes, and what Enterprise adds
Gatling open source, which Gatling calls its Community Edition, is free and, in its docs’ words, “licensed under the Apache License v2.0 and other specific licenses”. It ships SDKs in Java, Kotlin, Scala, JavaScript and TypeScript, a recorder and an HTML report per run; Gatling Enterprise adds managed load generators, run history and team dashboards.
Gatling Corp maintains the open-source project and sells Gatling Enterprise on top of it. The free edition, as Gatling’s documentation describes it, is the load engine, the SDKs, a recorder that captures browser actions into a script, Maven, Gradle and sbt plugins, a JavaScript CLI, and a static HTML report after each run.
Gatling vs JMeter comes down to code in a real language against a GUI plan; both give you an HTML report without extra tools. Gatling’s own comparison page sets out what Enterprise adds:
| Capability | Gatling open source (Community Edition) | Gatling Enterprise |
|---|---|---|
| Test code | All five SDKs | Runs “that same script” |
| Where the load comes from | ”Ideal for local testing” | Managed load generators, or self-hosted on AWS, Azure, GCP or Kubernetes |
| Reports | ”Static HTML reports per run" | "Dynamic, shareable dashboards” |
| History and comparison | ”No history, no comparison, no sharing" | "Full run history, SLO tracking, and run comparison” |
| Team access | ”No sharing or collaboration” | Users and teams management, role-based access control, single sign-on |
| CI | Maven, Gradle and sbt plugins and a JavaScript CLI; assertions fail the simulation | Native plugins for GitHub Actions, GitLab CI, Jenkins and Azure DevOps that “check SLOs on every build” |
Checked 2026-09-30 on Gatling’s documentation and its community-vs-enterprise page. Gatling offers Enterprise free “for 14 days, no credit card required.”
The JavaScript and TypeScript SDK matters to a small team: a team that writes TypeScript does not have to learn Scala or Java to use Gatling. The reason such a team might still pass lies beyond HTTP: the gRPC and MQTT components its JavaScript SDK can use are under the Gatling Enterprise Component License, and on the Community Edition their use is limited to “5 users maximum” and “5 minute duration tests”.
Why the four tools report different numbers for the same app
Load testing tools report different numbers for the same app for 4 main reasons: a closed workload sends less when the server slows while an open one keeps its rate, slow requests that were never sent are never timed, each tool times its own interval, and a load generator that runs short of resources can give wrong results.
Here is each reason in the tools’ own terms.
First, the workload model. In k6’s explanation of open and closed models, “in the closed model, VU iterations start only when the last iteration finishes”, and “Slower response times means longer iterations and a lower arrival rate of new iterations.” In an open model, “The response times of the target system no longer influence the load on the target system.”
Second, the requests that were never sent. The same k6 page names the effect: “In some testing literature, this problem is known as coordinated omission.” The HdrHistogram project’s README describes it without the name: response times may exceed the expected interval between requests, leading to “dropped” measurements that “would typically correlate with “bad” results”. Percentiles, why an average hides the slow end, and coordinated omission in depth belong to tail latencies.
Third, what “response time” covers. JMeter’s glossary times elapsed time “from just before sending the request to just after the last response has been received”, latency to “just after the first response has been received”, and connect time separately, noting that “connect time is not automatically subtracted from latency”. k6’s http_req_duration is the sending, waiting and receiving time “without the initial DNS lookup/connection times”. Those are not the same interval.
Fourth, the generator. Locust logs a warning when it gets close to running out of CPU. JMeter’s best practices warn that threads sized wrong for the machine lead to the “Coordinated Omission” problem “which can give you wrong or inaccurate results”.
| What differs | What it does to the number | What to do about it |
|---|---|---|
| Closed against open workload | A closed model sends fewer new iterations as responses slow, so the load drops just as the server struggles | Hold the request rate (an arrival-rate executor, injectOpen, JMeter’s Open Model Thread Group) when the question is requests per second |
| Requests never sent | The slow requests a closed model never sends are never timed, so the slowdown can look smaller | Use an open model, and read percentiles rather than the average |
| The interval each tool times | JMeter’s elapsed time and k6’s http_req_duration start and stop at different points | Compare runs of one tool against each other, never one tool’s number against another’s |
| The generator’s own load | A generator short of CPU or threads can skew its own timings | Watch the generator’s CPU during the run, and read the warning Locust logs when it nears its limit |
Take a small team that load tests its API with a fixed number of virtual users, each waiting for its answer before sending the next, and reads a steady latency. They rerun the same journey with a tool set to hold a request rate. The server slows, the new run shows queued requests and errors the first never produced, and the team blames the new tool. The first run had sent less load exactly when the server slowed, so the slow requests it never sent were never timed. In my reading, when two load tests of the same app disagree, check first what each one held constant, users or requests per second, before blaming the tool; compare runs of one tool, and hold the rate.
When to pick each, and which one to run against an AI-built app
The load-testing tool to pick is the one written in your team’s language: k6 for JavaScript and TypeScript, Locust for Python, Gatling for the JVM, and JMeter for testers who do not code or for protocols beyond HTTP. With no preference, my working rule is to take the one you can install in about five minutes.
These rules are mine. Read them top to bottom; the first row that matches wins.
| If this is true | Run this | Because |
|---|---|---|
| The team writes TypeScript or JavaScript and wants the test in CI | k6 | The test is a JS or TS file, and a failed threshold fails the run |
| The team writes Python, or the test must reuse the app’s Python code | Locust | The test is plain Python and imports Python libraries |
| The team is on the JVM, or wants code-first tests with an HTML report and no external dashboard | Gatling | Java, Kotlin and Scala SDKs, and a static HTML report per run |
Testers who do not code will own the tests, the protocol is not HTTP (JDBC, LDAP, JMS), or the company already has .jmx plans | JMeter | GUI, recorder, and those protocols on its home page list |
| The run will happen in Azure Load Testing | JMeter or Locust | Those are the two script types the service runs |
| Nobody on the team has a preference | Whichever you can install in about five minutes | The tool matters less than a reproducible script and dataset |
The Azure Load Testing row comes from Microsoft’s overview, which says the service “supports running Apache JMeter-based tests or Locust-based tests”. For k6 vs Locust, the team’s language decides, not speed. k6 vs Gatling is closer for a TypeScript team, since both take TypeScript; k6 runs as one binary, while Gatling’s JavaScript SDK needs Node.js v24 or later.
k6 fits when the code is JavaScript or TypeScript and CI is where tests live. It does not fit when the test must reuse app code that relies on Node.js module resolution without a bundling step.
Locust fits a Python team, or any test that wants to call the app’s own Python code. It does not fit a team that would install Python only for the load test.
Gatling fits a JVM shop, and a TypeScript team testing plain HTTP that wants an HTML report per run. It is a poor fit when the test needs gRPC or MQTT beyond 5 users or 5 minutes and the team will not move to Enterprise.
JMeter fits when testers who do not code own the tests, when the target speaks JDBC, LDAP or JMS, or when .jmx plans already exist. It fits least when the only reviewer reads code and would have to review XML.
If the app is TypeScript on a managed host, with a hosted Postgres behind an API, and the only reviewer is the founder or one contractor, the test has to live in the same repository as the code and read cleanly in review, which points at a text-script tool. The ceilings such an app hits first (connection limits, N+1 queries, missing indexes) and the controlled test that finds them are in why an AI-built app stalls at 100 concurrent users. Check the host’s terms before you point any generator at it, whichever tool you pick. If the target is an API with no front end, start from API load testing.
What it costs to run the tests: free tools, paid clouds
Load tests with JMeter, k6, Locust or Gatling cost nothing to run yourself beyond the machine or the CI minutes, and my working rule is that this covers a small app’s first test. A paid cloud adds managed load generators and a shared history of runs, and each one bills on its own meter.
The licenses are in the comparison table. The cloud columns below come from each maker’s pricing page: Grafana Cloud pricing, Gatling Enterprise pricing and Azure’s load testing pricing.
| Tool | Run it yourself | Maker’s or managed cloud | Free allowance | The meter |
|---|---|---|---|---|
| JMeter | Free (Apache-2.0) on your machine or CI runner | A maker’s cloud is not stated on Apache JMeter’s site; Azure Load Testing runs JMeter scripts | See the Azure row | See the Azure row |
| k6 | Free (AGPL-3.0) | Grafana Cloud k6 | 500 virtual user hours a month on the free tier | Pro: a $19 a month platform fee with 500 virtual user hours, then from $0.15 per virtual user hour |
| Locust | Free (MIT) | Locust’s docs name Azure Load Testing, “a Microsoft-managed load testing service” | See the Azure row | See the Azure row |
| Gatling | Free (Community Edition) | Gatling Enterprise | A 14-day free trial | Basic: €89 a month billed annually or €99 monthly, with 1 hour of testing and 1 load generator; 1 credit is a 1-minute test on 1 load generator |
| Azure Load Testing (JMeter and Locust scripts) | Not applicable | Microsoft-managed | Not stated on Azure’s pricing page | Virtual user hours, in two tiers, up to 10,000 and above 10,000; see the pricing page for the dollar rate. Since 1 March 2026, the minimum per run is 10 virtual users per engine for the run’s duration, or for 10 minutes if the run lasts less than 10 minutes |
Prices checked 2026-09-30.
A cloud run is worth paying for when you need load from several regions, more load than one machine can generate without becoming the bottleneck, or a history of runs the whole team can see. It is not worth it for the first test. In my reading, a small app’s own ceiling usually arrives well before one laptop’s does.
How to check the tool fits before you write a second script
A load-testing tool fits, by my working rule, when it passes 5 checks: a first result in about an hour, a test file a teammate can review, a CI job that fails on a missed target, a generator kept below its CPU warning, and a report of workload, duration, environment, concurrency, latency and error rate before and after a change.
Each check below says what to keep as evidence and what a fail looks like.
- 01 Install to first result in about an hour. Point the tool at staging, script one user journey, run about ten users, and keep the terminal output or report with the date. Fail: it took a day.
- 02 The test file sits in the repository, and a teammate can read the diff of a one-line change. Keep the pull request link. Fail: an XML diff nobody can review.
- 03 The test runs headless in CI, and a deliberately low fail line turns the job red: a k6 threshold, a Gatling assertion, or for Locust a process_exit_code check in the script, since by default it sets the exit code only when the results contain a failure or error. For JMeter, use one of the third-party CI libraries its home page points to. Keep the red run's link. Fail: the job stays green, and a tool that cannot turn the build red only gives you a dashboard.
- 04 The load generator's CPU stayed below the warning level its docs give: Locust warns above 90 percent. Where the docs give no level, as with JMeter's machine-sizing advice, keep it below about 80 percent. Keep the reading. Fail: the generator was the bottleneck.
- 05 In the Production Hardening Sprint, deliverable 9.2 is verified this way: "Report the workload, duration, environment, concurrency, latency, and error rate before and after changes." Hold your own report to the same line, with a run before and a run after one change side by side. Fail: a report that shows only a user count and an average.
A tool that passes all five is the right tool for your team, whatever a comparison post says. Those before-and-after reports are also the raw material for what a capacity plan is: one sentence on how much traffic the app takes, and under what conditions.
Where the sprint fits
In the Production Hardening Sprint, deliverable 9.2, load test and bottleneck fixes, is the load test itself: “Simulate concurrent users, identify the first bottlenecks, fix them, and rerun the workload”, verified by the report in check 5 above. Deliverable 9.6, the written capacity statement, documents “measured concurrent capacity on the current infrastructure and the changes needed to plan for five times that workload”, and we verify it by linking capacity claims to load-test evidence and identifying untested projections as projections. The production readiness report, deliverable 13.1, delivers the result for every scope item, the work completed and its verification evidence. Hosting, paid tools and API usage remain in your accounts, and any required third-party costs are explained before they are enabled. Every deliverable, with how we verify it, is in the published scope.
Common questions about load-testing tools
Is k6 better than JMeter?
No, not in general. k6 fits a team that writes JavaScript and runs its tests in CI; JMeter fits testers who do not code, protocols beyond HTTP and plans that already exist. The k6 section above sets out the documented differences.
Is JMeter still relevant?
Yes, where testers do not code, where protocols beyond HTTP matter, and where a company already has .jmx plans. It is still an Apache project, and its latest release, 5.6.3, is dated 9 January 2024 on the project’s GitHub releases page (checked 2026-09-30). For a new team that writes code, a text-script tool usually fits better.
Is coding required for JMeter?
No, not for a basic plan: you build it in the GUI or record it with the HTTP(S) Test Script Recorder. Once a test needs logic, you write scripts in JSR223 elements, and Apache’s best practices point to Groovy for them.
How long will it take to learn JMeter?
My working rule is an afternoon or so for a first plan, working through the manual’s getting-started page. The slow part comes later: pulling a value out of one response to use in the next request, and building test data that behaves like real users.
Which is better, LoadRunner or JMeter?
Neither in general; for a small SaaS the license and the team’s language decide it. LoadRunner Professional is now called OpenText Professional Performance Engineering, and OpenText sells it in license bundles with a free trial. JMeter is open source under the Apache-2.0 license, free to run.
How do I install Gatling on Windows?
Download the Maven-Java project ZIP from Gatling’s install page, extract it, and run mvnw.cmd gatling:test from the project folder; the Maven wrapper means you do not install Maven yourself. Java projects need a 64-bit OpenJDK LTS version from 11 to 25. For the JavaScript or TypeScript SDK, Gatling installation needs Node.js v24 or later and npm v11 or later, then npm install and npx gatling run. If you use the standalone bundle, the standard Windows zip tool will not extract it; Gatling recommends 7-zip.
If you have a working app built with these tools and need it ready for real customers, this is what we do.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase