JMeter is the Apache Software Foundation’s free, open-source load testing tool: a Java application that simulates heavy load on a server and measures how the server performs. It is one of 4 open-source load generators a small team should weigh, beside k6, Locust and Gatling, and the right pick is mostly the language your team already writes.

What we are comparing, and for whom: four open-source load generators and one small app

This is for a team of one to five people with a web app on a managed host such as Vercel, Render, Railway or Supabase, and no load test yet. Running one, from a user journey to a fixed bottleneck, is its own subject: load testing a web application. Here the only job is picking the tool.

The set is four open-source load generators, each free to run on a laptop or a CI runner under the license listed in the table below. Artillery load testing is the fifth tool a Node.js team tends to meet, and it stays out of this comparison. Commercial suites such as LoadRunner get one answer in the FAQ and nothing more. Every version, license and price below carries the date I checked it.

This is documented analysis from each project’s own documentation, repository and pricing page, not a benchmark I ran. It quotes no requests-per-second or users-per-machine figure from anyone’s blog. The tool matters less than whether the script and the dataset are reproducible, so the question is which tool your team will keep using, not which one is fastest.

Across the 21 third-party apps I audited in June and July 2026, the Performance and Scale pillar averages 53.3 out of 100, scored on all 21. Those apps are a selected set I chose to audit, not a random sample, so the average is no rate for AI-built apps in general. Load testing is one control among several in web performance optimization.

JMeter, k6, Locust and Gatling: the comparison, row by row

JMeter, k6, Locust and Gatling are 4 free, open-source load generators that differ mainly in how a test is written: JMeter builds an XML test plan in a GUI, k6 runs a JavaScript or TypeScript script, Locust runs Python, and Gatling runs code in Java, Kotlin, Scala, JavaScript or TypeScript. All four run headless from the command line.

The table sets k6 vs JMeter vs Gatling side by side, with Locust added. Row 5 and the last column are my reading, not a vendor fact: a test plan saved as XML reviews as a diff of markup, while the other three review as code, and for a team that reviews every change that row matters most.

JMeterk6LocustGatlingWhat a small team should care about
1. What it is, in its maker’s wordsA Java application that can “simulate a heavy load on a server, group of servers, network or object""an open source load and performance testing tool""an open source performance/load testing tool for HTTP and other protocols""a high-performance load testing tool built for efficiency, automation, and code-driven testing workflows”Same job; the rows below are where they part
2. Who maintains itThe Apache Software FoundationGrafana Labs, which acquired k6 in 2021Jonatan Heyman, Lars Holmberg and Andrew Baldwin, as its docs list its authorsGatling Corp, which also sells Gatling EnterpriseWhere you file the bug you hit
3. LicenseApache-2.0AGPL-3.0MITApache-2.0 in the repository; the docs add “other specific licenses”Free to run on your own machines
4. How a test is writtenA test plan built or recorded in the GUI, saved as a .jmx XML fileA JavaScript or TypeScript filePlain Python codeJava, JavaScript, TypeScript, Scala or Kotlin codeIs it a language the team already writes?
5. How a change reads in a pull requestA diff of XML markupA diff of codeA diff of codeA diff of codeCan the person who reviews your code read the test?
6. Runtime on the machineJava 8+ for release 5.6.3One standalone binaryPython 3.11 or later (Locust 2.46.6)A 64-bit OpenJDK LTS, 11 to 25; or Node.js v24+ and npm v11+ for JavaScript and TypeScriptOne more runtime on the CI image, or none
7. Workload model out of the boxThread groups, each thread running the whole plan; throughput timers; an experimental Open Model Thread GroupClosed executors such as constant-vus, and arrival-rate executors that hold a rateA peak user count and a spawn rate (-u, -r)injectOpen and injectClosed profilesCan it hold a request rate? See the next section
8. Pass or fail in CICI “through 3rd party Open Source libraries for Maven, Gradle and Jenkins”; an exit code for failed samples is not stated in Apache JMeter’s docsThresholds; a failed one makes k6 exit with a non-zero codeExit code 1 on any failed sample by default; process_exit_code set in the script for a latency or failure-ratio lineAssertions: “If at least one assertion fails, the simulation fails”Does a missed target turn the build red?
9. Report out of the boxAn HTML dashboard it can generate from CLI-mode resultsAn end-of-test summary printed to stdoutA live web UI, --csv files and an --html report fileA static HTML report per runIs there a file to keep beside the pull request?
10. Protocols beyond HTTPSOAP and REST, FTP, JDBC, LDAP, JMS, SMTP, POP3 and IMAP, TCP, shell commands, Java objectsWebSockets, gRPC, browser “and more”; extensions for othersAny system you write a client for; third-party extensionsJMS ships with it; its docs have guides for WebSocket and server-sent events; gRPC and MQTT are add-on components, limited to 5 users and 5-minute tests on the free Community EditionMatters only if the app speaks something other than HTTP

Checked 2026-09-30 against each project’s documentation, repository and download page.

Apache’s home page says it plainly under the heading “JMeter is not a browser”: “JMeter is not a browser, it works at protocol level”, and it does not execute the JavaScript found in HTML pages. k6 documents a browser API for browser-based performance tests. How fast a page renders in a real browser is a different test from the load tests compared here.

The four tools, one by one

Each tool below ends on a documented property that could make a small team pass on it.

What JMeter is used for, and what it asks of you

JMeter is used to load test web applications, APIs and other protocols, such as databases over JDBC, by simulating concurrent connections and timing the answers. It has two modes: a GUI for building and debugging the test plan, and a command-line mode that Apache’s own manual says to use for the real load test.

In Apache JMeter’s home page words, it is “a 100% pure Java application designed to load test functional behavior and measure performance”, first built for web applications and since “expanded to other test functions”. The getting-started guide is blunt about the modes: “Don’t run load test using GUI mode !” JMeter’s best practices put CLI mode first in their list of ways to cut resource use, and advise switching from BeanShell to JSR223 elements, with Groovy as the language they point to. The HTTP(S) Test Script Recorder captures a browser session into a test plan.

The current release is 5.6.3. Its download page lists it as requiring Java 8+, and the project’s GitHub releases page dates it 9 January 2024 (checked 2026-09-30).

What it asks of you: a Java install, a GUI-built XML file in your repository, and a thread per virtual user, since each thread “will execute the test plan in its entirety”. When one machine is not enough, the manual’s answer is to replicate the test across more computers. Data analysis and visualization plugins extend it.

The reason a code-first team might pass: the test lives in XML, so a reviewer reads each change as markup. The steps to a first JMeter run belong to load testing a web application, not to this comparison.

k6 load testing vs JMeter: a script in JavaScript against a test plan in a GUI

k6 and JMeter differ where a small team feels it most: the test file and the fail line. A k6 test is a JavaScript file, and a failed threshold makes k6 exit with a non-zero code for CI. A JMeter test is an XML plan built in a GUI, and its home page points to third-party libraries for CI.

Grafana Labs documents k6 as an open-source performance testing tool whose tests are JavaScript or TypeScript, with TypeScript on by default since k6 v0.57, run by a standalone binary (the k6 documentation). TypeScript support “strips the type information but doesn’t provide type safety”. Put k6 vs JMeter side by side and three documented differences stand out:

  • The test is a text script, not an XML plan, so it diffs like the rest of your code.
  • Thresholds are the pass or fail line, and a failed one ends the run with a non-zero exit code; JMeter’s home page sends you to “3rd party Open Source libraries for Maven, Gradle and Jenkins” instead.
  • Arrival-rate executors start iterations at a set rate, independent of the response time, as long as VUs are available. In JMeter, each thread in a thread group runs the whole plan, and holding a rate there takes a throughput timer or its experimental Open Model Thread Group.

JMeter still gives people who do not write code more: a full GUI, the recorder, and the protocol list in the table. k6 has a no-code path too. Grafana’s k6 Studio is a desktop app that records a user flow from browser interactions into a HAR file and generates a k6 script from it.

The reason a small team might pass is in how k6 loads code. Its require() “only handles loading of built-in k6 modules, scripts on the local filesystem, and remote scripts over HTTP(S), but it does not support the Node.js module resolution algorithm.” App code that depends on that algorithm will not load as it is, and the docs show Webpack and Rollup templates for bundling it first. Grafana’s own k6-against-JMeter post dates from January 2021 and is the maker’s view. The next step after picking k6 is a k6 load testing example with thresholds in the script.

Locust load testing vs JMeter: Python users against thread groups

Locust and JMeter split on language and on machine use. A Locust test is ordinary Python that can import the app’s own client code, and since Python cannot fully use more than one core per process, you run one worker per core to reach all of a machine’s computing power. JMeter needs no code for a basic plan.

Locust’s own description says it lets you “define your tests in regular Python code”. Every user runs in its own greenlet, the design is event-based so one process holds many users, and a web UI “shows the progress of your test in real-time”, with load you can change mid-run. Tests can also be spread over several machines.

The documented limit, in the docs’ words: “Because Python cannot fully utilize more than one core per process (see GIL ), you need to run one worker instance per processor core in order to have access to all your computing power.”

Locust vs JMeter on tooling: JMeter’s recorder and protocol list are built in. Locust’s docs point to a third-party extension that turns a browser recording (a HAR file) into a locustfile, and to writing your own client for other systems. The report is there too: --html stores an HTML report to a file.

What Locust has that JMeter does not is the plain-Python test. For a Python app, the test can call the same client code, build seed data and sign tokens the way the app does. The reason a small team might pass: with no Python in the stack, the test adds a runtime, Python 3.11 or later for the current release, just for itself. The next step after picking it is a runnable locustfile, the core of Locust load testing.

Gatling open source: what the free edition includes, and what Enterprise adds

Gatling open source, which Gatling calls its Community Edition, is free and, in its docs’ words, “licensed under the Apache License v2.0 and other specific licenses”. It ships SDKs in Java, Kotlin, Scala, JavaScript and TypeScript, a recorder and an HTML report per run; Gatling Enterprise adds managed load generators, run history and team dashboards.

Gatling Corp maintains the open-source project and sells Gatling Enterprise on top of it. The free edition, as Gatling’s documentation describes it, is the load engine, the SDKs, a recorder that captures browser actions into a script, Maven, Gradle and sbt plugins, a JavaScript CLI, and a static HTML report after each run.

Gatling vs JMeter comes down to code in a real language against a GUI plan; both give you an HTML report without extra tools. Gatling’s own comparison page sets out what Enterprise adds:

CapabilityGatling open source (Community Edition)Gatling Enterprise
Test codeAll five SDKsRuns “that same script”
Where the load comes from”Ideal for local testing”Managed load generators, or self-hosted on AWS, Azure, GCP or Kubernetes
Reports”Static HTML reports per run""Dynamic, shareable dashboards”
History and comparison”No history, no comparison, no sharing""Full run history, SLO tracking, and run comparison”
Team access”No sharing or collaboration”Users and teams management, role-based access control, single sign-on
CIMaven, Gradle and sbt plugins and a JavaScript CLI; assertions fail the simulationNative plugins for GitHub Actions, GitLab CI, Jenkins and Azure DevOps that “check SLOs on every build”

Checked 2026-09-30 on Gatling’s documentation and its community-vs-enterprise page. Gatling offers Enterprise free “for 14 days, no credit card required.”

The JavaScript and TypeScript SDK matters to a small team: a team that writes TypeScript does not have to learn Scala or Java to use Gatling. The reason such a team might still pass lies beyond HTTP: the gRPC and MQTT components its JavaScript SDK can use are under the Gatling Enterprise Component License, and on the Community Edition their use is limited to “5 users maximum” and “5 minute duration tests”.

Why the four tools report different numbers for the same app

Load testing tools report different numbers for the same app for 4 main reasons: a closed workload sends less when the server slows while an open one keeps its rate, slow requests that were never sent are never timed, each tool times its own interval, and a load generator that runs short of resources can give wrong results.

Here is each reason in the tools’ own terms.

First, the workload model. In k6’s explanation of open and closed models, “in the closed model, VU iterations start only when the last iteration finishes”, and “Slower response times means longer iterations and a lower arrival rate of new iterations.” In an open model, “The response times of the target system no longer influence the load on the target system.”

Second, the requests that were never sent. The same k6 page names the effect: “In some testing literature, this problem is known as coordinated omission.” The HdrHistogram project’s README describes it without the name: response times may exceed the expected interval between requests, leading to “dropped” measurements that “would typically correlate with “bad” results”. Percentiles, why an average hides the slow end, and coordinated omission in depth belong to tail latencies.

Third, what “response time” covers. JMeter’s glossary times elapsed time “from just before sending the request to just after the last response has been received”, latency to “just after the first response has been received”, and connect time separately, noting that “connect time is not automatically subtracted from latency”. k6’s http_req_duration is the sending, waiting and receiving time “without the initial DNS lookup/connection times”. Those are not the same interval.

Fourth, the generator. Locust logs a warning when it gets close to running out of CPU. JMeter’s best practices warn that threads sized wrong for the machine lead to the “Coordinated Omission” problem “which can give you wrong or inaccurate results”.

What differsWhat it does to the numberWhat to do about it
Closed against open workloadA closed model sends fewer new iterations as responses slow, so the load drops just as the server strugglesHold the request rate (an arrival-rate executor, injectOpen, JMeter’s Open Model Thread Group) when the question is requests per second
Requests never sentThe slow requests a closed model never sends are never timed, so the slowdown can look smallerUse an open model, and read percentiles rather than the average
The interval each tool timesJMeter’s elapsed time and k6’s http_req_duration start and stop at different pointsCompare runs of one tool against each other, never one tool’s number against another’s
The generator’s own loadA generator short of CPU or threads can skew its own timingsWatch the generator’s CPU during the run, and read the warning Locust logs when it nears its limit

Take a small team that load tests its API with a fixed number of virtual users, each waiting for its answer before sending the next, and reads a steady latency. They rerun the same journey with a tool set to hold a request rate. The server slows, the new run shows queued requests and errors the first never produced, and the team blames the new tool. The first run had sent less load exactly when the server slowed, so the slow requests it never sent were never timed. In my reading, when two load tests of the same app disagree, check first what each one held constant, users or requests per second, before blaming the tool; compare runs of one tool, and hold the rate.

When to pick each, and which one to run against an AI-built app

The load-testing tool to pick is the one written in your team’s language: k6 for JavaScript and TypeScript, Locust for Python, Gatling for the JVM, and JMeter for testers who do not code or for protocols beyond HTTP. With no preference, my working rule is to take the one you can install in about five minutes.

These rules are mine. Read them top to bottom; the first row that matches wins.

If this is trueRun thisBecause
The team writes TypeScript or JavaScript and wants the test in CIk6The test is a JS or TS file, and a failed threshold fails the run
The team writes Python, or the test must reuse the app’s Python codeLocustThe test is plain Python and imports Python libraries
The team is on the JVM, or wants code-first tests with an HTML report and no external dashboardGatlingJava, Kotlin and Scala SDKs, and a static HTML report per run
Testers who do not code will own the tests, the protocol is not HTTP (JDBC, LDAP, JMS), or the company already has .jmx plansJMeterGUI, recorder, and those protocols on its home page list
The run will happen in Azure Load TestingJMeter or LocustThose are the two script types the service runs
Nobody on the team has a preferenceWhichever you can install in about five minutesThe tool matters less than a reproducible script and dataset

The Azure Load Testing row comes from Microsoft’s overview, which says the service “supports running Apache JMeter-based tests or Locust-based tests”. For k6 vs Locust, the team’s language decides, not speed. k6 vs Gatling is closer for a TypeScript team, since both take TypeScript; k6 runs as one binary, while Gatling’s JavaScript SDK needs Node.js v24 or later.

k6 fits when the code is JavaScript or TypeScript and CI is where tests live. It does not fit when the test must reuse app code that relies on Node.js module resolution without a bundling step.

Locust fits a Python team, or any test that wants to call the app’s own Python code. It does not fit a team that would install Python only for the load test.

Gatling fits a JVM shop, and a TypeScript team testing plain HTTP that wants an HTML report per run. It is a poor fit when the test needs gRPC or MQTT beyond 5 users or 5 minutes and the team will not move to Enterprise.

JMeter fits when testers who do not code own the tests, when the target speaks JDBC, LDAP or JMS, or when .jmx plans already exist. It fits least when the only reviewer reads code and would have to review XML.

If the app is TypeScript on a managed host, with a hosted Postgres behind an API, and the only reviewer is the founder or one contractor, the test has to live in the same repository as the code and read cleanly in review, which points at a text-script tool. The ceilings such an app hits first (connection limits, N+1 queries, missing indexes) and the controlled test that finds them are in why an AI-built app stalls at 100 concurrent users. Check the host’s terms before you point any generator at it, whichever tool you pick. If the target is an API with no front end, start from API load testing.

What it costs to run the tests: free tools, paid clouds

Load tests with JMeter, k6, Locust or Gatling cost nothing to run yourself beyond the machine or the CI minutes, and my working rule is that this covers a small app’s first test. A paid cloud adds managed load generators and a shared history of runs, and each one bills on its own meter.

The licenses are in the comparison table. The cloud columns below come from each maker’s pricing page: Grafana Cloud pricing, Gatling Enterprise pricing and Azure’s load testing pricing.

ToolRun it yourselfMaker’s or managed cloudFree allowanceThe meter
JMeterFree (Apache-2.0) on your machine or CI runnerA maker’s cloud is not stated on Apache JMeter’s site; Azure Load Testing runs JMeter scriptsSee the Azure rowSee the Azure row
k6Free (AGPL-3.0)Grafana Cloud k6500 virtual user hours a month on the free tierPro: a $19 a month platform fee with 500 virtual user hours, then from $0.15 per virtual user hour
LocustFree (MIT)Locust’s docs name Azure Load Testing, “a Microsoft-managed load testing service”See the Azure rowSee the Azure row
GatlingFree (Community Edition)Gatling EnterpriseA 14-day free trialBasic: €89 a month billed annually or €99 monthly, with 1 hour of testing and 1 load generator; 1 credit is a 1-minute test on 1 load generator
Azure Load Testing (JMeter and Locust scripts)Not applicableMicrosoft-managedNot stated on Azure’s pricing pageVirtual user hours, in two tiers, up to 10,000 and above 10,000; see the pricing page for the dollar rate. Since 1 March 2026, the minimum per run is 10 virtual users per engine for the run’s duration, or for 10 minutes if the run lasts less than 10 minutes

Prices checked 2026-09-30.

A cloud run is worth paying for when you need load from several regions, more load than one machine can generate without becoming the bottleneck, or a history of runs the whole team can see. It is not worth it for the first test. In my reading, a small app’s own ceiling usually arrives well before one laptop’s does.

How to check the tool fits before you write a second script

A load-testing tool fits, by my working rule, when it passes 5 checks: a first result in about an hour, a test file a teammate can review, a CI job that fails on a missed target, a generator kept below its CPU warning, and a report of workload, duration, environment, concurrency, latency and error rate before and after a change.

Each check below says what to keep as evidence and what a fail looks like.

  1. 01 Install to first result in about an hour. Point the tool at staging, script one user journey, run about ten users, and keep the terminal output or report with the date. Fail: it took a day.
  2. 02 The test file sits in the repository, and a teammate can read the diff of a one-line change. Keep the pull request link. Fail: an XML diff nobody can review.
  3. 03 The test runs headless in CI, and a deliberately low fail line turns the job red: a k6 threshold, a Gatling assertion, or for Locust a process_exit_code check in the script, since by default it sets the exit code only when the results contain a failure or error. For JMeter, use one of the third-party CI libraries its home page points to. Keep the red run's link. Fail: the job stays green, and a tool that cannot turn the build red only gives you a dashboard.
  4. 04 The load generator's CPU stayed below the warning level its docs give: Locust warns above 90 percent. Where the docs give no level, as with JMeter's machine-sizing advice, keep it below about 80 percent. Keep the reading. Fail: the generator was the bottleneck.
  5. 05 In the Production Hardening Sprint, deliverable 9.2 is verified this way: "Report the workload, duration, environment, concurrency, latency, and error rate before and after changes." Hold your own report to the same line, with a run before and a run after one change side by side. Fail: a report that shows only a user count and an average.

A tool that passes all five is the right tool for your team, whatever a comparison post says. Those before-and-after reports are also the raw material for what a capacity plan is: one sentence on how much traffic the app takes, and under what conditions.

Where the sprint fits

In the Production Hardening Sprint, deliverable 9.2, load test and bottleneck fixes, is the load test itself: “Simulate concurrent users, identify the first bottlenecks, fix them, and rerun the workload”, verified by the report in check 5 above. Deliverable 9.6, the written capacity statement, documents “measured concurrent capacity on the current infrastructure and the changes needed to plan for five times that workload”, and we verify it by linking capacity claims to load-test evidence and identifying untested projections as projections. The production readiness report, deliverable 13.1, delivers the result for every scope item, the work completed and its verification evidence. Hosting, paid tools and API usage remain in your accounts, and any required third-party costs are explained before they are enabled. Every deliverable, with how we verify it, is in the published scope.

Common questions about load-testing tools

Is k6 better than JMeter?

No, not in general. k6 fits a team that writes JavaScript and runs its tests in CI; JMeter fits testers who do not code, protocols beyond HTTP and plans that already exist. The k6 section above sets out the documented differences.

Is JMeter still relevant?

Yes, where testers do not code, where protocols beyond HTTP matter, and where a company already has .jmx plans. It is still an Apache project, and its latest release, 5.6.3, is dated 9 January 2024 on the project’s GitHub releases page (checked 2026-09-30). For a new team that writes code, a text-script tool usually fits better.

Is coding required for JMeter?

No, not for a basic plan: you build it in the GUI or record it with the HTTP(S) Test Script Recorder. Once a test needs logic, you write scripts in JSR223 elements, and Apache’s best practices point to Groovy for them.

How long will it take to learn JMeter?

My working rule is an afternoon or so for a first plan, working through the manual’s getting-started page. The slow part comes later: pulling a value out of one response to use in the next request, and building test data that behaves like real users.

Which is better, LoadRunner or JMeter?

Neither in general; for a small SaaS the license and the team’s language decide it. LoadRunner Professional is now called OpenText Professional Performance Engineering, and OpenText sells it in license bundles with a free trial. JMeter is open source under the Apache-2.0 license, free to run.

How do I install Gatling on Windows?

Download the Maven-Java project ZIP from Gatling’s install page, extract it, and run mvnw.cmd gatling:test from the project folder; the Maven wrapper means you do not install Maven yourself. Java projects need a 64-bit OpenJDK LTS version from 11 to 25. For the JavaScript or TypeScript SDK, Gatling installation needs Node.js v24 or later and npm v11 or later, then npm install and npx gatling run. If you use the standalone bundle, the standard Windows zip tool will not extract it; Gatling recommends 7-zip.