Load testing with k6 starts before the first request, with the thresholds at the top of the script: http_req_failed under 1 percent, or the run fails. This k6 load testing example walks a real web app’s sign-in and one core journey with ramped stages, 3 thresholds and a GitHub Actions job, then reads the summary k6 prints.
The k6 load testing example: one script, annotated
A k6 load testing example worth copying has 4 parts in one JavaScript file: options with ramped stages, thresholds that fail the run, a setup step that signs in a test user, and a default function that walks a real user journey with checks. Everything else in k6 is optional for a first test.
I wrote the script below from k6’s documentation for v2.3.0, the release marked Latest on the grafana/k6 releases page on 2026-10-04, and I have not run it against a real app. It assumes a web app with a JSON API behind a sign-in.
A load test is one part of web performance optimization: it finds where the ceiling is, and the rest of that work moves it.
Which journey to test, and how many users to model, comes from the method for load testing a web application. This page is the runnable k6 version of that method.
// load-tests/journey.js: written from the k6 v2.3.0 docs, not run against a real app
import http from 'k6/http';
import { check, group, sleep } from 'k6';
import exec from 'k6/execution';
const BASE = __ENV.BASE_URL; // staging, never production by default
const PEAK = Number(__ENV.PEAK_VUS || 1); // your expected peak, in virtual users
// Stage profiles: durations are my placeholders, see the stages table
const PROFILES = {
smoke: [{ duration: '1m', target: 1 }],
load: [{ duration: '2m', target: PEAK }, { duration: '5m', target: PEAK }, { duration: '1m', target: 0 }],
stress: [{ duration: '2m', target: PEAK }, { duration: '5m', target: PEAK * 2 }, { duration: '5m', target: PEAK * 3 }, { duration: '1m', target: 0 }],
spike: [{ duration: '10s', target: PEAK * 3 }, { duration: '1m', target: PEAK * 3 }, { duration: '10s', target: 0 }],
};
export const options = {
stages: PROFILES[__ENV.PROFILE || 'smoke'],
thresholds: {
http_req_failed: ['rate<0.01'], // the docs' example: under 1%
'http_req_duration{type:read}': ['p(95)<500'], // placeholder from the docs' examples
'http_req_duration{type:write}': ['p(95)<800'], // placeholder from the docs' examples
checks: ['rate>0.99'], // my starting value
},
};
// Runs once, before any virtual user starts: sign in the dedicated test user
export function setup() {
const res = http.post(`${BASE}/api/auth/login`, JSON.stringify({
email: __ENV.TEST_EMAIL,
password: __ENV.TEST_PASSWORD,
}), { headers: { 'Content-Type': 'application/json' } });
if (res.status !== 200) exec.test.abort(`sign-in failed with status ${res.status}`);
return { token: res.json('token') };
}
// One iteration = one pass through the journey by one virtual user
export default function (data) {
const params = (type) => ({
headers: { Authorization: `Bearer ${data.token}`, 'Content-Type': 'application/json' },
tags: { type },
});
let id = null;
group('load dashboard', () => {
const res = http.get(`${BASE}/api/dashboard`, params('read'));
check(res, {
'dashboard status 200': (r) => r.status === 200,
'dashboard has items': (r) => r.status === 200 && Array.isArray(r.json('items')),
});
});
group('create record', () => {
const res = http.post(`${BASE}/api/records`, JSON.stringify({ title: 'k6 run' }), params('write'));
check(res, {
'create status 201': (r) => r.status === 201,
'create returns an id': (r) => r.status === 201 && r.json('id') !== undefined,
});
if (res.status === 201) id = res.json('id');
});
if (id !== null) {
group('read it back', () => {
const res = http.get(`${BASE}/api/records/${id}`, params('read'));
check(res, {
'read status 200': (r) => r.status === 200,
'read title matches': (r) => r.status === 200 && r.json('title') === 'k6 run',
});
});
}
sleep(1); // think time between journeys
}
The notes below are a short tutorial on that file: each part of the k6 script in the order k6 runs it, and what it does during a load testing run. The routes (/api/auth/login, /api/dashboard, /api/records) and the token, items and title fields are placeholders for your app’s own.
options.stages sets the shape of the run: ramp up, hold, ramp down. The script picks one of four stage profiles by name, and the stages section below says what each one answers. For finer control the docs point to a scenario with the ramping-vus executor, which “ramps the number of VUs according to your configured stages”.
options.thresholds is where a run becomes pass or fail. The first line is the docs’ own example, http_req_failed: ['rate<0.01'], commented there as “http errors should be less than 1%”. The two duration lines split reads from writes by tag, with values taken from the docs’ examples as placeholders until you have your own, as k6’s thresholds documentation shows for tagged requests.
setup() signs in once. k6 calls setup “only once per test”, before the virtual users start, and whatever it returns reaches the default function as data. So one dedicated test user signs in against the app’s own auth route, and every virtual user reuses that token instead of signing in on each loop. If the sign-in fails, exec.test.abort() stops the run with exit code 108 and the message you pass it. The order of these stages, and the rule that only JSON data passes from setup to the virtual users, is in the k6 test lifecycle.
The default function is one iteration of one journey: load the dashboard data, create a record, read it back. Each step sits in a group() and is followed by a check() on the status and on one field of the body, because, in the docs’ words, “Sometimes, even an HTTP 200 response contains an error message.” A status check alone would pass that response. The read and write tags feed the per-path thresholds, and the body checks feed the checks threshold; k6 checks covers the API.
sleep(1) is think time. Without it each virtual user fires its next journey the moment the last one returns, faster than a person clicks, and the run measures a tighter loop than real use. The one second follows the docs’ own examples.
The target and the test account come from environment variables, read through __ENV and passed with -e on the command line or from the shell, so no URL or password is ever written in the file. k6 environment variables lists both routes.
For API load testing with no front end at all, the k6 script keeps this shape and only the requests change.
What k6 testing is, and what k6 means
k6 is an open-source load testing tool from Grafana Labs. A test is one JavaScript file, run from the command line by an engine written in Go and measured with built-in metrics such as http_req_duration. Its HTTP tests measure the server side, the requests and their timings; a separate browser module drives a real browser.
k6 testing, in practice, means writing that file and running it against an app you own. The code lives in the grafana/k6 repository, which describes the project as “A modern load testing tool, using Go and JavaScript” and ships it under the AGPL-3.0 license. Grafana’s documentation calls k6 an “open-source, developer-friendly, and extensible performance testing tool”.
k6 scripts are JavaScript, but in the docs’ words “k6 isn’t Node.js or a browser”, and packages that rely on Node.js APIs such as os and fs “won’t work in k6”. An npm package your app imports may not load in a test; k6’s modules page covers bundling for the ones that can be made to work.
What an HTTP test measures is each request, its status and its timing: http_req_duration is the time the server took to process the request and respond, without the initial DNS lookup and connection times. It renders nothing. The browser module is the exception: it automates a Chromium-based browser and collects frontend performance metrics, which makes it a different kind of test from the one on this page.
What the name k6 stands for is not stated on the grafana/k6 repository page, its releases page or the k6 docs landing page, as read on 2026-10-04. I would not repeat an expansion those pages do not give.
The same journey can also be scripted for Locust load testing or Artillery load testing, and the trade-offs between k6 and JMeter are a separate decision from writing this script.
How to fill it in: install, the journey, the stages, the thresholds
Four decisions turn the example into your test, taken here in the order a first run meets them.
Install k6 and run the script
Before any terminal command, create the dedicated test user in staging and walk the journey once in the browser: sign in, open the dashboard, create a record, open that record again. If a step fails there, a load test will only repeat the failure faster.
Install k6 says “k6 has packages for Linux, Mac, and Windows” and offers a Docker container or a standalone binary as the alternatives. The commands as that page prints them:
brew install k6 # macOS (Homebrew)
winget install k6 --source winget # Windows (Windows Package Manager)
sudo apt-get install k6 # Debian/Ubuntu, after adding k6's apt repository and running apt-get update (both on the install page)
sudo dnf install https://dl.k6.io/rpm/repo.rpm # Fedora/CentOS: add the repository,
sudo dnf install k6 # then install
docker pull grafana/k6 # Docker image
The docs call the Chocolatey package unofficial, and they call the winget packages official, installed from k6 manifests “created by the community”. For a machine with no package manager, the k6 download for load testing is the standalone binary: in the install page’s words, the GitHub Releases page “has a standalone binary for all platforms”.
Run it only against a system you own or have written permission to load, in staging by default, after checking your host’s terms (the next section). Then prove the script and the account with a single journey before any load:
k6 run --once -e BASE_URL=https://staging.example.com load-tests/journey.js
k6 run -e BASE_URL=https://staging.example.com -e PROFILE=load -e PEAK_VUS=50 load-tests/journey.js
The --once flag, added in v2.3.0, runs a script that has at most one scenario, as this one does, with one virtual user and one iteration. TEST_EMAIL and TEST_PASSWORD can stay in the shell environment, which k6 also reads into __ENV. A run whose thresholds all pass exits with code 0.
The journey and the test account
Pick one journey that matters, usually sign in, the main read and the main write, rather than the home page. A home page served from a CDN tells you little about the database behind the dashboard.
Use a dedicated test user, or a small pool of them, in a test tenant, so every record the run creates can be found and deleted afterwards. On Supabase Auth or Firebase Authentication, sign in once in setup() and reuse the token: signing in on every iteration would put the auth provider’s own limits into your result, in my reading. If your app keeps the session in a cookie instead of a bearer token, send it as a header from data as well, because k6 clears cookies each time a virtual user starts its loop again. Keep the run shorter than the token’s lifetime, which your auth provider sets, or refresh the token inside the script. Otherwise expired-token errors in the last minutes of a long run read as the app failing (my reading).
Leave third parties out of the load. Payment, email and AI model providers are stubbed or pointed at their test modes: you are not authorized to load someone else’s service, and their limits would end up in your numbers. The target is staging sized like production, or production in a quiet hour with the owner’s written go-ahead. Check the host’s terms before generating load: Vercel’s fair use guidelines list “Load Testing without authorization” under “Never fair use”.
Stages: how many virtual users, for how long
How a count of real users becomes requests per second is worked through in turn users into requests before you size anything; this section takes your expected peak of simultaneous users as PEAK_VUS and gives it a shape. The table lists the four profiles in the script. The durations are my working rule, starting points rather than standards.
| Test type | Stages shape (the profile in the script) | What it answers |
|---|---|---|
| Smoke | one virtual user for about a minute, or --once for a single journey | Do the script, the test account and the target work at all? |
| Load | ramp to PEAK_VUS over about two minutes, hold about 5 to 10 minutes, ramp down | Does the app hold the expected peak inside the thresholds? |
| Stress | ramp past the peak in steps, to two and then three times it, until a threshold fails | Where is the first limit, and what breaks there? |
| Spike | jump to three times the peak in seconds, hold briefly, drop | Does the app survive a sudden burst and recover after it? |
One machine is enough for a small app. k6’s large-test guide says that, “Depending on the available resources, and with the guidelines described in this document”, a single instance “can run 30,000-40,000 simultaneous users (VUs)”. Read k6’s guide to running large tests before you go anywhere near that.
When the goal is a fixed request rate instead of a user count, the arrival-rate executors fit better: constant-arrival-rate “starts iterations at a constant rate”, and ramping-arrival-rate “ramps the iteration rate according to your configured stages”. Both are set as a scenario rather than with the top-level options.stages, described in k6 scenarios and executors.
k6 performance testing thresholds: the three to set, and where the numbers come from
Thresholds turn a k6 run into a pass or fail test. My starting three are http_req_failed below 1 percent, the 95th percentile of http_req_duration below the journey’s own target, and passed checks above 99 percent. If any one fails, k6 exits with a non-zero code and the CI job goes red.
| Threshold | Metric, as k6 names it | Starting value | Where your own value comes from |
|---|---|---|---|
| Errors | http_req_failed, a Rate: “The rate of failed requests according to setResponseCallback” | rate<0.01, the docs’ example | the error rate your capacity target allows |
| Latency | http_req_duration, a Trend, split by the type tag | p(95)<500 for reads and p(95)<800 for writes, placeholders from the docs’ examples | the response time the product promises for that journey |
| Correct answers | checks, a Rate: “The rate of successful checks” | rate>0.99, my working rule | how many wrong answers in a hundred the journey can tolerate |
A founder’s k6 script checks only the status code, and its thresholds cover http_req_failed and the p(95) of http_req_duration. Under load the app’s API starts answering with status 200 and an error message in the body; k6 treats statuses from 200 to 399 as expected by default, so http_req_failed stays at zero, and even a body check that fails does not fail the run on its own, so the run passes. In my reading, a check on one field of the body plus a threshold on the checks rate is the line that turns that wrong answer under load into a failed run; why an API answers 200 with an error is covered in when 200 is the protocol, not the outcome.
The starting values are placeholders. A real value comes from the target in your load test plan or capacity statement (what a capacity plan is), and the article on apps stalling at 100 users warns, in its answer on response time under load, against treating fixed numbers as a general baseline without evidence for the app.
Thresholds work on tags, so the write path can carry a different limit from the read path: the docs’ example is 'http_req_duration{type:API}': ['p(95)<500'] on requests tagged type: 'API'. Every built-in metric name and type is in k6’s built-in metrics reference.
A stress test should stop at the first breach instead of running on. The long threshold format adds abortOnFail, and delayAbortEval waits for some samples first:
thresholds: {
http_req_failed: [{ threshold: 'rate<0.01', abortOnFail: true }],
'http_req_duration{type:read}': [{ threshold: 'p(95)<500', abortOnFail: true, delayAbortEval: '10s' }],
'http_req_duration{type:write}': [{ threshold: 'p(95)<800', abortOnFail: true, delayAbortEval: '10s' }],
checks: [{ threshold: 'rate>0.99', abortOnFail: true }],
},
A failed threshold changes the exit code, which is what CI reads: the docs say “k6 would exit with a non-zero exit code”. Checks on their own do not do that. In the checks page’s words, “When a check fails, the script will continue executing successfully and will not return a ‘failed’ exit status” unless checks are combined with thresholds.
The k6 script itself: what goes in the repo, and running it in CI with k6 GitHub actions
The k6 script belongs in the app’s repository, not on a laptop: one folder with the script and a README that names the run command, the environment variable names and the last result, plus a CI workflow that runs it on demand against staging. The k6 GitHub repository holds the tool itself, its releases and its license.
The README holds variable names only, never their values; the values live in the CI provider’s secrets.
For k6 load testing on GitHub, your own repository needs three files:
| Path | What it holds |
|---|---|
load-tests/journey.js | the script above |
load-tests/README.md | the run commands, the variable names (BASE_URL, PROFILE, PEAK_VUS, TEST_EMAIL, TEST_PASSWORD), the stage profiles, and the last result with its date, environment and commit |
.github/workflows/load-test.yml | the workflow below |
Grafana provides two official actions for this: grafana/setup-k6-action installs k6 and grafana/run-k6-action runs the script, both at @v1 in Grafana’s guide to adding k6 to a CI pipeline. The workflow below starts only on demand, through workflow_dispatch, takes the target, the profile and the peak as inputs, and reads the test account from secrets:
name: Load test
on:
workflow_dispatch:
inputs:
target: { description: 'Staging base URL', required: true, type: string }
profile: { description: 'Stage profile', required: true, type: choice, options: [smoke, load, stress, spike] }
peak_vus: { description: 'Expected peak virtual users', required: true, type: number }
jobs:
k6:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: grafana/setup-k6-action@v1
- uses: grafana/run-k6-action@v1
env:
BASE_URL: ${{ inputs.target }}
PROFILE: ${{ inputs.profile }}
PEAK_VUS: ${{ inputs.peak_vus }}
TEST_EMAIL: ${{ secrets.LOADTEST_EMAIL }}
TEST_PASSWORD: ${{ secrets.LOADTEST_PASSWORD }}
with:
path: load-tests/journey.js
flags: --summary-export ${{ github.workspace }}/summary.json
- uses: actions/upload-artifact@v7
if: ${{ !cancelled() }}
with: { name: k6-summary, path: summary.json }
The inputs reach k6 through a step-level env: block, which Grafana’s guide offers as the alternative to --env in flags. --summary-export writes the end-of-test summary to a JSON file; k6’s options reference says the flag is “not deprecated yet” but discouraged, and points to handleSummary() once you want your own format. The upload step uses if: ${{ !cancelled() }}, GitHub’s recommended way to run a step whether the job passed or failed, so a failed run still keeps its summary. By default the run action prints only the summary to the job log; its debug input shows k6’s full output.
My working rule is never on every push and never against production by default: a load test on each commit spends runner minutes and loads staging for no new answer, so it runs on demand and before a release. If you want a cheap contract check on every pull request, a second workflow can run the same script with --once against staging.
A filled example: reading one run’s summary
A k6 summary is read in 3 places first: the thresholds block at the top, http_req_failed for errors, and the checks lines for wrong answers that still returned 200. Then iterations per second and http_reqs per second tell you the throughput the app held at that number of virtual users.
The block below is illustrative, not a measured run. I constructed it for this page on stated assumptions: the load profile at fifty virtual users, held for five minutes against staging, with round values chosen to show one threshold passing and one failing. Its layout follows the compact summary in k6’s end-of-test summary docs: thresholds first, then checks, then the HTTP, EXECUTION and NETWORK groups. The docs’ sample shows no tagged threshold and no failed check, so how the {type:...} lines and the ✗ check lines are laid out is my guess.
█ THRESHOLDS
checks
✓ 'rate>0.99' rate=99.71%
http_req_duration{type:read}
✓ 'p(95)<500' p(95)=412.6ms
http_req_duration{type:write}
✗ 'p(95)<800' p(95)=1.38s
http_req_failed
✓ 'rate<0.01' rate=0.18%
█ TOTAL RESULTS
checks_total.......................: 59940 124.9/s
checks_succeeded...................: 99.71% 59766 out of 59940
checks_failed......................: 0.29% 174 out of 59940
✗ dashboard status 200
✗ dashboard has items
✓ create status 201
✓ create returns an id
✓ read status 200
✓ read title matches
HTTP
http_req_duration..................: avg=318ms min=41ms med=176ms max=4.9s p(90)=861ms p(95)=1.07s
{ expected_response:true }.......: avg=316ms min=41ms med=175ms max=4.9s p(90)=858ms p(95)=1.06s
http_req_failed....................: 0.18% 54 out of 29971
http_reqs..........................: 29971 62.4/s
EXECUTION
iteration_duration.................: avg=1.95s min=1.12s med=1.6s max=7.3s p(90)=3.1s p(95)=3.6s
iterations.........................: 9990 20.8/s
vus................................: 1 min=1 max=50
vus_max............................: 50 min=50 max=50
NETWORK
data_received......................: 48 MB 100 kB/s
data_sent..........................: 9.7 MB 20 kB/s
| Line | What it says | What to do if it is bad |
|---|---|---|
| THRESHOLDS block | each threshold with ✓ or ✗; here the write path’s p(95) failed and the run failed with it | record which threshold failed; the run is a fail even when every other line looks healthy |
http_req_failed | the share of requests outside the expected statuses (200 to 399 by default) | rerun with --summary-mode full to see which group the failures came from |
http_req_duration avg, med, p(90), p(95) | how long the server took to answer | compare p(95) with the target; the average hides the slow tail |
checks_succeeded, checks_failed | answers that came back wrong, including a 200 with an error in the body | find the named check that failed and the request behind it |
iterations, http_reqs per second | the journeys and requests per second the app held at this load | record it beside the virtual-user count as the run’s throughput |
vus, vus_max | how many virtual users ran, and the most allocated | if the maximum is below PEAK_VUS, the run never reached the load you meant to test |
Why p(95) and not the average: in this block the average answer took about a third of a second, while one request in twenty took more than a second. A user who meets that slow tail on every save does not experience the average (my reading).
The pattern here is reads passing and the write path’s p(95) breaching its threshold during the hold stage. The full mode, --summary-mode full, adds per-group results, which tells you whether the slow part is the create or the read-back. What to do next is not on this page. Finding which limit was hit is the subject of the article on why apps stall at 100 users, and the order of fixes is covered under scaling web applications. An HTTP test does not see the browser, so a slow first paint needs a different kind of work, such as how to reduce JavaScript bundle size. Results can also stream to Grafana Cloud k6, whose free plan is “Limited to 500 virtual user hours per month” on Grafana’s pricing page, as read on 2026-10-04.
How to verify the result
A load test result is trustworthy when 5 things are true: the app’s own logs show the requests arrived, the checks rate is reported, the load generator kept idle CPU, the record names the environment, stages and thresholds, and the same script was rerun after the fix. A number without those is an anecdote.
- 01 The run hit the intended target. The app logs or the host request metrics show requests arriving in the run window at about the count the summary reports. Exact equality is not the test, since a host may log function calls but not cached or static responses. Evidence: both numbers, saved right after the run.
- 02 The checks lines are reported and thresholded, because a 200 that carries an error never raises http_req_failed (the example under the thresholds section). Evidence: the checks_failed line and the checks threshold result.
- 03 The load generator was not the bottleneck. k6 recommends sizing the machine for at least 20% idle cycles, up to 80% used by k6. Evidence: the CPU reading of the machine that ran k6, for the length of the run.
- 04 The record is complete: the environment and its size, the profile and stages, each threshold with its pass or fail, the date and the commit. Evidence: the last-result entry in load-tests/README.md.
- 05 The identical script was rerun after the fix, against the same environment and data. Evidence: two summaries from the same commit of the script, before and after.
The replay rule in the last check is also the final step of the 100-users article’s load test list. For every run, keep the saved summary, the workflow run URL and the commit of the script. In the Production Hardening Sprint, deliverable 9.2, load test and bottleneck fixes, is verified this way: “Report the workload, duration, environment, concurrency, latency, and error rate before and after changes.”
Where the sprint does this
The load test itself is deliverable 9.2 of the Production Hardening Sprint: “Simulate concurrent users, identify the first bottlenecks, fix them, and rerun the workload”, because “Launch capacity should be measured before real traffic becomes the test”. Deliverable 9.6, the written capacity statement, documents “measured concurrent capacity on the current infrastructure and the changes needed to plan for five times that workload”, and is verified this way: “Link capacity claims to load-test evidence and identify untested projections as projections.” Each result, the work completed and its verification evidence go into the production readiness report, deliverable 13.1, which accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. Hosting, paid tools, and API usage remain in your accounts. The full wording is in deliverables 9.2 and 9.6 in the published scope.
Common questions about k6
Is K6 free?
Yes, the k6 tool is free: it is open source under the AGPL-3.0 license, and you run it on your own machine or CI runner. The hosted Grafana Cloud k6 has a free plan capped at 500 virtual user hours a month, and its Pro plan “Starts at $0.15/ virtual user hour” with a “Platform fee of $19 per month” that includes 500 virtual user hours, on Grafana’s pricing page as read on 2026-10-04.
Is k6 open source?
Yes. The grafana/k6 repository names the AGPL-3.0 license. What that license asks of a company that modifies or redistributes k6 is a legal question, and this answer is not legal advice.
Is k6 Grafana?
k6 comes from Grafana Labs, but it is not the Grafana dashboard. Its code lives in the grafana/k6 repository and its docs sit on grafana.com, and results can be sent to Grafana Cloud k6, the hosted service. In my reading it is a separate tool from Grafana itself: you can run k6 with no Grafana product installed.
What is k6 vs K8?
k6 is a load testing tool, and “K8s” is shorthand for Kubernetes, which runs containers across a cluster (my reading, from common usage). They meet only when you run k6 inside a cluster: k6’s running guide lists a distributed mode where “the test execution is distributed across a Kubernetes cluster”.
What is the latest version of K6?
v2.3.0, released on 2026-09-21, was marked Latest on the grafana/k6 releases page on 2026-10-04. Check that page again before you pin a version in CI; k6’s install page shows the upgrade command for each package manager.
If this checklist left you with more open items than you expected, the sprint below works through all of them in ten working days.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase