A scan lists what might be wrong; web application penetration testing is a person, with written permission, trying to reach what an attacker could and proving each step. In 3 of the 21 third-party apps I audited in June and July 2026, scanners flagged 33 to 44 vulnerabilities and the audit traced exactly zero as reachable.

What web application penetration testing covers, and what it does not

Web application penetration testing is an authorized, time-boxed attempt by a person, working under agreed constraints, to reach what an attacker could reach in one app, with every finding proven and written up. It covers only the surfaces the signed scope names, and where a scan stops at possible problems, the test shows what can actually be reached.

Those 21 apps are the ones I audited in June and July 2026: 11 public third-party vibe-coded apps audited exhaustively across all 12 pillars, and 10 held-out third-party apps audited blind. They are a selected set, chosen for audit, so the number is neither a random sample nor a rate for AI-built apps in general. What it does show is that raw scanner output is not risk; it makes no claim about how a pen test would score. The test is one control inside web app security, and this page stays with that control.

The definition above is my plain-English reading of NIST’s definition of penetration testing: “A test methodology in which assessors, typically working under specific constraints, attempt to circumvent or defeat the security features of a system.”

Pen testing web applications, or penetration testing a website, covers what the signed scope names: the app’s routes and forms, its API, the sign-in flows, the hosting edge and the integrations the owner controls. Anything outside that list stays untouched, however close it sits. Software pen testing is the same work pointed at a wider target, such as a desktop program or a device, so a web application penetration test is the narrower job most SaaS founders are asked for.

The primary goal of penetration testing is proof: which weaknesses can be reached, and what reaching each one gives an attacker. For a small team, that is why penetration testing is important, and the benefit is concrete: a short, ranked list of real, reachable problems with the proof attached to each, where a scanner hands over a long export of possibilities.

Web pentesting on a SaaS usually means one production app plus its API, while website pentesting for a brochure site with no sign-in is a much smaller job. A pen test of an application with customer accounts takes more effort, because every role and every boundary between customers is one more thing to try. The same targets apply when penetration testing apps built with AI coding tools: the tester sees the running app, never the tool that wrote it.

A quote for web-app penetration testing should list its targets by hostname; if it lists none, ask. Whatever the label, web application security penetration testing is judged on one thing, whether each finding comes with proof. If a quote says application security penetration testing, ask which apps it covers: the web app, a mobile app, or both.

An OWASP penetration test is one that follows the OWASP Web Security Testing Guide, OWASP’s catalog of web security tests; its project page says “v4.2 is currently available as a web-hosted release and PDF” and that release version 5.0 is in development. VAPT testing is a vulnerability assessment plus a pen test in one engagement, and how the two halves differ is in pen test vs vulnerability assessment.

Three questions sit next to this one and have their own pages. Whether you need a test at all, and what SOC 2 and customers actually require, is answered in do I need a pen test. For the choice between a scan, a code audit and a pen test, read which kind of security review you are being asked for. And the checks you can run yourself before paying anyone are in the website security check.

Black box, grey box and white box

Black box gives the tester no knowledge of the inside, grey box gives some, such as a user account and the architecture, and white box gives the code and configuration too. A multi-tenant SaaS usually needs at least grey box, because a cross-customer failure only shows to a logged-in user.

In pen testing, white and black box describe how much the tester is told before starting; the terms have nothing to do with packaging. NIST’s glossary defines black box testing, gray box testing and white box testing by that knowledge. The last two columns below are my reading of how the knowledge is handed over on a small SaaS, which is where black box penetration testing and a white box pentest part ways in practice.

BoxWhat NIST’s glossary saysWhat the tester is given on a small SaaS (my reading)What it can find (my reading)
Black box”assumes no knowledge of the internal structure and implementation detail of the assessment object”The public URLs, no accountsWhat an anonymous visitor can reach: public pages, the sign-in form, headers, exposed routes
Grey box”assumes some knowledge of the internal structure and implementation detail of the assessment object”; “Also known as focused testing”Two user accounts in two separate customers, plus the architectureThe above, plus what a logged-in user can reach, including another customer’s data
White box”assumes explicit and substantial knowledge of the internal structure and implementation detail of the assessment object”, filed under comprehensive testingThe repository and configuration as wellThe above, plus flaws in code paths the tester reads before trying them

A black box pentest answers one question well: what can a stranger on the internet reach? Blackbox pentesters start from the public URL, where a real attacker starts too, so that view is worth having, though it misses everything behind sign-in. Gray box testing, the middle option, hands the tester part of the inside view: user accounts, the architecture, sometimes the API spec. White box security testing adds the repository and configuration, so a suspect code path can be read first and then tried against the running app.

For a multi-tenant SaaS, grey box penetration testing is the version I would buy first. In my June and July 2026 audits, 7 of the 21 third-party apps had confirmed cross-user or cross-tenant authorization failures, where a logged-in user could read or write another customer’s data. A tester with no account cannot be that logged-in user. Those 21 were apps picked for audit, so the 7 describe that group only, not AI-built apps at large. Choosing between black box and white box penetration testing therefore starts with one question: does the app hold customer accounts whose data must stay apart?

Take a founder with a multi-tenant SaaS who buys a black-box pen test and provides no accounts. The report covers the public pages, the sign-in form and the headers, and no route behind sign-in is tried, so nobody checks whether one customer’s account can read another’s. The box you buy sets what can be found. My take on the lesson: give the tester two accounts in two customers, because for a multi-tenant app the grey-box test is the one that can reach the failures that matter.

What goes wrong without it

Four failures follow when an app ships with no scoped test, or with the wrong kind.

FailureWhat the owner seesWhat a scoped test does about it
A privileged endpoint anyone can callNothing: the app works for every normal userA tester on the external scope calls it with no session and records the response
One customer reads another customer’s dataNothing in normal use, since each user sees their own screensA grey-box tester signs in as one customer and requests the other’s records (the box section above)
A scan export sent where a test was asked forA long list of findings nobody can rankEach finding is shown reachable, with proof, or dropped (the number at the top of this page)
A fix shipped with no retestA closed ticketThe tester replays the proof after the fix and records a dated result

The first row has a number behind it. In the same June and July 2026 audits, 11 of the 21 third-party apps had unauthenticated endpoints doing privileged work. That count comes from apps I chose to audit; it describes them, and it is no estimate of how often the problem occurs elsewhere. The last row is the quiet one: without a retest, a finding stays open on paper and in fact, whatever the ticket says. The third row is a tool comparison in its own right, covered in why scanner, audit and pentest answer different questions.

How a small scope is agreed, and how the test runs

A small web app pen test runs in three stages: agree the scope on one page, let the tester work through the phases on the agreed targets, then retest every fix. Most of what the test can find is decided in the first stage, by the targets and accounts the tester gets.

Scope: external, internal, network, infrastructure, cloud, API

External means everything reachable from the internet, and it is the default for a SaaS. Network and infrastructure are mostly the managed host’s layer. Cloud means the account configuration. API means the routes behind the front end, tested with the spec in hand.

Each “what it means” cell is my plain-English reading, and each “need it” cell is my working answer for an app on a managed host.

Scope wordWhat it meansDoes a small SaaS on a managed host need it?
ExternalEverything reachable from the internet: the app, its API, its sign-in, its public file linksYes, the default
InternalTesting from inside an office network or VPN, as an insider or a stolen laptop wouldRarely: an app on a managed host has no office network in its path
Network and infrastructurePorts, servers and the layer the host runsMostly the host’s side; test only what you run yourself
CloudThe account configuration: identity and access, storage, network rulesYes, if you run your own cloud account; read the provider’s testing policy first
APIEvery route the server exposes, including ones the front end never calls, with the spec providedYes; for GraphQL, hand over the schema

The first two rows settle internal vs external penetration testing for most SaaS teams: external first, internal only when an office network sits in front of the app. An external pentest starts where an attacker starts, which is why it comes first, and when you ask for external penetration testing services, ask for this scope before anything else. External network penetration testing, the port-and-server half of that view, matters once you run your own servers.

On a managed platform, network pen testing is mostly a test of the host’s side of the line. An infrastructure pen test makes sense once you run your own virtual machines or containers. If a firm sends a network penetration testing checklist for a SaaS on a managed host, strike the lines that belong to the provider and keep the app lines; the same goes for network security assessment services, which suit an office more than an app. Remote penetration testing is the norm here, since nobody needs to visit an office to reach an app on the internet.

An API pen test works from the spec, so the tester can reach routes the front end never calls, and API penetration testing without that spec leaves the tester guessing. For GraphQL pentesting, hand over the schema. GraphQL’s docs on GraphQL introspection say “It’s often useful to ask a GraphQL schema for information about what features it supports”, name introspection as the way to do it, and add that “Disabling introspection in production is common in order to reduce the API’s attack surface.” If it is off in production, the tester cannot pull the schema from the live API, so give it to them.

Cloud providers publish what customers may test. AWS’s penetration testing policy welcomes AWS penetration testing of customers’ own AWS infrastructure “without prior approval for the services listed” under Permitted Services, while an AWS pentest of any other service means working “directly with AWS Support or your account representative.” Microsoft’s Azure penetration testing page says “You don’t need Microsoft’s pre-approval to do it, but you do need to follow the published rules,” and Azure penetration testing by an outside firm needs “explicit written authorization from the resource owner.” The pen-test decision article linked in the first section has more on the AWS rules.

The scope document and the rules of engagement, on one page

A small pen test is agreed on one page: the targets by hostname and route, the accounts provided, what is out, the dates and hours, the emergency contact and stop word, the signed authorization, the evidence rules, the report and retest window, and the price and days.

Pen testing scope is settled on paper before anyone sends a request, and for a small app it fits on one page. This is the pen test plan I would copy, a penetration test plan template of nine lines. To scope a small pen test, fill it in before the first call with the tester.

  1. 01 Targets: each hostname and route in scope, and whether it is production or staging (if staging, loaded with production-shaped data)
  2. 02 Accounts: the logins provided and their roles, with at least two customers so cross-customer access can be tried
  3. 03 Out of scope: third-party services you do not own, denial of service, and social engineering
  4. 04 Dates and hours: when testing may run, in one stated time zone
  5. 05 Contacts: an emergency contact on each side, and a stop word that halts testing at once
  6. 06 Authorization: written permission signed by whoever owns the systems, before any testing starts
  7. 07 Evidence rules: what the tester may keep as proof, such as a row count, never a customer record
  8. 08 Report and retest: the report format, and the window for retesting fixes
  9. 09 Price and duration: the fee and the number of testing days

The second and seventh lines are my working rules rather than any standard’s. The first three lines are the penetration testing scope, what will be tested, so they also settle what the scope of a penetration test is for your app. The third line keeps the security pen test on your own systems: third parties you do not own are theirs to authorize. The penetration testing rules of engagement are the fourth to seventh lines: when, who to call, what counts as proof, and the signed permission.

The sixth line is the one no test starts without. The PTES pre-engagement section says “It is critical that testing does not begin until this document is signed by the customer,” and splits this page into two halves: “While the scope defines what will be tested, the rules of engagement defines how that testing is to occur.” NIST’s glossary defines rules of engagement as “Detailed guidelines and constraints regarding the execution of information security testing.”

Manual penetration testing means a person works the targets on this page, choosing each next request from what the last one returned. That judgment is the part a scanner does not have, and it is what you pay for in a manual pen test. A pre-launch penetration testing checklist is the same page with the go-live date as the deadline, and the rest of the launch list is in the go-live checklist. For an automated or AI tool run, the safety checklist is in the pen-test decision article linked in the first section.

What a test costs is covered in affordable penetration testing, and who sells which shape of test in web application security testing services. The contract and the report that go with this page are in penetration test report format.

Pentest steps, and the security methodologies behind them (PTES, NIST SP 800-115, OWASP WSTG)

A pen test, seen from the client’s side, runs in seven steps: the signed scope, reconnaissance, scanning, exploitation of what the scope allows, post-exploitation kept to proof, the report, and the retest. The report should name its method: PTES, NIST SP 800-115 or the OWASP WSTG.

Penetration testing step by step, as the client sees it, looks like this; the order and the right-hand column are my framing.

StepWhat happensWhat you see from the client side
Signed scopeThe one-page scope and the written authorization are agreedA signed page; no traffic yet
ReconnaissanceThe tester maps hostnames, routes and technologies from public sources and the app itselfLittle or nothing
ScanningTools and the tester probe the targets for candidate weaknessesA burst of requests in your logs, inside the agreed hours
ExploitationThe tester tries the candidates the scope allows, to show what each gives an attackerTest accounts doing odd things; the emergency contact may get a call
Post-exploitation, kept to proofThe tester stops at proof: a row count, a screenshot, a changed test recordNothing leaves the app beyond what the evidence rules allow
ReportFindings with severity, proof and fixThe report
RetestThe tester replays each proof after your fixesA dated pass or fail per finding

The exploitation phase begins once the earlier steps have found something the scope allows the tester to try. What the exploitation phase aims to achieve is proof of impact: what that weakness gives an attacker, shown with a request and its response. PTES has a section with the same name, Exploitation, followed by one called Post Exploitation. In practice the penetration testing life cycle loops, since something found during exploitation can send the tester back to reconnaissance.

Pen testing steps are grouped differently by each method; these are the three named at the top of this section. PTES, short for the Penetration Testing Execution Standard, “consists of seven (7) main sections” (the full list is in the questions at the end). The PTES methodology covers the paperwork as well as the testing: its stages of penetration testing start with Pre-engagement Interactions, before any traffic is sent, and end with Reporting. As a framework, PTES stays high level: its main page says the standard “does not provide any technical guidelines as far as how to execute an actual pentest”, and points to a separate technical guide.

NIST SP 800-115, the “Technical Guide to Information Security Testing and Assessment”, shows four phases: Planning, Discovery, Attack and Reporting. It calls them “an example of how the penetration process can be divided into phases” and adds that “There are many acceptable ways of grouping the actions involved in performing penetration testing.” NIST 800-115, published in September 2008, is the reference when a report says it follows NIST penetration testing guidance.

The OWASP WSTG is the web test catalog: its latest edition lists individual tests from information gathering through authorization, session management and API testing, including GraphQL. The WSTG’s list of penetration testing methodologies names PTES and NIST 800-115 among others, and it also lists one literally called the Penetration Testing Framework.

A pen test framework, in a report, is the method the tester says they followed. Ask which penetration testing methodology the report will follow before you sign, and write the answer into the eighth line of the scope page. A pentest methodology named in the report gives the customer’s reviewer something to check coverage against; without one, the reviewer has only the tester’s word. Whichever of these penetration test standards the report names, the penetration testing process still ends with the report and your retest, and the client-side table above maps onto the pen test phases each one uses. Testing each OWASP Top 10 category has its own walkthrough in how to test OWASP Top 10 vulnerabilities. Types of penetration testing come in two sorts on this page: by what the tester knows (the box section) and by where the tester stands (the scope section). Social engineering and red teaming, next, are the other types of pen testing people ask about.

Social engineering, red team and adversary simulation: adjacent, and why they are not this

In NIST’s glossary, social engineering is “An attempt to trick someone into revealing information (e.g., a password) that can be used to attack systems or networks.” Social engineering pen testing points that at your people, by email or phone, where a web app test points at your code. Social engineering penetration testing is out of this page’s scope, and out of the third line of the scope document.

In cyber security, NIST’s glossary defines a red team as “A group of people authorized and organized to emulate a potential adversary’s attack or exploitation capabilities against an enterprise’s security posture.” That red teaming definition puts the adversary first, not a list of targets. The blue team is “The group responsible for defending an enterprise’s use of information systems by maintaining its security posture against a group of mock attackers (i.e., the Red Team).” Red team against blue team, then, is attack against defense, run as an exercise.

In my reading, a red team operation tests people, network and app together toward one objective, and adversary simulation is the commercial name for the same idea. Red teaming and penetration testing overlap in technique, but a pen test has a target list and a red team has an objective, so the red team methodology may skip the app entirely if a phishing email gets there faster. A red team assessment then reports whether the objective was reached and what the defenders noticed. Red team testing earns its cost when there is a blue team to measure; without a security function, nobody is there to exercise. All three suit organizations that have one. If a firm offers you social engineering testing services or adversary simulation services before your first app test, a scoped web app test is still the first test a small SaaS needs.

Continuous penetration testing, automated scans, and where AI testing fits

Continuous penetration testing means automated scans on every change plus a person testing on a schedule and after big changes. The scans give breadth; the scheduled human test gives the proof a customer’s reviewer asks for.

My working rule for the schedule: a person tests before a big customer contract, after a large release and after an incident. It is important to continuously conduct penetration testing in this lighter sense because code changes, and a pen test report describes the app on the day it ran. When continuous pentesting is sold as a subscription, ask what the person does each month and what the tool does.

The scan half of continuous security testing, and its tools, is covered by the website security check linked in the first section; the ongoing program around it is in vulnerability management tools. AI pentesting, what it counts for and how to run it safely, is covered in the pen-test decision article linked in the first section.

How to verify it: the test ran on the agreed scope and the fixes held

A pen test report lists the scoped targets, names the methodology, and gives each finding a severity on a named scale, proof without customer data, a fix a developer can act on, and a dated retest. Fix by severity, retest, then show the one-page summary.

Six checks tell you the test covered what you paid for and the fixes held; each names where the evidence sits.

  1. 01 Every target on the scope page appears in the report's coverage list, with what was tried against it
  2. 02 Every finding carries its proof, a request and response or a screenshot, in its evidence field, with no customer data in it
  3. 03 Every severity names its scale and that scale's version, CVSS or the tester's own
  4. 04 Every fix is written so a developer can act on it, in the finding's fix field
  5. 05 Retest findings after remediation: the tester replays each proof after the fix, as the same role and the same account the finding used, and records a dated pass in the retest section
  6. 06 You replay one fixed finding yourself from its proof steps, it no longer returns what the proof showed, and you keep a dated note of it

On the third check, CVSS “is currently at version 4.0”, and a score can be translated into a rating “such as low, medium, high, and critical”. On the sixth, judge the replay by what comes back: the other customer’s data, or the privileged action, should be gone, whatever status code the app returns. The report’s sections, its executive summary and how to read the results in full are in the report-format article linked in the scope document section.

In the Production Hardening Sprint, deliverable 3.10, targeted external penetration testing, is verified this way: record the five surfaces, authorized tests, findings, fixes, and retest outcomes. AxonBuild performs this targeted test. What to show a customer’s reviewer afterwards is a separate question: technical documentation for investors.

Where the sprint does this

Deliverable 3.10 tests the five highest-risk externally reachable attack surfaces, remediates findings, and delivers the methods and evidence. The package also includes an OWASP review and a controls checklist with evidence. Results go into the production readiness report, deliverable 13.1, which accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. Formal third-party certifications and independent audit opinions are separate from the sprint deliverables. Every deliverable is listed in the published scope.

Common questions about pen testing a small SaaS

What are the 7 steps of pen testing?

Seen from the client, the seven steps are: agree and sign the scope, reconnaissance, scanning, exploitation of what the scope allows, post-exploitation kept to proof, the report, and the retest of every fix. Other lists group the same work into fewer or more steps; PTES’s seven sections, listed two answers down, are one such grouping.

What are the three main types of penetration testing?

By what the tester knows, the three are black box (no knowledge of the inside), grey box (some knowledge, such as user accounts) and white box (explicit and substantial knowledge, including the code). Other lists sort by target instead and name external, internal and web application tests.

What are the 7 phases of PTES?

The Penetration Testing Execution Standard lists seven main sections: Pre-engagement Interactions, Intelligence Gathering, Threat Modeling, Vulnerability Analysis, Exploitation, Post Exploitation and Reporting.

Does NIST 800-53 require penetration testing?

For some systems. Control CA-8 in NIST SP 800-53 Revision 5 reads “Conduct penetration testing [Assignment: organization-defined frequency] on [Assignment: organization-defined systems or system components],” and the publication is a catalog of controls “for information systems and organizations”. SP 800-53B selects CA-8 in the high baseline only, not the low or moderate one. In my reading, a small SaaS has to meet it only when a customer or contract asks for that framework; this is not legal advice.

Can you pentest a website?

Yes. A website is the usual target of a web application pen test: its pages, forms, API and sign-in flows, as far as the signed scope names them, and only after the owner’s written authorization.