A web application security audit is a review of one app’s code and configuration against a fixed list of controls, with every finding checked against the running app and written up with proof, severity and a fix. My June and July 2026 audits produced 958 confirmed findings across 21 third-party apps, about 46 per app.
Those 21 apps are a selected set of third-party apps I audited, not a random sample, so the count is not a rate for AI-built apps in general, and each of the 958 is a human-verified finding rather than a scanner line. The audit is one piece of web app security: the piece that tells you, with evidence, which controls hold on one app today.
What a web application security assessment or audit is
A web application security assessment or audit reviews one app’s code and configuration against a fixed control list; each finding is checked on the running app and written up with proof, severity and a fix. A review is the same work under a softer name; a penetration test is a different product.
On a web app security assessment, the auditor does four things: reads the code, tests the controls, tries each candidate finding against the live app, and writes up what held and what failed. Assessment, review and auditing are names for that same work, so web application security auditing follows the steps below whatever the contract calls it, and the same holds for a security review: what it checks and who can run it depend on the app, not on the name. A test in the narrow sense is web application penetration testing: someone attacks the running app from outside and reports what got through, with its own method and its own report.
An information security audit, in the organization-wide sense, covers networks, policies and compliance across a whole company, and it is a different discipline. NIST’s glossary definition gives that general meaning: “Independent review and examination of records and activities to assess the adequacy of system controls, to ensure compliance with established policies and operational procedures, and to recommend necessary changes in controls, policies, or procedures.” Application auditing narrows that to one piece of software and the data it touches.
For one app, the purpose of a security audit is proof that its controls hold, kept in a form the owner can hand to a customer or an investor. For an app built with an AI tool, whether you need an AI app security audit at all is its own decision, with its own costs and its own checks.
Why it matters for an app you did not write
When you inherit an app, or buy one, or sign off on one a contractor built, you do not know which shortcuts were taken. The person who took them is often gone, and the code does not say where a check was skipped.
The numbers from my own work point the same way. Of the 26 apps I audited, 22 had at least one confirmed-critical finding. Across all 26, the readiness scores range from 29 to 81 out of 100, with a mean of 52.1 and a median of 51. Sorted by band, 22 of the 26 landed in the RED band, 4 in AMBER and 0 in GREEN.
Those figures come from my June and July 2026 audits of 26 apps that I picked, not apps drawn at random. They show where problems turn up in that group; they do not measure how often AI-built apps fail across the market.
My reading of them: a scanner reports patterns, and an audit’s value is the checked path from input to consequence. A scanner can flag a missing header or a package with a known flaw. It cannot tell you that a signed-in customer can open another customer’s invoices because the ownership check sits in the browser, since that takes following one request from the form to the database and back. For an app someone made with an AI builder, the vibe coding security audit covers what changes in that kind of codebase.
How a web application security audit is run
A web application security audit runs in seven steps: the scope in writing, an inventory of entry points and data, the control walk with a test per item, code reading along the risky paths, verification of each finding against the running app, ranking, and a report with a retest. Steps one, three and five are the minimum.
To perform a security audit on one web app, take the steps in order, because each one feeds the next. The list is what occurs during a security audit that is done properly:
- 01 Write the scope down: which repository, which environments, which surfaces
- 02 Inventory the entry points, roles, data stores and third-party calls
- 03 Walk the control list, with a test for every item
- 04 Read the code along the risky paths: sign-in, payments, uploads and the AI feature
- 05 Check every candidate finding against the running app
- 06 Rank what was confirmed, so the fixes go in the right order
- 07 Write the report, fix, and retest
The scope comes first and goes in writing: which repository, which environments (staging, production or both), and which surfaces, such as the web app, its API, the admin panel and any mobile client. Anything the scope does not name is not audited, which is why the report ends with a list of what was left out.
The inventory lists every entry point (pages, API routes, webhooks, scheduled jobs), every role, every data store and every call that leaves the app for a third party. It is the map the later steps check against, and the architecture review below draws the same map as a diagram.
The control walk takes a list and tests each item: the application security checklist with a test per item, and the OWASP Top 10 categories, where knowing how to test the OWASP Top 10 vulnerabilities one at a time matters more than naming them. For a web application security testing guide to walk the controls against, one free reference is the OWASP Web Security Testing Guide, which its project page calls “the premier cybersecurity testing resource for web application developers and security professionals.” Its project page says “v4.2 is currently available as a web-hosted release and PDF” and that release version 5.0 is being developed. Web security testing tools cover part of this step: a website security check proves some controls from outside, never the ones that live only in the code.
Code reading follows the risky paths, not every line: sign-in and session handling, payments, file uploads, and the AI feature if the app has one. The reviewer traces each path from the request to the database and back, looking for the place where a check should be and is not.
Verification is the step that separates an audit from a scan. Every candidate finding, whether it came from a tool, the checklist or the code, is tried against the running app; one that cannot be shown is dropped or marked unconfirmed.
Ranking sets the order of the fixes. In the same audits, the 958 confirmed findings across the 21 third-party apps break down as 58 critical, 362 medium, 403 smell and 135 hygiene. Critical items are a small part of that list, so a report that marks everything high gives you no order to work in.
The report writes each finding up with proof and a fix, and the retest re-runs each fixed finding’s proof to show it no longer works.
Written down like this, the list is the security audit procedure you can ask a firm to show before the work starts, and whatever web application security testing methodology a firm names should map onto it; one with no verification step leaves you with unconfirmed tool output. The security audit process stays the same whether an outside firm or your own engineer runs it. For an app built with an AI builder, the vibe coding security audit article walks the same ground stage by stage, from boundary to priority; the seven steps here are the general version.
If time is short, my working rule is to make a minimum security audit out of steps one, three and five: the written scope, the control walk and the verification. A web application security questionnaire from a customer asks for the results of steps three and five in its own format, and a security questionnaire is far quicker to answer once those results exist.
In the Production Hardening Sprint, two deliverables carry this kind of work. Deliverable 3.7, the OWASP Top 10 review, is verified this way: we deliver category-level results and supporting test evidence, with justified non-applicable cases identified. Deliverable 3.10, targeted external penetration testing, is verified this way: we record the five surfaces, authorized tests, findings, fixes, and retest outcomes, and AxonBuild performs this targeted test.
What a finding has to contain
A finding has six fields: where it is, the proof, the consequence in plain words, the severity with its reason, the fix, and the retest result. A finding without proof is a guess, and a report of guesses is a scanner export with a cover page.
In a security test audit report, every finding carries all six. The example column follows one constructed finding, not one from a real app:
| Field | What goes in it | Example (constructed) |
|---|---|---|
| Location | The file and line, or the route | The invoices route handler, at the line that loads the record, and the route GET /api/invoices/:id |
| Proof | The request and response, or the query, that shows it | Signed in as user B, a request for user A’s invoice id returns user A’s invoice |
| Consequence | What an attacker gets, in plain words | Any signed-in customer can read any other customer’s invoices |
| Severity, with the reason | The rating and why it earned it | Critical: customer data exposed to any account, with no special access needed |
| Fix | The change, in the place it belongs | Load the invoice only where its owner matches the session user, on the server |
| Retest result | The proof re-run after the fix | The same request as user B now returns not found; user A still sees the invoice |
The standard I hold my own audits to is one example of what “verified” can mean in practice. A held-out validation run produced 0 cry-wolf false positives across 10 apps, recall of 1.0 against human-verified ground truth, and 10 out of 10 correct N/A gating, the call on which checks do not apply to an app. The recall counts confirmed findings, and the 10 held-out apps are a selected set, not a sample of apps in general. Separately, the 11 apps in my deep audits went through 36 independent re-verify runs across the 12 pillars of the audit, with zero regressions and zero new false positives. What a verified finding’s evidence looks like for a builder-made app is set out in the vibe coding security audit article.
The architecture security review
An architecture security review is the audit’s inventory drawn as a diagram: where the trust boundaries are, where authorization is enforced, where secrets live, and which calls leave the app. For one small app it comes down to eight questions and, by my working rule, an hour or so with the diagram.
The review needs a diagram first, and knowing how to create an architecture diagram with boxes, arrows and boundaries is most of the preparation. As a security architecture review checklist, the eight questions fit in one table:
| Question | What a reviewer looks for |
|---|---|
| Where are the trust boundaries? | Every point where data from a user, a browser or another service enters code you trust |
| Where is authorization enforced? | On the server, on every read and write, never only in the browser |
| Where do secrets live? | In server-side settings, never in the client bundle or the repository history |
| Which calls leave the app? | Each third-party API, webhook and AI provider call, and what data it sends |
| Which data is personal? | The tables and fields that hold personal data, and who can read each |
| Where do logs go? | Where logs are stored, who can read them, and whether secrets or personal data end up in them |
| What runs with admin rights? | Service keys, admin roles, background jobs and database functions that skip normal checks |
| What is the rollback? | How a bad release or a bad migration is undone, and whether anyone has done it |
A security architecture review questionnaire asks the same eight in writing, for a customer or a new engineer to answer. Run early, as a secure architecture review of a design that is not built yet, the questions catch a missing boundary before anyone codes around it. An application security risk assessment then ranks what the review found, and a security risk assessment template, at its simplest, is a table with three columns: the risk, its likelihood and its impact. The deeper version of that ranking is threat modeling.
What the report has to contain, and how to judge one you are handed
An application security audit report has eight sections: scope and dates, method, inventory, findings with proof, a severity summary, the fix list with owners, retest results, and what was not tested. A report with no proof column, no retest and no list of what was left out is a scanner export.
| Report section | What it must hold | Red flag |
|---|---|---|
| Scope and dates | The repository and commit, the environments, the surfaces, and the dates tested | No commit or dates, so nobody can tell which version was checked |
| Method | The steps followed, the standard walked, the tools used and what each covered | An automated scan given as the whole method |
| Inventory | Every entry point and user role, where data is stored, and which outside services the app calls | Missing, or copied from the owner’s own description |
| Findings with proof | Every finding with its six fields | No proof column |
| Severity summary | Counts by severity, each rating with its reason | Ratings with no reason, or everything marked high |
| Fix list with owners | What to change, who owns it, in what order | Generic advice with no file or route |
| Retest results | Each fixed finding re-run, with the new result | No retest, or a fix taken on the owner’s word |
| Not tested | Surfaces, roles and environments left out, with the reason | Missing, so the report reads as if it covered everything |
If the method names a standard, the OWASP ASVS is one to look for; its latest stable version is 5.0.0, from May 2025, and it defines three verification levels, saying of Level 2 that “Most applications should be striving to achieve this level of security.” The standard leaves the choice to the owner: an organization “should analyze its risks and decide what level it believes it should be at”.
Copy the table and it works as a web application security audit report template, one row per section, filled in by whoever ran the audit. A firm that calls its work an assessment should fill an application security assessment report template with the same sections, even under other headings, and a sample security assessment report is this table filled in for one app. When you receive a report, the same eight rows work as an application audit checklist: tick each section, then read its red flag. A customer’s application security assessment questionnaire, in my reading, covers much of the same ground in its own questions, so answer it from the report rather than from memory.
Read as application security review questions, the table also tells you what to ask a firm before you sign: what each section will hold. For an application security risk assessment, the checklist and the template are simpler: the three-column table of risk, likelihood and impact described under the architecture review above.
Five red flags tell you an audit you paid for is weak. There is no proof column. Severity comes with no reason. Scanner output is pasted in as findings. Nothing was retested. And there is no list of what was not tested. When a firm sends a security assessment report example before you buy, look for the proof column first.
A founder buys a website security audit to answer a customer’s security questionnaire. The report covers the marketing site’s web server and its HTTP security headers, the kind of checks ImmuniWeb’s public website test lists when it offers to “Test your website and web server for security, privacy, encryption, protection from data scraping, and compliance with GDPR and PCI DSS”, while the app and its API live on a subdomain the written scope never named. The questionnaire asks about the app, and the report says nothing about it. The lesson I take from it: an audit proves only what its written scope names, which is why the list of what was not tested matters as much as the findings.
If what you are buying is a code audit service rather than an audit of a running app, judge its sample report against that buyer’s questions instead. Comparing the firms that sell web application security testing services is a separate job from judging the report they hand back.
How to check your own app before the audit
A pre-audit self-check is the application security checklist run in order by the founder, with the evidence kept for each item. The audit then starts from a known state, and the auditor’s days go to the paths a checklist cannot reach.
How to do web application security testing yourself, before an auditor starts, comes down to that checklist with a test per item, run against staging with two test accounts and the evidence saved as you go. The list is not repeated here, because the application security checklist already holds a test for each item. For an app built with an AI tool, the AI app security audit article gives seven checks to run before you pay anyone.
For a builder-made app, a vibe code security check shows what the build tool covers and what is still yours. After the audit, the ongoing work is vulnerability management, where vulnerability management tools keep watch on new flaws in the packages you ship.
No single tool is the audit. A scanner covers part of the control walk and none of the code reading or verification, and tools are best judged against the step they cover.
What it costs and how long it takes
Website security audit cost is usually set per scope or per hour; that is my reading of how firms quote, not a published survey. One seller’s published figure, checked on 2026-09-30: Astra lists its Pentest Expert plan at $5,999 a year for one target, and counts “One web or SaaS app” as one target, “including all APIs consumed.” That plan includes “Manual Pentest by certified experts in OWASP, APTS, SANS, PTES standards” and “2 Re-scans by experts to verify fixes”.
Three things move the price most: how much the scope covers, whether fixes are included or left to you, and whether a retest is included. A quote that leaves out the retest leaves you to prove the fixes yourself.
One app with one repository in scope takes days, not weeks, by my working rule; several apps, a mobile client or a large API add to that. Comparing sellers on price is part of choosing among web application security testing services, a separate decision from how the audit is run.
Where the sprint does this
In the sprint, deliverable 13.1, the production readiness report, delivers the result for every scope item, the work completed, and its verification evidence, and it is verified this way: we account for all 123 IDs, keep failures visible until resolved and explain genuine non-applicable items. Formal third-party certifications and independent audit opinions are separate from the engineering deliverables. The OWASP Top 10 review and the targeted test named under the audit steps are two of the items that report accounts for, and every item is listed in the published scope.
Common questions about auditing a web app’s security
What happens during a security audit?
A reviewer writes down what is covered, maps the app’s entry points and data, tests each control on a list, reads the code along the risky paths, confirms each finding on the running app, ranks the results, and hands over a report followed by a retest.
How to perform a security audit?
Start with a written scope, test every control on your list, and confirm each finding against the live app; those three are the least an audit can be. Then rank what held up, write each finding with its proof and fix, and re-run the proof once the fix is in.
What does a security audit consist of?
A security audit consists of the work and the report, and the report is the part you can check. It should say what was covered and when, how it was tested, what the app exposes, each finding with its proof, a tally by severity, who fixes what, what the retest showed, and what was left out.
What should a security assessment include?
A security assessment should include, for every finding, its location, the proof, the consequence, a severity with its reason, the fix and a retest result. As a whole, it should also say which parts of the app it did not test.
How to write a security assessment report?
Write it in a fixed order: scope and dates, method, inventory, findings, severity summary, fixes with owners, retest, and exclusions. Put proof under every finding so a reader can reproduce it without asking you.
What is the main purpose of a security audit?
The main purpose of a security audit is evidence: a written record that an app’s controls hold, detailed enough that an owner, a customer or an investor can check each claim for themselves.
The checks in this guide show you where the app is open. The sprint below closes those gaps, tests the result and writes the evidence down.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase