OWASP’s Secure Code Review Cheat Sheet lists the steps of a baseline review, from architecture to configuration, and none of it is wrong. It says nothing about code an AI assistant wrote. My audit found 11 of 21 apps from other teams letting a request with no login check reach privileged actions, so this secure code review checklist starts there.
The secure code review checklist: eight passes, and the order to run them
A secure code review checklist for an AI-built app runs 8 passes: entry points, trust in the client, authorization per object and tenant, cost and abuse limits, input and prompts, money paths, secrets, then dependencies. The first three decide who can do what, and a missing check leaves no pattern for a tool to match, so they come first.
The counts on this page come from the 21 third-party apps I audited in June and July 2026: 11 public vibe-coded apps audited exhaustively across all 12 pillars, and 10 held-out apps, audited blind. Those audits produced 958 confirmed findings across the 21 apps, about 46 per app, every finding adversarially verified. They are a selected set, not a random sample, so no count below is a rate for AI-built apps in general. The 14 AI apps in passes 4 and 5 are the third-party AI apps inside that set, and the 26 in pass 8 is every app I audited, my own five production apps included. Of the engineering standards for AI-assisted teams, this is the one written for whoever reads the code.
Read the table as a code review security checklist: one row per pass, run from the top. The order is my working rule, not a ranking by count: access passes go first because an absent check is invisible to pattern matching, and the last two go last because scanners and published advisories catch most of what they look for, which is why secrets sit seventh despite their count.
| Pass | What to open first | What to look for | How often it showed up in my audits | Reading enough, or a test needed |
|---|---|---|---|---|
| 1. Entry points | The route inventory: every route, server action, function and webhook | Anything that changes data, reads private data or spends money with no session check on the server | 11 of the 21 third-party apps had unauthenticated endpoints doing privileged work | Test needed |
| 2. Trust in the client | The handlers that read the request body | A user id, role, price, plan or outbound URL taken from the request instead of decided by the server | 10 of the 21 third-party apps trusted the client: the server accepted whatever the browser asserted | Reading usually finds it; the fix gets a test |
| 3. Authorization per object and per tenant | Every query that takes an id, and every table’s row-level security policy | A lookup by id with no owner or tenant condition; a policy that allows more than its name says | 7 of the 21 had confirmed cross-user or cross-tenant authorization failures, where a logged-in user could read or write another customer’s data; 9 of the 21 had row-level security gaps | Test needed |
| 4. Cost and abuse | The most expensive endpoint: model calls, sandboxes, email, uploads | What limits it, keyed to which identity, and what happens when the limiter’s own settings are missing | 13 of the 21 had no rate limiting on their most expensive endpoint; 12 of the 14 AI apps had a confirmed denial-of-wallet path, where a stranger or free account can burn the owner’s paid AI or compute bill without limit | Test needed |
| 5. Input, output and the model | Validation at the server boundary, raw SQL, HTML rendering, file uploads, prompt assembly | A request value reaching a query, a page, a file path or a model’s instructions unchecked | 8 of the 14 AI apps had a live prompt-injection path, with untrusted text flowing straight into the model’s instructions | Reading finds the path; a test confirms the rejection |
| 6. Money paths | The payment webhook handler and the code that grants a plan | The provider’s signature checked before the payload is used; the plan decided on the server from the provider’s event | Not counted on this page | Test needed |
| 7. Secrets and configuration | The built bundle, the repository history, the environment mapping | A private key in anything the browser downloads or the history keeps | 6 of the 21 third-party apps shipped a real secret; the Secrets & Credentials pillar still averages 84.4 out of 100, scored on all 21, the best of the 12 pillars | Test needed: a scan of the build |
| 8. Dependencies and framework version | The lockfile and the framework’s version | That version against its published advisories | 9 of the 26 audited apps ran a framework version with a publicly known, reachable RCE or auth bypass, and the fix was often a one-line version bump | Reading matches the version; reachability decides the priority |
The first pass needs the full list of ways in, which the next section builds before any reading starts. On each route, the session check has to sit inside the handler, not in the page that calls it.
Trust in the client means the server believes a value it should have decided itself: the user id in the body, the price in the cart, the plan in a hidden field. The fix is to recompute trusted values on the server from the session and the database.
Row-level security policies are read against what the table’s name promises. A policy on an orders table that lets any signed-in user select every row says more than the name does, and a lookup by id needs an owner or tenant condition in the same query. Reading cannot prove isolation, which is why this row’s test sits in the reading-stops table further down.
Cost and abuse is about money the app pays out, not money it takes in. Follow the limiter from the expensive endpoint: which identity it counts (an IP address, a user id, an API key), and what it does when its configuration is absent. A limiter that fails open passes every read-through and stops nothing. The control itself is rate limiting in an API.
For input and output, follow each request value to the places it can do harm, and in an AI app the prompt is one of those places. Untrusted text that lands inside a system prompt is the case to look for, and knowing what prompt injection is tells you which model outputs to distrust afterwards.
Money paths get a pass of their own because a payment webhook is a public endpoint that grants access. Two things to find: the signature check before anything reads the payload, and the plan decided from the provider’s event, never from a page the browser lands on after checkout. Both belong to webhooks security.
Secrets are a build question as much as a source question. A key can be absent from every file in the repository and still sit in the JavaScript bundle the browser downloads, or in a commit from months ago. Secret scanning covers both, which is why this pass leans on a tool.
The last pass is a lookup: read the framework and its version from the lockfile, then check that version against the project’s published advisories. Whether the vulnerable code path is reachable in this app decides how urgent the upgrade is. On a Node project, the npm audit command sends a description of the project’s dependencies to the registry and asks for a report of known vulnerabilities, and by default npm requires a package-lock or shrinkwrap to run it.
This security code review checklist is for the person who reads the code. An application security code review of an app you cannot read yourself starts with how to secure a vibe-coded app when you can’t read the code, and the same risks as twelve checks with a test each, written for a founder, are in the vibe coding security checklist. The starter secure code review checklist softwaresecured keeps on GitHub is a list of categories, with headings from Information Gathering to Log Management, and it carries no severity or prioritization.
How to run a code review for security, pass by pass
A code review for security runs in the same order every time: build the route inventory, run the scanners and treat their output as leads, choose the riskiest paths, read each one end to end, then test whatever reading cannot prove. The route inventory is the spine the rest of the review ticks off.
I give no fixed duration. Reading a whole app has no published number, and the walk-through of how long a code review takes explains why. I set the time box from the size of the route inventory instead, and because the passes run in order, a review that stops early has still covered who can do what.
Before you read: the route inventory, a scanner run and the riskiest paths
- 01 Build the route inventory: every route, server action, function and webhook through which a request reaches the backend. Mark each one that touches money, touches personal data or calls a paid API.
- 02 Run the scanners and file every flag as a lead to check, never as a finding.
- 03 Choose the riskiest paths first. My working rule is about ten, taken from the routes marked for money, personal data or a paid API.
- 04 Trace each of those paths from the request to the database and back: from sources such as user inputs, file uploads, API calls, database reads and environment variables, to sinks such as database queries, file writes, output rendering, logging and external APIs.
- 05 Run the eight passes over those paths, then over the rest of the inventory as the time box allows, and give every route a result.
The route inventory is the same list that work on broken access control starts from; the three marks per route are my working rule.
Scanner output goes in as leads for a reason. In 3 of the 21 third-party apps, scanners flagged 33 to 44 vulnerabilities and the audit traced exactly zero as reachable. Those apps come from the same June and July audits, a chosen set rather than a sample of all AI-built apps. OWASP’s cheat sheet makes the point in its own terms: its SAST triage line has automated findings set the priorities for where a reviewer reads first. Which scanner to run, and what static analysis can see, belong to static code analysis tools and what is a code scan; what a scanner report proves and cannot prove is covered in the audit article linked in the next section. The tracing step is that same static analysis technique, done by a person reading.
An AI reviewer is one more source of leads. Picking the best AI code reviewer, or a CodeRabbit alternative is its own question, and so is whether ChatGPT can review your code and which prompt to give it.
The OWASP code review guide and the Top 10: what each one is for
The OWASP Code Review Guide is OWASP’s comprehensive guide to reviewing code for security vulnerabilities and security bugs, the Secure Code Review Cheat Sheet is the working reference, and the OWASP Top 10 describes itself as a standard awareness document. In a short review, the Top 10 works best last, as a coverage check on the passes.
| OWASP resource | What it is | How to use it in a short review |
|---|---|---|
| Code Review Guide | OWASP’s long-form guide; its own repository calls it the Secure Code Review Guide. The published edition is version 2.0, from 2017; the repository that announced a 3.0 release was archived, read-only, in April 2025. | The long reference, when a finding needs the method behind it |
| Secure Code Review Cheat Sheet | Baseline reviews of the whole codebase and diff-based reviews of changes, with review steps, checklists, and templates for a finding report and a review summary | The working reference beside the eight passes |
| Secure Coding Practices checklist | Archived: OWASP says the project “has now been archived” and its checklists “have also been migrated to the Developer Guide”, which says checklists “are also useful during code reviews and design activities” and “are not meant to be followed in their entirety” | The fix reference, read in its Developer Guide home rather than the old pages |
| OWASP Top 10 (2025 edition) | In OWASP’s words, “a standard awareness document for developers and web application security”; ten categories, from A01 Broken Access Control to A10 Mishandling of Exceptional Conditions | The coverage check at the end |
All four sit on OWASP’s own sites: the OWASP Code Review Guide, OWASP’s Secure Code Review Cheat Sheet, the OWASP Developer Guide’s web application checklist and the OWASP Top 10:2025.
Used as the running order, an OWASP Top 10 code review works through risk categories, while the passes follow the path a request takes through one app. My method puts the Top 10 last: once the passes are done, an OWASP Top 10 source code review becomes a coverage check, with every category marked either covered by a pass or not applicable, and the reason written down. Where AI-built apps landed on six of the ten 2025 categories is already tabled under where AI-built apps land on the OWASP Top 10, and the same article holds the scanner section step 2 pointed to. Testing each category is a separate job, how to test OWASP Top 10 vulnerabilities, and so is the background on what OWASP is and what it publishes.
Reviewing code for quality and security in the same sitting
Reviewing code for quality and security in one sitting works when the security passes run first and quality notes go on a separate list. In AI-built code, duplicated handlers and unchecked API boundaries are where security defects hide.
The two reviews read the same files and ask different questions. Duplicated handlers can mean an auth check fixed in one copy and not the other, which is a security reason to know how to find unused code in a repo. An API boundary nobody typed means the server can trust a shape it never checked, the same untyped seam that leaves a team saying the frontend broke after a backend change. Why two names for one thing usually mean two versions of the same logic, and why consistency turns into a security question, is in the code review walk-through linked above.
The broader list and its report belong to the general code review checklist. My one rule for doing both at once: a quality note never rides inside a security finding, so each finding keeps one owner and one retest.
Where reading stops and a test has to run
Reading code shows what it says, not what it does. Six claims need a test before a review can call them done: authorization, tenant isolation, rate limits, webhook signatures, input validation and secret exposure. Each test runs against a system you own or are authorized to test, and leaves dated evidence.
| The claim | Why reading cannot settle it | The test that does | The evidence to keep |
|---|---|---|---|
| Authorization | A check that exists is not a check every path runs | Call each protected action as a signed-out user and as a lower-role user; each is refused or gets nothing back | The refused requests, with status and date |
| Tenant isolation | A policy’s text is not the rows it returns | Read and write as each identity (signed out, each role, a user in a second tenant); each sees and changes only what it should | Each identity’s result |
| Rate limits | A limiter that exists can be keyed on the wrong value or fail open | Send a burst past the limit, then one ordinary request from a second account | The burst’s refusals and the second account’s request that went through |
| Webhook signatures | Verification code can sit off the path the provider actually calls | Have the provider send one genuine test event, then send one altered copy of an event | Both results |
| Input validation | A declared schema is not proof that every writer uses it | Send invalid and hostile values to each endpoint that writes data | The rejected requests |
| Secret exposure | Clean source can still build into a bundle that holds a key | Scan the built bundle as well as the source, and watch the browser’s network traffic | The scan output and the traffic check |
In the Production Hardening Sprint, the six tests above map to six deliverables, each with a written verify step. Deliverable 1.3, server-side authorization: call protected actions directly as unauthorized and underprivileged users; confirm rejection. For deliverable 1.4, data isolation across every table: run read and write tests as anonymous users, different roles, and separate tenants. Deliverable 3.6, public-endpoint rate limits: simulate bursts and verify enforcement without blocking ordinary use. For deliverable 5.1, webhook signature verification: confirm valid events succeed and invalid or altered payloads are rejected. Deliverable 3.1, server-side request validation: submit invalid and hostile payloads and confirm safe rejection or handling. For deliverable 2.1, frontend secret removal: scan source and built assets and inspect browser traffic for private credentials.
My working rule is to run these in staging where you can, and on someone else’s app only with the owner’s written permission; the same rule is why no payloads or bypass strings appear on this page.
A filled example: one row of the checklist, written up
One finding from an app I audited, an AI coding workspace, fits this checklist’s row fields as written below. The finding is the report’s; the fix and the retest are my recommendation.
Pass: 2, trust in the client
What I opened: the server code that sends AI requests
What I looked for: values the server takes from the request
What I found: the AI request's base URL came from the client, while the
API key fell back to the owner's server-side key, with no
allow-list on the URL. One anonymous request pointing the
app at an outside server would deliver the owner's AI key
to that server's logs.
Reading or test: reading shows the path
Fix: the server sets the base URL itself, or checks it against
an allow-list. After the fact, rotation is the only remedy.
Retest: a request naming an outside host is refused
File, line and commit are left out here because they would identify the repository; the review record keeps them. My working rule is never to prove a finding like this by sending a real key anywhere: the code path is the evidence, and the retest proves the fix.
This is the class pass 2 exists for, and its count sits in the table above; the app comes from the same June and July 2026 audits of 21 third-party apps, a set I chose rather than sampled. The evidence a verified finding carries beyond this one row is laid out under what evidence a verified finding should contain.
A value the server should decide, taken from the request instead, is a finding that reading can show on its own. The test belongs to the fix.
How to verify the review happened
A secure code review is verifiable when someone else can check five things: a result for every route in the inventory, findings that carry their evidence, a recorded run for every test the checklist asked for, the scanner output marked up, and a retest for every fix, all tied to one commit.
These five checks are my working rule, and each is one that someone other than the reviewer can run and fail.
- 01 Every route in the inventory has a pass result or a written reason it was skipped. Evidence: the inventory with its result column.
- 02 Every finding carries the evidence fields the filled example pointed to, plus a location someone else can open. Evidence: a second person reproduces one finding from the record alone.
- 03 Every row the checklist marked as needing a test has a recorded run. Evidence: the request and the response (for the secrets pass, the scan output), with the date.
- 04 The scanner output sits beside the findings, with each flag marked reachable, not reachable or not checked. Evidence: the marked-up report.
- 05 Every fixed finding has a retest that passes after the fix, and a run that failed before it wherever running it before the fix was safe. A finding like the filled example, where the unfixed path would hand a real key to an outside server, is shown by its code path instead. Evidence: the results, with their dates.
OWASP’s Review Summary Template already asks for the version or commit hash and the scope of what was reviewed; the five checks are what I add on top. From the buyer’s side, what the written report should contain is set out in the audit article the filled example links to. A founder who cannot read the code and wants to check a review they were handed should start with code review for non-developers.
In the Production Hardening Sprint, deliverable 3.7, the OWASP Top 10 review, is verified this way: deliver category-level results and supporting test evidence, with justified non-applicable cases identified. The targeted security tests cover the five priority attack surfaces documented in the report. Formal third-party certifications and independent audit opinions are separate from these engineering deliverables. Deliverable 3.7 sits with the rest on the published list of all 123 deliverables.
The record names the commit it covers, because the next prompt can change the files it describes. Keeping fixed things fixed after that is the work of guardrails for AI coding agents.
Owning an app means being able to run it, change it and recover it without guessing. The sprint below leaves you with the runbooks and documentation to do that.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase