The first enterprise customer’s IT team asks for the date of your last security review, and there is no date, because nobody has read the app against a written list. 22 of the 26 AI-built apps I audited had at least one confirmed-critical finding. A security review is that read: the list, the evidence for each line, and a date.
What a security review is, and what it checks
A security review is a person reading the app against a written list and keeping the evidence for each line, with a date. A tool-run review, such as Claude Code’s security review command, checks code for common vulnerability classes; Anthropic’s own guide says its reviews should complement, not replace, existing security practices and manual code reviews.
The 26 apps behind the opener’s number are every app I audited in June and July 2026: 21 public and held-out third-party apps, plus five of my own. I chose them, so they are a selected set, not a random sample, and 22 of 26 describes those apps, not AI-built apps in general. A security review is the part of web app security that proves the other parts are in place.
The reader can be you, a teammate who did not build the app, or an independent third party; once it is someone independent, it is an audit, and the web application security audit is that version. The kinds of review people name (internal, customer, vendor and code review) differ mostly by who asks or who reads: your own team on a schedule, a prospect’s IT team before signing, you checking a supplier before sharing data, or an engineer reading the code and the design. For a security review at a startup, the internal kind comes first, because it is the one you control. If the app is already live, the security review is the same one, run against production settings rather than a staging copy, and a review prompt from a Reddit thread can feed into it but cannot stand in for the written list and the evidence.
Tools help. Anthropic’s guide to automated security reviews says Claude Code’s review checks for “SQL injection risks, cross-site scripting (XSS) vulnerabilities, authentication flaws, insecure data handling, and dependency vulnerabilities.” What the command covers and misses is set out in what Claude Code’s own security review does and does not cover, and the prompts to paste into a builder before launch are in security prompts for your AI builder. The review is what checks whether the prompt’s output holds. My pointers for an internal security review are short: use a written list, keep evidence per line, and date it. The twelve lines below are the minimum I’d make a security review out of.
Why it matters before the first customers
The first customers ask about app security before they sign, and the answer they want is a date and a list of findings. Here are four failures from the same audits, with the baseline line that catches each one.
| Finding | How many | The baseline line that catches it |
|---|---|---|
| Unauthenticated endpoints doing privileged work | 11 of the 21 third-party apps | Line 1, server-side session checks |
| Confirmed cross-user or cross-tenant authorization failures, where a logged-in user could read or write another customer’s data | 7 of the 21 third-party apps | Lines 2 and 3, server-side roles and tenant isolation |
| A real secret shipped | 6 of the 21 third-party apps | Lines 4 and 5, no secrets in the repository, separate keys |
| A framework version with a publicly known, reachable RCE or auth bypass; the fix was often a one-line version bump | 9 of the 26 audited apps | Line 8, dependencies scanned and the framework current |
The 21 are public and held-out third-party apps that I picked and audited in June and July 2026; the framework row’s 26 adds my own five. Nobody drew them at random, so the counts say what turned up in that set, not how every startup’s app is built.
A regulator’s case shows the same gap from the other side. The complaint in the FTC’s case against Verkada, a security camera company, filed on August 30, 2024, says the company failed to develop adequate security vulnerability management practices, such as “testing, auditing, assessing, or reviewing its products’ or applications’ security features” and “conducting regular risk assessments, vulnerability scans, and penetration testing of its networks and databases”. It says a consulting firm’s posture assessment, shared in February 2021, identified critical and high-level gaps, that the company “failed to address known security gaps”, and that on March 8, 2021 an intruder gained access to a support account with administrative-level privileges and, through it, had access to over 150,000 live customer cameras. The lesson I take from it, in my reading and not the regulator’s words: a written review with a date is the cheapest control a small team can show, because its absence is what a regulator or a customer names first.
Handling security before your first launch means running that review while nobody is waiting on it. After launch, the same question tends to arrive as a questionnaire, and answering it row by row is covered in a software security assessment questionnaire from a big customer. Someone now responsible for security at a scaleup runs the same baseline and puts the next review’s date on the calendar.
The baseline: twelve lines an early-stage startup needs, with the evidence for each
An early-stage startup’s security baseline is twelve lines: server-side session checks, server-side roles, isolation on tenant tables, no secrets in the repository or its history, separate keys per environment, validated input, rate limits, scanned dependencies, security headers, alerts that reach a person, a restored backup, and a deploy gate. Each line needs evidence.
NIST’s glossary defines a control baseline as “The set of controls that are applicable to information or an information system to meet legal, regulatory, or policy requirements, as well as address protection needs for the purpose of managing risk.” For one early-stage app, that set is the table below, and the best security baseline for early-stage startups is the one you can show evidence for, line by line.
| # | Line | What must be true | The evidence a reviewer accepts |
|---|---|---|---|
| 1 | Every route checks the session on the server | Permissions are enforced on every protected route, API endpoint and server action | Call protected actions directly as unauthorized and underprivileged users; confirm rejection |
| 2 | Roles enforced on the server | Admin routes and operations are limited to explicit, verified roles | Test administrative operations using both authorized and ordinary accounts |
| 3 | Isolation on every tenant table | Row-Level Security or equivalent server-side scoping on all tables, with explicit rules for intentionally public data | Run read and write tests as anonymous users, different roles, and separate tenants |
| 4 | No secret in the repository or its history | Git history scanned for committed secrets, and every exposed credential rotated | Retain scan results and confirm exposed credentials have been invalidated and replacements work |
| 5 | Separate keys per environment | Development, staging and production have their own variables, databases and keys | Inspect environment mappings and confirm staging actions stay within staging services |
| 6 | Input validated on the server | Every data-writing endpoint uses explicit schemas and safe input and output handling | Submit invalid and hostile payloads and confirm safe rejection or handling |
| 7 | Rate limits on public endpoints | Public forms, signup flows and resource-consuming endpoints have appropriate limits | Simulate bursts and verify enforcement without blocking ordinary use |
| 8 | Dependencies scanned, framework current | Known-vulnerable packages upgraded or removed, unused packages removed | Record the dependency scan, remediation, and regression checks after changes |
| 9 | Security headers set | CSP, HSTS, framing restrictions and other appropriate browser security headers configured | Inspect response headers and exercise the application to confirm intended protections without broken flows |
| 10 | Alerts that reach a person | Alerts go to a designated Slack or email destination | Send test alerts and verify their destination, context, and response instructions |
| 11 | A backup restored once | Automated backups verified, retention set, and a real restore performed | Restore a backup into an isolated environment and check representative data integrity |
| 12 | A deploy gate | The main branch is protected with required checks and controlled merge permissions | Attempt a failing merge and confirm the protection prevents it |
Line 12 is often the missing one: in the same audits, at least 17 of the 21 third-party apps had no deploy gate, and so did all 5 of my own apps, so every push ships straight to production with nothing checking it first. It also depends on your GitHub plan. GitHub says protected branches “are available in public repositories with GitHub Free and GitHub Free for organizations” and “in public and private repositories with GitHub Pro, GitHub Team, GitHub Enterprise Cloud, and GitHub Enterprise Server”, and rulesets follow nearly the same list, so a private repository on GitHub Free cannot use either one and needs one of the other plans for this line.
As a startup security checklist, the table is worked top to bottom. A security checklist for CTOs at an AI startup is the same twelve lines, because the checks do not change with who or what wrote the code. Out of all the best security practices for startups, these twelve are the minimum I’d want evidence for before the first enterprise customer. The IT security checklist for startups on the company side (laptops, email, access to SaaS tools) is a separate list; this one covers the app. Two lines, 4 and 8, are scans; the other ten are read and tested, and together they make the security audit checklist a reviewer works through.
Lists like this have prior art. The SaaS CTO Security Checklist from goldfiglabs.com was discussed on Hacker News as “The SaaS CTO Security Checklist Redux” in June 2021, and Security4Startups publishes “a checklist of the security controls you should consider implementing in your startup.” I built this one around the evidence column, because the evidence is what a customer’s reviewer asks to see. For a startup on AWS, AWS’s Startup Security Baseline is the counterpart one layer down: controls for early-stage startups on a single AWS account, organized as account controls and workload controls. A startup security checklist for accelerators is, in my reading, this table under another name, and the free copy is the table itself, with nothing to download.
Keep the security baseline for your early-stage startup as a Markdown file in the app’s GitHub repository, so every change to it has a date and an author. If your early-stage startup’s security baseline started as a list from a Reddit thread, add the evidence column; the lines alone are not what a reviewer asks for. The twelve checks in tiers, each with a test and a builder prompt, are in the vibe coding security checklist; the version arranged by area for a SaaS product is the SaaS security checklist; and a self-check for an app made in a builder is the vibe code security check.
Assessment, audit, review: what each word means
A security assessment is the read against the list. An audit is the same read by someone independent, with a report. A review is what engineers call either one when it is internal. Assessment and testing adds tools and a person trying to break the lines that passed.
For the formal security assessment meaning, NIST’s definition of a security assessment reads: “The testing and/or evaluation of the management, operational, and technical security controls in an information system to determine the extent to which the controls are implemented correctly, operating as intended, and producing the desired outcome with respect to meeting the security requirements for the system.” Applied to one application, in my reading, what is included in a security assessment is the twelve lines above, each checked for whether it is implemented correctly and working as intended. The security assessment report is the filled table with dates and artifacts, and the security assessment plan is which lines it covers plus the date of the next one. Security assessment tools, meaning scanners and checklists, are inputs to the read, never the read itself.
| Word | Who does it | What you get | When a customer asks for it |
|---|---|---|---|
| Assessment | You or your team, against a written list | The filled list, with evidence and a date | In a security questionnaire, before a contract |
| Review | An engineer on the team, ideally one who did not build the feature | Findings and the fixes they led to | When they ask for the date and findings of your last one |
| Audit | An independent party | A report with the auditor’s findings | When their policy requires a third party |
| Assessment and testing | Your team, tools, and a tester trying to break what passed | Test results beside the evidence | For sensitive data or a larger contract |
The table is my reading of how the words get used, not a standard’s definitions. A security audit in cybersecurity is the read done by an independent party with a written report, which is also what an external security audit means. In my own grouping, not a standard’s, security audit types come down to internal, external and compliance. Once testing is added, the step up to a penetration test is the pen test vs vulnerability assessment question. Prices, and the split between a scan, a code audit, a pen test and a certification, are in what an AI app security audit costs.
Posture, risk and readiness: what the words mean on a team of two
Security posture is how many of the twelve lines are true today, with evidence. Risk is the chance and the cost of a line staying false. Readiness is the same read before a specific event: a customer’s review, an audit, an investor’s diligence.
Improving security posture means turning false lines true, and I’d work in this order: identity, then secrets, then the request surface. That order is my reading, not a measured ranking. Each line you turn true, with evidence, is an increase in security posture you can show someone. Cyber security posture at company scale also covers people, devices and vendors; for two people shipping one app, the twelve lines are most of it, in my reading.
A cyber security readiness assessment is the baseline read run before a named event, such as a first enterprise contract, an audit or a funding round’s diligence. Cyber readiness is the state that read measures on the day you run it.
NIST’s definition of cyber risk is “The risk of depending on cyber resources (i.e., the risk of depending on a system or system elements that exist in or intermittently have a presence in cyberspace).” For a small app, that is where a vulnerability sits in risk management: it is one concrete way a false line gets used, the opening that turns it into a loss. A security gap analysis template is the baseline table with two extra columns, today and required, in the same file.
SOC 2 uses these words in a narrower sense, and the three jobs it separates are laid out in gap, readiness and risk assessment. Security risk as a general question about apps built with AI tools is answered in are vibe-coded apps safe.
How to check your own app: the minimum internal review, in an afternoon
The minimum internal security review fits in an afternoon: print the twelve lines, produce the evidence or write down a no for each, run the dependency and secrets scans, read the failures with someone who did not build the app, fix them, sign-in and access first, and write the date on one page.
- 01 Print the twelve lines of the baseline table
- 02 For each line, produce the evidence from the last column, or write no
- 03 Run the two scans: dependencies for line eight, and secrets in the repository and its history for line four
- 04 Read the failures with someone who did not build the app
- 05 Fix them with sign-in and access first, keys and secrets next, and the request surface last
- 06 Write the date and the findings on one page, and set the date of the next review
The fifth step follows the same working order as the posture section. My working rule is to repeat the review at least before each enterprise customer’s review and after any change to sign-in, keys or the schema.
On a Production Hardening Sprint, the record a reviewer opens is deliverable 13.1, the production readiness report, and it is verified this way: account for all 123 IDs; keep failures visible until resolved and explain genuine non-applicable items.
For a founder who cannot read code, how do I know my app was built properly is the plain-language starting point, and the other sense of a security review, whether an agent tool is safe on a live codebase, is in is Google Antigravity safe on production code.
Where the sprint does this
Area 03 of the sprint scope, Input, output & application security, contains 12 of the 123 deliverables. Deliverable 3.7 reviews the application against the OWASP Top 10 and records findings, fixes, and evidence by category. Deliverable 3.10 tests the five highest-risk externally reachable attack surfaces, remediates findings, and delivers the methods and evidence, and I perform this targeted test myself. Formal third-party certifications and independent audit opinions are separate from the sprint deliverables. Every deliverable, with how each one is verified, is listed in the published scope, and buying the test from an outside firm instead is web application security testing services.
Common questions about reviewing an early-stage app
What is an example of a security baseline?
An example of a security baseline is the twelve-line table on this page, including server-side session checks, isolation on every tenant table and a restored backup. For a startup on AWS, AWS’s Startup Security Baseline is a cloud example, with its controls split into account controls and workload controls.
What is a security baseline document?
A security baseline document is the baseline table written down with a date: each line, the evidence for it, and who checked it. It fits on one page, and it is the record you hand a customer’s reviewer.
What does “security posture” mean?
Security posture means the state of your security controls today, backed by evidence. For an early-stage app, it is the count of baseline lines that are true and proven, out of twelve.
How to increase security posture?
Increase security posture by turning false baseline lines true: start with who can reach what, then the keys, then the inputs and public endpoints. That order is my own judgment. Then set a date for the next review, because a line that is true today can go false after a change.
If this checklist left you with more open items than you expected, the sprint below works through all of them in ten working days.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase