Of the 26 AI-built apps I audited in June and July 2026, 22 had at least one confirmed critical finding, none scored green, and at least 23 had no working automated test. What investors look for in code is whether a company is one of those and how fast its founder can prove otherwise. Eight red flags decide it.

Those 26 are 21 apps other people built and five of my own, each report made of human-verified findings. They are a selected set of audited apps, not a random sample, so none of the numbers on this page is a rate for all AI-built apps. This page sits in the reviewer’s seat of production hardening; the founder across the table can start with how to prepare for technical due diligence.

What investors look for in code: what the reviewer is trying to get

A technical reviewer wants three answers from a startup’s code: can someone other than the founder maintain it, is there a landmine, and how much of the next round goes to rework. Eight red flags answer those three, and each one is cleared by an artifact, not a promise.

Maintainability comes first. The reviewer wants to know whether an engineer hired after the round could open the repository, run it and change it without the founder in the room. The landmine is anything that turns into a cost the new money inherits: a critical security finding, a dependency under a license the company can’t comply with, or a provider account that sits in someone else’s name. Rework is the arithmetic that follows. If half the app has to be rebuilt before it can grow, part of the round pays for the past.

A codebase is investable when those three answers come with evidence behind them. Tidy style, a fashionable framework and a neat commit log change none of them, and a reviewer who grades style is answering a different question.

A code audit is a review of an app’s source code and setup against named checks that ends in a written report of findings. A code audit before investment points that review at the three questions: a web application security audit for the landmines, plus the ownership and maintainability checks that decide whether the company can run the thing it paid for. If you are choosing who does that work, code audit services sorts the kinds of seller.

The founder’s side of the same process is startup technical due diligence, and buying a whole company built with AI is a wider review, covered in due diligence when buying an AI-built app. One limit holds in every seat: the review settles the code, not the team, the market or the round’s terms.

The eight red flags in a startup codebase, and what clears each one

Each flag is a check a reviewer can run with read access to the repository and the provider dashboards, and each has one piece of evidence that clears it. An accelerator can run the same table across a whole batch: it is a fair list for technical reviewers to ask cohort companies to answer before demo day.

#Red flagHow the reviewer checksRate in my auditsEvidence that clears itThe page on it
1No working testsRun the suite, break one line of business logic, run it againAt least 23 of 26; at least 18 of the 21 that other people builtA suite that fails on a deliberate regressionhow to write end-to-end smoke tests
2Every push goes to productionOn GitHub: Settings, then Branches. Look for a branch protection rule on the production branch with “Require status checks before merging” turned onAt least 17 of the 21 that other people built, and 5 of 5 of my ownA protected branch and a promotion gateGitHub branch protection
3A critical security findingAsk for the latest security audit report and the retest that closed its criticals22 of 26 had at least one confirmed criticalAn audit report with a retestThe web application security audit
4Secrets in the repository or the bundleGitHub secret scanning “scans your entire Git history on all branches”; its alerts appear on the repository’s Security and quality tab. Then search the built browser bundle6 of the 21 that other people built shipped a real secret; three had one permanently in git historyA full-history scan and a rotation recordsecret scanning
5Provider accounts in a contractor’s nameSupabase: who holds Owner, in the organization’s team settings. Vercel: who holds Owner (a role on Pro and Enterprise plans), under Settings, then Members. Stripe: who is the Account Owner, from the Team tab in the DashboardNot counted in my auditsThe ownership inventorythe provider inventory a founder should hold
6No backup ever restoredAsk for the date of the last test restore and where the data was restored toNot counted in my auditsA restore recorddatabase backup checklist for startups
7A copyleft or unknown license in the treeRun a license scanner over the full dependency tree and read every flagNot counted in my auditsThe license report, with every flag resolvedlicense scanning
8A bus factor of one, where the one is a chat historyHand the README to someone who has never seen the app and watch them try to run itNot counted in my auditsA README a stranger can follow and a recorded tourwhat belongs in a README; what a codebase walkthrough is

The rates come from my June and July 2026 audits of 26 apps, 21 built by other people plus 5 of my own: a chosen set rather than a sample of the whole field, so read them as what I found, not as odds for any given app.

Tests are the flag I’d check first, because every other fix leans on them. A test folder that exists proves little. The reviewer’s move is to change one line of business logic, a price calculation or a permission check, and run the suite again. If it stays green, the tests aren’t testing anything. What clears the flag is a suite that goes red on that deliberate break, green again when the line is put back, and runs on every change. A handful of end-to-end smoke tests through sign-up, the main action and payment is the smallest version that counts.

When a push to the main branch goes straight to customers, a typo travels as fast as a fix, and that empty deploy path is the second flag. On GitHub the reviewer looks for a branch protection rule with required status checks, which GitHub says must all pass before collaborators can merge into the protected branch. Availability depends on the plan: GitHub lists protected branches for public repositories on GitHub Free, and for public and private repositories on GitHub Pro, Team, Enterprise Cloud and Enterprise Server. The same page also points to its guide on converting branch protections to rulesets. The hosting side then needs a promotion gate, so a deploy waits for those checks instead of firing on every merge.

A confirmed critical finding settles a report on its own. In my audits’ scoring, any confirmed critical made the band red whatever the score, and the four apps without one were exactly the four that landed amber. The reviewer doesn’t have to find the critical personally. Ask for the most recent security audit, then ask for its retest. A report that lists criticals with no retest shows the problems were found. It does not show they were fixed.

The obvious fix for a leaked secret fails. Deleting a key from today’s code leaves it in every earlier commit, and anything bundled for the browser can be read by every visitor. GitHub’s secret scanning reads the whole history; it runs automatically, for free, on public repositories, while an organization’s private repositories need GitHub Secret Protection enabled on GitHub Team or GitHub Enterprise Cloud. Where it isn’t on, any scanner that reads full history does the job. The evidence is the scan output plus a rotation record: which key, when it was replaced, and proof the old one stopped working.

A login hides the fifth flag. A founder who can sign in to the database, the host and the payment account may still be a guest in each, because the owner role can sit with whoever set the account up. That role has the last word on each account; at Stripe there can only be one Account Owner, and it can close the account. So the check is the role, not the login, and the table says where each provider shows it. The ownership inventory clears it: each provider, whose account it is, and who holds that role today.

Backups look cleared from the dashboard, which is why they need a second question. A schedule that says backups run proves a schedule exists. The reviewer asks when someone last restored one, into what, and whether the app’s data was all there afterwards. A restore record answers all three: the date, the target (a fresh project or a scratch database, never production), and a check that the rows and files the app depends on came back.

License problems never show in the product. A generated app can pull in a package to make one feature work, and a copyleft license deep in the tree can carry obligations the business never meant to take on, such as providing source code with software it distributes. The reviewer runs a license scanner over the full dependency tree, and a report with each copyleft or unknown license replaced or approved in writing clears the flag. Ownership of the generated code is a separate question with a US answer: the US Copyright Office’s report of 29 January 2025 concludes that generative AI outputs can be protected by copyright only where a human author has determined sufficient expressive elements, which the mere provision of prompts does not meet (the US Copyright Office’s guidance on AI and copyright). That is the office’s view of US law, and this is not legal advice.

The license risk has a court record behind it. In Software Freedom Conservancy’s suit against Vizio, an order of the Orange County Superior Court dated 29 December 2023 records the allegation that Vizio distributed smart TVs with GPL-licensed software without providing the source code, and denies Vizio’s motion for summary judgment. That is an allegation the order let go forward, not a finding of breach, and the order does not decide the case. In my reading, the case shows why the license report is on a reviewer’s list: a copyleft license in shipped code can turn into a legal claim, and the order records the claim, not its outcome.

The last flag is a bus factor of one, and in an AI-built app the one can be a chat history. The reasons behind the code sit in conversations with a coding assistant that don’t travel with the repository, so the founder can explain the app and nobody else can. The reviewer’s test is plain: hand the README to someone who has never seen the app and see whether they can install it, run it and find where the payment logic lives. A README that passes, plus a recorded tour of the codebase, clears it.

What good looks like: the evidence pack

The evidence pack a reviewer accepts is eight artifacts: a readiness report with failures visible, an architecture diagram, a data model, the security checklist with proof, a tested capacity statement, the license report, the provider inventory, and runbooks. The risk band is the posture; the score is not a grade.

Each artifact answers one or more of the flags with something the reviewer can open instead of taking on trust.

ArtifactWhat it provesWho produces it
Readiness report, a row per check, failures kept visibleWhere the app stands on every check, including the ones it failsThe engineer who did the work
Architecture diagramHow the pieces connect, so a stranger can find their way aroundThe founder’s team
Data modelWhat data the app holds and how the tables relateThe founder’s team
Security checklist with evidenceEach control, with the test, log line or screenshot that shows it worksThe engineer who did the work
Tested capacity statementThe load the app handled in a real test, as opposed to a guessThe engineer who ran the test
License reportEvery dependency’s license, with copyleft and unknown ones resolvedThe engineer who did the work
Provider inventoryWhich accounts the company owns, and who holds the owner role in eachThe founder
Runbook setHow to deploy, roll back, rotate a key and restore a backupThe engineer who did the work

Technical readiness for demo day is the same pack with a single page on top that answers the eight flags, one line each, with a link to the evidence behind it. Check an accelerator’s technical requirements against this list before the batch starts; anything it asks for that is missing here belongs in the pack as well. For a startup heading into an accelerator, the security checklist on GitHub titled Startup Security Checklist covers part of this: its Technology section has two headings, Asset Ownership (no items under it yet) and Backups (one of its two items is restoring backups to a fresh environment), and the rest covers Customer Trust and Legal and Compliance.

Scores need the same care as flags. In the scoring I used for my June and July 2026 audits, a selected set of apps, the band measures risk and the score measures how much is left to do, so the two can disagree: in the illustrative case from my method notes, a pristine app with one critical can score 76 and still be red, while a 59-scoring amber app has no landmines but 21 things to finish.

In the Production Hardening Sprint, this pack is deliverable 13.6, the technical due diligence pack: the readiness report, architecture diagram, data model, security checklist and capacity statement bundled into one PDF, checked for completeness, consistent version references and readable linked evidence. Its readiness report, deliverable 13.1, is verified this way: account for all 123 IDs; keep failures visible until resolved and explain genuine non-applicable items. A control that does not apply to the product is marked with a written reason, and the final report records every item as verified or not applicable with a reason. The open-source license audit, deliverable 10.11, puts the license report into the same pack with zero unresolved flags. The provider inventory is deliverable 2.8’s evidence: each provider, the owning account, the owner role holder, and the date the previous builder’s access was removed. The operating runbooks, deliverable 13.3, cover deployment, rollback, key rotation, backup restoration and the response to each operational alert, and they are part of the same handover.

What it costs and how long it takes

A fund has three ways to check a portfolio company’s code: its own reviewer works through the eight flags, for the cost of the reviewer’s time; a security firm audits one codebase, priced per scope; or a fixed-scope hardening pass fixes what the flags find and leaves the evidence pack behind.

OptionWho does itTimeCost basis
The fund’s own reviewer works through the eight flagsA technical partner or advisor to the fundAbout an afternoon, my working ruleThe reviewer’s time
A security audit or pen test of one codebaseAn audit or pen-test firm10 to 15 working days for the manual pentest, on Astra’s pricing pagePer target per year: $5,999 a year for Astra’s Pentest Expert plan, where one web or SaaS app counts as one target
A fixed-scope hardening pass that fixes what the flags find and leaves the pack behindAn engineering teamFixed per codebase, stated in the seller’s published scopeFixed per codebase, stated in the seller’s published scope

Astra’s figures are from its pricing page, checked on 27 September 2026. Who else sells the second option, and how their offers differ, is in web application security testing services.

The Production Hardening Sprint is one such pass: it runs for 10 working days on one codebase, and it ends in the due diligence pack described above. It leaves verification evidence for every scope item in its readiness report, under the rule quoted in the evidence pack section, and formal third-party certifications and independent audit opinions are separate from its deliverables.

A post-investment technical review is the same eight flags run again, and my working rule is to run it at the first board meeting after the round closes. The reason is timing. The ICO’s final penalty notice to Marriott, dated 30 October 2020, records that Starwood’s systems were compromised in 2014, with a web shell installed on 29 July 2014, that Marriott acquired Starwood in 2016, and that Marriott did not detect the attack until September 2018. It also records Marriott’s own statement that it was only able to carry out limited due diligence on Starwood’s systems during the acquisition; that is Marriott’s account, not an ICO finding, because the notice says the Commissioner has not determined whether such diligence was possible, and the decision relates solely to Marriott’s failures after 25 May 2018. Its general line is that a controller’s due diligence on its data operations “is not time-limited or a ‘one-off’ requirement.” This was a large hotel acquisition, not a seed round. My reading, not the ICO’s finding: a problem already in the systems before the money moves is inherited with them, which is why the eight flags belong before the investment closes and again at that first board meeting.

A portfolio company security baseline is the same pack asked of every company in the fund, on one template, so a partner compares like with like and sees at a glance which flags are open where. For the companies selling to enterprises, the next step up is SOC 2, and SOC 2 compliance consultants covers who helps with that.

Where the sprint fits

Y Combinator (YC) keeps a due diligence list for the Series A in its Startup Library. YC’s Series A diligence checklist lists the information a company needs ready once it signs a term sheet, all of it company paperwork: corporate records, the business plan and financials, intellectual property, securities, material agreements, disputes, and employees. None of its items asks for source code or technical documentation; the nearest are the list of trademarks, patents, copyrights and domain names, and the agreements covering proprietary information and technology. The code half of diligence is what the eight flags and the pack above cover.

The sprint’s handover is that code half: the updated codebase, tests and deployment configuration, a production readiness report accounting for all 123 deliverables under the verify rule quoted earlier, the technical due diligence pack in one PDF, operating runbooks, guidance and automated checks for future AI-assisted changes, and a recorded 60-minute handover walkthrough with the client team plus a 20 to 30-minute codebase tour. A fund can ask a portfolio founder for that pack by name, and every deliverable behind it is listed in the published scope.

Common questions about technical reviews of startup code

What are red flags in due diligence?

In a technical review of a startup’s code, the eight red flags to check are no working tests, every push going straight to production, a critical security finding, secrets in the repository or the browser bundle, provider accounts in a contractor’s name, backups never restored, a copyleft or unknown license, and a bus factor of one. Three deserve the first look. Missing tests come first, since nothing else can be changed safely without them. A deploy path with no gate comes next, because it lets any mistake reach customers unchecked. A confirmed critical finding is the third, because it is the one flag a reviewer can’t weigh against the others.

What does “technical readiness” mean?

For a startup facing demo day or a first round, technical readiness means the code clears the eight red flags with evidence a reviewer can open: the readiness report, architecture diagram, data model, security checklist, capacity statement, license report, provider inventory and runbooks described under “What good looks like”. It has nothing to do with NASA’s technology readiness levels, which use the same words for a different thing.

Is 100,000 lines of code a lot?

Size on its own isn’t one of the flags, and a reviewer shouldn’t treat it as one. A large codebase with tests that catch a deliberate break, a gated deploy path and a README a stranger can follow is in better shape than a small one with none of those. An AI-built app of that size with no tests has the first red flag, not a size problem, and the fix is the same as for a small app: tests that go red when something breaks.

What are common mistakes investors make?

In a technical review, the four I’d watch for are reading the paperwork instead of the code, taking a readiness score as a grade when one critical finding outweighs any score, taking a founder’s login as proof the company owns the account, and checking once at the deal and never again. Each has a cheap fix. Open the repository, read the band and the criticals before the number, ask who holds the owner role, and run the same flags at the first board meeting.