Run the tests the code claims to have, before anything else. Of 21 audited apps that other teams built, no more than 3 had a test that worked, and in one retail app the checkout tests never touched the real checkout code. To verify AI generated code before production without reading it, you collect evidence that it behaves.

What it means to verify AI generated code before production when you cannot read it

Verifying AI-generated code you cannot read means collecting evidence that it behaves, not reading it. Seven pieces cover a small SaaS: a test run that has failed on purpose once, a dependency scan, a secrets scan of the history, a permission test, a tenant rule check, a strict type check, and one human read of the riskiest paths.

The opener’s figure comes from 21 apps other people built, public and held-out, that I audited in June and July 2026. I picked them; they were not drawn at random, so the figure is not a rate for AI-built apps in general. This page sits in area 10, code health and AI development guardrails, of the engineering standards for AI assisted teams.

Whether you review AI generated code, evaluate it or audit it, you are reading the same evidence with a different question in mind; the vocabulary table further down maps each word to its pieces. An LLM’s code you can’t read gets its audit from what it does when it runs: black-box behavior, not line-by-line reading.

This page leaves four neighboring questions to other pages. Whether a tool wrote the code at all is a different question: how to check if code is AI generated. Ten tests you can do from a browser tab are in checks you can run without reading code. Asking the assistant to grade its own work is covered under whether AI can do code reviews. Writing behavior contracts and tests for code you did not write is how to test a vibe-coded app.

Is this code safe? The five-minute version

The five-minute version of “is this code safe” is five checks: search the repository and its history for live secret keys, run the package audit and count what it finds, run the test command, on a Supabase-style app list the tables with row-level security off, and ask a second account for the first account’s record. Any fail is today’s answer.

  1. 01 Scan the repository and its whole history with gitleaks, a secrets scanner. Its default rules cover Stripe keys, OpenAI and Anthropic keys, and JWTs, the form Supabase's legacy keys take. They name no rule for Supabase's newer secret keys, so also search the history for the text sb_secret_. Fail: a key starting sk_live_ or rk_live_, a Supabase secret or service_role key, or an AI provider key, in any commit, even if the file is gone now.
  2. 02 Run the package manager's audit and count what it reports. npm and pnpm grade each advisory by severity, so note how many sit at the highest level shown. For pip-audit, count the vulnerable packages it lists. Fail: no result at all (by default npm needs a lockfile), or a high or critical advisory.
  3. 03 Run the project's test command and read the last lines it prints. Fail: no tests, a missing test script, or a pass that ran zero tests.
  4. 04 If the browser talks to the database directly, as a Supabase app does, run the query below in your database's SQL editor to list the tables with row level security off. Fail: any table holding user data in the list. If your server makes every database call, check 5 does this job.
  5. 05 Create a second test account in your own app, sign in with it, and ask for a record that belongs to the first account. Fail: any of the first account's data comes back. An empty answer or a refusal is a pass.
npm audit                    # or: pnpm audit, or: pip-audit -r ./requirements.txt
gitleaks git -v .
git log --all -S sb_secret_
select relname from pg_class
where relkind = 'r' and not relrowsecurity
  and relnamespace = (select oid from pg_namespace where nspname = 'public');

The secrets check leans on gitleaks because a plain text search misses keys. The gitleaks git command scans a repository’s history through git log -p, and its default rules include Stripe keys with sk_ or rk_ and test, live or prod, OpenAI keys starting sk-, Anthropic keys and JWTs. Stripe’s live mode secret and restricted keys start with sk_live_ or rk_live_, so a search for those alone misses an OpenAI or Anthropic key, which starts sk-. Supabase’s legacy anon and service_role keys are long JWTs that begin eyJ, so searching for the word service_role finds variable names, not the key. A JWT hit can be the anon key, which Supabase calls the legacy version of the publishable key, the one safe to expose; compare the hit with the keys on your project’s Settings > API Keys page, and only the service_role key is a fail. Gitleaks’ default config has no Supabase rule, which is why the git log line searches for sb_secret_ across every branch: -S finds commits that add or remove that text, and --all covers every ref. A search of the built bundle the browser receives is a separate check, covered in the article on why AI coding tools ship security holes.

On check 2, npm audit grades advisories as info, low, moderate, high or critical, and by default it “requires a package-lock in order to run the audit”, so a project with no lockfile gets no npm result until it has one. pnpm audit uses low, moderate, high and critical. The pip-audit README shows a report of name, version, advisory ID and fix versions and names no severity field. The high-or-critical line in check 2 is where I draw it for a first pass, not a standard.

On check 4, relrowsecurity is PostgreSQL’s catalog flag, “True if table has row-level security enabled”, and relkind = 'r' keeps the list to ordinary tables. The query checks the public schema; change the name if your app exposes another one. Supabase’s publishable key is safe to expose in a web page and “only reaches what Row Level Security allows”, while a secret key “bypasses every Row Level Security policy you have”. Supabase’s row level security guide adds that a table in an exposed schema without RLS “is readable and writable by any role with a grant on it”, which is why check 4 lists those tables. Supabase also says it is deprecating the anon and service_role keys by the end of 2026; Supabase’s API keys guide shows both generations.

Check 5 is the same test as check 6 in the browser checks, which walks through it step by step. Run it only on an app you own, with two accounts you made. The seven pieces below are what to collect once these five are done. Separately, “is this safe” sometimes means letting a coding agent run without asking permission, which is Claude Code YOLO mode and not this page’s question.

Why it matters: what my audits found when the evidence was missing

Each row pairs a missing piece of evidence with what I found in the apps I audited.

Missing evidenceWhat I foundThe piece that catches it
A test run that fails when something breaksNo working test anywhere in at least 18 of the 21 third-party apps (the opener’s figure)1
A permission testUnauthenticated endpoints doing privileged work in 11 of the 21 third-party apps4
A secrets scan of the historyA real secret shipped in 6 of the 21 third-party apps3
A dependency scanA framework version with a publicly known, reachable RCE or auth bypass in 9 of the 26 audited apps, the only row counted across all 26 apps including my own; the fix was often a one-line version bump2

These counts come from the same June and July 2026 audits of 21 third-party apps, with my own apps added only in the dependency row. I chose the apps, so they show what turns up when the evidence is absent, not how common each failure is across all AI-built code.

The published research on AI-generated code security, and why its headline rates disagree, is covered in why AI coding tools ship security holes. Fixing what these checks find belongs to web app security.

The seven pieces of evidence, and how to get each without reading a line

Each piece of evidence is one command or one request and one output to keep: the failing test run, the audit summary, the secrets scan report, the refused request, the per-table test record, the clean type check, and the review note. An output you cannot produce is a gap, not a pass.

EvidenceHow to get it without readingThe output to keepWhere it goes deeper
1. A suite that has been made to fail onceRun the suite in CI and demonstrate that an intentional regression fails itThe red run and the green runhow to write end to end smoke tests
2. A dependency scanRecord the dependency scan, remediation, and regression checks after changesThe audit summary with its datethe npm audit command
3. A secrets scan of the repository and its historyRetain scan results and confirm exposed credentials have been invalidated and replacements workThe scan report and a note of each key rotatedsecrets management best practices
4. The permission testCall protected actions directly as unauthorized and underprivileged users; confirm rejectionThe refused responsesbroken access control
5. The tenant rule checkRun read and write tests as anonymous users, different roles, and separate tenantsThe test record per tabledata consistency checklist for SaaS
6. The typed boundary and strict type runIntroduce a deliberate contract mismatch and confirm type or contract checks catch it; record the strict configuration and a clean checking run without suppressing the errors being fixedThe caught mismatch and the clean runcode quality checks
7. A human read of the riskiest pathsA person who reads code goes through sign-in, billing and the AI prompt path (my choice of paths)The review notecode review best practices, with the code review checklist

Row 1 is where a green tick misleads most. In one of my own apps, the CI step named “npm install, build, and test” passed on every push without running a test: there was no test script and no test file in the tree. A green CI step proves only that the step ran; the evidence is a run that went red when something broke.

Row 4 runs with the second account’s own session, never the first’s. Where row level security guards a table, as on a Supabase app, a refused read comes back as an empty list rather than an error. PostgreSQL’s manual says that with no policy on a table, “no rows are visible or can be modified”, and rows a policy does not allow “will not be processed”. The empty response is the pass, so keep it as the evidence.

Row 7 is one I insist on: pay for a person to read sign-in, billing and the AI prompt path even when every other row passes, because the other six prove behavior, not intent. What that review hands back to a founder who cannot read code is explained in code review for non-developers.

On a Production Hardening Sprint, deliverable 13.1, the production readiness report, delivers the result for every scope item, the work completed, and its verification evidence, and it is verified this way: account for all 123 IDs, keep failures visible until resolved and explain genuine non-applicable items.

Which word means which evidence: test, review, evaluate, audit

The wordThe evidence it asks forThe pieces above
TestDoes the app do what it should when run, and refuse what it should refuseThe failing test run, the permission test, the tenant rule check
ReviewHas a person who reads code checked the paths where a mistake costs mostThe human read
EvaluateIs what the code stands on sound: its packages, its secrets, its typesThe dependency scan, the secrets scan, the strict type run
AuditEvery piece, each with a dated record someone else can inspectAll seven

The mapping reflects how I hear founders and developers use the four words, not a standard definition. A fifth word, detect, is the odd one out: it asks who wrote the code, not whether the code works, and the detector article answers it. For the human read, the best AI code review tools, and who says so compares the tools that can give a second read alongside the person.

How to check your own app

A routine for your own app is the five-minute version today and the seven pieces over the next week, each output kept with its date in one folder. A piece you cannot produce is the next thing to build or buy. Then add guardrails so the evidence stays true after the next AI change.

Make one folder per date and drop the seven outputs into it: the red and green runs, the audit summary, the scan report, the refused responses, the per-table test record, the clean type run and the review note. Next month’s folder should hold the same seven, so a gap shows up as a missing file rather than a feeling.

Treat missing evidence as unknown, not as a pass. The evidence also goes stale the moment an assistant edits the code again, which is what guardrails for AI coding agents are for: CI checks and repository rules that rerun the proof on every change. What the code itself should look like once it is ready is a separate question: how to write production ready code.

Where the sprint does this

In the sprint, deliverable 10.8, AI repository guardrails, provides CLAUDE.md, AGENTS.md, Cursor rules or equivalents describing conventions and protected patterns, and adds CI checks for enforceable rules; it is verified by reviewing the guidance and demonstrating CI catching a representative forbidden regression. Deliverable 3.7 reviews the application against the OWASP Top 10 and records findings, fixes and evidence by category. Both are listed, with how each is verified, in the published scope.

Common questions about checking code an AI tool wrote

How do you validate AI generated code?

By running it and keeping what it produces, not by reading it. Here validate and verify mean the same thing: collect a red-then-green test run, a dated audit summary, a history scan report, refused requests from a second account, a per-table tenant test, a clean strict type run and a person’s review note, and treat any output you cannot produce as a gap.

Can AI generate source code?

Yes. AI code generation produces source code from natural language instructions, in tools such as Claude Code, Cursor and Lovable. AI generated source code arrives without any evidence that it behaves; that evidence comes from running it and keeping the outputs, which is what the seven pieces are for.

How do you review an AI code?

Run it first, then have a person read the sign-in, billing and AI prompt paths. I’d pay for that read even when the automated checks pass. An AI reviewer is useful as a second read, never the only one; the review tools article compares the options.

How to check if a software is safe?

Run the five checks from the five-minute version: secrets in the history, the package audit, the test command, tables with row level security off, and a second account asking for the first account’s record. For software you are about to ship, run them before launch; for software someone hands you, run the same five before it touches real data.