Coding standards used to be a document people read. When an assistant writes most of the lines, the standard has to be files the assistant reads and checks that CI runs. In a selected set of 21 third-party apps I audited in June and July 2026, at least 18 had no working test anywhere. So engineering standards for AI-assisted teams come down to 11 controls, each with a test.
The engineering standards for AI-assisted teams: 11 controls
Engineering standards for AI-assisted teams are 11 controls: dead code and duplicate logic removed, one source of truth for state, typed API boundaries, a documented repository, critical-path smoke tests, strict typing, auth and billing unit tests, guardrails the assistant reads and CI enforces, a recorded codebase tour, a clone-to-running setup guide, and a license audit.
Together they make up one area of production hardening, 11 of the 123 checks, and one of the areas that keep the rest of that work true after the next change lands. The table gives each control a short label in my words, the reason it matters, and the page that goes deep on it.
| Control | What it is | Why it matters |
|---|---|---|
| 1. Dead code and duplicate logic | Unused code deleted, repeated logic merged into one path (how to find unused code in a repo) | Repeated implementations make it easy to fix one path and leave another broken. |
| 2. One source of truth for state | Each piece of state has one owner, and the pattern is written down (two screens show different values) | Competing sources of truth can show different values in different screens. |
| 3. Typed API boundaries | Request and response shapes shared by frontend and backend (the frontend broke after a backend change) | Uncoordinated schema changes can break clients silently. |
| 4. Documented repository structure | A consistent layout, explained in the README (what belongs in a README) | Future developers and AI tools need to know where responsibilities belong. |
| 5. Critical-path smoke tests | Signup, login, the main action and payment, run automatically (how to write end-to-end smoke tests) | A small change can break a business-critical journey elsewhere. |
| 6. Strict typing | The type checker at its strictest setting, with the errors fixed (method named under control 6) | Stronger checks catch classes of mistakes before execution. |
| 7. Auth and billing unit tests | Tests on who may do what and on payment state, including the refusals (same smoke-test page as control 5) | Sensitive decisions need a reliable safety net when the implementation changes. |
| 8. AI repository guardrails | A rules file the assistant loads, plus CI checks for the rules a machine can test (guardrails for AI coding agents) | Future AI-assisted changes can undo hardening unless the workflow checks them. |
| 9. Recorded codebase tour | A video walk through the repo, the data model and the big decisions (method named under control 9) | A future engineer should not have to rediscover the system from scratch. |
| 10. Reproducible local setup | One command from a fresh clone to a running app (linked under What goes wrong without it) | The next engineer’s first hour decides whether they can work on the product at all. Generated apps rarely have this. |
| 11. Open-source license audit | Every dependency’s license known, the risky ones resolved (linked under What goes wrong without it) | License questions are on every technical due diligence list, and generated apps pull packages without checking. |
Each control turns coding best practices and coding standards into a file that exists or a check that fails, two things an assistant can’t quietly drift away from between sessions. As a practice, AI-assisted software engineering here means the things a two-person team can show on request: files, configs, and CI runs that pass or fail.
The formal layer sits above all of this: ISO standards for software development, and NIST software development standards such as the SSDF. NIST’s SSDF (SP 800-218) is the NIST Secure Software Development Framework, Version 1.1, published February 2022, and it describes itself as “a core set of high-level secure software development practices that can be integrated into each SDLC implementation”.
Coding standards best practices: the rules of coding an assistant reads before it writes
Coding standards for an AI-assisted team are three files the assistant reads before it writes: a rules file, a lint and format configuration, and a strict type configuration. NASA/JPL’s Power of Ten was written for safety-critical C, yet its author calls rules 4 to 7 “fairly broadly accepted as standards for good coding style.”
The rules file is where the conventions and the protected patterns live, in whatever file the team’s tool loads (control 8 below names them, the guardrails for AI coding agents page goes deeper, and CLAUDE.md best practices covers what to write in one). The lint and format configuration turns style rules for coding into a failing check instead of a request: ESLint and Prettier on a TypeScript app, for example, and the case for linters is a topic of its own. The type configuration is the strict setting in control 6, which makes the compiler refuse a whole class of guesses.
Programming rules only work when a tool can check them, and the NASA coding rules make that case well. Holzmann’s Power of Ten paper, “The Power of Ten: Rules for Developing Safety Critical Code”, comes from Gerard J. Holzmann at the NASA/JPL Laboratory for Reliable Software, and he wrote these NASA coding guidelines as an answer to long ones: most existing guidelines, he says, “contain well over a hundred rules”. The NASA Power of 10 limits itself to “no more than ten rules”, each specific enough that it “can be checked mechanically”, and the rules “primarily target C”.
Rules 4 and 7, two of the four the author files under good coding style, show why this NASA coding standard reads well outside C. Rule 4 caps a function at what fits on one printed sheet: “Typically, this means no more than about 60 lines of code per function.” Rule 7 says “The return value of non-void functions must be checked by each calling function, and the validity of parameters must be checked inside each function”.
My reading of the other rules: among the NASA code standards in the paper, rule 10 transfers most directly to a web app. It says “All code must compile with these setting without any warnings” and asks for a daily run of at least one static analyzer, and Holzmann adds that this “should be considered routine practice, even for non-critical code development.” On a web app, the nearest thing is strict typing plus a linter that fails CI. Rules 1 to 3, 8 and 9 (recursion, loop bounds, memory allocation, the preprocessor, pointers) are NASA coding requirements for C in safety-critical systems, and a TypeScript app has little use for them.
The paper makes no claim that code written this way cannot fail. Its claim is narrower: a well-chosen rule set “could” make critical components “more thoroughly analyzable”, and “These ten rules are being used experimentally at JPL in the writing of mission critical software, with encouraging results.”
Over-engineering and under-engineering: what an assistant does to a small codebase
Over-engineering is building something more complicated than the problem needs: an abstraction used once, a plugin system with one plugin. Under-engineering is the opposite failure, a copy where a call belonged. An assistant can do both in the same codebase, and removing dead code and duplicate logic clears the evidence of each.
Cambridge gives the meaning of over-engineer as “to create, design, or build something to be more complicated or perform more actions than is necessary or helpful” (Cambridge’s definition). The word “helpful” is the part that matters for a founder: a thing is over engineered when the extra structure helps nobody who uses or maintains it.
Over engineering examples in generated code follow two habits, in my reading. The first is an abstraction layer for a thing used once: an interface, a factory and a registry wrapped around the one email sender the app has. The second runs the other way: the assistant copies a handler into a new file instead of calling the one that exists, so the same logic now lives twice and drifts. Overengineering and underengineering land in one repo because the assistant solves each prompt fresh.
The fix is control 1 (the table’s first row says why), and whatever survives that cleanup is tech debt, and sizing it is a separate job. My working rule for what is over engineering and what is fine: it’s bad when the extra structure has one user and nobody asked for it.
Separation of concerns and DRY, and why generated code breaks both
Separation of concerns means each part of the code deals with one concern; don’t repeat yourself means each piece of knowledge lives in one place. Generated code can break both the same way: the same validation copied into three handlers, and a business rule written inside a UI component.
The separation of concerns principle goes back to Dijkstra’s note EWD447, “On the role of scientific thought”, dated 30 August 1974, where he writes that nothing is gained “by tackling these various aspects simultaneously” and names the habit of studying one aspect at a time “the separation of concerns”. DRY, the don’t repeat yourself principle, is tip 15 in The Pragmatic Programmer, and The Pragmatic Programmer’s DRY tip reads “Every piece of knowledge must have a single, unambiguous, authoritative representation within a system.”
In day-to-day programming, separation of concerns in software engineering means the component renders, a service holds the business rules, and one module validates input, whether the app is React or Java. In my reading, generated code tends to collapse those layers. The discount rule sits inside the checkout button, the email check is written again in every form with a slightly different pattern each time, and the cart lives in a client store and in the database at once, which is control 2’s problem (the table’s second row). Code that repeats itself like this is sometimes called WET, the opposite of DRY.
Bus factor: when one person or one model holds the codebase together
The bus factor is the number of people who would have to leave before nobody could maintain the codebase. On an AI-built app it can be one, and that one may be a chat history rather than a person. A README, a recorded codebase tour and a clone-to-running guide raise it.
A bus factor of 1 on a generated app can mean the reasons behind the schema, the order the migrations ran in and the steps of the last deploy sit in one conversation. Nobody can replay that conversation and get the same code back, including the person who had it.
To put a number on the bus factor of your own app, I use a working rule, not a published formula: list the critical areas (auth, billing, the database, the deploy) and count, for each, the people who could change it safely today; the lowest count is the number. Controls 4, 9 and 10 raise it, and area 13 covers what a developer handoff looks like. For the opposite seat, inheriting a codebase someone else built, legacy code takeover is a separate topic.
What is a bug in software, and what testers mean by defect and regression
A bug in software is behavior that differs from what the code was meant to do. A defect is the name a test report gives it, and a regression is a bug the previous version did not have, which is what critical-path smoke tests are there to catch.
That definition of a bug in software is mine, not a standard’s. A bug in code can sit unnoticed from the first release; a regression is the one a recent change introduced, and that makes it the kind a test can catch before a user does. Controls 5 and 7 exist for that. Software defects found in an audit are counted per app on how many bugs is normal.
What goes wrong without it
Each symptom below is what a founder or a new developer sees first. The cause is one missing control, and the page to open is named in each.
Two screens show different values
The dashboard shows one balance, the account page shows another, and a refresh changes which one looks right. Under the hood, two parts of the app each keep their own copy of the same data, and after an update only one copy changes. The table’s second row gives the reason this bites. Start at the two-screens page named in the table to find the second copy and pick one owner.
The frontend broke after a backend change
A field gets renamed in the API, the deploy goes green, and the settings page now renders blank for every user. Nothing tied the frontend’s idea of the response to the backend’s, so no check had anything to compare, and the break can reach users before anyone notices (the table’s third row). The backend-change page named in the table is the one to open.
A new developer cannot run it locally
A contractor clones the repo and spends the first day hunting for environment variables, a database with some data in it, and the one command that starts everything. The app runs on one laptop and in production, and nowhere in between. Nobody wrote down the one command, the environment example or the seed data, and generated apps rarely come with them (the table’s tenth row). Open cannot run the project locally.
Nobody can explain why this file exists
Someone asks why there are two date helpers and a folder called utils2, and the only honest answer is that the assistant made them. Dead files and duplicate helpers make any change a guess about which copy is live, and an unexplained layout leaves the next reader, human or assistant, putting new code wherever it lands (rows 1 and 4 of the table). Start with the unused-code page, then the README page, both named in the table; if the codebase is new to you, treat it as a legacy code takeover first.
A change to login shipped with nothing to catch it
A small edit to the sign-in form goes live on a Friday, and on Monday the support inbox says nobody can log in. No test ran the login journey, so the change could break it and still ship (the table’s fifth row). In my audits of June and July 2026, at least 18 of the 21 third-party apps had no working test anywhere, at least 23 of all 26 audited apps had zero working automated tests, and all 5 of my own apps had no tests. Those apps were picked for audit, not drawn at random, so the counts describe that set and not AI-built apps in general. In one of my own apps, the CI step named “npm install, build, and test” passed on every push without running a test, because there was no test script and no test file, which is the reason to verify AI-generated code before production. My reading: a CI step’s name is a claim, and only a run that fails when it should proves the suite exists. The smoke-test page linked in the table is where to start.
A copyleft license inside a commercial app
An investor’s due diligence list asks for a license report, and it turns up a copyleft package pulled in through another dependency. The assistant added the parent package because it solved the prompt, and nobody read the license before it shipped; the table’s last row says why buyers ask. What copyleft means for a closed-source product, and how to scan for it, is on license scanning.
The 11 controls, one by one
Each control below says what is in place when it’s done, the test that proves it, and the page with the method.
1. Dead code and duplicate logic are gone
The control: dead code is removed and duplicated logic is consolidated while preserving required behavior. A finder for unused files and exports gives the candidate list; each removal is a small, revertible commit. The test: record removed or consolidated code and run regression checks on affected flows. The second half matters, because a deleted “unused” function sometimes turns out to be called by a cron job or a webhook. The method is on the unused-code page named in the table.
2. State has one source of truth
The control: conflicting state ownership is consolidated into a documented, consistent pattern. In practice that means deciding, per piece of data, whether the server or the client owns it, and deleting the other copy. The test: exercise refreshes, mutations, and navigation and verify consistent displayed and stored state. Compare what the screen shows with what the database holds after each step, not only with what the previous screen showed. The two-screens page named in the table has the method.
3. Every API boundary is typed
The control: types are defined at every API boundary and contracts are shared between frontend and backend using the stack’s appropriate tooling. That might be a shared TypeScript package, a schema library that validates at runtime, or types generated from an API spec; the choice follows the stack. The test: introduce a deliberate contract mismatch and confirm type or contract checks catch it. If renaming one field on the server doesn’t fail anything on the client, the boundary isn’t typed yet. The backend-change page named in the table covers the options.
4. The repository structure is documented
The control: the codebase is organized consistently and its layout is explained in the README. A README that says where routes, business logic, database access and tests live is also the first thing an assistant reads when it looks for a place to put new code. The test: follow the README from a clean checkout and trace a core feature through its documented modules. Any step where the README and the folders disagree is a finding. The README page named in the table lists what belongs in it.
5. Critical-path smoke tests run in CI
The control: smoke tests are automated for signup, login, the core product action, and payment flows. CI here means whatever runs on every push; GitHub Actions is one example of a CI runner. For payment, the test drives the app up to the hand-off to the payment provider and proves the payment in test code, because Stripe’s automated-testing guide says its “Frontend interfaces, like Stripe Checkout or the Payment Element, have security measures in place that prevent automated testing” and that “you can simulate the output of our interfaces and API requests using mock data”. The test: run the suite in CI and demonstrate that an intentional regression fails it. Smoke tests cover behavior; code quality checks and what a code scan is cover what can be read from the code without running it. The smoke-test page linked in the table has the method.
6. Typing is strict
The control: TypeScript strict mode is enabled and its errors are resolved in TypeScript applications, with an equivalent strict-checking approach for other supported stacks. TypeScript’s strict option says turning it on “is equivalent to enabling all of the strict mode family options”, and warns that “upgrades of TypeScript might result in new type errors in your program”; the current reference lists its default as true, and writing "strict": true into tsconfig.json keeps the setting visible whatever version the project runs. The test: record the strict configuration and a clean checking run without suppressing the errors being fixed. Run the same type-check command the build uses, since a check pointed at the wrong config can pass without checking anything. The switch itself is a separate topic: how to enable TypeScript strict mode.
7. Auth and billing have unit tests
The control: unit tests sit around authorization and payment logic, including negative and edge cases. The negative cases carry the weight: a user from another account asking for a record, an expired plan asking for a paid feature, a canceled subscription moving back to active without a payment. The test: run the tests and show rejection of unauthorized access and incorrect billing transitions. A suite where every test expects success hasn’t tested this control. The smoke-test page from control 5 shows where these sit beside the end-to-end tests.
8. The assistant has guardrails and CI enforces them
The control: CLAUDE.md, AGENTS.md, Cursor rules, or equivalents describe conventions and protected patterns, and CI checks are added for enforceable rules. Claude Code reads project instructions from ./CLAUDE.md or ./.claude/CLAUDE.md, and by default reads AGENTS.md “only when you have no CLAUDE.md in your working directory or above it”; its docs also say “Claude treats them as context, not enforced configuration” (Claude Code’s memory docs). In Cursor, “Project rules live in .cursor/rules as .mdc files and are version-controlled”, and a plain .md file in that folder is ignored (Cursor’s rules docs). So a rule binds only where something enforces it, and a CI check on every change is the enforcement this control asks for. The test: review the guidance and demonstrate CI catching a representative forbidden regression. An AI review bot such as CodeRabbit adds another reviewer to each pull request; compare the options on a CodeRabbit alternative, and see an AI code check for spotting generated code. The guardrails for AI coding agents page has the method.
9. There is a recorded codebase tour
The control: a recorded walkthrough of the repository, data model, and key architectural decisions. The point of the video is the reasons, which the code can’t show: why the queue exists, why billing state lives where it does, which folder nobody should touch without a test. The test: deliver an accessible recording with chapter markers or a short contents list. A tour nobody can navigate gets watched once. What to cover follows from what a codebase walkthrough is for.
10. A new machine goes from clone to running by the guide alone
The control: one documented command brings the application up locally, with an environment example and seed data. The environment example lists every variable with a safe placeholder, and the seed data gives a fresh database enough rows to click through the main journey. The test: go from clone to running application on a new machine by following the guide alone. “By the guide alone” means no message to the person who wrote it. The page on a project you cannot run locally has the setup steps.
11. Every dependency’s license has been audited
The control: the licenses of every dependency are inventoried, copyleft or unlicensed packages are flagged, and each one is replaced or approved. Transitive packages count too, since nobody chose them on purpose. The test: include the license report in the due diligence pack with zero unresolved flags. An approved exception is a resolved flag if the reason is written down. The license scanning page is where to find the tools.
Reviewing code you cannot fully read
A review of code an assistant wrote asks four questions in order: does it run, does it break a critical path, does it touch auth or billing, and does it add structure nobody asked for. The pages below give the checklist, the security version and the method for each question.
The order is my own choice: a change that doesn’t run needs no further review, and a change that touches auth or billing needs more than a skim. The fourth question is the over-engineering test from earlier on this page.
| What you are asked for | Page |
|---|---|
| A checklist and template for any pull request | code review checklist |
| The security version, for auth, data access and secrets | secure code review checklist |
| How a review runs, from request to merge | code review best practices |
| What “production ready” means for generated code | how to write production-ready code |
| Checking assistant output before it ships | verify AI-generated code before production (named in the login section above) |
How to verify the whole area in an afternoon
Verifying this area takes 11 tests and, as my working estimate for a small app, about an afternoon: clone onto a fresh machine, walk the README, run the strict type check, break an API contract on purpose, break login and watch CI fail, and read the license report. Record each result.
The order below is the one a one-person team can run top to bottom, because most later tests need the running copy the first one produces. Each item names the control, the condition that makes the test count, and what to write down; the full test sits under the control’s own heading above.
- 01 Control 10, local setup: a new machine and the guide alone. Write down each step the guide left out.
- 02 Control 4, structure: a clean checkout, one core feature traced through the folders. Write down where the README and the code disagree.
- 03 Control 6, strict typing: a clean run with no suppressed errors. Write down the config and the command output.
- 04 Control 3, typed boundaries: a planted mismatch between the two sides, caught by a check. Write down which check failed.
- 05 Control 5, smoke tests: a planted break in a critical journey, failing the suite in CI. Write down the link to the red run.
- 06 Control 7, auth and billing: an unauthorized request and a wrong billing transition, both refused by tests. Write down the test names.
- 07 Control 2, state: refresh, change and move between screens, comparing the screen with the database. Write down any mismatch.
- 08 Control 1, dead code: the removal list, with the affected flows checked again. Write down what was removed.
- 09 Control 8, guardrails: one forbidden change, caught by CI. Write down the rule and the failed run.
- 10 Control 9, the tour: a recording with chapter markers or a contents list. Write down where it lives.
- 11 Control 11, licenses: the license report with nothing left unresolved. Write down each approved exception and its reason.
This area is 11 lines of a longer list. Every area together makes up the full production readiness checklist.
Where the sprint stops
In the Production Hardening Sprint, the work happens inside the existing codebase, and the app’s current framework and hosting setup are the starting point. New features that change the product’s core capabilities are separate work. Formal third-party certifications and independent audit opinions are separate from the sprint deliverables.
Where the sprint does this
Area 10 of the sprint is these 11 controls. Each one is delivered and then verified by the test written under its heading above. The result for every scope item, the work completed and its verification evidence go into the production readiness report, where failures stay visible until resolved. The deliverables and their tests are listed under area 10 of the published scope.
Common questions about engineering standards
Is over engineering good or bad?
Bad, once the added layers serve a single caller that nobody requested; that’s my working rule, explained in the over-engineering section above. Structure earns its place when a second real use already exists or is on this month’s plan, and the cost of adding it later would be high.
What are some examples of overengineering?
Three that show up in generated code, as my examples: a repository pattern wrapped around a single table, a plugin system with one plugin, and a configuration service that serves two values. Each adds files a reader has to understand before changing anything, and none of them is used a second time.
What are NASA’s 10 coding rules?
They are Gerard J. Holzmann’s ten rules for safety-critical code from the NASA/JPL Laboratory for Reliable Software: simple control flow with no goto or recursion, a fixed upper bound on every loop, no dynamic memory allocation after initialization, functions of about 60 lines at most, at least two assertions per function on average, data declared at the smallest scope, return values and parameters always checked, a limited preprocessor, restricted pointers, and zero warnings from the compiler and a daily static analyzer. Holzmann counts rules 4 to 7 as broadly accepted standards for good coding style, and the coding standards section above covers which rules travel to a web app.
What does bus factor mean?
Bus factor means how many people a project can lose before no one left can keep it running. A low number is the risk; a bus factor of one means a single person, or a single chat history, holds what the next change needs.
What is an example of separation of concerns?
A signup form where one validation module checks the email and password, a service creates the account, and the component only renders the form and the errors it’s given. If the rules change, one module changes, and the component doesn’t.
What are basic coding standards?
For a team that works with an assistant, the basics are three files in the repo: instructions the assistant loads at the start of each session, a linter and formatter configuration that fails the build on a violation, and a type-checker configuration at its strictest setting.
Owning an app means being able to run it, change it and recover it without guessing. The sprint below leaves you with the runbooks and documentation to do that.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase