Code quality checks are the tests, run by tools or by a reviewer, that tell you whether software can be changed without breaking: types, lint, tests, duplication, complexity, dead code and dependency health. None needs you to read code. Across the 21 third-party apps I audited in June and July 2026, the Maintainability & Evolvability pillar averages 61.1 out of 100.

What code quality checks are, and the seven worth asking for

Code quality checks are automated or reviewer-run tests of how safely code can be changed. A small team needs seven: a type check, a linter, tests on the critical paths, a duplication scan, a complexity scan, a dead code scan and a dependency audit. Each returns a pass, a count or a score you can read without reading code.

That 61.1 comes from a chosen group: 11 public vibe-coded apps I audited across all 12 pillars, plus a held-out set of 10 other third-party apps I audited blind, all in June and July 2026. It is not a random sample, and not a rate for every app an AI helped build.

My definition of code quality is how cheaply and safely someone who did not write the code can change it. Whether the app works today is a different question, answered by testing the running product. The standards these checks belong to are set out in engineering standards for AI-assisted teams.

A few words come up whenever a developer talks about this. Maintainability, in my working definition, is how easily code can be understood, changed and tested; the longer answer, with how tools score it, is in why an AI-built app gets harder to change. Structural quality is the term CISQ’s page on ISO 5055 uses for the quality of the code’s internal structure, as opposed to what the software does when it runs. Code QA (quality assurance) is the habit of running these checks on every change; it describes a practice, not a job title.

To check if code is good quality without opening a single file, ask for the output of each check below and read one line of it.

CheckWhat it tells youThe output to ask forWhat it cannot tell you
Type checkErrors the compiler can prove, such as a value passed where another type was expectedThe error count with strict settings onWhether the logic does what the business needs
LintKnown bad patterns and likely bugs that the chosen rules describeThe error count, with the rule set namedWhether those rules matter for this app
Tests on the critical pathsWhether signup, login, the core action and payment still work after a changePass and fail totals, and the user flows coveredAnything no test exercises; a coverage percentage alone says little
DuplicationHow much logic is copied between filesThe percentage of duplicated lines and the largest clonesWhich copy is the correct one
ComplexityWhich functions hold too many decisions to test easilyThe functions over the limit, by nameWhether a complex function is wrong
Dead codeFiles, exports and dependencies nothing importsThe unused listWhether something is loaded in a way the tool cannot follow
Dependency auditKnown vulnerabilities in the packages you installnpm: findings with a severity each; pip-audit: the list of known-vulnerable packagesWhether a package is abandoned, which is a separate look

None of the seven knows whether the business rules are right, or whether user A can read user B’s data. Those need a person who reads the code, or tests written for exactly that.

What is a code smell, and how it differs from a bug

A code smell is “a surface indication that usually corresponds to a deeper problem in the system”, in Martin Fowler’s words; he credits the term to Kent Beck. A smell is quick to spot and not always a problem: long functions, duplicated blocks, huge files and long parameter lists are common ones.

A bug is behavior that is wrong. A smell is a property of the code that makes the next bug more likely or the next fix harder, and the program can run perfectly well with it. Martin Fowler on code smells makes the two points behind that: a smell is by definition quick to spot, and smells do not always indicate a problem (“Some long methods are just fine”). He also notes that inexperienced people can spot smells even when they do not know enough to judge them, and the table below relies on exactly that.

SmellWhat it looks like from outsideThe check that flags it
Long functionOne function name keeps appearing in bug reports and review commentsComplexity scan, or a lines-per-function rule
Huge fileOne file is far longer than the rest in a file listing sorted by sizeA file-size sort or a lines-per-file rule
Duplicated blockThe same fix has to be made in several placesDuplication scan
Long parameter listEvery call to a function passes a long row of valuesA lint rule on parameter count
Dead codeFiles and exports that nothing importsDead code scan
A “utils” file everything importsAlmost every change touches the same shared fileA dependency graph, or the churn list

In my reading, two that turn up often in AI-built code are duplication (the same fetch written into several components) and dead code left behind by prompts that were abandoned halfway. Removing them safely is its own job: how to find unused code in a repo. A smell is the plain-language version of the numbers in the sections that follow.

Why it matters for an app you cannot read

For a startup, the cost of low code quality shows up as the price and risk of the next change. Every fix breaks something else, a new developer needs weeks before they are useful, and an investor’s technical reviewer reads the repository and prices what they find into the round. That is my reading of where the money goes, and the table sorts it by the check that would have warned you.

What low quality costsHow it shows upWhich check would have warned
Fixes that break other thingsA bug fixed on one screen reappears on anotherTests on the critical paths; duplication scan
Slow onboardingA new developer cannot run the app or find where things liveType check; dead code scan
Slower featuresEstimates grow every month for changes of the same sizeComplexity scan, read with the churn list
A harder due diligenceThe reviewer’s findings list has items you could have fixed firstAll seven outputs, kept current
Security holes in old codeA package with a known vulnerability is still installedDependency audit

The 61.1 average from the opener, scored on all 21 third-party apps, ranks 10th of the 12 pillars, counting from the weakest. All 5 of my own production apps, which went through the same audit, had no tests. At least 17 of the 21 third-party apps had no deploy gate, so every push ships straight to production with nothing checking it first, and several apps disable their own type and lint checks at build time.

Those 21 are the 11 public and 10 held-out third-party apps from my June and July 2026 audits, and the 5 are my own; together they are a selected set, not a measure of how AI-built apps usually turn out. What I take from it: the checks exist in most stacks, and in these apps they were switched off or never required. What the pillar averages measure, and how the debt builds up, is covered in technical debt in AI-generated code; the term itself is a separate question: what tech debt means in an app you did not write. None of this says your code is bad. The outputs of the checks decide that.

How it works: the checks, the numbers, and what each is worth

Each section below takes one family of numbers: what it counts, what it predicts, and what it does not.

How to check code quality without reading code: ask for seven outputs

Checking code quality without reading code means asking for 7 outputs and reading the summary line of each: the type checker’s error count, the linter’s error count, the test run’s pass and fail totals, the duplication percentage, the list of functions over the complexity limit, the unused files list and the dependency audit.

Start in the browser. On GitHub, open the repository and click Actions under the repository name; other hosts call the same page CI or pipelines. If there are no runs, or the only runs are deploys, then nothing checks a change before it is merged, and that is your first finding. A host’s build may still run the type check when it deploys, so ask one more question: does a type error stop the deploy?

These seven are the code quality metrics I would use to measure code quality on a small team. Each ask below is a message you can send a developer as written, with the line to read and what a good answer looks like in my working rules.

  1. 01 Type check. Send: "Run the project's own type-check command with strict settings on and send me the last line." Read the error count. Good: zero errors, with no suppression comments added to get there.
  2. 02 Lint. Send: "Run the linter and send me the summary line, and tell me which rule set it uses." Good: zero errors, and a rule set someone can name.
  3. 03 Tests. Send: "Run the tests, send me the totals line, and list the user flows they cover." Good: everything passes, the pass count is above zero, and signup, login and payment are on the list.
  4. 04 Duplication. Send: "Run a duplication scan such as jscpd and send me the percentage and the largest clones." Good, by my working line: duplication in single digits.
  5. 05 Complexity. Send: "List every function over the complexity limit, by name, with its score." Good: none over the limit without a written reason.
  6. 06 Dead code. Send: "Run an unused-code scan such as knip and send me the list." Good: the list exists and someone has gone through it.
  7. 07 Dependency audit. Send: "Run npm audit (or pip-audit on Python) and send me the summary." Good: no high-severity finding on npm; on Python, no known-vulnerable package.

Two of those asks need care. On Vite’s React TypeScript template, the root tsconfig.json lists no files ("files": []) and only points to two other configs, and the build runs tsc -b, so, by my reading of that empty list, a bare tsc --noEmit at the root has nothing to check there: ask for the project’s own command. And npm audit needs a lockfile by default (“By default npm requires a package-lock or shrinkwrap in order to run the audit”), while it prints a severity for each advisory; pip-audit’s README documents no severity field, and its exit code says only whether known vulnerabilities were found, which is why the Python answer is a yes or no.

In my June and July 2026 audits, a voice-AI SDK had a test command and an installed test runner and not one test in the repository. The lesson I take from it: a test command proves nothing until it runs something, so ask for the totals line, never just whether a test command exists.

A paste-in “check my code” tool reads a snippet, not the system around it, and whether any tool can even tell AI-written code apart is covered in how to check if code is AI generated. A three-command version for a founder who has a terminal is the “How to measure it in ten minutes” section of the maintainability article linked above. Checks on the running app that need no code access at all are in ten checks on an AI-built app you can run without reading code.

Code complexity metrics: cyclomatic, cognitive, and what a complexity score means

A code complexity metric counts how hard a piece of code is to follow. Cyclomatic complexity counts the independent paths through a function, and NIST’s structured testing report gives 10 as the limit McCabe proposed. Cognitive complexity weights nesting. Both are scored per function, so a repository-wide average hides the functions that matter.

Cyclomatic complexity was defined by Thomas McCabe in “A Complexity Measure”, IEEE Transactions on Software Engineering, December 1976. The source I link for the definition and the limit is NIST’s structured testing report, which McCabe co-wrote in 1996. It speaks of modules, and says a module “corresponds to a single function or subroutine in typical languages”.

MetricWhat it countsScored perWhere the threshold comes from
Cyclomatic complexityThe decision logic in the code: the minimum number of paths that, combined, generate every path through itFunction (a module, in NIST’s wording)McCabe’s original limit of 10, as NIST’s report states it; the report adds that limits as high as 15 have been used successfully
Cognitive complexityBreaks in the linear flow of the code, with an extra increment when they are nestedFunctionNot stated in SonarSource’s white paper
Halstead measuresDistinct and total operators and operands, turned into volume, difficulty and effortFunction or moduleNot stated in radon’s docs
Maintainability indexA formula over lines of code, cyclomatic complexity and Halstead volumeModuleradon’s own ranks: A from 20 to 100, B from 10 to 19, C from 9 to 0; its docs call it “still a very experimental metric”
Lines of codeLines per function and per fileFunction, fileNone cited on this page
CouplingHow many other modules a file depends onFile or moduleNone cited on this page

Composite scores such as the maintainability index are hard to act on, in my reading: when the number drops, it does not say which function to open. Lines of code are crude and still useful as a smell. On whether a codebase is “a lot” of code, size matters less, in my reading, than how much of it is duplicated or unused.

A complexity score belongs to one function, and so do the complexity levels a tool prints for it, such as radon’s letter ranks; my rule is that a repository-wide average of them is noise. In this structural sense, you measure code complexity by taking each measure in the table per function or per file, and code complexity measurement tools do that counting for you: radon for Python, ESLint’s complexity rule for JavaScript and TypeScript, and escomplex, whose repository describes it as “Software complexity analysis of JavaScript-family abstract syntax trees”. escomplex’s last commit on its main branch is from July 2017 and its latest npm release, 2.0.0-alpha, from February 2016; the repository is not archived.

Time complexity, written in Big-O notation, is a different subject: it describes how running time grows with the size of the input, and the FAQ below covers it.

Cyclomatic complexity too high: what the number means and how to reduce it

Cyclomatic complexity is too high when one function holds more decisions than a person can test. NIST’s report says McCabe’s original limit of 10 has significant supporting evidence and that limits as high as 15 have been used successfully. Reduce it by extracting functions, returning early instead of nesting, and replacing long condition chains with a lookup table.

The warning comes from a rule with a limit. ESLint’s complexity rule warns when a function crosses its configured threshold, which defaults to 20, and counts classic McCabe complexity unless you choose its modified variant. In Python, radon ranks each block’s cyclomatic complexity from A to F: 1 to 5 is A, 6 to 10 is B, 11 to 20 is C, 21 to 30 is D, 31 to 40 is E and 41 or more is F. radon cc -s --min C . lists every block ranked C or worse, with its score.

NIST’s report adds a condition: limits over 10 “should be reserved for projects that have several operational advantages over typical projects”, such as experienced staff, formal design, code walkthroughs and a comprehensive test plan. A cyclomatic complexity of 20 is high against NIST’s figures: it is twice McCabe’s original limit and above the 15 the report says has been used successfully, while it sits exactly at ESLint’s default threshold, which a function has to cross before the rule warns, and radon ranks it C, “moderate - slightly complex block”. My working rule follows the policy the report recommends: keep a limit, and write down the reason wherever a function goes over it.

The five moves, in the order that carries the least risk in my working rules:

  1. 01 Return early instead of nesting: handle the error and edge cases at the top, so the main path is not buried inside other conditions.
  2. 02 Extract a well-named function for each branch that does real work.
  3. 03 Replace a long if-else or switch chain with a lookup table when the branches only choose a value.
  4. 04 Separate validation from work: one function checks the input, another acts on it.
  5. 05 Delete branches nothing reaches, once a test or a coverage report shows they are dead.

Every move goes behind a test first, because reducing complexity is a refactor and a refactor must not change behavior; how to write end-to-end smoke tests is where to start. What not to do: split one function into five that only make sense together, just to get under the number. NIST’s report warns that splitting a control dependency across a module boundary risks “introducing control coupling between the modules”.

A constructed illustration of move 3, not code from any audited app:

// Before: one branch per plan; a new plan means a new branch
function seatLimit(plan: string) {
  if (plan === 'free') {
    return FREE_SEATS;
  } else if (plan === 'team') {
    return TEAM_SEATS;
  } else if (plan === 'business') {
    return BUSINESS_SEATS;
  }
  throw new Error(`Unknown plan: ${plan}`);
}

// After: the branches become data
const SEAT_LIMITS = new Map([['free', FREE_SEATS], ['team', TEAM_SEATS], ['business', BUSINESS_SEATS]]);
function seatLimitAfter(plan: string) {
  const limit = SEAT_LIMITS.get(plan);
  if (limit === undefined) throw new Error(`Unknown plan: ${plan}`);
  return limit;
}

Code churn, and what the research found it predicts

Code churn is how much a file changes over a period, counted from version control history. On its own it predicts little: Microsoft Research’s study of Windows Server 2003 found absolute churn a poor predictor of defect density and churn relative to other variables, such as size and time, highly predictive. My working list pairs churn with complexity.

The command that ranks files by how often they changed, and the list of files near the top of both the size and the churn rankings, are in the “How to measure it in ten minutes” section of the maintainability article; what churn looks like in AI-built repos is its “The doom loop, measured” section. This page adds the research. Microsoft Research’s relative code churn study, by Nagappan and Ball (May 2005), relates the amount of churn to “other variables such as component size and the temporal extent of churn”, and reports that its churn metric suite is able to tell fault-prone binaries from the rest with an accuracy of 89.0 percent.

My working rule, not the study’s: I pair churn with the complexity list to pick the first files to test and simplify. The study measured relative churn, not that pairing.

Developer productivity metrics: what to measure in engineering, and what backfires

Developer productivity metrics measure a delivery system, not a person. DORA’s software delivery metrics include deployment frequency and change fail rate, the two I’d have a team of one to three read from its own deploy log. Counting lines, commits or tickets per developer measures activity and rewards the wrong work.

DORA’s software delivery metrics are five today: change lead time, deployment frequency and failed deployment recovery time for throughput, and change fail rate and deployment rework rate for instability. DORA’s own list of pitfalls includes setting a metric as a goal, which it says increases the likelihood that teams will try to game the metrics, and “Having one metric to rule them all”. The SPACE framework from Forsgren and colleagues (ACM Queue, 2021) makes the same point from another side: developer productivity “cannot be measured by a single metric or dimension”.

For a team of one to three, an AI agent included, my working rule is to track two numbers from the deploy log: deployment frequency, and change fail rate, which DORA defines as the ratio of deployments that need immediate intervention afterwards. At this size, that is how to measure developer productivity: count what reached users and how often it broke. Measuring engineering by the person, with lines of code, commit counts or tickets closed, measures activity and invites gaming, and in my reading AI assistance inflates all three. The quality checks on this page keep the delivery numbers honest, because a fast deploy that fails its checks was never fast.

Code quality standards: ISO 5055 and CISQ, and what they are worth to a small team

ISO/IEC 5055 is the standard for automated source code quality measures, adopted by ISO in 2021 from the OMG standard built on CISQ’s measures. It measures four factors: reliability, security, performance efficiency and maintainability. For a team of two it is worth knowing as vocabulary; running the checks matters more than formal measurement against it.

CISQ developed the measures for the four factors; they were approved as OMG standards and combined into OMG’s Automated Source Code Quality Measures, the standard ISO adopted as ISO/IEC 5055:2021. CISQ itself was co-founded by the Object Management Group and the Software Engineering Institute at Carnegie Mellon University. The measures work by “detecting and counting severe weaknesses in source code”, in OMG’s words; each weakness is an entry in the Common Weakness Enumeration, and CISQ’s page says the standard is implemented by vendors of static analysis technology.

What that is worth to a team of two, in my reading: the vocabulary, and the weakness lists as a reading list of mistakes to avoid. A procurement questionnaire from a larger customer may ask about ISO 5055; CISQ’s page says the measures can be written into contracts with third-party suppliers. The useful answer from a small team is to show that these checks run on every change, with the outputs attached.

The Production Hardening Sprint draws the same line: formal third-party certifications and independent audit opinions are separate from the sprint deliverables.

Code quality checkers by type: what each looks at, and the free ones

Code quality checkers fall into 7 types by what they check: linters, type checkers, test runners, duplication detectors, complexity analyzers, dead code finders and dependency auditors. Each type has an open-source example. Commercial platforms bundle them behind one dashboard and a grade, which adds convenience rather than extra evidence.

How jscpd, the ESLint complexity rule, radon and knip are used to track a codebase over time is the “Tools that go past grep” section of the technical debt article linked above. This table only names the type and open-source examples of each; no tool here was tested for this page, and none is ranked.

Checker typeWhat it checksAn open-source exampleWhere it runs
LinterKnown bad patterns and likely bugsESLint for JavaScript and TypeScript; Ruff for Python; SQLFluff for SQLThe command line, so any CI step
Type checkerType errors the compiler can proveThe TypeScript compiler, tscThe command line and the build
Test runnerWhether the tests passWhichever runner the project’s test script callsThe command line and CI
Duplication detectorCopied blocks across filesjscpd, a “Copy/paste detector for source code” (MIT license)The command line, a GitHub Action or a pre-commit hook
Complexity analyzerFunctions over a limitradon for Python; ESLint’s complexity rule; escomplex for JavaScriptThe command line and CI
Dead code finderUnused files, exports and dependenciesknip, for JavaScript and TypeScript projectsThe command line and CI
Dependency auditorKnown vulnerabilities in installed packagesnpm audit (needs a lockfile by default); pip-audit for PythonThe command line and CI

Two related jobs are separate topics: linters for a codebase an AI wrote, and how to enable TypeScript strict mode. The docs for the three smaller tools are jscpd, knip and SQLFluff’s getting started guide. To install SQLFluff, the guide’s command is pip install sqlfluff, and sqlfluff lint test.sql --dialect ansi lints one file with the ansi dialect.

A free online code checker has the same limit as the paste-in tools above, in my reading: it sees the lines you paste and none of the files they depend on. Comparing named products, and the lists of tools that go with that, is the job of static code analysis tools; security scanners and what a scan finds are in what a code scan is; AI review bots, and who ranks them, are in who says which AI code reviewer is best. SonarQube, Codacy and similar commercial platforms are automated code quality tools that roll several of these checks into one report with a grade. That grade is their formula, in my reading, so read the findings under it.

How to improve code quality, and how to maintain it once it is there

Improving code quality on an inherited app follows one order in my working rules: tests on the paths that make money first, then strict types, then lint with the noisy rules off, then all three required in CI, and only then the complexity hotspots. Maintaining it means the checks block a merge, so quality stops depending on memory.

This is how to check your own app before you change anything. The eight code quality assessment questions below are the ones I would put to a developer, an agency or an acquirer’s reviewer, each with the evidence that answers it.

  • Which checks run on every change, and which of them block a merge? Evidence: the CI configuration and a pull request that a failing check held back.
  • How many type errors are there with strict settings on? Evidence: the last line of the type-check run.
  • What do the tests cover, named by user flow? Evidence: the list of flows and the latest totals line.
  • When did a test last catch a real bug? Evidence: the failing run and the fix that followed it.
  • Which ten files change most, and are they tested? Evidence: the churn list next to the test list.
  • What is unused? Evidence: the dead code scan’s output, and what was done about it.
  • How long does a new developer need to run the app locally? Evidence: the setup instructions, timed by someone who has not seen the app.
  • What is an AI agent forbidden to touch? Evidence: the rules file and the check that enforces it.

The improvement order, in my working rules:

  1. 01 Tests on the money paths: signup, login, the core action and payment.
  2. 02 Strict types, with the errors fixed rather than suppressed.
  3. 03 Lint with the noisy rules off, so every remaining error is worth fixing.
  4. 04 All three required before a merge.
  5. 05 The hotspots: files high on both the complexity list and the churn list.

Testing comes first because improving code quality without tests means every later change is a guess; the smoke tests guide linked in the cyclomatic complexity section is that step. For step 4, the gate is GitHub branch protection or a ruleset with required status checks, and after enabling them “all required status checks must pass before collaborators can merge changes into the protected branch”. GitHub’s docs say protected branches and rulesets are available in public repositories with GitHub Free, and in private repositories with GitHub Pro, GitHub Team or GitHub Enterprise Cloud.

To maintain code quality once it is there, the checks run in CI on every pull request, a failing check blocks the merge, and an AI agent works under the same gate as a person; the agent’s side is a separate topic: how to verify AI-generated code before production. A human review still reads what no tool can, and the code review checklist lists what that review covers.

Where the sprint fits

In the Production Hardening Sprint, deliverable 10.1 removes dead code and consolidates duplicated logic while preserving required behavior ; 10.5 automates smoke tests for signup, login, the core product action and payment flows ; 10.6 enables TypeScript strict mode and resolves errors in TypeScript applications, with an equivalent strict-checking approach for other supported stacks ; 10.7 adds unit tests around authorization and payment logic, including negative and edge cases ; and 10.8 provides CLAUDE.md, AGENTS.md, Cursor rules or equivalents describing conventions and protected patterns, and adds CI checks for enforceable rules. Deliverable 7.3 runs linting, type checks, builds and tests on every pull request. The result for each goes into the production readiness report, which accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. Building new product features or modules sits outside the sprint. Each item, with how it is verified, is listed in the code health checks in the published scope.

Common questions about measuring code quality

What is the 80/20 rule in coding?

The 80/20 rule in coding is the Pareto principle applied loosely: the idea that a small share of causes produces most of the effects, such as a few files taking most of the edits. The principle itself says that for many outcomes roughly 80% of consequences come from 20% of causes; Joseph Juran developed it for quality control after reading Vilfredo Pareto, who showed that about 80% of the land in Italy was owned by 20% of the population. In coding it works as a rule of thumb, not a measured ratio, and a churn list is where it shows up in a repository.

How to measure time complexity of code?

Time complexity is algorithmic complexity, written in Big-O notation: how running time grows as the input grows. You work it out from the loops (one pass over the input grows in step with it, a loop inside that loop grows with its square) and check it by timing the code on inputs of increasing size. It is a separate subject from the structural complexity on this page, which counts the decisions inside a function rather than the work done per input.

Can AI detect smells?

Yes, for the pattern kinds, in my reading: long functions, duplicated blocks and deep nesting are patterns, and linters already flag many of them without AI. An AI reviewer’s findings need the same checking as AI-written code, because either can be wrong. Who claims which AI reviewer is best is the subject of the AI code review article linked in the checkers section.

Does AI actually boost developer productivity?

DORA’s 2025 report on AI-assisted software development describes AI’s primary role as an amplifier, “magnifying an organization’s existing strengths and weaknesses”, and says the greatest returns come from the underlying organizational system rather than the tools themselves. My line: the checks on this page are what make the extra output safe to merge.