A change to the pricing page can ship on a Friday and leave nobody able to pay until a customer emails on Monday, because nothing ran checkout before the deploy. This is how to write end to end smoke tests for signup, login, the core action and a Stripe test payment, and the unit tests behind them.

How to write end to end smoke tests, and the unit tests beside them: what these controls are, together

The smallest test suite that protects a SaaS, by my working rule, is four end-to-end smoke tests, for signup, login, the core action and payment, run against a deployed build before every release, plus unit tests around the permission and billing decisions with the negative cases written out. Everything else can wait.

Both controls belong to the code health area of the engineering standards for AI assisted teams: the habits that keep an app safe to change once an AI tool writes much of its code.

ControlWhat it isWhy it matters
Critical-path smoke testsThe four journeys (signup, login, the core product action, payment) driven in a real browser against the deployed build, in CI on every deploy; the deploy waits for them where the host and plan allow it (step 6 below)A change in one file can break checkout in another, and nobody finds out until a customer does
Auth and billing unit testsTests on the decisions behind login, permissions and money, each with its negative and edge cases written outA permission or billing mistake costs a customer’s data or money, and any later edit can bring it back

The four journeys are your critical user journeys, and for a first suite they are the whole critical user journey testing checklist: a stranger signs up, a customer logs in, the customer does the one thing the product exists for, and the customer pays. End to end testing automation here means a real browser driving the deployed app the way a customer would, with nobody clicking. Among e2e testing frameworks, Playwright and Cypress both drive the browser; Playwright describes its test runner as “an end-to-end test framework for modern web apps”. Jest, Vitest and pytest run the unit tests. None of the choices matters as much as which journeys get tested.

In the Production Hardening Sprint these are two deliverables: 10.5 automates smoke tests for signup, login, the core product action, and payment flows, and 10.7 adds unit tests around authorization and payment logic, including negative and edge cases.

What a smoke test is

A smoke test is the smallest check that a build is not broken in an obvious way, run first and fast. For a web app it is the four journeys a customer must be able to complete, driven in a real browser against the deployed build.

It runs before anything slower. If login is down, there is no point waiting for the rest of the suite to report the same outage again and again. The definition is mine, written for a small web app rather than taken from a testing glossary.

A full regression suite is about breadth: every feature checked against its last known good behavior. A smoke test is the opposite, the few paths that must never break, checked on every build. Smoke testing in software engineering names the same idea for any kind of build, a compiled binary as much as a web app: prove the thing starts and does its main job before anyone spends time on the details.

Unit, integration and end to end: which layer catches which failure

Integration testing checks that separately built parts work together. A unit test catches the wrong permission decision and misses the missing route. An integration test catches the query that returns another tenant’s rows and misses the broken button. An end-to-end test catches the broken button and misses the reason. Use each where it is cheapest.

LayerWhat it exercisesThe failure it catchesThe failure it misses
UnitOne function, such as canReadInvoice(user, invoice), with no database or browserA permission check that returns true for the wrong userA route that was never wired to that function
IntegrationYour code plus a real database, queue or API, without a browserA query that returns another tenant’s rowsA button that no longer submits the form
End to endThe deployed app, driven in a browser from the first page to the stored resultA signup, login or checkout that no longer completesWhich line of code caused it

The table is my reading of where each layer earns its cost, not a standard. The e2e vs integration testing question comes down to what each can see: the integration test knows which query leaked a row, while the end-to-end test only knows the page showed the wrong thing. Software integration and test is the enterprise name for the middle row, and system testing is the usual name for the end-to-end row, so unit, integration and system testing describes the same three layers.

One integration test best practice, my working rule: run it against a real database with rows for two tenants seeded, then ask as the first tenant for the second tenant’s data and expect nothing back. A mocked database cannot fail that test, which is why it needs a real one.

What goes wrong without them

Three failures, each one a missing test.

A change to pricing ships and nobody can pay

Say you run a small SaaS. You change the pricing page and deploy on a Friday. The change breaks the path from the pricing page to payment, and because no check in your deploy exercised payment, your first sign is a customer’s email days later. That is how a deploy broke the checkout flow while every existing check stayed green. The point of the case: one end-to-end payment test in the deploy gate is the cheapest place to catch a change that stops customers paying, because it fails that change before it ships.

In the apps I audited, at least 18 of the 21 third-party apps had no working test anywhere: 17 with literally none, plus a retail POS whose checkout “test suite” never executed the actual checkout code. Those 21 come from my June and July 2026 audits, 11 public apps and a held-out set of 10, a selected set rather than a random sample, so the count describes those apps and is not a rate for AI-built apps in general.

Why a suite that exists can still prove nothing is its own subject, covered in when tests pass but the app is still broken.

A permission check is loosened and another customer’s data shows

In the same audits, 7 of the 21 third-party apps had confirmed cross-user or cross-tenant authorization failures, where a logged-in user could read or write another customer’s data. That is the same hand-picked group of apps, so treat it as what those apps showed, not a forecast for yours.

The unit test that catches a loosened check is one negative case: user B asks for user A’s record with user B’s own credentials, and the test expects a refusal. Where row-level security filters the query instead of raising an error, a correct build returns no rows, so the test expects an empty result rather than an error code. Either way, the test fails the day someone relaxes the check.

The demo passed, so the site was “tested”

Testing a website before launching is not clicking through it once on the demo account. A click-through proves one path worked once, on one machine, with data the builder already set up. The manual checks worth doing, in order of consequence, are in how to test a vibe-coded app. A beta with real users comes after the suite exists, not instead of it; beta testing an app you built with AI covers that stage.

How to set them up

Write the four journeys first, against a deployed preview, each asserting one visible outcome and one stored one. Then write my six unit cases for auth and billing, three of them negative. Then make the deploy wait for all of them.

The four end-to-end smoke tests

  1. 01 Point the browser runner at a deployed preview, not localhost, and set baseURL from the preview address. Evidence: the CI log shows the preview URL the run used.
  2. 02 Signup: a fresh email address, the confirmation step, then the first screen a new user sees. Evidence: the new user row exists and the first screen rendered.
  3. 03 Login: the right password reaches the dashboard, then the wrong one is rejected. Evidence: the rejection message and no session.
  4. 04 The core action: the one thing the product is for, asserting the result appears on screen and is stored. Evidence: the saved record, read back.
  5. 05 Payment: drive the app as far as its own hand-off to Stripe in test mode, then prove the paid state in test code. Evidence: the payment recorded against the test user.
  6. 06 Make the deploy wait for the four. Evidence: a deploy that was held because one of them failed.

On step 1, Playwright’s CI guide has an “On deployment” workflow that it describes this way: “This will start the tests after a GitHub Deployment went into the success state. Services like Vercel use this pattern so you can run your end-to-end tests on their deployed environment.” That workflow hands the deployed address to the tests as PLAYWRIGHT_TEST_BASE_URL; point baseURL in the Playwright config at it. Where the host puts previews behind a login, the runner sends the host’s automation bypass: Vercel says “Protection Bypass for Automation is available on all plans”, sent as the x-vercel-protection-bypass header or query parameter.

On step 2, the confirmation needs either an inbox the test can read or the auth provider’s own way to confirm a test user. Pick one before writing the test, or signup will be the journey that always gets skipped.

Step 3 is the one worth copying first, because the wrong-password case is the half people leave out. One Playwright test, with the visible outcome and the stored one (no session) both asserted:

import { test, expect } from '@playwright/test'

test('login rejects a wrong password', async ({ page }) => {
  await page.goto('/login')
  await page.getByLabel('Email').fill(process.env.SMOKE_USER_EMAIL!)
  await page.getByLabel('Password').fill('not-the-password')
  await page.getByRole('button', { name: 'Sign in' }).click()
  // Visible: use the exact error text your app shows
  await expect(page.getByText('Wrong email or password')).toBeVisible()
  // Stored: no session, so a protected page sends you back to login
  await page.goto('/dashboard')
  await expect(page).toHaveURL(/\/login/)
})

Step 5 has a limit worth knowing before you write it. Stripe says its “Frontend interfaces, like Stripe Checkout or the Payment Element, have security measures in place that prevent automated testing”. So the browser test stops at the app’s own hand-off: the checkout session is created and Stripe’s page or form appears. The paid state is then proved in test code without Stripe’s form. Stripe’s testing docs say: “When writing test code, use a PaymentMethod such as pm_card_visa instead of a card number”, so a test can make a payment on the test user’s own Stripe customer with pm_card_visa. The other route is the one Stripe’s webhook-testing page shows: post an event signed with a test secret (its example uses whsec_test_secret) to the handler, and expect 200 for a valid signature and 400 for a bad or missing one. For that route, the event carries the test user’s ids and the handler under test is configured with the same test secret. Either way, the test then asserts that the handler recorded the payment for that user.

A bare stripe trigger payment_intent.succeeded does not do this job. Stripe’s webhook docs show it printing “Running fixture for: payment_intent”, and that fixture’s payment belongs to no user of your app, so a correct handler has nothing to record for the test user. On a deployed preview the event also has to reach the build: through a webhook endpoint registered for that preview, or through the Stripe CLI’s stripe listen --forward-to, which Stripe says must run “with Stripe CLI in a terminal”. Without one of those, a correct build fails the check. Before release, a person pays once in the real form with a test card such as 4242 4242 4242 4242, which Stripe lists under “Testing interactively”.

On step 6, the promotion gate belongs to the release readiness checklist, including whether your host can hold a deploy and on which plan. Two facts shape it. A required status check on the branch stops a merge, but GitHub says “Protected branches are available in public repositories with GitHub Free and GitHub Free for organizations” and, for private repositories, with GitHub Pro, GitHub Team, GitHub Enterprise Cloud and GitHub Enterprise Server. And a host that deploys every merge needs its own gate: Vercel, for one, promotes the latest successful production build by default, and its Deployment Checks “hold each production deployment until all required checks pass”.

My working rule for all four: each test asserts one visible outcome and one stored one. A green page with nothing saved behind it is how a broken signup passes.

The auth and billing unit tests

These six cases are my list for any app that has logins and charges money:

  1. 01 The wrong user cannot read the record.
  2. 02 The wrong role cannot call the admin action.
  3. 03 An expired session is rejected.
  4. 04 A plan change moves the subscription to the right state, and no other.
  5. 05 A failed payment moves the subscription to the state the app defines for unpaid, not to canceled.
  6. 06 A refund never leaves the ledger negative.

Here is the first case in Vitest, in a file named permissions.test.ts, since Vitest by default only picks up files with .test. or .spec. in the name. The same shape works in Jest or pytest with their own syntax.

import { expect, test } from 'vitest'
import { canReadInvoice } from './permissions'

// Why: a loosened ownership check shows one customer another's invoices
test('refuses user B reading user A invoice', () => {
  const invoiceOfA = { id: 'inv_1', ownerId: 'user_a' }
  const userB = { id: 'user_b', role: 'member' }
  expect(canReadInvoice(userB, invoiceOfA)).toBe(false)
})

Unit test documentation, for a suite this size, is two things: a test name that says what it proves, and a one-line comment that says why the case exists. The name ends up in every CI log, so “refuses user B reading user A invoice” tells whoever reads a red run what broke without opening the file. Where authorization should be enforced in the first place is a web app security question; these tests pin the decision wherever it lives.

Writing the cases: happy path, boundary and negative

Structural testing writes cases from the code’s own branches. Every case is one of three kinds: the happy path, the boundary (the last day of a trial, a zero invoice), or the negative (a downgrade the plan forbids). Name the outcome, write the assertion, then the setup.

Structural software testing sounds academic, but for a billing module it means reading each if and writing a case for each side of it. The happy path is a valid upgrade that should just work. Boundary testing puts the input exactly on an edge: the trial’s last day rather than its middle, an invoice that totals nothing rather than a normal one. A negative test case feeds the code something it must refuse and passes only when it refuses, such as a downgrade to a plan the customer’s current usage does not fit, expecting a rejection.

Inside out testing means writing the unit tests first and the journeys last. On an app that already exists I would go the other way round, because four journeys cover more ground on day one than any unit test can. As a testing procedure, starting from the assertion keeps each case honest: you decide what “correct” looks like before the setup code can bend it.

Testing checklists you can copy

A checklist in software testing is only useful if every line is something a test can pass or fail, so every line below is.

The journey list, which also serves as the UI testing checklist for the screens that matter:

  • A new visitor signs up, confirms, and reaches the first screen.
  • A customer logs in with the right password; the wrong password is refused.
  • A customer completes the core action and the result is stored.
  • A customer reaches the Stripe hand-off and the paid state is recorded.
  • Password reset works end to end, when the app has one.
  • Account deletion removes the account, when the app has one.

The billing logic test coverage checklist, which is the six unit cases above plus four that bite later:

  • The six auth and billing cases, each with its negative written out.
  • Proration on a mid-cycle plan change lands on the amount you expect.
  • Trial end moves the account to the right paid or unpaid state.
  • The same event delivered twice is processed once, by its event id.
  • A replay test signs its payload fresh at run time.

The duplicate case follows Stripe’s own webhook advice, “Track event IDs to identify duplicate deliveries”, and the replay case has to sign fresh because Stripe says “Our libraries have a default tolerance of 5 minutes between the timestamp and the current time”.

The developer testing checklist, run before any merge:

  • The whole suite ran, not a filtered subset.
  • Coverage on the auth and billing modules did not drop.
  • No test was skipped or marked to be fixed later.

Hand these lists to whoever does QA and they are the software QA checklist for the app. An auditor can use the same lists as a software QA audit checklist, asking for the red run as the evidence behind each line. For an integration testing checklist template, add one line to the billing list: seeded rows for two tenants, with every query run as each of them. If you want one testing checklist for the whole application, it is these three lists in this order. None of them needs a download; copy the lines into your repo’s README or pull request template.

TDD and BDD on an app that already exists

Test-driven development cannot put the test first on code that already exists. The honest version of TDD on an AI-built app is: write the four journeys as the safety net now, then write the test first for every change from here on.

TDD writes a failing test, then the smallest code that passes it, then tidies both. The advantages of test-driven development come down to two benefits of TDD: a design that has to be testable, and a regression net that grows with every change. The cost of TDD software development is a slower first week, while both the tests and the habit get built.

BDD writes each case as a sentence a non-engineer can read, which is what a founder-written journey test already is: “a new visitor can sign up and reach the first screen”. Behavior-driven development (BDD, spelled behaviour driven development in British English) fits agile development well because each user story already names who wants what, and the story becomes the scenario. Cucumber, a BDD tool, calls BDD “the software development process that Cucumber was built to support”, and its Gherkin syntax “uses a set of special keywords to give structure and meaning to executable specifications”. You do not need either tool to get the benefit; plain test names written as sentences do most of the work.

How to verify each one

The proof that a test suite works is a red run: break the login redirect or remove the webhook handler on purpose, watch CI fail, revert, watch it pass. Keep both runs. A suite that has never been red proves nothing.

  1. 01 Break something on purpose on a branch: change the login redirect, or remove the webhook handler.
  2. 02 Push and watch CI go red on the journey that covers it. Evidence: the red run link and its time.
  3. 03 Run the unit tests and read the refusal cases pass by name, then loosen one permission check and one billing transition and watch those cases turn red. Evidence: both outputs.
  4. 04 Revert and watch CI go green. Evidence: the green run.
  5. 05 Keep the diff, the red run and the green run together, dated.

To verify tests fail on an intentional bug, the bug has to sit inside something the suite claims to cover; breaking a page no journey visits proves nothing about the suite. Software verification asks whether the product was built right, and this is the part of that question a small app can answer with evidence: these journeys and these cases, shown failing and passing.

In the Production Hardening Sprint, deliverable 10.5 is verified this way: run the suite in CI and demonstrate that an intentional regression fails it; deliverable 10.7 is verified this way: run the tests and show rejection of unauthorized access and incorrect billing transitions.

Coverage: what the number means, and where to stop

Coverage counts the lines a test ran, not the lines a test checked. My working rule: push the four journeys and the auth and billing modules high, leave the rest, and list every excluded branch, because an exclusion comment hides it from the number.

Code coverage analysis tells you which lines executed while the tests ran. It cannot tell you whether any assertion looked at the result; what a green run and full coverage cannot see is a section of the tests-passed article linked above. To analyze code coverage usefully, read it per module, not as one number for the whole repo.

My working rule is also where to stop: past those modules, chasing full coverage buys tests of getters and layout code while the billing branches wait.

The trap is the exclusion comment. In coverage.py, a line carrying # pragma: no cover is left out, and “coverage.py excludes it from the list of missing code” in its reports, according to coverage.py’s exclusion docs. Istanbul’s comments such as /* istanbul ignore next */ do the same job, which its nyc README describes as ways to “exclude from coverage tracking”. Neither tool’s exclusion docs say the report lists the excluded lines, so my rule is that the one-page test report lists every exclusion, with its reason. A code coverage analyzer is whatever produces the report: coverage.py for Python, Istanbul’s nyc for JavaScript, or your stack’s own.

Verification, validation and the test report you hand over

Verification and validation in software engineering split one question in two. Verification asks whether it was built right, which this suite answers. Validation asks whether it was the right thing to build, which only real users answer, in the beta mentioned earlier.

A testing report sample for a small app fits on one page: what the suite covers, what is excluded and why, the link to the last red run and the last green run, and the date. Software testing and quality assurance at this size is that page kept current, not a department.

The ROI of automation testing on a small software product is not a calculator figure. It is the list of regressions the suite caught before a customer did, so the report records each one with its red run. That list is the only test automation ROI figure worth keeping.

Where the sprint does this

Critical-path smoke tests (10.5) and auth and billing unit tests (10.7) are the two controls described above. Both results go into deliverable 13.1, the production readiness report, which is verified this way: account for all 123 IDs; keep failures visible until resolved and explain genuine non-applicable items. Building new product features or modules sits outside the sprint and is separate work. The review that reads these tests is in the code review checklist, and the automated checks that run beside them are in code quality checks. Every deliverable, with how it is verified, is in the published scope.

Common questions about smoke tests and a first test suite

Which tool is used for smoke testing?

For a web app, a browser runner such as Playwright or Cypress runs the smoke tests; this page runs them from CI against a deployed preview. The tool matters less than the journeys: a suite that covers signup, login, the core action and payment in either tool beats a larger suite that skips payment.

What is the difference between smoke and sanity testing?

Smoke testing asks whether the build works at all, by running the few journeys that must never break. Sanity testing asks whether one specific fix worked, by checking that change and what sits right next to it. The wording is mine; both are quick checks, pointed at different questions.

Is 80% code coverage good?

A coverage percentage means little without the list of what was excluded from it. My working rule is to cover the four journeys and the auth and billing modules as far as they go and not set a whole-repo target at all.

Which is better, BDD or TDD?

Neither replaces the other. TDD is how a developer writes the next change, test first. BDD is how a founder states a journey in a sentence anyone can read, and Cucumber’s Gherkin syntax is one way to write it. On an existing app, use BDD-style names for the journeys and TDD for each change after.

What are the drawbacks of using TDD?

The main drawback is speed at the start: the early days go slower because every change now comes with its test. On code that already exists there is a second one: the tests cannot come first, so TDD only starts with the next change, and the journeys have to cover what is already there.

What is a critical user journey?

Google’s SRE workbook defines it: “A critical user journey is a sequence of tasks that is a core part of a given user’s experience and an essential aspect of the service.” For a SaaS, that means the path a customer must complete for the product to be worth paying for: signing up, logging in, doing the core action and paying.