Write one seed script and run it against an empty staging database before anyone tests anything. To generate realistic fake data for staging, the script creates three kinds of rows: accounts you log in as, a few hundred records from a Faker library with a fixed seed (my working rule for volume), and the awkward cases your app must survive. It never copies production rows.
Generate realistic fake data for staging: what a seed script is and what goes in it
Realistic fake data for staging comes from one seed script with four layers: reference data that ships with migrations, fixture accounts for every role and at least two tenants, generated rows from a Faker library, and hand-written edge cases. The script refuses production, runs after migrations, and produces the same rows every run.
Staging seed data sits alongside the other controls in the data consistency checklist for SaaS. The table splits the four layers by job, size and home.
| Layer | What it is for | Rough size (my working rule) | Where it lives | Also needed in production? |
|---|---|---|---|---|
| Reference data | Plans, roles, feature flags, countries: rows the app needs before anyone signs up | Whatever the app needs to start | Migrations, not the seed | Yes |
| Fixture accounts | One login per role, spread over at least two tenants, so every permission path can be walked | One account per role in each tenant | The seed script; passwords in the staging secret store | No |
| Generated volume | Enough rows that pagination, search and slow queries show up | A few hundred to a few thousand rows per main table | The seed script, through a Faker library with a fixed seed | No |
| Edge cases | Hand-written rows for the states that break things | One row per case | The seed script, in its own named block | No |
The size and production columns are my working rules, not a standard. Reference data belongs in migrations because production needs it too, and Supabase’s docs advise that seed files hold only data insertions and avoid schema statements.
The script needs three properties: safe (it refuses to run anywhere but staging or a local machine), repeatable (the same input gives the same rows) and ordered (parents go in before children). It lives in the repository beside the migrations and runs with one command. Supabase’s seeding guide describes seeding as a way to create “reproducible environments” for local development, staging, and production, and says seeding “occurs after all database migrations have been completed”.
What is seed data in development
Seed data in development is the set of rows a script inserts so a fresh database is usable. The word has a second meaning here: the starting value of a random generator, which is fixed, with the library version and a base date pinned, so generated names, dates and amounts repeat exactly.
On a developer’s machine, seed data lets the app run on its first start. In staging, it lets people test real workflows instead of empty screens. Supabase defines seeding as “the process of populating a database with initial data, typically used to provide sample or default records for testing and development purposes”. A seed in a database, then, is that file or script plus the rows it writes.
In programming, seed can also mean the starting value of a pseudorandom generator. Python’s Faker states the rule plainly: “A Seed produces the same result when the same methods with the same version of faker are called.” The condition in that sentence, the same version, is why the seed script pins its library.
Three neighboring words get mixed up. Django’s docs say “A fixture is a collection of files that contain the serialized contents of the database”. A factory, in my reading, builds rows inside one test and throws them away. A seed fills a whole environment and stays there.
What goes wrong without it: a staging environment has no test data, or has the wrong data
An empty staging system gives the team little opportunity to test realistic workflows. Four habits produce that problem, or its cousin, a staging database full of the wrong rows.
| What the team does | What goes wrong | The seed layer that prevents it |
|---|---|---|
| Leaves the staging database empty | Nobody can check a dashboard, a report or a permission rule against zero rows, so testing drifts to production | Fixture accounts and generated volume |
| Types three rows by hand | Every list fits on one page, so pagination, slow queries and empty-state bugs never appear | Generated volume and edge cases |
| Copies production into staging | Customer data now sits in a second environment that needs production’s protections | The whole seed, which makes the copy unnecessary |
| Runs a seed with no fixed seed or base date | A bug seen on Tuesday cannot be reproduced on Wednesday | The repeatability rules in step 4 below |
When you cannot test features with an empty database, the testing moves to production, where every mistake lands on a real customer. Whether staging should hold a copy of production data at all is answered “Usually no” in one database and no staging, which also sets out what staging must isolate. The legal side sits with GDPR for a small SaaS.
A wrong-environment run is the other way seed data goes wrong. Picture a developer whose shell still holds the production connection string in its database variable from an earlier task, running the seed command to refresh staging. The seed’s first step truncates the seeded tables, and TRUNCATE ... CASCADE removes all rows from those tables and from every table that references them, in production. The guard that checks the environment and the database host before connecting is the step that stops it, which is why it comes first below; my reading is that where a seed may run is decided by the guard in the script, not by whichever connection string the shell happens to hold.
How to do it on the common stacks
Four parts follow: the shape every seed script shares, the seed command your stack already has, how to make the rows realistic, and what to say when someone suggests copying production.
The shape of a seed script: guarded, repeatable, ordered
A seed script runs seven steps: check the environment and host against an allowlist, run after migrations, reset or upsert so a second run changes nothing, fix the random seed, the library version and the base date, insert parents before children, create auth users through the provider’s admin interface, and exit non-zero on error.
The steps are my method; the tool facts under them come from each tool’s own docs.
- 01 Guard: read the environment name and the database host, and exit with an error unless both are on an allowlist (
staging,local). Never trust a default. - 02 Schema first: migrations run before the seed, and reference data arrives with them.
- 03 Reset or upsert: truncate the seeded tables, or upsert on stable keys, so a second run adds no duplicates.
- 04 Repeatability: fix the random seed, pin the faker library to an exact version, and set a fixed base date for relative dates.
- 05 Order: insert parents before children, one transaction per group of related tables.
- 06 Auth users: create login accounts through the auth provider's admin interface where it documents one, from a server process only.
- 07 Report: print rows per table and the fixture login emails, and exit non-zero on any error.
Step 1 checks two values because either can be wrong on its own: an environment variable can say staging while the connection string still points at production. Laravel’s seeding docs show a guard the framework ships: “you will be prompted for confirmation before the seeders are executed in the production environment”, and the --force flag runs them without the prompt. A CI job that passes --force needs its own check, and the docs describe that prompt for the production environment without a word about the database host. The guard below uses placeholders for your own names.
// scripts/seed.ts: the guard runs before any database connection opens
const env = process.env.APP_ENV ?? "";
const host = new URL(process.env.DATABASE_URL ?? "postgres://unset").hostname;
const allowedEnvs = ["staging", "local"];
const allowedHosts = ["<your-staging-db-host>", "localhost"];
if (!allowedEnvs.includes(env) || !allowedHosts.includes(host)) {
console.error(`Seed refused: env="${env}" host="${host}"`);
process.exit(1);
}
Step 2 keeps schema work in migrations, which run first; shipping those safely is its own subject, the database migration checklist. Step 3 is what makes a second run safe. The Rails guide’s rule for db/seeds.rb is that the code “should be idempotent so that it can be executed at any point in every environment”. If you reset instead of upserting, know what the reset does: PostgreSQL’s TRUNCATE reference says TRUNCATE “quickly removes all rows from a set of tables”, and CASCADE also truncates every table with a foreign-key reference to the named ones. Drizzle’s reset helper generates those TRUNCATE ... CASCADE statements on PostgreSQL.
Step 4 comes from the faker libraries’ own docs. Faker’s guide to reproducible results warns that “When upgrading to a new version of Faker, you may get different values for the same seed”, and that for relative-date methods such as faker.date.past, “setting a random seed is not sufficient to have reproducible results”; its fix is a fixed reference date, set once with faker.setDefaultRefDate. Python’s Faker says results “are not guaranteed to be consistent across patch versions”, and, if you hardcode results in tests, to pin the version “down to the patch number”.
Step 6 matters most on Supabase. The call is Supabase’s admin createUser, and its reference says: “This function should only be called on a server. Never expose your service_role key in the browser.” It adds that createUser() “will not send a confirmation email”, and that setting email_confirm to true confirms the address, so a fixture account does not wait on an inbox.
Step 7 makes failure loud. My working rule is to run the same script in CI against a throwaway database on every pull request, so a schema change that breaks the seed fails there, not on the next staging refresh.
The seed command your stack already has
Most frameworks already ship a seed entry point, and using it beats inventing one. Each cell below comes from the tool’s own docs.
| Stack | Where the seed lives | The command | What it does about resets |
|---|---|---|---|
| Supabase CLI | supabase/seed.sql by default | Runs on the first supabase start and on every supabase db reset | supabase db reset re-runs the seed after all migrations |
| Prisma ORM 7 | The command in the seed key of migrations in prisma.config.ts | npx prisma db seed | Seeding is only triggered explicitly, never by migrate dev or migrate reset |
| Prisma ORM 8 (release candidate) | A plain script, for example src/prisma/seed.ts | npx tsx src/prisma/seed.ts; there is no seed command | Drop and recreate the database with your database tools, then prisma db init |
| Drizzle | Your own script, using drizzle-seed | Not stated in Drizzle’s seeding docs | reset(db, schema) issues TRUNCATE ... CASCADE on PostgreSQL |
| Rails | db/seeds.rb | bin/rails db:seed | bin/rails db:seed:replant reloads the seed data |
| Django | Fixtures in each app’s fixtures directory or in FIXTURE_DIRS | django-admin loaddata <fixture label> | Not stated in Django’s fixtures docs |
| Laravel | database/seeders, starting from DatabaseSeeder | php artisan db:seed | php artisan migrate:fresh --seed drops all tables and re-runs every migration first |
| App with no framework seed command | scripts/seed.ts plus a seed entry in package.json | npm run seed | The script’s own reset step |
The last row is my working rule for an app generated without a framework seed command. Prisma is mid-change: its docs now default to Prisma ORM 8, a release candidate, and the page for people coming from version 7 says “there is no seed command”; you write a script and run it like any TypeScript file. On version 7, Prisma’s seeding guide keeps prisma db seed. Drizzle’s package, drizzle-seed, generates “deterministic, yet realistic, fake data” from a seedable pseudorandom number generator. Rails documents both of its commands in the Rails migrations guide’s seeding section, and Django’s loader is in Django’s fixtures docs.
For a hosted staging project, the Supabase staging retrofit guide covers the remote reset and its warning that the command erases the linked project’s data.
Making fake data realistic: faker libraries and the cases that matter
Fake data is realistic when its shape matches production, not only its names: a few large tenants and many small ones, months of activity, and edge cases such as a tenant at a plan limit, look-alike data in a second tenant, and a timestamp on a daylight saving boundary.
Faker libraries do the generating. Faker for JavaScript describes its job as “Generate massive amounts of fake (but realistic) data for testing and development”, with generators for names, addresses and dates. Python’s Faker “generates fake data for you”, and both take a seed, as step 4 set out.
Realism is mostly shape, and the shape in that opening line is my working rule, not a measured standard: tenant sizes skewed the way production’s are, activity spread over many months, and a believable mix of active and churned accounts. Those are the shapes that make reports, plan limits and slow queries behave the way they will in production.
Every generated email uses a reserved domain. RFC 2606 lists example.com, example.net and example.org as reserved second level domain names that can be used as examples. Faker’s API notes that its ordinary emails “could coincidentally be real email addresses” and points to exampleEmail() for that case.
Payment objects are real sandbox objects created through the provider’s API with test keys, never invented ids. Stripe’s API keys docs say each mode has its own keys and “objects in one mode aren’t accessible to the other”. Keeping those modes apart across environments is part of the payment go-live checklist.
The core awkward shapes, from the empty state to a name with an apostrophe, are listed in the one-database staging article. These eight add to them:
- A tenant sitting exactly at its plan limit.
- A second tenant with look-alike data (the same display names) for isolation tests.
- A very long name, to test truncation and layout.
- Non-Latin text and emoji in names and free-text fields.
- An optional field left empty on every row of a table.
- Timestamps on a month end and on both daylight saving boundaries.
- A soft-deleted parent that still has live children.
- A record at the maximum size the app allows.
The month-end and daylight saving rows follow the time zone handling checklist. The soft-deleted parent needs a clear answer to what soft delete is in your app before you can say what its children should show. A random data generator such as Mockaroo is fine for a one-off CSV; the script stays the source of truth.
If someone says “just copy production”: mask, subset, or do not
The default answer is no, and the seed script is the reason no copy is needed; the case against copying is made in the two staging articles already named, and it is not repeated here. If the ask is really for production’s structure, take the schema without rows and seed it; the Supabase staging article lists the commands. If a production-shaped data set is truly required, for a bug that only shows with real data or a migration rehearsal, my working rule is a masked subset produced by a documented step, run by someone with production access, recorded, and deleted afterwards. The masking happens before the export leaves production, the rule the one-database article sets. One open-source option is PostgreSQL Anonymizer, “an extension to mask or replace personally identifiable information (PII) or commercially sensitive data”.
How to verify it, and verify staging does not touch live data
A staging seed is verified in six checks: it runs clean on an empty database, a second run changes nothing, every fixture role completes the core flows, it refuses a production connection string, staging’s database, payment keys, email and webhooks are all its own, and the seed runs in CI on every pull request.
Deliverable 4.10 of the Production Hardening Sprint is verified this way: “Run the seed process on a clean staging database and exercise core flows.” The checks below turn that line into steps you can run on your own app.
- 01 Create a clean staging database, run the migrations, then run the seed. Pass: exit code 0 and row counts that match the printed summary.
- 02 Run the seed a second time. Pass: identical counts and no duplicate fixture accounts.
- 03 Log in as each fixture role and walk the core flows: sign up, the main task, a sandbox checkout, and an email that lands in a test inbox. Pass: every flow completes for every role.
- 04 Point the seed at the production host on purpose, from a fresh shell, with no working password in the connection string. Pass: the guard's refusal message and a non-zero exit before any connection opens.
- 05 Check isolation from the staging side: a database host that differs from production's, a payment secret key that starts
sk_test_(orrk_test_for a restricted key), the email provider in its test or sandbox setting where it has one, and webhook endpoints that belong to staging. Pass: one staging action, such as creating an order, leaves no trace in production. - 06 Confirm the seed job runs in CI on every pull request. Pass: the latest pull request shows it green.
Check 4 leaves the password out, so even if the guard is broken, the test holds no credentials that could change production data. Check 5 is how you verify staging does not touch live data; the full list of what staging must isolate, and why, lives in the one-database article’s “What staging must isolate” section, and this check only confirms it.
If the seed fails with cannot execute INSERT in a read-only transaction, one hosted cause is size. Supabase’s database size docs say “Free Plan projects enter read-only mode when your database size exceeds 500 MB”, and give a paid-plan example: “uploading more than 1.5x the current size of your database storage will put your database into read-only mode”. Size is not the only cause of cannot execute INSERT in a read-only transaction, and the others need different fixes.
Keep the evidence: the output of both seed runs, the guard’s refusal, and the isolation checks written up as a dated table.
Where the sprint does this
In the Production Hardening Sprint, deliverable 4.10 is to “Provide a repeatable seed script with realistic non-sensitive test records.” The environments it runs in are deliverable 7.1, which provides “development, staging, and production with separate databases, credentials, and configuration” , and payment separation is deliverable 5.5: “Verify test and live credentials are isolated by environment.” The production readiness report, deliverable 13.1, is verified this way: “Account for all 123 IDs; keep failures visible until resolved and explain genuine non-applicable items.” Hosting, paid tools and API usage remain in your accounts, and any required third-party costs are explained before they are enabled. Each deliverable, with how it is verified, is in the published scope.
Common questions about seed data and test data
What are the three types of test data?
One common split is valid, invalid and boundary data; that is my reading of how the term is used, not a formal standard. In a seed script, fixture accounts and generated rows are the valid data, and the hand-written edge cases hold the boundary and invalid ones, such as a tenant exactly at its plan limit or a field left empty.
How can I test a database?
Test a database by checking four things: constraints reject bad rows, migrations apply cleanly to a fresh copy, the seed loads twice with the same result, and your core queries return the rows you expect. The schema design itself is a separate review, the job of a database review checklist.
How do I empty a database?
Empty a staging or local database, never production, with TRUNCATE on the seeded tables: CASCADE also empties every table that references them, and RESTART IDENTITY restarts their sequences. A stack’s reset command does the same job, such as supabase db reset on a local Supabase project, which runs the seed file again afterwards. Check which host you are connected to before running either; inside the seed script, the guard makes that check for you.
PostgreSQL’s reference adds that TRUNCATE takes an ACCESS EXCLUSIVE lock on each table, which blocks all other concurrent operations on it.
Which tool is used for database testing?
No single tool: in my reading, database testing uses the migration tool, the seed script and the test runner your stack already has. The migration tool proves the schema applies, the seed proves the data loads and repeats, and the test runner checks the queries and rules against those rows.
What does seed mean in AI?
In AI, as in a faker library, a seed is the starting value for a pseudorandom generator, set so that a run can be repeated. Drizzle’s docs describe the same mechanism: the generator is “based on an initial value called a seed”, so “you can control its randomness”. It has nothing to do with seed data in a database.
The checks in this guide show you where the app is open. The sprint below closes those gaps, tests the result and writes the evidence down.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase