Run the migration as its own release step, after the build and before the new code takes traffic, and let a failed migration cancel the deploy. That is how to run database migrations on deploy on any host. While the new version rolls out, two versions of the app share one schema, so every change has to work with both.
What a migration-aware deploy is
A migration-aware deploy is a release that runs its schema change exactly once, in a fixed place between the build and the traffic switch, and stops if the change fails. The pipeline or the host’s release hook runs it, never the app at boot.
The step belongs to the release path, next to the build and the smoke check; that whole path, from one repository to a team, is DevOps for startups: the release path. I keep it out of the app’s start command and off every laptop pointed at production: one place, owned by the pipeline, with its own log. For the word itself in plain language, start with what deploying means.
Three properties carry the rest of this page. Ordering: the schema change goes first when it adds something and last when it removes something the old code reads. Compatibility: the old code and the new code both work against the schema in between. Failure handling: a failed migration stops the release and leaves the old version serving.
All of this assumes your migrations already exist as ordered files in the repository, as in the database migration checklist.
What goes wrong when migrations and code deploy out of order
Here are five ways the order goes wrong, each with the symptom you see first. The “first check” column is my own order of what to look at; what to do after that first look lives in the rollback guide linked further down.
| Symptom | What happened | First check (mine) |
|---|---|---|
New code is live and requests fail with column "x" does not exist (SQLSTATE 42703) or relation "x" does not exist (42P01) | The code shipped and its migration did not reach this database | The production database’s migration history, then the deploy log for the migration step’s exit code |
| The old version, still serving, starts failing on a column the migration dropped or renamed | The migration ran first and removed something the old code still reads | Which version is serving right now, and what the migration removed |
| Two instances both ran the migration at boot; one crashed or sat waiting on a lock | The start command migrates, and more than one instance started at once | The instance logs from the deploy, looking for two migration runs |
| The migration stopped partway through its statements | One statement failed after earlier ones had already run | The tool’s history table and the actual columns of the tables it touched |
| The deploy hangs on the migration step and never finishes | A statement is waiting on a lock held by a long query or open transaction | The step’s live log, then whichever session holds the lock on that table |
The error text comes from PostgreSQL’s source, which prints a missing column as column "x" does not exist, or as column users.x does not exist when the query names the table, and a missing table as relation "x" does not exist. Their SQLSTATE codes are 42703, undefined_column, and 42P01, undefined_table, so a log search on the code finds them whatever the column is called.
Buildkite published one of these failures on its status page. On May 1, 2026 a migration that renamed a column on its users table ran before the new application code was deployed, so the code already running still expected the original column name and queries that loaded a user record began failing across the product until the column was restored, 19 minutes of customer-facing impact. Buildkite’s postmortem, posted May 8, 2026, adds that an LLM-assisted workflow had generated the migration and that it shipped in the same pull request as the code, so the tests only ever ran the new code against the new schema. The schema and the code are two releases wearing one version number, and someone has to order them on purpose.
In my June and July 2026 audits, at least 17 of the 21 third-party apps had no deploy gate, and so did all 5 founder apps: every push ships straight to production with nothing checking it first. Those 21 are 11 public third-party apps I audited exhaustively across all 12 pillars and 10 held-out third-party apps audited blind, and the 5 are my own production apps; I chose all of them, so the counts hold for these apps and are no rate for AI-built apps in general. In the same audits the Deployment & Operations pillar averages 37.0 out of 100 across the 21 third-party apps scored on it.
App crashed after deploy, missing column
An app that crashes after a deploy with a missing column error shipped its code without its migration. PostgreSQL reports it as SQLSTATE 42703, undefined_column. Check the migration history on the production database first, then the deploy log for the migration step’s exit code.
The code version and the schema version disagree, and my list of causes has three entries: the migration was never run against production, it ran against a different database (a staging URL left in the production job’s secrets), or it failed and the code went out anyway. Heroku documents one path to that third case: when a release phase fails after a successful build, a later release triggered by an add-on’s config var change uses the built slug and does not run the release phase, so the code can deploy without the migration it depends on. The codes are listed in PostgreSQL’s error codes appendix if your log shows a different one.
Look at the production migration history before anything else; that check has its own step in the full write-up of the migration that ran in dev and never ran in production. Then open the deploy log and find the migration step’s exit code. If the pending migration only adds things, the fast path is to run it now. If it removes or changes something, redeploy the code that matches the current schema instead, the case the rollback guide’s recovery table calls “Only the deployed code is incompatible”.
The migration ran before the code deployed
A migration that ran before the code deployed is a problem only when it removed or tightened something. Additive changes are safe to run first. A dropped or renamed column breaks the old version for as long as it keeps serving, so removals ship one deploy after the code stops using them.
Running first is the right order for an added column and exactly the failure for a dropped column, a rename, a new NOT NULL column with no default, or a tightened type. The old code keeps serving until the rollout finishes, or indefinitely if the deploy then fails. On Fly.io’s default rolling strategy, each running Machine is taken down and replaced one by one. I take that to mean some requests reach the old version and some the new one until the last Machine is replaced, and that is the window where the old code meets the new schema.
The first check is which version is serving and what the migration removed. Rolling the code back will not put the column back, which is the point of does a rollback undo a database migration. The rule that prevents this sits in the next section: never remove something in the same deploy that stops using it.
How to run database migrations on deploy: one step, in order, that can stop the release
Database migrations on deploy run as a separate step: after the build, before the new code takes traffic, once, with its exit code checked. A non-zero exit stops the release and leaves the old version serving. The app never migrates itself at boot when two instances can start together.
A CI/CD database migration step has the same shape on every host:
- Build the artifact the new code will run from.
- Run the migration once, as its own step, with production database credentials scoped to that step only.
- Check the step’s exit code; anything other than zero ends the release here.
- Switch traffic to the new code.
- Run the smoke check against the live version.
For me, the app’s start command is the wrong place to migrate when more than one instance can start at the same time, because each instance would try. Two migration runs also must not overlap, and some hosts and tools document how they handle it. Fly.io runs release_command once per deploy attempt. On PostgreSQL, Prisma’s docs say that when two db migrate runs start against the same database, one of them waits for the other. Prisma’s PgBouncer guide, kept in its ORM 7 docs, puts CLI commands on a direct connection because “Prisma Migrate uses database transactions” and its schema engine uses a single connection that does not support pooling with PgBouncer. PgBouncer’s own feature map marks session-level advisory locks “Never” under transaction pooling. So on a host that hands you a pooled connection string, I would give the migration step a direct connection instead; Supabase’s docs list the direct connection for migrations and the session-mode pooler string for an IPv4-only network. EF Core 9 and later use migration locking. Whatever the tool, keep the step’s log with the release it belongs to.
The host behavior below comes from each host’s own docs, read on October 4, 2026; it is what those pages state, not a run of my own. Flyway and Liquibase are two more tools that do the same job as the ones named here.
Where the migration step goes on GitHub Actions, Fly.io, Heroku, Render, Railway and Vercel
| Host | Where the step lives | What a failure does | The catch |
|---|---|---|---|
| GitHub Actions | A migrate job that the deploy job needs, bound to environment: production so only that job gets the environment’s secrets | The jobs that need it are skipped, unless they use a conditional expression that makes them continue | Environments and environment secrets in private or internal repositories need GitHub Pro, GitHub Team or GitHub Enterprise; required reviewers on Free, Pro or Team are only for public repositories |
| Fly.io | release_command in [deploy], run in a temporary Machine from the newly built image before any Machines are created or updated | A non-zero exit status stops the deployment | The temporary Machine has no volumes attached |
| Heroku | A release process type in the Procfile, run in a one-off dyno whenever a new release is created, unless an add-on’s config var change caused it; app dynos don’t boot for the new release until it finishes successfully | The new release is not deployed and the current release is unaffected | A config var change also creates a release (unless the var belongs to an add-on), and the var stays changed even if the release command fails |
| Render | The pre-deploy command, run after the build finishes and before that build is deployed, on a separate instance | The deploy fails and the service keeps running its most recent successful deploy (if any) | Available for paid web services, private services and background workers |
| Railway | The pre-deploy command, run between building and deploying, in a separate container on the private network with the app’s environment variables | It is not retried and the deployment does not proceed | Volumes are not mounted |
| Vercel | A release-phase migration hook is not stated in Vercel’s docs; a deployment is the result of a successful build | Not stated in Vercel’s docs | Each merge to the production branch creates a production deployment, so a CI migrate job alone stops nothing; Deployment Checks hold the production deployment until required checks pass, and Force Promote bypasses them |
| Supabase (database only) | supabase db push from a GitHub Actions workflow, which Supabase recommends for production over deploying from your local machine | Not stated in Supabase’s docs | The app’s host deploys the code separately, so I’d have the push finish before the code takes traffic |
| A single server or VM | The deploy script, after the new code is on disk and before the process restart | I’d make the script stop on a non-zero exit | Not a vendor fact: the script is yours to write |
| Kubernetes with Helm | A Job annotated as a pre-upgrade hook, run after templates are rendered and before any resources are updated | Helm waits for the Job to complete; if the hook fails, the release fails | No walkthrough here: one row |
The sources, in table order: GitHub’s deployment environments and GitHub’s workflow syntax for needs, Fly.io’s release_command, Heroku’s release phase, Render’s pre-deploy command, Railway’s pre-deploy command, Vercel’s deployments docs with Vercel’s Deployment Checks, Supabase’s guide to managing environments and Helm’s chart hooks. On a GitHub Free private repository, where environments are not available, drop the environment line and let the migrate job read a repository secret instead.
On GitHub Actions, the deploy job waits for the migrate job and is skipped when it fails. That holds the release only when this deploy job is what deploys; a host that deploys every push on its own is, as far as I can tell, not held by it, which is the reason for the Vercel row.
jobs:
migrate:
runs-on: ubuntu-latest
environment: production # this job alone gets the production secrets
steps:
- uses: actions/checkout@v7
- run: <your migration tool's apply command>
env:
DATABASE_URL: ${{ secrets.DATABASE_URL }}
deploy:
needs: migrate # skipped if migrate fails
runs-on: ubuntu-latest
steps:
- run: <your host's deploy command>
On Fly.io the same step is one line in fly.toml, and the timeout line is optional:
[deploy]
release_command = "<your migration tool's apply command>"
release_command_timeout = "10m" # default is 5 minutes
On Vercel, make the migrate job a required Deployment Check, imported from GitHub Actions, so the production deployment waits for it; I would never run migrations from the build command. If the database is Supabase and the code is on Vercel, I’d use the job that runs supabase db push as that required check, so the push finishes before the code goes live. The full staging-then-production version, one GitHub Actions workflow per environment, is the Supabase staging environment release workflow.
Two tool notes. Prisma ORM 8’s npx prisma db migrate runs your migrations against a database and replaces Prisma ORM 7’s prisma migrate deploy; it works on PostgreSQL and MongoDB, with SQLite still experimental, and in CI and production you never plan a migration (per Prisma’s guide to applying a migration). EF Core says to use a migration bundle for automated deployment and lists runtime migration for “Applications that accept startup migration tradeoffs” in EF Core’s guide to applying migrations. A preview build that runs migrations against the production database is a separate hazard that follows from what preview deployments are.
What is a backward compatible migration?
A backward compatible migration is a schema change the currently running code can ignore. Adding a nullable column, a table or an index qualifies. Renames, drops and type changes do not, and ship over several deploys: add the new shape, move the code, remove the old.
| Schema change | Safe in one deploy? | How to ship it |
|---|---|---|
| Add a nullable column | Yes | One deploy; no table rewrite |
| Add a table | Yes | One deploy |
| Add an index | Yes | Build it concurrently on PostgreSQL so inserts, updates and deletes are not locked out; a failed concurrent build leaves an invalid index that still costs update overhead, so drop it and try again |
| Add a NOT NULL column | Only with a non-volatile default | A non-volatile default is stored without rewriting the table; otherwise add it nullable, backfill, then tighten in a later deploy |
| Rename a column | No | Several deploys, mapped in the next section |
| Drop a column | No | Stop reading and writing it first, then drop it in a later deploy |
| Change a column’s type | No | Add a new column, backfill, switch the code, drop the old one later |
| Add a foreign key or check constraint | In steps | Add it NOT VALID, which skips the scan of existing rows, then VALIDATE CONSTRAINT, which takes only a SHARE UPDATE EXCLUSIVE lock on the table (plus ROW SHARE on the referenced table for a foreign key) |
The “safe in one deploy?” column is my rule; PostgreSQL’s behavior in the last column comes from PostgreSQL’s CREATE INDEX and PostgreSQL’s ALTER TABLE pages. Safe here means both code versions keep working, not lock-free: ALTER TABLE states “An ACCESS EXCLUSIVE lock is acquired unless explicitly noted”, so I’d assume even the quick forms queue behind a long-running query. That wait is what the lock timeout in the fails-halfway section is for. Backward compatible database changes are the rows marked yes, and every other row becomes one once it is split into deploys that each work with the code around them. Before you put a concurrent index build in a migration file, read the rollback guide’s section on how PostgreSQL transaction rollback differs from migration rollback.
Zero downtime deployment with database changes: which deploy runs each migration
Zero downtime deployment with database changes means no release step removes anything the code still serving uses. A column rename takes 4 deploys: the first adds the new column, a separate job backfills it, the next two move the code over, and the fourth drops the old column.
The pattern behind this, its full sequence and why it avoids emergency rollbacks, is in expand-contract in the database migration rollback guide, and I won’t repeat it. What this page adds is which release runs what, for a rename of users.name to users.full_name:
| Deploy | Migration the release step runs | What the code does | Safe to roll the code back to |
|---|---|---|---|
| 1 | Add full_name, nullable | Writes both columns, reads name | The previous release |
| Between 1 and 2 | None (a one-off job copies name into full_name in batches) | Unchanged | Not a code change |
| 2 | None | Reads full_name, still writes both | Deploy 1 |
| 3 | If name is NOT NULL, ALTER COLUMN name DROP NOT NULL first | Stops writing name | Deploy 2 |
| 4 | Drop name | Nothing reads or writes name | Deploy 3 |
Deploy 3 needs the DROP NOT NULL because that form changes whether a column is marked to allow or reject null values, and my inference is that a row the new code inserts without name would otherwise be rejected. The “safe to roll the code back to” column is mine. I leave the backfill out of the release step because the release step holds the deploy, and Fly.io, Heroku and Render cap that step while Railway by default waits on it; the limits are in the next section.
-- Deploy 1, release step
ALTER TABLE users ADD COLUMN full_name text;
-- Between deploys, one-off job, repeated per batch of ids
UPDATE users SET full_name = name WHERE id BETWEEN $1 AND $2 AND full_name IS NULL;
-- Deploy 3, release step, only if name is NOT NULL
ALTER TABLE users ALTER COLUMN name DROP NOT NULL;
-- Deploy 4, release step
ALTER TABLE users DROP COLUMN name;
Blue-green and canary deployments don’t change this; Fly.io lists both canary and bluegreen strategies, and in each the old and new versions use the same database during the switch, so in my view the same schema rule applies. For a small app, zero downtime honestly means no failed requests caused by the schema. It does not mean the deploy carries no risk.
When a migration fails halfway
A migration that fails halfway should stop the deploy while the old version keeps serving, and the release steps on Fly.io, Heroku, Render and Railway all stop on a failed command. The harder case is a migration waiting on a lock. A lock timeout on the migration session turns that wait into a failure.
The deploy half of a failure is the “what a failure does” column of the host table above. What the database kept after a failed migration depends on the transaction boundary, which the rollback guide’s section on PostgreSQL transaction rollback versus migration rollback covers.
A migration that waits instead of failing runs into each host’s release-step limit:
- Fly.io: 5 minutes by default before the command is killed, changed with
release_command_timeout. - Heroku: a 1-hour timeout that cannot be extended.
- Render: 30 minutes for the pre-deploy command.
- Railway: no time limit by default, so a command waiting on a lock holds the deployment in progress rather than failing it; a Pre-deploy Timeout of 1 to 3,600 seconds can cap it.
The fix sits in the migration’s own session: set lock_timeout, which aborts any statement that waits longer than the limit while trying to acquire a lock, and in practice that makes the migration step fail instead of hang. The limit applies to each lock acquisition attempt separately, a value without units is taken as milliseconds, and zero, the default, disables it; setting it in postgresql.conf is not recommended because it would affect all sessions. The wording is in PostgreSQL’s lock_timeout setting. I start at a few seconds on the migration session and raise it only for a statement you know needs longer.
What to do next, stopping retries, reading the history table and choosing between rollback and fixing forward, is the rollback guide’s section on what to do immediately after a database migration fails. The release side of that recovery belongs in a rollback plan you have rehearsed.
How to verify it: test a schema change through the pipeline
A migration-aware deploy is verified with two test changes. The first, a harmless new column, must run once before the traffic switch and appear in the history table. The second, broken on purpose, must stop the deploy and leave the old version serving an unchanged schema.
Run this on a staging database first; the article on rollbacks and staging, linked in the section on a migration that ran before the code, explains why. Each check is something a correct pipeline passes and a broken one fails, and each leaves evidence you can keep.
- 01 Add a harmless migration, such as a nullable column on a low-traffic table, and merge it. Evidence: the commit id.
- 02 Open the host's deploy log page and confirm the migration step ran once, before the traffic switch, with exit code zero. Evidence: the log, with its date.
- 03 Open the table view in your database dashboard and confirm the migration tool's history table lists the migration and the new column exists. Evidence: a screenshot of both.
- 04 Add a second migration with one statement that references a table that does not exist, merge it, and confirm the deploy stopped, the old version still answers requests and the schema did not change. Evidence: the failed log and one response from the old version. Then delete the broken migration before the next deploy.
- 05 Restart the app, or scale it to two instances where your plan allows it, and read the instance logs from the restart. Evidence: no log line shows a migration run. A host with no start command skips this check.
- 06 Time the migration step and note it against your deploy budget and the host's release-step limit. Evidence: the step's duration from the log.
Check five needs the instance logs because the history table cannot show a boot-time run: a migration tool runs only pending migrations, so a second run should find nothing to do and leave no new row. Scaling has plan limits: Heroku’s docs say you can’t run more than one Eco or one Basic dyno per process type, and Render’s docs mark its free instance type “1 (scaling unavailable)”, while Vercel has no start command to check. For check six, a slow migration step spends deploy budget, and keeping that budget small is part of how to speed up CI build times. The checks that run before merge, a migration lint among them, are CI/CD best practices, and a migration file that skips review is a job for GitHub branch protection.
In the Production Hardening Sprint, deliverable 7.5 is verified this way: deploy a representative schema change and verify sequencing and failure handling.
Where the sprint does this
Deliverable 7.5 of the Production Hardening Sprint integrates database migrations into the deployment process with ordering and compatibility checks; 4.6 moves schema changes into ordered migrations committed to the repository; 7.6 documents and rehearses release rollback, including how database changes are handled safely; and the production readiness report, 13.1, delivers the result for every scope item, the work completed and its verification evidence. All four are among the 123 deliverables in the sprint.
Common questions about migrations and deploys
What does “db migration” mean?
A db migration is a versioned, ordered change to a database’s schema, kept as a file and applied once to each database, as opposed to moving data from one system to another. Each file adds, changes or removes tables, columns, indexes or constraints, and the migration tool records which files a database has already applied.
What does “code migration” mean?
Code migration means moving an app’s code to a new language, framework, library version or host. It changes the code, not the structure of the database, and it is a separate job from the schema migrations a deploy runs.
What’s the difference between forward and backward compatibility?
Backward compatibility means the new schema still works with the old code; forward compatibility means the old schema still works with the new code. A deploy that adds a column first needs the first, and a deploy that ships code before removing a column needs the second, because the new code runs for a while against a schema that still has it.
How do you handle schema changes?
Handle schema changes as migration files that are reviewed, run once by the release step, and split so additions ship first and removals ship in a later deploy. Each deploy should work with the schema before and after it.
If you have a working app built with these tools and need it ready for real customers, this is what we do.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase