Run the migration as its own release step, after the build and before the new code takes traffic, and let a failed migration cancel the deploy. That is how to run database migrations on deploy on any host. While the new version rolls out, two versions of the app share one schema, so every change has to work with both.

What a migration-aware deploy is

A migration-aware deploy is a release that runs its schema change exactly once, in a fixed place between the build and the traffic switch, and stops if the change fails. The pipeline or the host’s release hook runs it, never the app at boot.

The step belongs to the release path, next to the build and the smoke check; that whole path, from one repository to a team, is DevOps for startups: the release path. I keep it out of the app’s start command and off every laptop pointed at production: one place, owned by the pipeline, with its own log. For the word itself in plain language, start with what deploying means.

Three properties carry the rest of this page. Ordering: the schema change goes first when it adds something and last when it removes something the old code reads. Compatibility: the old code and the new code both work against the schema in between. Failure handling: a failed migration stops the release and leaves the old version serving.

All of this assumes your migrations already exist as ordered files in the repository, as in the database migration checklist.

What goes wrong when migrations and code deploy out of order

Here are five ways the order goes wrong, each with the symptom you see first. The “first check” column is my own order of what to look at; what to do after that first look lives in the rollback guide linked further down.

SymptomWhat happenedFirst check (mine)
New code is live and requests fail with column "x" does not exist (SQLSTATE 42703) or relation "x" does not exist (42P01)The code shipped and its migration did not reach this databaseThe production database’s migration history, then the deploy log for the migration step’s exit code
The old version, still serving, starts failing on a column the migration dropped or renamedThe migration ran first and removed something the old code still readsWhich version is serving right now, and what the migration removed
Two instances both ran the migration at boot; one crashed or sat waiting on a lockThe start command migrates, and more than one instance started at onceThe instance logs from the deploy, looking for two migration runs
The migration stopped partway through its statementsOne statement failed after earlier ones had already runThe tool’s history table and the actual columns of the tables it touched
The deploy hangs on the migration step and never finishesA statement is waiting on a lock held by a long query or open transactionThe step’s live log, then whichever session holds the lock on that table

The error text comes from PostgreSQL’s source, which prints a missing column as column "x" does not exist, or as column users.x does not exist when the query names the table, and a missing table as relation "x" does not exist. Their SQLSTATE codes are 42703, undefined_column, and 42P01, undefined_table, so a log search on the code finds them whatever the column is called.

Buildkite published one of these failures on its status page. On May 1, 2026 a migration that renamed a column on its users table ran before the new application code was deployed, so the code already running still expected the original column name and queries that loaded a user record began failing across the product until the column was restored, 19 minutes of customer-facing impact. Buildkite’s postmortem, posted May 8, 2026, adds that an LLM-assisted workflow had generated the migration and that it shipped in the same pull request as the code, so the tests only ever ran the new code against the new schema. The schema and the code are two releases wearing one version number, and someone has to order them on purpose.

In my June and July 2026 audits, at least 17 of the 21 third-party apps had no deploy gate, and so did all 5 founder apps: every push ships straight to production with nothing checking it first. Those 21 are 11 public third-party apps I audited exhaustively across all 12 pillars and 10 held-out third-party apps audited blind, and the 5 are my own production apps; I chose all of them, so the counts hold for these apps and are no rate for AI-built apps in general. In the same audits the Deployment & Operations pillar averages 37.0 out of 100 across the 21 third-party apps scored on it.

App crashed after deploy, missing column

An app that crashes after a deploy with a missing column error shipped its code without its migration. PostgreSQL reports it as SQLSTATE 42703, undefined_column. Check the migration history on the production database first, then the deploy log for the migration step’s exit code.

The code version and the schema version disagree, and my list of causes has three entries: the migration was never run against production, it ran against a different database (a staging URL left in the production job’s secrets), or it failed and the code went out anyway. Heroku documents one path to that third case: when a release phase fails after a successful build, a later release triggered by an add-on’s config var change uses the built slug and does not run the release phase, so the code can deploy without the migration it depends on. The codes are listed in PostgreSQL’s error codes appendix if your log shows a different one.

Look at the production migration history before anything else; that check has its own step in the full write-up of the migration that ran in dev and never ran in production. Then open the deploy log and find the migration step’s exit code. If the pending migration only adds things, the fast path is to run it now. If it removes or changes something, redeploy the code that matches the current schema instead, the case the rollback guide’s recovery table calls “Only the deployed code is incompatible”.

The migration ran before the code deployed

A migration that ran before the code deployed is a problem only when it removed or tightened something. Additive changes are safe to run first. A dropped or renamed column breaks the old version for as long as it keeps serving, so removals ship one deploy after the code stops using them.

Running first is the right order for an added column and exactly the failure for a dropped column, a rename, a new NOT NULL column with no default, or a tightened type. The old code keeps serving until the rollout finishes, or indefinitely if the deploy then fails. On Fly.io’s default rolling strategy, each running Machine is taken down and replaced one by one. I take that to mean some requests reach the old version and some the new one until the last Machine is replaced, and that is the window where the old code meets the new schema.

The first check is which version is serving and what the migration removed. Rolling the code back will not put the column back, which is the point of does a rollback undo a database migration. The rule that prevents this sits in the next section: never remove something in the same deploy that stops using it.

How to run database migrations on deploy: one step, in order, that can stop the release

Database migrations on deploy run as a separate step: after the build, before the new code takes traffic, once, with its exit code checked. A non-zero exit stops the release and leaves the old version serving. The app never migrates itself at boot when two instances can start together.

A CI/CD database migration step has the same shape on every host:

  1. Build the artifact the new code will run from.
  2. Run the migration once, as its own step, with production database credentials scoped to that step only.
  3. Check the step’s exit code; anything other than zero ends the release here.
  4. Switch traffic to the new code.
  5. Run the smoke check against the live version.

For me, the app’s start command is the wrong place to migrate when more than one instance can start at the same time, because each instance would try. Two migration runs also must not overlap, and some hosts and tools document how they handle it. Fly.io runs release_command once per deploy attempt. On PostgreSQL, Prisma’s docs say that when two db migrate runs start against the same database, one of them waits for the other. Prisma’s PgBouncer guide, kept in its ORM 7 docs, puts CLI commands on a direct connection because “Prisma Migrate uses database transactions” and its schema engine uses a single connection that does not support pooling with PgBouncer. PgBouncer’s own feature map marks session-level advisory locks “Never” under transaction pooling. So on a host that hands you a pooled connection string, I would give the migration step a direct connection instead; Supabase’s docs list the direct connection for migrations and the session-mode pooler string for an IPv4-only network. EF Core 9 and later use migration locking. Whatever the tool, keep the step’s log with the release it belongs to.

The host behavior below comes from each host’s own docs, read on October 4, 2026; it is what those pages state, not a run of my own. Flyway and Liquibase are two more tools that do the same job as the ones named here.

Where the migration step goes on GitHub Actions, Fly.io, Heroku, Render, Railway and Vercel

HostWhere the step livesWhat a failure doesThe catch
GitHub ActionsA migrate job that the deploy job needs, bound to environment: production so only that job gets the environment’s secretsThe jobs that need it are skipped, unless they use a conditional expression that makes them continueEnvironments and environment secrets in private or internal repositories need GitHub Pro, GitHub Team or GitHub Enterprise; required reviewers on Free, Pro or Team are only for public repositories
Fly.iorelease_command in [deploy], run in a temporary Machine from the newly built image before any Machines are created or updatedA non-zero exit status stops the deploymentThe temporary Machine has no volumes attached
HerokuA release process type in the Procfile, run in a one-off dyno whenever a new release is created, unless an add-on’s config var change caused it; app dynos don’t boot for the new release until it finishes successfullyThe new release is not deployed and the current release is unaffectedA config var change also creates a release (unless the var belongs to an add-on), and the var stays changed even if the release command fails
RenderThe pre-deploy command, run after the build finishes and before that build is deployed, on a separate instanceThe deploy fails and the service keeps running its most recent successful deploy (if any)Available for paid web services, private services and background workers
RailwayThe pre-deploy command, run between building and deploying, in a separate container on the private network with the app’s environment variablesIt is not retried and the deployment does not proceedVolumes are not mounted
VercelA release-phase migration hook is not stated in Vercel’s docs; a deployment is the result of a successful buildNot stated in Vercel’s docsEach merge to the production branch creates a production deployment, so a CI migrate job alone stops nothing; Deployment Checks hold the production deployment until required checks pass, and Force Promote bypasses them
Supabase (database only)supabase db push from a GitHub Actions workflow, which Supabase recommends for production over deploying from your local machineNot stated in Supabase’s docsThe app’s host deploys the code separately, so I’d have the push finish before the code takes traffic
A single server or VMThe deploy script, after the new code is on disk and before the process restartI’d make the script stop on a non-zero exitNot a vendor fact: the script is yours to write
Kubernetes with HelmA Job annotated as a pre-upgrade hook, run after templates are rendered and before any resources are updatedHelm waits for the Job to complete; if the hook fails, the release failsNo walkthrough here: one row

The sources, in table order: GitHub’s deployment environments and GitHub’s workflow syntax for needs, Fly.io’s release_command, Heroku’s release phase, Render’s pre-deploy command, Railway’s pre-deploy command, Vercel’s deployments docs with Vercel’s Deployment Checks, Supabase’s guide to managing environments and Helm’s chart hooks. On a GitHub Free private repository, where environments are not available, drop the environment line and let the migrate job read a repository secret instead.

On GitHub Actions, the deploy job waits for the migrate job and is skipped when it fails. That holds the release only when this deploy job is what deploys; a host that deploys every push on its own is, as far as I can tell, not held by it, which is the reason for the Vercel row.

jobs:
  migrate:
    runs-on: ubuntu-latest
    environment: production        # this job alone gets the production secrets
    steps:
      - uses: actions/checkout@v7
      - run: <your migration tool's apply command>
        env:
          DATABASE_URL: ${{ secrets.DATABASE_URL }}
  deploy:
    needs: migrate                 # skipped if migrate fails
    runs-on: ubuntu-latest
    steps:
      - run: <your host's deploy command>

On Fly.io the same step is one line in fly.toml, and the timeout line is optional:

[deploy]
  release_command = "<your migration tool's apply command>"
  release_command_timeout = "10m"   # default is 5 minutes

On Vercel, make the migrate job a required Deployment Check, imported from GitHub Actions, so the production deployment waits for it; I would never run migrations from the build command. If the database is Supabase and the code is on Vercel, I’d use the job that runs supabase db push as that required check, so the push finishes before the code goes live. The full staging-then-production version, one GitHub Actions workflow per environment, is the Supabase staging environment release workflow.

Two tool notes. Prisma ORM 8’s npx prisma db migrate runs your migrations against a database and replaces Prisma ORM 7’s prisma migrate deploy; it works on PostgreSQL and MongoDB, with SQLite still experimental, and in CI and production you never plan a migration (per Prisma’s guide to applying a migration). EF Core says to use a migration bundle for automated deployment and lists runtime migration for “Applications that accept startup migration tradeoffs” in EF Core’s guide to applying migrations. A preview build that runs migrations against the production database is a separate hazard that follows from what preview deployments are.

What is a backward compatible migration?

A backward compatible migration is a schema change the currently running code can ignore. Adding a nullable column, a table or an index qualifies. Renames, drops and type changes do not, and ship over several deploys: add the new shape, move the code, remove the old.

Schema changeSafe in one deploy?How to ship it
Add a nullable columnYesOne deploy; no table rewrite
Add a tableYesOne deploy
Add an indexYesBuild it concurrently on PostgreSQL so inserts, updates and deletes are not locked out; a failed concurrent build leaves an invalid index that still costs update overhead, so drop it and try again
Add a NOT NULL columnOnly with a non-volatile defaultA non-volatile default is stored without rewriting the table; otherwise add it nullable, backfill, then tighten in a later deploy
Rename a columnNoSeveral deploys, mapped in the next section
Drop a columnNoStop reading and writing it first, then drop it in a later deploy
Change a column’s typeNoAdd a new column, backfill, switch the code, drop the old one later
Add a foreign key or check constraintIn stepsAdd it NOT VALID, which skips the scan of existing rows, then VALIDATE CONSTRAINT, which takes only a SHARE UPDATE EXCLUSIVE lock on the table (plus ROW SHARE on the referenced table for a foreign key)

The “safe in one deploy?” column is my rule; PostgreSQL’s behavior in the last column comes from PostgreSQL’s CREATE INDEX and PostgreSQL’s ALTER TABLE pages. Safe here means both code versions keep working, not lock-free: ALTER TABLE states “An ACCESS EXCLUSIVE lock is acquired unless explicitly noted”, so I’d assume even the quick forms queue behind a long-running query. That wait is what the lock timeout in the fails-halfway section is for. Backward compatible database changes are the rows marked yes, and every other row becomes one once it is split into deploys that each work with the code around them. Before you put a concurrent index build in a migration file, read the rollback guide’s section on how PostgreSQL transaction rollback differs from migration rollback.

Zero downtime deployment with database changes: which deploy runs each migration

Zero downtime deployment with database changes means no release step removes anything the code still serving uses. A column rename takes 4 deploys: the first adds the new column, a separate job backfills it, the next two move the code over, and the fourth drops the old column.

The pattern behind this, its full sequence and why it avoids emergency rollbacks, is in expand-contract in the database migration rollback guide, and I won’t repeat it. What this page adds is which release runs what, for a rename of users.name to users.full_name:

DeployMigration the release step runsWhat the code doesSafe to roll the code back to
1Add full_name, nullableWrites both columns, reads nameThe previous release
Between 1 and 2None (a one-off job copies name into full_name in batches)UnchangedNot a code change
2NoneReads full_name, still writes bothDeploy 1
3If name is NOT NULL, ALTER COLUMN name DROP NOT NULL firstStops writing nameDeploy 2
4Drop nameNothing reads or writes nameDeploy 3

Deploy 3 needs the DROP NOT NULL because that form changes whether a column is marked to allow or reject null values, and my inference is that a row the new code inserts without name would otherwise be rejected. The “safe to roll the code back to” column is mine. I leave the backfill out of the release step because the release step holds the deploy, and Fly.io, Heroku and Render cap that step while Railway by default waits on it; the limits are in the next section.

-- Deploy 1, release step
ALTER TABLE users ADD COLUMN full_name text;
-- Between deploys, one-off job, repeated per batch of ids
UPDATE users SET full_name = name WHERE id BETWEEN $1 AND $2 AND full_name IS NULL;
-- Deploy 3, release step, only if name is NOT NULL
ALTER TABLE users ALTER COLUMN name DROP NOT NULL;
-- Deploy 4, release step
ALTER TABLE users DROP COLUMN name;

Blue-green and canary deployments don’t change this; Fly.io lists both canary and bluegreen strategies, and in each the old and new versions use the same database during the switch, so in my view the same schema rule applies. For a small app, zero downtime honestly means no failed requests caused by the schema. It does not mean the deploy carries no risk.

When a migration fails halfway

A migration that fails halfway should stop the deploy while the old version keeps serving, and the release steps on Fly.io, Heroku, Render and Railway all stop on a failed command. The harder case is a migration waiting on a lock. A lock timeout on the migration session turns that wait into a failure.

The deploy half of a failure is the “what a failure does” column of the host table above. What the database kept after a failed migration depends on the transaction boundary, which the rollback guide’s section on PostgreSQL transaction rollback versus migration rollback covers.

A migration that waits instead of failing runs into each host’s release-step limit:

  • Fly.io: 5 minutes by default before the command is killed, changed with release_command_timeout.
  • Heroku: a 1-hour timeout that cannot be extended.
  • Render: 30 minutes for the pre-deploy command.
  • Railway: no time limit by default, so a command waiting on a lock holds the deployment in progress rather than failing it; a Pre-deploy Timeout of 1 to 3,600 seconds can cap it.

The fix sits in the migration’s own session: set lock_timeout, which aborts any statement that waits longer than the limit while trying to acquire a lock, and in practice that makes the migration step fail instead of hang. The limit applies to each lock acquisition attempt separately, a value without units is taken as milliseconds, and zero, the default, disables it; setting it in postgresql.conf is not recommended because it would affect all sessions. The wording is in PostgreSQL’s lock_timeout setting. I start at a few seconds on the migration session and raise it only for a statement you know needs longer.

What to do next, stopping retries, reading the history table and choosing between rollback and fixing forward, is the rollback guide’s section on what to do immediately after a database migration fails. The release side of that recovery belongs in a rollback plan you have rehearsed.

How to verify it: test a schema change through the pipeline

A migration-aware deploy is verified with two test changes. The first, a harmless new column, must run once before the traffic switch and appear in the history table. The second, broken on purpose, must stop the deploy and leave the old version serving an unchanged schema.

Run this on a staging database first; the article on rollbacks and staging, linked in the section on a migration that ran before the code, explains why. Each check is something a correct pipeline passes and a broken one fails, and each leaves evidence you can keep.

  1. 01 Add a harmless migration, such as a nullable column on a low-traffic table, and merge it. Evidence: the commit id.
  2. 02 Open the host's deploy log page and confirm the migration step ran once, before the traffic switch, with exit code zero. Evidence: the log, with its date.
  3. 03 Open the table view in your database dashboard and confirm the migration tool's history table lists the migration and the new column exists. Evidence: a screenshot of both.
  4. 04 Add a second migration with one statement that references a table that does not exist, merge it, and confirm the deploy stopped, the old version still answers requests and the schema did not change. Evidence: the failed log and one response from the old version. Then delete the broken migration before the next deploy.
  5. 05 Restart the app, or scale it to two instances where your plan allows it, and read the instance logs from the restart. Evidence: no log line shows a migration run. A host with no start command skips this check.
  6. 06 Time the migration step and note it against your deploy budget and the host's release-step limit. Evidence: the step's duration from the log.

Check five needs the instance logs because the history table cannot show a boot-time run: a migration tool runs only pending migrations, so a second run should find nothing to do and leave no new row. Scaling has plan limits: Heroku’s docs say you can’t run more than one Eco or one Basic dyno per process type, and Render’s docs mark its free instance type “1 (scaling unavailable)”, while Vercel has no start command to check. For check six, a slow migration step spends deploy budget, and keeping that budget small is part of how to speed up CI build times. The checks that run before merge, a migration lint among them, are CI/CD best practices, and a migration file that skips review is a job for GitHub branch protection.

In the Production Hardening Sprint, deliverable 7.5 is verified this way: deploy a representative schema change and verify sequencing and failure handling.

Where the sprint does this

Deliverable 7.5 of the Production Hardening Sprint integrates database migrations into the deployment process with ordering and compatibility checks; 4.6 moves schema changes into ordered migrations committed to the repository; 7.6 documents and rehearses release rollback, including how database changes are handled safely; and the production readiness report, 13.1, delivers the result for every scope item, the work completed and its verification evidence. All four are among the 123 deliverables in the sprint.

Common questions about migrations and deploys

What does “db migration” mean?

A db migration is a versioned, ordered change to a database’s schema, kept as a file and applied once to each database, as opposed to moving data from one system to another. Each file adds, changes or removes tables, columns, indexes or constraints, and the migration tool records which files a database has already applied.

What does “code migration” mean?

Code migration means moving an app’s code to a new language, framework, library version or host. It changes the code, not the structure of the database, and it is a separate job from the schema migrations a deploy runs.

What’s the difference between forward and backward compatibility?

Backward compatibility means the new schema still works with the old code; forward compatibility means the old schema still works with the new code. A deploy that adds a column first needs the first, and a deploy that ships code before removing a column needs the second, because the new code runs for a while against a schema that still has it.

How do you handle schema changes?

Handle schema changes as migration files that are reviewed, run once by the release step, and split so additions ship first and removals ship in a later deploy. Each deploy should work with the schema before and after it.