How long should a deploy take? My working target for a small web app is under 10 minutes from merge to live. How to speed up CI build times starts with a stopwatch on each stage, then two usual fixes: cache what every run redoes, and run in parallel the jobs that now wait in a line.

What a timed deployment is, and how long a deploy should take

A timed deployment is the release path measured end to end: from the merge to the new version answering on production, with every step in between listed and timed. For a small web app, my working target is under ten minutes with every check included, so an urgent fix ships in one sitting.

Timing the release path is one control on the deployment side of DevOps for startups, and it rests on one rule: one start event, one finish event, and every step between them written down. A number with no named start, finish or step list cannot be compared with next month’s.

The ten-minute line is my working target, not an industry standard. It counts only a representative release, which in my working rule is a normal code change that goes through every check the team relies on: install, lint, types, tests, build, migrations, the provider’s rollout and the health check. A README edit that skips half of those says nothing about the path.

The reason to time it at all is short: a slow release process delays delivery of urgent fixes.

Keep this apart from the metric DORA calls change lead time. DORA’s metrics guide starts that clock when a change is committed to version control and stops it when the change is deployed in production, so it also counts the hours a pull request sits waiting for review. The DORA section below covers that metric; this page shortens the pipeline part of it.

What goes wrong without it

A slow release path changes how people ship long before anyone measures it. The table is my reading of the four patterns, with the stage I would time first for each.

What you seeWhat it costsThe stage usually behind it
The deploy takes an hour to ship a fixA production bug stays live for an hour after it is understoodTests and build running one after the other, nothing cached
People batch changes to avoid the waitEach release is bigger and harder to roll backEvery stage: each merge pays the full wait
Someone deploys from a laptop or edits in the hosting dashboardThe repository no longer matches productionWhichever stage made the pipeline feel too slow to wait for
Checks get deleted to save timeA slow pipeline is traded for an untested oneUsually the test stage, the longest one

The second row is the one that hurts later. A batched release has more in it to undo, and what a rollback plan is depends on knowing exactly which changes went out together.

The gap my audits turned up was a different one: missing checks. At least 18 of the 21 third-party apps had no working test anywhere: 17 with literally none, plus a retail POS whose checkout “test suite” never executed the actual checkout code. Across those apps, the Deployment & Operations pillar averages 37.0 out of 100, scored on 21 of the 21 third-party apps.

The 21 apps are third-party apps I audited in June and July 2026, chosen rather than drawn at random, so neither number is a rate for AI-built apps in general. My reading of them: many of those pipelines were fast only because they checked nothing. The goal here is a fast pipeline with the checks still in it.

My own code is not exempt. In one of my own apps, the CI step named “npm install, build, and test” passed on every push without running a test. A green tick there proved only that the step ran, which is why the way to verify AI-generated code before production starts with reading what each step executes.

What the pipeline should run, and in what order, belongs to CI/CD best practices. This page times whatever it runs.

How to measure end to end deploy time

End to end deploy time is measured from one start event, the moment the change merges to main, to one finish event, the first successful health check on the new version. Queue time, install, lint, tests, build, migrations and the provider’s rollout each get their own line, because the slow stage is rarely the one people guess.

These are the eight stages I time, as my working rule, with where each timestamp comes from on GitHub Actions and Vercel.

StageWhere the timestamp comes fromUsual cause when slowFirst fix
QueueGitHub: the run’s createdAt against the first job’s startedAt, both in gh run view --json; Vercel: a build waits while all build slots are busyEvery slot or runner busy, or several pushes in a rowCancel superseded runs; on Vercel, a newer commit on the same branch already skips older queued builds
Checkout and installThe install step’s startedAt and completedAt in the run’s jobs fieldNo dependency cacheCache the package manager’s downloads, keyed on the lockfile
Lint and type checkIts job’s execution time, shown under the job summaryRunning in line with the testsA separate job that runs in parallel
TestsThe test job’s time, one figure per shardOne long suite on one machineShards across parallel jobs
BuildThe build job on GitHub; on Vercel, the build logs in the Deployments sectionThe framework’s build cache is not kept between runsPersist the build cache (Vercel does this for Next.js)
MigrationsThe migration step’s startedAt and completedAt in the run’s jobs fieldMore work in the migration than the schema changeKeep data changes out of the deploy step
Upload and rolloutThe deploy job’s time; on Vercel, the deployment’s build details in the Deployments sectionA held or slow promotionRead the provider’s log before changing anything
Health checkYour own request to the health endpoint after the rolloutCold start, or a check that does too muchKeep the check light, and have it return the version or commit hash

Queue time on Vercel has its own causes and settings, covered in a Vercel build stuck on queued. A slow migration step is a design question for how to run database migrations on deploy, not a caching one.

The first pass needs no terminal. Open the repository, click Actions, pick the workflow in the left sidebar, open a run from the list, and read each job’s execution time under the job summary. Do that for the last five runs and mark the longest job in each. My working rule is to time five runs rather than one and keep the median and the worst, because a run that starts with a cold cache looks nothing like the warm run after it.

GitHub’s page on job execution time also shows where billable time sits: under “Run details”, click Usage. Billable minutes are only shown for jobs on private repositories that use GitHub-hosted runners, rounded up to the next minute.

From a terminal, the GitHub CLI prints the same run as JSON:

gh run list --branch main --limit 5
gh run view <run-id> --json createdAt,startedAt,updatedAt,jobs
gh run view <run-id> --verbose

The gh run view manual lists createdAt, startedAt, updatedAt and jobs among a run’s JSON fields, and --verbose shows job steps. The manual does not list the fields inside jobs; the GitHub CLI’s source code exports a startedAt and a completedAt for each job and for each step in it. Read one run’s output before you script anything on top of it.

On Vercel, logs and build details are in the Deployments section of the dashboard, as Vercel’s builds docs put it. The finish is a request you send to the health endpoint once the rollout completes; a version or commit hash in the response tells you the new build is the one answering.

Two clocks get mixed up here. The pipeline’s duration is the one this page shortens; change lead time adds the wait for review, and a team that reports one as the other argues about the wrong number. The record you keep of a timed run is in “How to verify it” below.

How to speed up CI build times: caching and parallelism, the two reasons a pipeline got slow

Caching and parallelism are the two levers that shorten most pipelines. Caching keeps downloaded packages and build outputs between runs, keyed on the lockfile. Parallelism runs lint, type checks and test shards as separate jobs, so the pipeline takes as long as its slowest job, not the sum of them all.

The settings below come from each tool’s documentation, linked in the sections that follow; the trap column is my reading.

Slow stageThe fixThe settingThe trap
Dependency installCache the package manager’s downloadsactions/setup-node with cache: 'npm' (npm, yarn and pnpm supported)It does not cache node_modules, so the install step still runs
Framework buildKeep the build cache between runsactions/cache on .next/cacheA cache that never restores goes unnoticed unless someone reads the build log
Container image buildCache image layerscache-from: type=gha and cache-to: type=gha,mode=maxThe inline cache only supports min cache mode
Lint, types and tests in one lineMake them separate jobsJobs run in parallel by default; needs only on the deploy jobA needs chain on every job puts them back in a line
One long test suiteSplit it into shardsPlaywright --shard=x/y, Jest --shard, a matrix of jobsShards that share one database or one port
Runs nobody needs any moreCancel superseded runsconcurrency with cancel-in-progressOn a deploy or migration job, it can stop a deploy halfway
A docs-only change runs everythingSkip the workflow for those pathspaths-ignoreA filtered workflow made a required check leaves pull requests “Pending”

Caching: stop redoing work that did not change

Three caches pay off, and in my working rule they pay off in this order: the package manager’s download cache, the framework’s build cache, then image layers if the pipeline builds a container.

The first is one input on a step you already have. actions/setup-node has built-in caching for npm, yarn and pnpm; it uses actions/cache under the hood for global package data, by default uses the hash of the lockfile in the repository root as part of the cache key, and does not cache node_modules. GitHub’s dependency caching guide says the setup actions need minimal configuration and “will create and restore dependency caches for you”.

The second belongs to the framework. Per Next.js CI build caching, Next.js keeps its build cache in .next/cache, and a CI that does not persist that folder between builds may show a No Cache Detected error; on Vercel the cache is configured for you. Both caches fit in one job; the actions/cache path and keys follow the Next.js guide, and the action versions follow each action’s README:

steps:
  - uses: actions/checkout@v7
  - uses: actions/setup-node@v7
    with:
      node-version: 24
      cache: 'npm'
  - run: npm ci
  - uses: actions/cache@v6
    with:
      path: ${{ github.workspace }}/.next/cache
      # new cache whenever packages or source files change
      key: ${{ runner.os }}-nextjs-${{ hashFiles('**/package-lock.json') }}-${{ hashFiles('**/*.js', '**/*.jsx', '**/*.ts', '**/*.tsx') }}
      # source changed, packages did not: start from a prior cache
      restore-keys: |
        ${{ runner.os }}-nextjs-${{ hashFiles('**/package-lock.json') }}-
  - run: npm run build

The third only matters if the pipeline builds an image. Docker’s cache backends for GitHub Actions cover inline, registry, GitHub and local caches plus cache mounts. Docker says that in most cases you want the inline cache exporter, and notes that it only supports min cache mode. The GitHub cache backend uses GitHub’s cache service, and its cache backend API is marked experimental.

GitHub’s cache has limits worth knowing before you add more. Entries not accessed in over 7 days are removed, the default limit is 10 GB per repository, and once that is full the oldest-accessed caches are deleted first, which may cause cache thrashing.

Hosted platforms do part of this for you. Vercel’s build cache can be up to 1 GB and is retained for one month, and by default Vercel makes a shallow clone (git clone --depth=10) to speed up builds. A monorepo adds task caching on top: Turborepo caches to the local filesystem by default, and its remote caching shares those artifacts across machines.

The trap, in my reading: a cache keyed too loosely serves stale output, so every key carries the lockfile hash, and a cache is never the fix for a flaky test.

Parallelism: run side by side what runs in a line

Jobs in a GitHub Actions workflow run in parallel by default, and needs is what lines them up: a job with needs waits for the named jobs to complete successfully. So lint, type check and the test shards become separate jobs, only the deploy job waits on them, and the pipeline lasts as long as its slowest job.

The workflow below is for a pipeline that deploys from CI, built from GitHub’s matrix and job docs. Each job also starts with the checkout, setup-node and install steps from the caching block; they are cut here so the shape stays visible.

on:
  pull_request:
  push: { branches: [main] }
concurrency:
  group: ci-${{ github.ref }}
  cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
jobs:
  lint:
    runs-on: ubuntu-latest
    steps: # checkout, setup-node and npm ci go first
      - run: npm run lint
  typecheck:
    runs-on: ubuntu-latest
    steps:
      - run: npm run typecheck
  test:
    runs-on: ubuntu-latest
    strategy: { matrix: { shard: [1, 2, 3, 4] } }
    steps:
      - run: npx playwright test --shard=${{ matrix.shard }}/4
  deploy:
    needs: [lint, typecheck, test]
    if: github.ref == 'refs/heads/main'
    runs-on: ubuntu-latest
    steps:
      - run: npm run deploy

On a Vercel project deployed by its Git integration, Vercel creates a production deployment each time you merge to the production branch, so a needs chain in this file gates nothing there. Vercel’s gate is Deployment Checks: Vercel “will hold each production deployment until all required checks pass before assigning it to your custom production domains”, and its GitHub Checks source can “Import GitHub Actions workflow results as Deployment Checks”.

Playwright’s test sharding splits a suite with --shard=x/y, one shard per job, and can only shard tests that can run in parallel, which by default means whole test files. Each shard writes its own report; the blob reporter plus npx playwright merge-reports turns them into one. Jest takes a --shard option in the same index/count form, provided its test sequencer implements a shard method.

The concurrency block cancels superseded runs. By default a new run replaces any run still pending in the same group, and cancel-in-progress also cancels one already running. The expression above switches that off on main, where at most one run in the group runs at a time and the next one waits as pending. In my working rule, a deploy or a migration stopped halfway is worse than a short wait.

Path filters let a documentation-only change skip the test matrix: with paths-ignore, the workflow does not run when every changed path matches the patterns. GitHub keeps one condition on that: the checks of a workflow skipped by path filtering stay “Pending”, and “A pull request that requires those checks to be successful will be blocked from merging.” So a filtered workflow is never a required check, in my reading; a representative change still runs the full matrix.

Parallel jobs spend more runner minutes for less waiting. On private repositories, GitHub-hosted runner minutes come from a free quota set by the account’s plan: GitHub’s billing page lists 2,000 minutes a month on GitHub Free and 3,000 on Pro and Team. The trap, in my reading: test shards that write to the same database or bind the same port pass alone and fail together.

Bazel test no cache: rerunning tests Bazel reports as cached

Bazel test no cache means asking Bazel to rerun tests it would report as cached. The flag is --cache_test_results=no, which executes all tests unconditionally. It reruns tests without deleting the build outputs, which is why it suits a flaky test, while bazel clean deletes the output directories and resets caches.

By default the option is auto, and Bazel “will only rerun a test if any of the following conditions applies”: it detects changes in the test or its dependencies, the test is marked as external, multiple runs were requested with --runs_per_test, or the test failed. With yes, it behaves like auto but may also cache test failures and runs made with --runs_per_test. The short form is -t, and Bazel’s user manual notes that the option changes only whether Bazel uses previously saved results, not whether it saves the current run’s.

bazel test --cache_test_results=no //...

# .bazelrc: a named config, used only when asked for
test:nocache --cache_test_results=no
# then: bazel test --config=nocache //...

For a flaky test, --runs_per_test “specifies the number of times each test should be executed”, which turns one lucky pass into a count of passes and failures.

bazel clean is the heavy tool. It deletes the output directories for all build configurations and resets internal caches, and with --expunge it removes the entire output base. The manual says clean is provided “primarily as a means of reclaiming disk space for workspaces that are no longer needed”; reaching for it to rerun one test throws away the outputs the next build would have reused, which in my reading is how a slow Bazel pipeline gets slower.

A no-cache tag on a Bazel test target is a different switch: it keeps that action or test out of the local and remote caches, while Skyframe and the persistent action cache are not affected. In Bazel issue 19909 on GitHub, a user reported that the tag did not disable caching of a genrule; a maintainer said the team would treat it as a documentation bug for no-cache, and the issue is closed as not planned.

A small JavaScript app does not need Bazel. This section is for whoever inherited one.

CircleCI Chunk, test splitting, and other tool-specific fixes

CircleCI Chunk is not a test-splitting feature. In CircleCI’s words, “Chunk by CircleCI assists with CI/CD related tasks through a natural language chat interface and task scheduling”, and “Chunk is currently in beta.” CircleCI’s Chunk docs list six skills: fix flaky tests, extend test coverage, fix bugs, refactor code, improve documentation and optimize build configs. CircleCI describes it as an AI agent for CI/CD tasks, and I make no claim here about how well it does that work.

If what you wanted was CircleCI’s way to split a slow suite, that is test splitting with parallelism: the parallelism key sets how many independent executors run the job, and circleci tests split divides the tests between them. CircleCI’s test splitting docs call that the legacy circleci tests command family; the preferred way is now circleci testsuite, available on CircleCI Cloud only, and CircleCI Server keeps circleci tests glob and circleci tests split.

Whatever the CI, the same two levers do the work: a dependency cache keyed on the lockfile, and test containers running in parallel.

The five DORA metrics, and the two a small team can actually move

The five DORA metrics, as DORA’s guide lists them, are change lead time, deployment frequency, failed deployment recovery time, change fail rate and deployment rework rate. A small team can move two of them this month, in my reading: lead time, by shortening the pipeline, and change fail rate, by counting which deploys needed immediate intervention.

The second column copies DORA’s guide, last updated January 5, 2026. The last two columns are my reading for a team of three with no metrics platform.

MetricDORA’s definitionHow a team of three counts it without a platformCan you move it this month?
Change lead time”The amount of time it takes for a change to go from committed to version control to deployed in production.”First commit on the pull request to the finish time in the release recordYes, by shortening the pipeline and the review wait
Deployment frequency”The number of deployments over a given period or the time between deployments.”Count production deploys in the release log each weekIt follows from a fast pipeline; not a goal by itself
Failed deployment recovery time”The time it takes to recover from a deployment that fails and requires immediate intervention.”From the alert to the health check passing on the fixed versionThrough the rollback drill, not through this page
Change fail rate”The ratio of deployments that require immediate intervention following a deployment. Likely resulting in a rollback of the changes or a “hotfix” to quickly remediate any issues.”Deploys labeled as failed, divided by all deploys, monthlyYes, by counting it the same way every month
Deployment rework rate”The ratio of deployments that are unplanned but happen as a result of an incident in production.”Deploys made because of an incident, divided by all deploysIndirectly: it falls as fewer deploys fail

Deployment frequency rises on its own once the pipeline is short, so I would not chase it. Failed deployment recovery time belongs to the rollback drill. That leaves change lead time and change fail rate as the two worth a month of attention.

Change failure rate: what it is, the formula, and how to count it by hand

Change failure rate, which DORA’s guide calls change fail rate, is the ratio of deployments that require immediate intervention after they ship, likely a rollback or a hotfix. Count failed deployments and divide by all deployments in the same period: in my example, 4 interventions in 40 deploys is a rate of 10 percent.

DORA change failure rate and DORA change fail rate name the same metric; the guide’s exact definition is in the table above. Failure rate in reliability engineering, the one tied to MTBF, is a different measure that happens to share the words.

Written out, my example is 4 failed deploys ÷ 40 deploys = 0.10, a change fail rate of 10 percent; the numbers are invented to show the arithmetic, not taken from any team.

My working rule is to decide once what counts as a failure and write it down: a rollback, a hotfix for that deploy, or a feature flag turned off to stop harm. The rate means nothing if the definition drifts from one month to the next. No tool is needed to count it, either. Put a label such as deploy-failed on the pull request whose deploy needed intervention, or add a line to the release log, and tally both columns once a month.

DORA’s guide prints no benchmark levels, and the 2025 DORA report download asks for a name, a business email and a phone number, so this page quotes no levels. DORA’s Quick Check scores your five metrics and compares the results with benchmarks derived from its 2025 research.

With 10 deploys a month, one failure moves the rate by 10 points, so I read the trend over a quarter rather than any single month.

DORA capabilities: the list behind the metrics

DORA capabilities are the practices DORA’s research connects to delivery performance. DORA’s capabilities catalog lists 34 of them, tags 19 as core and 7 as AI, with two carrying both tags. DORA describes its Core Model as the capabilities, metrics and outcomes that represent its most firmly established findings.

A small team already touches several: continuous integration, test automation, deployment automation, database change management, working in small batches, and monitoring and observability are all in the catalog. Small batches matter most for this page, since DORA’s guide names reducing the batch size of changes as a common approach to improving all five metrics.

Engineering velocity: what people mean, and the version you can defend

Engineering velocity is a loose name for how fast a team ships working software. It is not one of the five metric names in DORA’s guide, and it is not the story-point velocity that Scrum teams use for planning.

For a board or a client report, the version I would defend is the pair this page measures: how long a change takes to reach production, and how often a deploy needs immediate intervention. One speed number on its own invites gaming. Two of the pitfalls in DORA’s guide are “Setting metrics as a goal”, which it says makes it more likely that teams will try to game the metrics, and “Having one metric to rule them all”.

How to verify it

A timed deployment is verified by releasing a representative change through the normal path with a stopwatch on it: record the start, the finish and every included step, and keep the run’s URL. Under ten minutes with every check included passes my working target; a run that skipped tests or migrations does not count.

Six steps, each one a run can fail:

  1. 01 Agree the deployment path in one sentence, for example: from the merge to main to the health check passing on production
  2. 02 Pick a representative change: a normal code change that runs every check, not a docs edit
  3. 03 Record the start as the moment the change reached main: the pull request's merge time, or for a direct push the workflow run's createdAt, never the commit's own timestamp
  4. 04 Let the normal pipeline run with nothing skipped
  5. 05 Record the finish at the first successful health check on the new version, with a version or commit hash in the response so there is no doubt which build answered
  6. 06 Fill in the release record: every included step and its duration, every skipped step with its reason, and the run URL

Step three reads the run’s createdAt field, which the gh run view manual lists, for a direct push. A commit’s own timestamp can be older than the merge, so a clock started there runs long.

Pass: under ten minutes with every step included, my working target. Fail: over ten minutes, or under ten because tests or migrations were skipped. Repeat the run after any change to the pipeline and keep the last three records.

The release record, ready to copy:

DateCommitChangeStart (merge)Finish (first healthy check)TotalIncluded stepsSkipped steps, with reasonRun URL
(date)(short hash)(one line)(time)(time)(minutes)install, lint, types, tests, build, migrations, rollout, health checknone(link)

This is how we verify deliverable 7.10, Timed deployment, in the Production Hardening Sprint: time the agreed deployment path and record its start, finish, and included steps.

Where the sprint does this

In the Production Hardening Sprint, deliverable 7.10, Timed deployment, is to optimize and document the deployment pipeline to complete a representative release in under ten minutes, verified as the section above describes. Deliverable 13.1, the production readiness report, delivers the result for every scope item, the work completed and its verification evidence, accounting for all 123 IDs, keeping failures visible until resolved and explaining genuine non-applicable items. Your app’s current framework and hosting setup are our starting point, and we refactor or replace components where the production work requires it. Hosting, paid tools, and API usage remain in your accounts. The 7.10 wording comes from deliverable 7.10 in the published scope, which lists all 123 deliverables.

Common questions about deploy time and DORA metrics

What are the 5 DORA metrics?

Change lead time, deployment frequency, failed deployment recovery time, change fail rate and deployment rework rate, as DORA’s guide, last updated January 5, 2026, names them. DORA counts the first three as throughput and the last two as instability. Here DORA is the software delivery research program run by Google Cloud, not the EU’s Digital Operational Resilience Act, a financial-sector regulation that shares the acronym.

What is considered a good change failure rate?

For a small team, a good change failure rate is one that falls over a quarter while being counted the same way each month. DORA’s guide gives no single number; DORA’s Quick Check is where you compare your rate with benchmarks from its 2025 research.

What is a good deployment frequency?

As often as a change is ready. For a small app that meets my ten-minute target, that can mean several deploys a day, and nothing forces a wait between deployments except the pipeline itself and the host’s build slots. DORA’s guide warns against targets such as “Every application must deploy multiple times per day by year’s end”, and its Quick Check gives the 2025 benchmarks if you need a comparison.

Are DORA metrics still relevant?

Yes. DORA still publishes and revises them: its guide, last updated January 5, 2026, moved from the original four keys to the current five-metric model. The guide also says the metrics are best suited for measuring one application or service at a time, which is how a small team should use them.

Where is the Bazel cache located?

Under Bazel’s output base, which sits inside an output root that defaults to ~/.cache/bazel on Linux, ~/Library/Caches/bazel on macOS from Bazel 9, and %HOME% or else %USERPROFILE% on Windows. Bazel 8.x and earlier on macOS used /private/var/tmp, and a set $XDG_CACHE_HOME moves the root to ${XDG_CACHE_HOME}/bazel. Run bazel info output_base to print the exact path for your workspace.