A container can be running and useless: the process is up, the app inside is wedged, and the status column still says Up 3 days. A docker compose health check changes that label. After 3 failed probes it marks the container unhealthy, and no restart follows, because Docker’s restart policies act when a container exits or Docker restarts.

Docker Compose health check: what it is, and the example to copy

A Docker Compose health check is a command Docker runs inside a container on a timer to tell running apart from working. Exit code 0 means healthy and 1 means unhealthy. It sets the container’s health status, which docker ps shows and which other services can wait on before they start.

In a Compose file the check is the healthcheck attribute of a service. Docker’s Compose healthcheck reference says it works the same way, with the same defaults, as the HEALTHCHECK instruction in a Dockerfile, and that your Compose file can override the values the image sets. The route the probe calls needs its own care, starting with what a health check endpoint is and what it must never expose. Container health is one layer of hardening SaaS applications for resilience.

Guides also write the tool as docker-compose, with a hyphen, and run health check together into one word; every spelling points at this same setting.

The docker compose healthcheck example below is for a small SaaS: a web app probed on its health route, Postgres probed with its own readiness tool, the web service held back until the database is healthy, and a restart policy on both.

services:
  web:
    build: .
    restart: unless-stopped              # start again after an exit or a Docker restart
    depends_on:
      db:
        condition: service_healthy       # create web only once db passes its check
    healthcheck:
      test: ["CMD-SHELL", "curl -f http://localhost:3000/health || exit 1"]
      interval: 15s                      # Docker's default is 30s
      timeout: 5s                        # a probe slower than this counts as a failure
      retries: 3                         # consecutive failures before "unhealthy"
      start_period: 30s                  # failures here do not count; size it to your cold start

  db:
    image: postgres:18
    restart: unless-stopped
    environment:
      POSTGRES_USER: app
      POSTGRES_DB: app
      POSTGRES_PASSWORD: ${POSTGRES_PASSWORD}   # required by the image; keep it in .env
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U $${POSTGRES_USER} -d $${POSTGRES_DB} || exit 1"]
      interval: 10s
      timeout: 5s
      retries: 5

The port and the /health path stand in for your own app’s values. The $$ makes Compose pass a literal $ through, so the shell inside the container reads the variables, the way Docker’s startup order example writes the same check.

Six options control the check. Docker’s defaults below come from the Dockerfile reference, read on 2026-09-30. The last column is my working rule for a small SaaS, not Docker’s advice.

OptionWhat it doesDocker’s defaultMy starting value for a small SaaS
testThe command. A list can start with CMD (run the program directly), CMD-SHELL (run a string in the container’s shell) or NONE (no check); a plain string means CMD-SHELLNone: without a check the container has no health statusHit the app’s health route
intervalTime between probes; the first runs one interval after the container starts30sAbout 10 to 15 seconds
timeoutA probe that runs longer counts as a failure, and Docker stops it with SIGKILL30sAbout a third of the interval
retriesConsecutive failures before the status turns unhealthy33
start_periodStart-up grace: failures inside it do not count, and one success ends it early0sThe slowest cold start you have measured, plus about half again
start_intervalTime between probes during the start period (needs Docker Engine 25.0 or later)5sLeave it

The time from a wedged app to unhealthy is roughly interval times retries. With the example’s values that is 3 probes 15 seconds apart, about 45 seconds. If each probe hangs until its 5-second timeout, it is closer to a minute, because Docker waits a full interval after each check completes before the next one.

Point the web probe at the app’s health route. Never the home page, which can still answer while the parts that matter are failing, and never a route that calls a paid API, because the probe runs every interval for as long as the container lives.

For the database, use the database’s own readiness tool. PostgreSQL’s pg_isready docs say it returns 0 “if the server is accepting connections normally, 1 if the server is rejecting connections (for example during startup), 2 if there was no response to the connection attempt, and 3 if no attempt was made”. Docker’s reference marks exit code 2 as reserved (“don’t use this exit code”), so the example ends with || exit 1 to turn every failure into a plain 1. MySQL, Redis and the rest ship their own readiness commands, and they go in the same place.

Docker health checks in the Dockerfile, and the health check command

Docker health checks can live in the image as a HEALTHCHECK instruction or in the Compose file, which can override it. In Compose the test comes in 2 forms: CMD runs a program directly and CMD-SHELL runs a string in the container’s shell. The probe program must exist inside the image, and distroless images ship no shell.

The Dockerfile version takes the same options as flags:

HEALTHCHECK --interval=15s --timeout=5s --start-period=30s --retries=3 \
  CMD curl -f http://localhost:3000/health || exit 1

A Dockerfile has one health check: if it lists more than one HEALTHCHECK, only the last takes effect, as Docker’s HEALTHCHECK instruction reference puts it. My working rule on where to put it: in the Dockerfile when you ship the image for other people to run, so the check travels with it, and in the Compose file when the right check depends on how you deploy.

The || exit 1 is there for a reason. With -f, curl fails “with error code 22” for HTTP response codes of 400 or greater, and without it curl does not treat an error status as a failure at all. Docker’s Engine API spec lists 0 as healthy, 1 as unhealthy, 2 as reserved and “(considered unhealthy)”, and any other value as “error running probe”, so a probe that ends in exactly 0 or 1 is the one to write.

A common reason a new check fails, in my reading, is that the program it calls is not in the image. The distroless images README says those images “contain only your application and its runtime dependencies” and no package managers or shells, so neither curl nor CMD-SHELL works there. Which docker health check command to use depends on what the image contains:

What the image hasThe command to use
curl and a shellThe CMD-SHELL line from the example file above: curl with -f, ending in || exit 1
BusyBox wget and a shell, no curl["CMD-SHELL", "wget -q --spider http://localhost:3000/health || exit 1"]; BusyBox documents spider mode as “only check file existence” and does not state its exit status on an HTTP error, so run it once against a failing route
Node.js, no curl and no shell["CMD", "node", "-e", "fetch('http://localhost:3000/health').then(r => process.exit(r.ok ? 0 : 1), () => process.exit(1))"]; fetch is global from Node 18 without a flag; if node is not on the image’s PATH, give the full path to the binary
Python, no curlA small script copied into the image, run as ["CMD", "python3", "/app/healthcheck.py"]: it calls the route with urllib.request and ends with sys.exit(0), or sys.exit(1) on any exception (urlopen turns an error status into an HTTPError)

Two more rules apply to every row. Probe localhost on the port the app listens on inside the container, not the port you published on the host, because the command runs inside the container. And keep it cheap: it runs every interval, for the whole life of the container.

Why it matters: up is not the same as working

Docker’s own reference names the case a health check exists for: it “can detect cases such as a web server stuck in an infinite loop and unable to handle new connections, even though the server process is still running”. Without a check, the STATUS column of docker ps shows Up and a duration for any live process. These three rows are illustrations, my reading of how the gap shows up, not incidents:

What happenedWhat docker ps showedWhat a health check would have shown
The app deadlocks or runs out of database connections and stops answering, while the process lives onUp 3 days(unhealthy) after the configured failures
The API starts before Postgres accepts connections and fails its first requests after every docker compose upUpPostgres as starting, with the API not created until it is healthy
A deploy script waits for “container started” and reports success while the app is still failing its migrationsUpstarting, then unhealthy

In one app I audited, a voice-AI SDK’s token server had no health endpoint and its container had no health check, so a stuck instance kept taking traffic. My reading of it: a container with no check can only ever look up, and nothing on that host could tell the stuck instance from a working one.

The same blind spot shows up one layer out. In my June and July 2026 audits, 17 of the 21 third-party apps had no error tracking or alerting: when a user hits an error, nothing records it. Those 21 apps are a selected set I audited, not a random sample, and the count is not a rate for AI-built apps in general.

A health status lives on the host. A person hears about it only if an outside check is watching, which is the job of uptime monitoring for founders; if the app is down as you read this, start with what to do first when your app is down. On a managed host, the platform’s own health check setting takes the place of the Compose one.

How it works: the states, startup order, and restart policies

Four pieces make the whole picture: reading the status, debugging an unhealthy one, using health to order startup, and what actually restarts a container.

Starting, docker healthy, unhealthy: the three states

Docker reports 3 health states. A container is starting at first, healthy after a passing probe, and unhealthy after the configured run of consecutive failures. The status shows in docker ps, and docker inspect prints the output of recent probes, which is where debugging starts.

A single passing probe makes a container healthy “whatever state it was previously in”, so a flapping app can swing back and forth. Three commands read the status; the first is read-only and safe to run on any host:

docker compose ps                                   # every running service, with its health
docker ps --filter health=unhealthy                 # only the unhealthy containers
docker inspect --format '{{json .State.Health}}' <container>

The inspect output has three fields: Status, FailingStreak (the number of consecutive failures) and Log, which “contains the last few results (oldest first)”, each with its start and end time, exit code and output. Docker keeps only the first 4096 bytes of a probe’s output, so a probe that prints a short reason on failure is easier to debug.

A container with no health check has no health status at all. The docker ps filter calls that state none, and it is not the same as healthy. Every change of status is also written to docker events as a health_status event, which a watcher script on the host can read.

Docker container unhealthy: causes and first checks

A Docker container marked unhealthy has failed its probe several times in a row, 3 by default. The causes I check first are a probe program missing from the image, the wrong port or address, a start_period shorter than the real cold start, and a health route that fails when an optional dependency does.

Docker stores whatever your probe command wrote to stdout or stderr as its output, so there is no fixed error text to quote. So the table describes what you will see rather than quoting it:

SymptomLikely causeFirst check
The probe output says the command was not foundcurl, wget or the shell is not in the imagedocker exec <container> curl --version; an error means the program is missing
The probe output shows a refused connectionThe app listens on another port, or only on an address the probe does not useCompare the port in test with the one the app logs at start-up; if the app binds IPv4 only, try 127.0.0.1 in place of localhost
Healthy on your laptop, unhealthy on the serverThe server’s cold start is slower than start_period allowsTime a cold start on the server and raise start_period
Flips between healthy and unhealthy under loadtimeout is too tight, or the health route does real workCompare each probe’s start and end time in docker inspect with timeout
Unhealthy whenever a third-party provider is downThe route treats an optional dependency as requiredMake that dependency report degraded, not failed
Unhealthy again after restarts, in a loopThe container may be getting killed for memorydocker inspect --format '{{.State.OOMKilled}}' <container>

The first move is the same every time. Read the probe’s own output in docker inspect, then run the same command yourself the way Docker runs it: docker exec <container> sh -c '<the test command>' for a CMD-SHELL test, since Docker runs that string in the container’s shell, or the program and its arguments without sh -c for a CMD test. Then check the exit code with echo $?. The app’s own logs come after that, and if they have grown large, how to read Docker logs and stop them filling the disk is the place to start.

Two rows lead elsewhere. A health route that fails when a payment or email provider is down turns every provider outage into an unhealthy container; an optional dependency belongs in a degraded status, which is what graceful degradation is about. For the memory row, .State.OOMKilled answers yes or no only for the time since the container last started, so after a restart also look for an oom event in docker events; reading an OOM kill is the next step when either says yes.

Startup order: depends_on with service_healthy

Docker Compose startup order is controlled by depends_on. The short form waits only for the other container to start. The long form with condition: service_healthy waits until that service’s health check passes, so the API no longer starts before the database accepts connections.

Docker’s startup order guide puts the problem in one line: “On startup, Compose does not wait until a container is “ready”, only until it’s running.” The long form fixes that per dependency:

  web:
    depends_on:
      db:
        condition: service_healthy
        restart: true
      migrate:
        condition: service_completed_successfully

There are three conditions. service_started is the same as the short form, service_healthy waits for the health check, and service_completed_successfully waits for a one-shot job, such as a migration, to finish with success. restart: true makes Compose restart web after it updates db, but only for “an explicit restart controlled by a Compose operation”; it “excludes automated restart by the container runtime after the container dies”. Setting required: false makes Compose only warn when that dependency is not started.

What this does not do is my reading, since neither Docker page covers it. depends_on orders docker compose up and the other Compose commands, not the Docker daemon bringing containers back after a reboot. And it does nothing for a running app when the database drops an hour later, so the app still needs to retry its database connection with backoff.

Docker Compose restart policy: no, always, on-failure, unless-stopped

A Docker Compose restart policy has 4 values: no, always, on-failure and unless-stopped. Docker’s restart-policy page names two triggers, a container exiting and Docker restarting, and health status is not one of them, so an unhealthy container keeps running unless the app exits, a supervisor restarts it, or an orchestrator replaces it.

Docker’s restart policy page says restart policies exist “to control whether your containers start automatically when they exit, or when Docker restarts.” Here is what each one does, from that page, with my working rule in the last column:

PolicyRestarts whenAfter you stop it by handAfter the Docker daemon restartsUse it for (my working rule)
no (the default)NeverStays stoppedNot startedOne-shot jobs such as migrations
on-failure[:max-retries]It exits with a non-zero code, up to the optional retry capStays stoppedNot restartedWorkers that can crash-loop, capped at about 5
alwaysIt stops, for any reasonRestarted only when the daemon restarts or you restart itRestartedServices you want back even after a manual stop, once the daemon restarts
unless-stoppedIt stops, unless you stopped itStays stoppedRestarted, unless you had stopped itLong-running services such as the web app and the database

In a Compose file the value no is written in quotes, restart: "no", as Docker’s Compose reference shows it. Docker’s Compose Deploy Specification also describes deploy.restart_policy, a separate setting with its own conditions (none, on-failure, any); when it is not set, Compose uses the service’s restart field.

Two documented details catch people out. A policy takes effect only after the container “starts successfully”, which the page defines as up for at least 10 seconds with Docker monitoring it. And if you stop a container yourself, “the restart policy is ignored until the Docker daemon restarts or the container is manually restarted”.

That page, read on 2026-09-30, never mentions health status. On plain Docker or Compose, an unhealthy container keeps running until something else acts on it. There are three ways to close that gap:

WayHow it worksWhat it costs (my reading)
Make the app exit when it cannot recoverA watchdog inside the app ends the process when its own self-check fails, and the restart policy starts it againCode in your app; the cheapest option, and the restart policy’s own rules still apply
Run a supervisor that restarts unhealthy containersdocker-autoheal watches labeled containers, or all of them, and restarts the unhealthy onesThe Docker socket mounted into a container, which hands that container control of the daemon
Move to an orchestratorSwarm may reschedule unhealthy service containers; Kubernetes restarts a container whose liveness probe failsA new platform to learn and run

docker-autoheal describes itself as a tool to “Monitor and restart unhealthy docker containers”, MIT licensed; it watches containers labeled autoheal=true, or every container when AUTOHEAL_CONTAINER_LABEL=all, and its Unix-socket example mounts /var/run/docker.sock. Weigh that mount against Docker’s security page: the daemon “requires root privileges unless you opt-in to Rootless mode”, and “only trusted users should be allowed to control your Docker daemon”. Docker’s Swarm docs say its scheduler “may reschedule your running service containers at any time if they become unhealthy or unreachable”. Kubernetes liveness probes go further: if the probe command fails, “the kubelet kills the container and restarts it”. Whichever you pick, a restart is not a fix. Alert on it and find the cause.

The open-source iwiki-mcp project hit this gap in September 2026 and wrote it up in its GitHub issue #93. Its Compose file had a health check and restart: unless-stopped; the container reported Up 2 hours (unhealthy) and kept answering 502 while nothing restarted it, for roughly two hours on 2026-09-20. The maintainer noted that the cause was external, the database it depended on was down, and that a health-driven restart with no cap “would have added container flapping on top of that”. Their fix, merged as iwiki-mcp#101, was a watchdog timer that restarts the service owning the container only after 3 consecutive unhealthy probes and at most 3 times an hour. In their check on the host, the container went from unhealthy to healthy again in three and a half minutes with nobody touching it. The lesson I take from it: a health status is a label until something acts on it, so decide what acts, and how often it may act, before the night it matters.

How to check your own app

Container recovery is proven with 5 tests on staging: a cold start waits for the database, stopping the database changes the web status as designed, a frozen process turns unhealthy in the expected time, a killed process is restarted by the policy, and services return after a daemon restart.

Run them on a staging host with the same Compose file as production, never on production itself. Write down the Docker Engine version from docker version with each result, since behavior can change between versions.

  1. 01 Cold start. From nothing, run docker compose up -d, then docker compose ps every few seconds. Pass: db shows starting, then healthy; web is created only after db is healthy and then turns healthy itself; every long-running service shows a health status, and none has a blank one. Evidence: the ps output.
  2. 02 Stop the database with docker compose stop db. Pass: the health route answers with an error status because the database is required, and web turns unhealthy within about interval times retries, around a minute with the example's values. An optional provider being down should show as degraded with a success status and leave web healthy. Evidence: the docker inspect health output and the timing. Start db again afterwards.
  3. 03 Freeze the app without stopping it: docker kill --signal=SIGSTOP on the web container pauses its main process. Pass: web turns unhealthy in about interval times retries, because a probe that gets no answer runs into its timeout and counts as a failure. Then wait five minutes and write down whether anything restarted it. Resume with --signal=SIGCONT. If the image starts the app through a shell-form CMD, the signal reaches the shell, not the app, and the app keeps answering. Evidence: the timing and whether a restart happened.
  4. 04 Crash the main process. Wait until web has been up for more than 10 seconds, read its PID with docker inspect --format '{{.State.Pid}}', and end that PID from the host with kill -9, as root. Pass: the restart policy starts the container again and its RestartCount in docker inspect rises by one. Do not use docker stop or docker kill for this: Docker ignores the policy for a container you stopped yourself, and its docs do not say whether docker kill counts. Evidence: the count before and after.
  5. 05 Restart the Docker daemon, or reboot the staging host. Pass: every long-running service under always or unless-stopped comes back on its own. The web app may start before the database, since depends_on ordering belongs to Compose and not to the daemon, so the pass is web turning healthy once db is healthy. Evidence: the ps output after the restart, with the date.

In the Production Hardening Sprint, deliverable 6.9, Health endpoint, is verified this way: check healthy and degraded responses and confirm sensitive details are not public.

Where the sprint fits

No deliverable is named “Docker health checks”. Deliverable 6.9 provides a safe health endpoint reporting the readiness of required services without exposing secrets, and deliverable 8.2 monitors the production URL and health endpoint with outage alerts. The app’s current framework and hosting setup are the starting point; components are refactored or replaced where the production work requires it. Post-handover support is 14 calendar days of fixes for defects in the delivered sprint work. Every deliverable, with how each is verified, is in the published scope.

Common questions about container health and restarts

How do I disable healthcheck in a Docker Compose file?

Set disable: true under the service’s healthcheck, or set its test to ["NONE"]. Either one turns off a health check the image defines, which helps when the image’s built-in check is wrong for your setup.

How do I restart a Docker Compose service?

Run docker compose restart <service>. It restarts the service’s containers as they are, and changes you made to the Compose file are not reflected; to apply those, run docker compose up -d <service>, which stops and recreates the container when its configuration or image changed.

How do I restart all running containers using docker compose?

Run docker compose restart with no service name; Docker’s reference says it “Restarts all stopped and running services”. The same limit applies, so after editing the Compose file use docker compose up -d instead, which picks up the changes.

What is an HTTP health check?

An HTTP health check is a request to a route that answers with a success status when the app can serve and an error status when it cannot. In Docker the probe makes that request from inside the container, and curl -f turns a response code of 400 or greater into a failed probe. What the route itself should test belongs to the health endpoint’s design, not to Docker.