A health check endpoint should return 200 only when the app can serve a real request, and report degraded, not down, when an optional provider such as Stripe or Auth0 fails. So what is a health check endpoint? One unauthenticated route, usually /health or /healthz, that a load balancer or uptime monitor polls, returning a status and nothing secret.

What is a health check endpoint, and what should it report

A health check endpoint is an unauthenticated HTTP route, usually /health or /healthz, that reports whether the app can serve requests right now. It answers one question for machines: a load balancer, a container platform or an uptime monitor polls it and acts on the status code, 200 for ready and 503 for not.

The route lives in the app itself, registered before or outside the authentication middleware, so a request with no session cookie still reaches it. Two kinds of caller read it. The first decides whether traffic should reach this instance at all: a load balancer, a container orchestrator, a managed platform. The second is the uptime monitor outside your infrastructure, which tells a person when the answer changes. It is one of the controls in hardening a SaaS application for resilience, and the one a monitor can read from outside.

Kubernetes splits the job into three probes, and two of its words, liveness and readiness, come back in App Engine, NestJS and Spring Boot later on this page. The table uses only what Kubernetes’ liveness, readiness and startup probes page says.

ProbeThe question it answersWhat a failure causes in Kubernetes
LivenessIs the container still making progress, or stuck, for example in a deadlock?After more failures than the configured tolerance, the kubelet restarts the container
ReadinessIs the container ready to accept traffic?The Pod’s IP address is removed from the EndpointSlices of every Service that matches it, so Services stop routing to it
StartupHas the application inside the container finished starting?The kubelet kills the container and its restart policy applies; liveness and readiness do not run until startup succeeds

The same Kubernetes page draws the line this article builds on: when an app has a strict dependency on back-end services, the liveness probe passes when the app itself is healthy, and the readiness probe additionally checks that each required back-end service is available. My working rule: a small app on one platform usually needs one readiness-style route and nothing more.

Two neighbors are easy to confuse with it. A status page is for people, and whether a small SaaS needs a status page is its own decision. The monitor is the thing that calls the route: the guide to uptime monitoring for founders tells you to point it at a route that can fail meaningfully, and this page is how to build that route.

In the Production Hardening Sprint this endpoint is deliverable 6.9, and the reason the published scope gives for it is that monitoring needs a reliable signal for whether the application can serve requests.

What goes wrong without it

Three failures are already covered in the uptime article’s section “Point the monitor at a route that can fail meaningfully”: the route that always answers ok, the route that checks every dependency, and the route that turns into load of its own. The table adds three more.

What the endpoint doesWhat you see in productionThe fix
Treats Stripe or the AI provider as requiredOne provider’s blip fails the check on every instance at the same moment, and what happens next depends on the platformReport the provider as degraded; fail only on what a request cannot work without
Sits behind the loginThe monitor gets a 401 or a redirect to the sign-in page, and a monitor that accepts any 200 or redirect reports the app as upRegister the route outside the authentication middleware
Prints everything it knowsFramework and dependency versions, hostnames, the environment name, or a raw connection error with the database URL in it, readable by anyoneReturn the status only; move detail to a second, protected route

The first row plays out differently per platform. Kubernetes stops routing to a Pod that fails its readiness probe, and App Engine flexible leaves an instance that fails its readiness check out of the pool of available instances. On those two, in my reading, a Stripe outage that fails readiness everywhere can take the whole app out of rotation while the app itself is fine. AWS does the opposite: when every target in a group is unhealthy at once, the load balancer fails open and sends traffic to all of them anyway. Azure App Service by default excludes no more than half of an app’s instances at a time and none when all are unhealthy, but it replaces an instance that stays unhealthy for one hour; on the Free and Shared plans, which cannot scale out, unhealthy instances are not replaced automatically. Whichever platform you are on, the health signal is then false: it says down when customers can still sign in and use most of the app, or it gets ignored. That last part is my reading, not a platform rule.

The second row is the quiet one. In my June and July 2026 audits, an ops SaaS had one page that reported datastore health, and it sat behind the login, so no uptime monitor could reach it. In my reading, a health check nobody outside can reach is a health check nobody runs.

The third row matters because the route is public by design: whatever it prints, anyone can read. For a sense of where operations work stands in the apps I audit: across the 21 third-party apps I audited in June and July 2026, the Deployment & Operations pillar averages 37.0 out of 100, third of 12 pillars counting from the weakest. Those 21 are a set I chose, 11 public vibe-coded apps audited across all 12 pillars plus a held-out group of 10 more audited blind, so the number describes them and is not a rate for every AI-built app. Whether your own route has any of these three problems is what the six checks at the end of this page decide.

How to do it: the checks, the response, and the platform settings

The work is a route with three possible answers, a short list of what it probes, and one setting per platform. The platform settings below follow each platform’s own documentation as published on 4 October 2026; they are not a test I ran.

Start in a browser, before any code. Open the route in a private window. If it asks you to log in, returns a page of HTML, or prints more than a status, there is work to do.

The handler below is a Next.js route handlers example, and apart from the connection() lines it would work in any framework that answers a request with a Response. It extends the readiness handler in the uptime article with the degraded state. The connection() call tells Next.js to wait for a real request instead of prerendering the route, which matters because a prerendered health answer would be frozen at build time.

// app/api/health/route.ts
import { connection } from 'next/server';
import { db } from '@/lib/db';                       // the pool the app already uses
import { recentProviderErrors } from '@/lib/metrics'; // error counters the app already keeps

const withTimeout = <T,>(p: Promise<T>, ms: number) =>
  Promise.race([p, new Promise<never>((_, reject) => setTimeout(() => reject(new Error('timeout')), ms))]);

let last: { at: number; code: number; state: string } | null = null;

export async function GET() {
  await connection(); // answer each real request; never prerender this route
  if (!last || Date.now() - last.at > 5_000) { // cache a few seconds: my working rule
    let code = 200, state = 'ok';
    try { await withTimeout(db.query('SELECT 1'), 1_000); } catch { code = 503; state = 'unhealthy'; }
    if (code === 200 && recentProviderErrors(['payments', 'auth', 'ai', 'email']) > 0) state = 'degraded';
    last = { at: Date.now(), code, state };
  }
  return Response.json({ status: last.state }, {
    status: last.code,
    headers: { 'Cache-Control': 'no-store' },
  });
}

Health check design checklist: what to probe, and what to leave out

A health check design checklist has 8 items: check only what a request needs, give every dependency check a short timeout, separate liveness from readiness, mark optional providers as degraded, cache the result for a few seconds, return a status code machines can act on, expose no detail publicly, and keep the route off the authentication middleware.

This is my working list, one line each.

  • Check only what a request needs to succeed.
  • Give every dependency check a short timeout.
  • Keep liveness separate from readiness.
  • Expose no detail on the public route.
  • Keep the route outside authentication.
  • Mark optional providers as degraded, never unhealthy.
  • Cache the result for a few seconds.
  • Return status codes a load balancer acts on.

The first five restate the advice in the uptime article’s section “Point the monitor at a route that can fail meaningfully”, and the reasons are there. Knowing how to set a timeout on fetch covers the outbound HTTP calls; give the database query the same kind of deadline.

The last three are this page’s own. Optional providers get the degraded state because their failure is real but survivable; the next section sets out the three responses. The cache keeps polling cheap: every AWS load balancer node checks each target, an App Engine readiness check and an uptime monitor poll on their own clocks, and my working rule of a few seconds of cache turns all of that into one database query per window instead of one per poll. Status codes matter because each platform has its own rule for healthy: AWS target groups accept only 200 unless you change the success codes, Azure App Service accepts anything from 200 to 299, and App Engine counts a 200 OK as success. A 204 from your route passes on Azure and fails on AWS’s default.

Healthy, degraded, unhealthy: the three responses and their status codes

A health response has three states. Healthy returns 200. Unhealthy returns 503 because a required dependency, usually the database, cannot be reached, so the instance should leave rotation. Degraded returns 200 with a status of degraded because an optional provider is failing, so the app stays up and a human is told.

StateHTTP statusBodyWhat the load balancer doesWhat the monitor does
Healthy200{"status":"ok"}Keeps the instance in rotationReports up
Degraded200{"status":"degraded"}Keeps the instance in rotationAlerts on the body, if it checks for a keyword
Unhealthy503{"status":"unhealthy"}Takes the instance out once its failure count is reached, within each platform’s limitsAlerts on the status code

Degraded is a 200 by my design rule: the instance can still serve most requests, so pulling it from rotation would turn a partial problem into a full one. The monitor still sees it, because a keyword check on the body catches the word; how that check works is in the uptime article’s section “Use a keyword check so a cached page cannot pass”. What the user sees while a provider is down, the fallback screen or the queued email, is a separate design question, and answering it starts with what graceful degradation is.

You do not have to invent the vocabulary. The IETF health check response format draft uses pass, warn and fail, where warn means healthy with some concerns and travels with a 2xx or 3xx code, the same place degraded sits here. The page itself marks it an expired individual Internet-Draft from October 2021, not endorsed by the IETF, so treat it as an option to borrow from rather than a standard. Whichever words you pick, send Cache-Control: no-store with every answer; the uptime article gives the reason.

Checking the services you depend on: the Auth0 health check API, Stripe and the database

A dependency check asks a required service the cheapest question that proves it answers: the database gets a SELECT 1 with a short timeout. Optional providers such as Auth0 and Stripe are read from their status pages or from the app’s own recent calls, and a failure there marks the app degraded, never unhealthy.

DependencyRequired or optionalThe cheapest honest checkTimeoutEffect on status
DatabaseRequiredSELECT 1 through the same pool the app usesShort, inside the prober’s own timeoutUnhealthy on failure
Cache or queueRequired only if requests fail without itOne ping through the app’s clientShortUnhealthy if required, else degraded
Auth provider (Auth0 and similar)Optional for users already signed in, required for new sign-insThe app’s recent sign-in errorsNo live callDegraded
Payment provider (Stripe)OptionalThe app’s recent payment API errorsNo live callDegraded
AI providerOptionalThe app’s recent model call errorsNo live callDegraded
Email providerOptionalThe app’s recent send errorsNo live callDegraded

NestJS’s TypeORM indicator makes the same choice for the database row: it runs a SELECT 1 behind the scenes.

For Auth0, no health endpoint is stated in Auth0’s docs: neither its Check Auth0 Status page nor its Monitor Auth0 page names one (as read on 4 October 2026). What Auth0’s guide to checking its status offers instead is the public cloud status page, where you pick your Region to see Core Services and Supporting Services, and an Atom or RSS feed “to get status updates that affect your tenant”. Auth0’s monitoring docs also mention watching it with “any tool that supports synthetic transactions”. In my reading, that means a scripted sign-in on a schedule, which is a job for the monitor rather than for this route.

For Stripe, no health or ping endpoint is stated in Stripe’s docs either (docs home and the health alerts page, checked the same day). Stripe does document health alerts: all Stripe users receive alerts for 402 errors from partner outages and 500 errors from Stripe outages, by email and in Workbench, and webhook delivery of those alerts needs a Growth, Premium or Enterprise Support plan. Stripe’s status page is there for a person to read.

Either way, my rule is that a live call to a payment provider on every poll is the wrong check. I read the outcome of recent real calls instead, from an error counter the app already has. A blip at the provider then shows up as degraded on every instance, which is true, instead of unhealthy on every instance, which on Kubernetes or App Engine can take you offline for something your app could have ridden out.

A provider’s blip must never fail a readiness probe, because every instance fails it at the same moment.

NestJS Terminus: the health module for a Nest app

NestJS Terminus is the health check module in NestJS’s own documentation, installed as @nestjs/terminus. It gives a controller one decorated route that runs a list of health indicators, such as a database ping, an HTTP ping, memory and disk, and returns 200 or 503 with a JSON body listing each indicator’s result.

Install the nestjs/terminus package with npm install --save @nestjs/terminus, add TerminusModule to a health module, and write a controller whose @Get() route carries @HealthCheck() and calls HealthCheckService.check(). NestJS’s Terminus recipe lists ten built-in indicators: HttpHealthIndicator, TypeOrmHealthIndicator, MongooseHealthIndicator, SequelizeHealthIndicator, MikroOrmHealthIndicator, PrismaHealthIndicator, MicroserviceHealthIndicator, GRPCHealthIndicator, MemoryHealthIndicator and DiskHealthIndicator. For anything else, such as Redis or a provider’s error counter, you write a custom indicator with HealthIndicatorService.

Terminus has a degraded state. An indicator can return degraded(), the overall status becomes 'degraded', and the HTTP status stays 200; any indicator that is 'down' makes the status 'error'. That arrived in version 12.0.0, released on 31 August 2026, so an older install has only up and down. One warning from the docs matters for custom indicators: an indicator built with up(), down() and degraded() that throws aborts the whole check with a 500 instead of a 503, so wrap anything that can throw in attempt().

The example below follows the TypeORM example in NestJS’s docs, with the timeout and cache it documents.

// Follows NestJS's Terminus recipe: a TypeORM ping with withTimeout() and cacheFor()
import { Controller, Get } from '@nestjs/common';
import { HealthCheck, HealthCheckService, TypeOrmHealthIndicator } from '@nestjs/terminus';

@Controller('health')
export class HealthController {
  constructor(private health: HealthCheckService, private db: TypeOrmHealthIndicator) {}

  @Get()
  @HealthCheck()
  check() {
    return this.health.check([() => this.db.pingCheck('database').withTimeout(1500).cacheFor(5000)]);
  }
}

The default body names every indicator under info, error and details. That is fine on an internal port; for the public route, my rule is to return the overall status only. ASP.NET Core has the same three answers built in: its health checks middleware returns 200 for Healthy and Degraded and 503 for Unhealthy by default. Spring Boot’s Actuator health endpoint has no degraded state among its built-in statuses; it maps DOWN and OUT_OF_SERVICE to 503, leaves UP at 200, and offers liveness and readiness groups.

A health endpoint for load balancer checks: AWS, App Service health check and the App Engine readiness check

A health endpoint for load balancer checks must be fast, unauthenticated and honest, because the platform takes an instance out, or restarts it, after a set number of failed polls. AWS target groups, Azure App Service Health check and App Engine readiness checks all work this way: know the path, the polling interval and the failure count.

PlatformWhere the setting isPathInterval and thresholds (defaults)What happens on failureDocs checked on
AWS Application Load BalancerTarget group health check settings/ unless you change itEvery 30 seconds for instance or IP targets, 5-second timeout, unhealthy after 2 failures in a row, healthy again after 5 successes, success code 200The target is taken out of service; if every target is unhealthy, the load balancer fails open and routes to all of them2026-10-04
Azure App ServicePortal: your app, then Monitoring, then Health checkNo default; you set it, for example /healthA ping every minute; removed after 10 failed pings by default (configurable from 2 to 10); healthy means 200 to 299Removed from the load balancer, no more than half of the instances by default and none if all fail; replaced after one hour unhealthy on Basic tier or higher2026-10-04
App Engine flexible, readinessreadiness_check in app.yamlNone by default; set one so the check reaches your containerEvery 5 seconds, 4-second timeout, 2 failures to fail, 2 successes to pass; success is 200 OKThe instance is not added to the pool of available instances; the all-instances case is not on Google’s app.yaml reference or instance management pages (checked 2026-10-04)2026-10-04
App Engine flexible, livenessliveness_check in app.yamlNone by defaultEvery 30 seconds, 4-second timeout, 4 failures to fail, 300-second initial delayUnhealthy instances are restarted2026-10-04

On AWS, the settings live on the target group, and AWS target group health checks lists every field. Change the default path of / to your health route; a homepage that renders without the database tells the load balancer nothing.

The App Service health check needs the most care about tier and size. Azure App Service Health check asks for the Basic tier or higher and at least two active replicas to get the full benefit, does not follow 302 redirects, and requires the path to allow anonymous access if you use your own authentication system. An app that redirects its default domain to a custom domain gets a 301 back, and Azure marks that worker unhealthy.

The App Engine readiness check never reaches your code until you give it a path: by default, health check requests are not forwarded to your application container. Set path under readiness_check in App Engine flexible’s app.yaml reference to your route. My rule for liveness_check is a second path that skips the database, so a database outage never gets instances restarted.

Containers run the same idea one level down, as a Docker Compose health check. On Vercel, Netlify or a single VPS, in my reading, no load balancer polls the route, so its only readers are the uptime monitor and the deploy script that checks a release. The endpoint is still worth having for those two.

What the endpoint must never expose

A public health endpoint should return the status and nothing else. It must never include five things: environment variables, connection strings or internal addresses, framework and dependency versions, stack traces or raw error messages, and build paths or commit hashes. Detail belongs on a second route behind authentication or on an internal port.

  • Environment variables and the environment name: they name the services you use and sometimes hold keys.
  • Connection strings, hostnames and internal IP addresses: they map your network for anyone who wants to probe it.
  • Framework and dependency versions: a version number is a lookup key for known vulnerabilities.
  • Stack traces and raw error messages: a database error can carry the connection URL, user name and file paths.
  • Build paths and commit hashes: they describe your build machine and let anyone match a release to its source.

The version item is the one with numbers behind it. 9 of the 26 apps I audited in June and July 2026 ran a framework version with a publicly known, reachable RCE or auth bypass, and the fix was often a one-line version bump. The 26 are the 21 third-party apps from earlier on this page plus my own 5 production apps, which went through the same audit, and because I picked them, the 9 says what turned up in that group, not how common the problem is everywhere. A health route that prints its framework version hands that lookup to anyone.

On commit hashes, my rule is softer: a short hash is tolerable for checking which release is live, if the repository is private. Two routes solve the rest. The public one returns the status only; a detailed one, with each dependency’s result and timing, sits behind authentication or is bound to an internal port the load balancer cannot route to.

Work that runs outside requests, such as workers and queues, reports through its own check, with queue depth and the dead-letter count, not through the web route. A stuck queue does not stop a page from loading, and the job side has its own rules, starting with an idempotency key for safe retries.

How to verify it

A health endpoint is verified with 6 checks in staging: it returns 200 when healthy, 503 when the database is unreachable, 200 with degraded when an optional provider is blocked, nothing sensitive in any of the three bodies, a response in well under a second, and an alert from the monitor that polls it.

Run them in staging, in this order. Check 3 is a test that health returns a degraded status, and it is the one that proves the optional-provider rule works.

  1. 01 Healthy: call the route with no session cookie. Pass: 200 and the ok body. Evidence: the response with its timestamp.
  2. 02 Unhealthy: stop the staging database or block its port from the app. Revoking its credentials may not be enough, in my reading, because connections already open in the pool can stay signed in. Pass: a 503 within the check timeout plus the cache window, not a hang and not a 200. Evidence: the response and the time it took.
  3. 03 Degraded: block the optional provider with a wrong API key in staging or a blocked host, then send one request that calls it so the error counter records the failure. Pass: a 200 with the degraded status, and the instance stays in rotation. Evidence: the response and the platform view showing the instance still in service.
  4. 04 Exposure: read all three bodies and the response headers. Pass: no version, hostname, path, variable or error text anywhere. Evidence: the three responses saved as text.
  5. 05 Cost: time the route and watch the database. Pass: an answer in well under a second, and one cheap query per instance per cache window, not one per poll. Evidence: the timings and the query count.
  6. 06 The reader reacts: where a load balancer or platform polls the route, it marks the instance unhealthy during check 2; on a single VPS or a serverless host nothing polls it, so skip that half. The uptime monitor alerts on checks 2 and 3 and reports recovery after each. Evidence: the platform health event and the monitor alerts.

Wait for each platform’s own failure count before reading check 6. Azure’s default of 10 failed one-minute pings means about ten minutes before an instance shows as unhealthy, while AWS’s defaults of 2 failures 30 seconds apart take about a minute. Keep the three responses with their status codes and timestamps, the platform’s health event and the monitor’s alert in one place; together they show the route tells the truth in all three states.

When we run this on a sprint, we verify deliverable 6.9 by checking healthy and degraded responses and confirming sensitive details are not public, and deliverable 8.2 by triggering a controlled check failure and verifying notification and recovery reporting.

Simulating a whole provider outage is the wider drill, and knowing what chaos testing is helps you plan one that goes past these six checks.

Where the sprint does this

Deliverable 6.9 provides a safe health endpoint reporting the readiness of required services without exposing secrets, verified as described under How to verify it. Deliverable 8.2 monitors the production URL and health endpoint with outage alerts, and 6.10 simulates AI or payment-provider failure in staging and verifies the expected recovery behavior. The result for each goes into the production readiness report, which accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. Hosting, paid tools, and API usage remain in your accounts. The full list is in the reliability checks in the published scope.

Common questions about health checks

What is a Kubernetes health check endpoint?

A Kubernetes health check endpoint is the HTTP path a liveness, readiness or startup probe calls on a container. A failed liveness probe gets the container restarted, a failed readiness probe takes the Pod out of its Services’ endpoints, and a failed startup probe gets the container killed and handled by its restart policy. A common pattern is to point liveness at the same low-cost route as readiness, with a higher failure threshold.

Why is the elb health check failing?

An ELB health check fails when the target answers with a code outside its success codes, does not answer within the timeout, or fails for a reason AWS reports as Target.FailedHealthChecks or Elb.InternalError. The console shows “Health checks failed with these codes: [code]” for the first case and “Request timed out” for the second. My addition, the two causes I would check first: a route behind login, which answers 401 or a redirect while the default success code is 200, and the default path of / left pointing at a page that redirects.

What is a health check software?

Health check software means two unrelated things. For a web app, it is an uptime monitor that polls your health endpoint from outside and alerts when the answer changes; the uptime monitoring for founders guide covers choosing one. For a computer, it is a desktop tool that checks a PC’s hardware and updates, which has nothing to do with web apps. That split is my reading of how the term is used.

What is the difference between Azure Health and Azure Monitor?

Azure Service Health tells you about Azure’s own problems: outages, planned maintenance and advisories for the Azure services and regions you use. Azure Monitor is Microsoft’s observability service for collecting, analyzing and acting on telemetry from your own resources and apps. App Service Health check is a third, separate feature: it pings your app’s health path every minute, and its status metric can be viewed and alerted on through Azure Monitor.