Production hardening is the engineering work that prepares an app you already have for strangers to rely on: who can reach the data, what happens when a payment retries, how it behaves under load, how fast it recovers from a bad release. For a web app that is 123 checks in 13 areas, from access control to handover.
What is production hardening? The 13 areas it covers, and what it does not
Production hardening for a web app covers 13 engineering areas: authentication and authorization; secrets, keys and configuration; input, output and application security; database and data; payments and billing; error handling and reliability; environments, CI/CD and deployment; logging, monitoring and alerting; performance; code health and AI guardrails; email; privacy and compliance; and handover. Together they hold 123 checks.
Every one of those checks, each with the step that proves it, adds up to the production readiness checklist, all 123 checks with how to verify each. This page is the map above that list. If you already know your situation (about to launch, raising a round, handing an app over, or a customer just asked for SOC 2), skip to “Where to start, by situation” further down: it names the first page to open.
The question as people type it, “what is hardening in IT”, has a one-line answer in NIST’s glossary definition of hardening: “A process intended to eliminate a means of attack by patching vulnerabilities and turning off nonessential services,” with NIST SP 800-152 as its source. That is the server and operating-system sense. Puppet’s guide to system hardening runs from server and OS hardening to physical measures, TuxCare’s guide is written for Linux, and HashiCorp’s “Production hardening” page covers its own Vault product. This page answers for a different system: one web application on managed hosting, built by a small team, often with an AI tool.
Hardening in software development, for one web app, reaches past the server. The attack surface is the data model, the API, the payment path and the release path as much as the machine underneath, so the 13 areas below cover configuration and a good deal besides. Prod hardening, the short form, means the same work.
The table below works as a software hardening checklist at the level of areas: one row per area, with a few of the checks it holds and how many there are. There is nothing to download; the table is the template.
| Area | What it covers | Checks |
|---|---|---|
| 1. Authentication and authorization | Session and token controls, server-side authorization, data isolation across every table, two-factor authentication for owner and admin accounts | 11 |
| 2. Secrets, keys and configuration | Frontend secret removal, repository history scan and rotation, environment configuration separation, metered-service spend alerts | 9 |
| 3. Input, output and application security | Server-side request validation, public-endpoint rate limits, an OWASP Top 10 review, AI feature hardening | 12 |
| 4. Database and data | Schema integrity review, connection pooling, version-controlled migrations, verified backups and restore | 14 |
| 5. Payments and billing | Webhook signature verification, idempotent webhook handling, the complete billing event lifecycle, server-side entitlements | 10 |
| 6. Error handling and reliability | Consistent exception handling, error tracking, external-call timeouts, a health endpoint | 10 |
| 7. Environments, CI/CD and deployment | Three isolated environments, a protected production branch, pull-request CI, a tested rollback | 13 |
| 8. Logging, monitoring and alerting | Structured, sanitized logs, uptime monitoring, operational threshold alerts, actionable alert routing | 8 |
| 9. Performance | Appropriate caching, a load test and bottleneck fixes, a written capacity statement, an accessibility baseline | 7 |
| 10. Code health and AI development guardrails | Critical-path smoke tests, auth and billing unit tests, AI repository guardrails, an open-source license audit | 11 |
| 11. Email and communications | Sending-domain authentication, transactional delivery tracking, bounce and complaint handling | 3 |
| 12. Privacy and compliance foundations | A personal-data inventory, data export and deletion paths, a SOC 2-style controls checklist, a subprocessor list | 8 |
| 13. Handover | A production readiness report, a system architecture diagram, operating runbooks, a technical due diligence pack | 7 |
| Total | 13 areas | 123 |
Why AI-built apps need it
AI-built apps need production hardening because a working demo does not exercise access rules, failure handling or the release path. Audits of 21 third-party apps confirmed 958 findings, about 46 per app, 58 of them critical. In 7 of the 21, a logged-in user could read or write another customer’s data, and 17 had no error tracking or alerting.
Those numbers come from the 26 real applications I audited between June and July 2026, in three cohorts: the 21 third-party apps above, plus five of my own. They are a selected set of audited apps, not a random sample, so none of the counts below is a rate for all AI-built apps. Every count rests on human-verified findings, verified against the code, not pattern-matched. The statistics page explains how the 26 apps were selected and scored and carries the full numbers; the table here keeps the rows that map to an area.
| What failed | How many of the audited apps | Area |
|---|---|---|
| Confirmed cross-user or cross-tenant authorization failures | 7 of the 21 third-party apps | 1 |
| Row-level security gaps | 9 of the 21 third-party apps | 1 |
| Unauthenticated endpoints doing privileged work | 11 of the 21 third-party apps | 1 |
| A real secret shipped | 6 of the 21 third-party apps | 2 |
| A confirmed denial-of-wallet path: a stranger or free account can burn the owner’s paid AI or compute bill without limit | 12 of the 14 third-party apps with an AI feature | 2 and 3 |
| No rate limiting on the most expensive endpoint | 13 of the 21 third-party apps | 3 |
| When a user hits an error, nothing records it | 17 of the 21 third-party apps | 6 and 8 |
| No deploy gate: every push ships straight to production with nothing checking it first | At least 17 of the 21 third-party apps, and all 5 of my own | 7 |
| Zero working automated tests | At least 23 of all 26 apps | 10 |
In the same audits, readiness scores ran from 29 to 81 out of 100 across all 26, with a mean of 52.1 and a median of 51, and 22 of the 26 landed in the red band, 4 in amber and none in green.
The tests row hides the sharpest case. One point-of-sale app I audited had an 818-line test suite, green on every run, that never called the real sale-creation code; I tell it in full in a test suite that stayed green without touching checkout. The lesson I take from it: a green suite proves only what it calls, and a suite that never touches checkout says nothing about the path that takes money.
The recurring patterns behind these rows, sorted by what it takes an owner to notice them, are the 14 recurring patterns a demo never shows.
Production hardening for vibe-coded apps also depends on the builder, because each platform handles some of this for you and leaves the rest: what a passing check can’t see on a Lovable app, the pre-launch gates for a Bolt.new app, and what Replit secures and what your app still needs. Using one of these tools was never the mistake; the gaps sit in the parts a prompt rarely asks for.
What vibe coding is, and what it leaves for someone else to fix
Vibe coding is building software by describing it to an AI tool in plain language and accepting the code it writes, often without reading it. It gets a working demo fast. What it leaves for someone else is the part no demo exercises: access rules, retries, backups and a safe release path.
The name has a known author. Andrej Karpathy coined vibe coding in a post on X in February 2025: “There’s a new kind of coding I call ‘vibe coding’, where you fully give in to the vibes, embrace exponentials, and forget that the code even exists.” That post, the Andrej Karpathy vibe coding tweet, is still the one people look up when they ask where the term came from.
In practice, vibe coding an app runs as a loop: describe the screen, accept the diff, run it, describe the next change. People do it in Lovable, Bolt.new, Replit, v0, Base44, Cursor and Claude Code, among others. On vibe coding vs AI-assisted development, my own reading (not a standard) is that the line falls on whether a person reads and owns the code before it ships, whichever tool wrote it.
What that loop leaves behind is the failure table above, row by row. If your next question is can AI write production-ready code, that page answers it. Eight vibe coding examples read at the code level show the same gaps in real apps, how to secure a vibe-coded app gives the order to close them in, and what vibe coding security covers takes the risk question on its own.
The four questions every hardening pass answers
As a production hardening checklist, the 123 checks answer four questions, and the hardening process in software, for one app, means answering them in this order.
Who can reach the data
Areas 1 and 2 answer this, with area 12 writing down what personal data exists in the first place. The proof for server-side authorization is to call protected actions directly as unauthorized and as underprivileged users and confirm they are rejected. Data isolation across every table is proved with read and write tests as anonymous users, different roles and separate tenants. Frontend secret removal is proved by scanning the source and built assets, and watching browser traffic, for private credentials. The personal-data inventory is checked by tracing a sample user through storage, logs and providers and comparing the result with the list. The first four rows of the failure table are this question answered badly.
What happens when a payment retries
Area 5 answers this, with area 6 for the retry itself. The app verifies the payment provider’s signatures before processing webhook payloads, because forged events can wrongly grant paid access. Every webhook handler is made safe to repeat, and the proof is to replay events and deliver them concurrently, then confirm one intended business effect. Paid-plan permissions are enforced in backend actions, tested with free, expired and valid paid accounts. Transient failures get bounded retries with backoff, only where repeating the operation is safe. The failure idempotent handling prevents, in plain words: a repeated payment event can duplicate the fulfillment or corrupt the subscription state if processed unsafely.
How the app behaves under load
Areas 3, 4, 6 and 9 answer this. Rate limits on public endpoints are proved by simulating bursts and checking they are enforced without blocking ordinary use. Connection pooling is tested with concurrent requests while recording connection usage and errors. Timeouts on external calls are proved with a slow dependency: the waits stay bounded and the failure state is a useful one. A load test reports the workload, duration, environment, concurrency, latency and error rate, before and after the fixes, and the written capacity statement links its claims to that evidence and labels untested projections as projections. The rate-limit and denial-of-wallet rows in the failure table are what an unanswered version of this question costs.
How fast it recovers from a failed release
Areas 7 and 8 answer this, with area 4 supplying the restore and the drill. Google’s Site Reliability Engineering book gives the reason to start at the release: “SRE has found that roughly 70% of outages are due to changes in a live system.” A protected production branch is proved by attempting a failing merge and watching the protection stop it. A tested rollback is proved by rolling back a test release and checking behavior and retained data. Uptime monitoring is proved by a controlled check failure that produces a notification and a recovery report, and each configured alert condition is triggered and its threshold and behavior recorded. Backups count only after a restore into an isolated environment, and the disaster recovery drill records its timing, the services recovered and the limits it hit. The deploy-gate and error-tracking rows in the failure table are this question left open.
Areas 10 to 13 (code health, email, privacy, handover) are what keep those four answers true after the next change; each has its own entry in “The 13 areas” below.
The 13 areas
Each area below gives security hardening examples for a web app rather than a server: what the area covers, a problem people search for in it, and the page that goes deeper. The areas stay the same whatever the stack. Hardening Node.js for production, a Python API or a Supabase app runs through the same 13.
1. Authentication and authorization
Area 1 decides who is who and what each person may touch. It covers session and token controls, so an old session stops working after logout or a password change; permissions enforced on every protected route, API endpoint and server action, because a hidden button does not stop a direct request; Row-Level Security or equivalent server-side scoping across all tables, with explicit rules for data meant to be public; and a second factor on every owner and admin account. The problem people search for here is broken access control. The authentication checklist covers the whole area.
2. Secrets, keys and configuration
Area 2 keeps credentials where only the server can use them. It removes private secrets from frontend code and downloadable bundles, scans Git history for committed secrets and rotates every exposed credential, separates development, staging and production variables, databases and keys, and puts spend or usage alerts on every metered API. API keys exposed on the frontend has its own page, and secrets management best practices covers the rest of the area.
3. Input, output and application security
Area 3 is how the app treats what strangers send it. Every data-writing endpoint gets explicit schemas and injection defenses; public forms, signup flows and resource-consuming endpoints get appropriate rate limits; the app is reviewed against the OWASP Top 10, whose current release is the OWASP Top 10 2025, with findings, fixes and evidence recorded by category; and every model-backed feature gets prompt-injection defenses, output validation, per-user limits and a timeout with a fallback. The web app security guide defines hardening in its application-security sense, and how to test for the OWASP Top 10 vulnerabilities covers the testing itself.
4. Database and data
Area 4 keeps the records correct and recoverable. It reviews types, nullability, foreign keys, uniqueness rules and cascade behavior; moves schema changes into ordered migrations committed to the repository; verifies automated backups, sets retention and performs a real restore test; and restores the whole application into a fresh environment to time a disaster recovery. The data consistency checklist for SaaS covers the area, and the database backup checklist for startups takes backups on their own.
5. Payments and billing
Area 5 makes billing state match what customers paid for. It covers signature verification on webhooks, handlers that are safe to repeat, handling for creation, updates, cancellation, payment failure, refunds and disputes (delayed or out-of-order events included), and test and live credentials isolated by environment. Stripe is the common case. The whole billing flow is a subject of its own, SaaS billing process best practices, and so is the webhook itself: webhooks security.
6. Error handling and reliability
Area 6 is what happens when something fails. Swallowed failures and empty catch blocks are replaced with deliberate recovery or useful error reporting; error tracking such as Sentry goes in with protected source maps, environment labels and alerts; AI, email, payment and other external calls get deliberate timeouts; and a safe health endpoint reports whether required services are ready without exposing secrets. The area as a whole is hardening SaaS applications for resilience, and the API side of it is API error handling best practices.
7. Environments, CI/CD and deployment
Area 7 is the path a change takes to customers. Development, staging and production get separate databases, credentials and configuration; the main branch is protected with required checks and controlled merge permissions; linting, type checks, builds and tests run on every pull request; and release rollback is documented and rehearsed, database changes included. DevOps for startups covers the release path, and GitHub branch protection takes the protected branch on its own.
8. Logging, monitoring and alerting
Area 8 makes problems visible before a customer reports them. Request logs are structured, with correlation IDs and appropriate user references, and exclude passwords, tokens and unnecessary personal data; the production URL and health endpoint are monitored with outage alerts; error spikes, latency, connection pressure and queue backlog raise alerts; and those alerts go to a named Slack or email destination with tuned thresholds. Logging and monitoring covers the area; if the app logs nothing useful yet, begin with how to do logging.
9. Performance
Area 9 measures capacity instead of guessing it. Repeated reads are cached where it is safe, with explicit invalidation and customer-data boundaries; a load test simulates concurrent users, finds the first bottlenecks, fixes them and reruns; a written statement documents measured concurrent capacity and what five times that workload would need; and the core flows reach WCAG AA for keyboard navigation, contrast, labels and focus order. Start with web performance optimization, then a k6 load testing example.
10. Code health and AI development guardrails
Area 10 keeps the next change from undoing the rest. Smoke tests cover signup, login, the core product action and payment flows; unit tests guard authorization and payment logic, negative and edge cases included; CLAUDE.md, AGENTS.md, Cursor rules or equivalents describe protected patterns, with CI checks for the enforceable ones; and every dependency’s license is inventoried, with copyleft or unlicensed packages flagged. The tests row of the failure table is this area left empty. The area as a whole is in engineering standards for AI-assisted teams, and the smoke tests have their own page, how to write end-to-end smoke tests.
11. Email and communications
Area 11 makes sure account and transaction emails arrive. SPF, DKIM and DMARC are configured for the transactional sending domain, a dedicated transactional provider tracks delivery of account and transaction messages, and bounces and complaints are processed, with further sending suppressed where appropriate. How to improve email deliverability is the question for the area, and spoofing of your sending domain is a separate one: how to stop email spoofing.
12. Privacy and compliance foundations
Area 12 is the record a customer or reviewer asks for. It documents what personal data is stored, where, why and who can access it; adds authenticated export and deletion workflows; delivers a SOC 2-style controls checklist with supporting evidence, which is not a SOC 2 audit report; and publishes the list of vendors that process customer data. The area is how to manage SaaS data compliance, and the buyer side of it shows up in security questionnaire examples.
13. Handover
Area 13 turns the work into artifacts that let someone who was not there run and judge the app: a production readiness report with the result and evidence for every check, a system architecture diagram of components, data flows and authentication boundaries, operating runbooks for deployment, rollback, key rotation and restores, and a technical due diligence pack that bundles the report, the diagram, the data model, the security checklist and the capacity statement into one PDF. The area is what a developer handoff looks like, and the place to start is a runbook template.
Where to start, by situation
If you are looking for a production readiness consultant, or deciding whether to do this work yourself, start from your situation. The table sorts the situations and names the first page to open for each.
| Your situation | Open first | Why |
|---|---|---|
| About to take paying customers | the go live checklist | What has to be true on launch day, in order |
| More users are coming than the app has seen | a scaling readiness checklist for startups | Capacity, limits and the parts that break first under traffic |
| Raising a round | how to prepare for technical due diligence | What a reviewer looks for in AI-built code |
| Raising a round and want the investor’s view | what investors look for in code | The questions behind an investor’s technical review |
| Delivering an app to a client | how to hand over an app to a client | The accounts, documents and access that change hands |
| Choosing between a fractional CTO, an agency or a fixed scope | fractional CTO vs agency | What each arrangement includes and when it fits |
| Buying diligence on an app | what a technical due diligence consultant delivers | What the report should contain |
| Wanting a code audit | what a code audit service should check | How to choose one and what to ask for |
| Wanting a person to review the code | source code review services | Who sells review of AI-written code |
| Wanting outside security testing | web application security testing services | The kinds of testing and what each finds |
| A penetration test on a small budget | affordable penetration testing | What a small-budget test can and cannot cover |
| A customer asked for SOC 2 | SOC 2 compliance consultants | Who does the readiness work and who does the audit |
| Wanting QA or test automation | software testing services | What outside testing includes |
| Wanting tools before services | web application testing tools | The tools a small team can run itself |
| Worried mainly about customer data | data security for a small SaaS | The data controls that matter first |
No page here ranks anyone as the best production hardening service. The hiring rows point to pages that say what each option includes and when it fits, from an agency or a fractional CTO to a fixed-scope security sprint for startups.
How to know you are done
Production hardening is done when every check has a recorded result, the evidence behind it and a date, in a record someone who was not there can read. A failed check stays on the list until it is fixed, and a check that does not apply says why. A clean scan or a green badge is not that record.
My rule for any hardening pass, whoever does it: one line per check, and the evidence is something a stranger can open, such as a test run, a screenshot, a query output or a log line. A reviewer will also ask how the system fits together and how to operate it, so the record travels with an architecture diagram and runbooks, and an investor’s reviewer will want the report and the diagram bundled as a due diligence pack; the readiness report and the due diligence pack between them cover both. The per-check list to work from is the production readiness checklist linked near the top of this page.
In the Production Hardening Sprint, that record is deliverable 13.1, the production readiness report: the result for every scope item, the work completed and its verification evidence. Deliverable 13.1 is verified this way: account for all 123 IDs, keep failures visible until resolved and explain genuine non-applicable items. Every deliverable on the scope page, from 1.1 to 13.7, states the same kind of verify line, so you can read every check with how it is verified for yourself.
Where the sprint does this
The Production Hardening Sprint is our software hardening service with a published scope. It covers one codebase, runs 10 working days, and includes 123 deliverables across the 13 areas above. At handover you receive the updated codebase, tests and deployment configuration, a readiness report and technical due diligence pack, instructions for releases, backups, recovery and incident response, guidance and automated checks for future AI-assisted changes, and a recorded 60-minute handover with a codebase tour. After handover come 14 calendar days of fixes for defects in the delivered sprint work and 30 calendar days of async access for questions about the handover and architecture. Where it stops: formal third-party certifications and independent audit opinions are separate from the sprint deliverables; legal advice and certification are separate services, while we implement and document the technical data-handling controls; building new product features or modules, completing unfinished core features or business workflows, and rebuilding core functionality that does not yet perform its intended job are separate work; and the app’s current framework and hosting setup are our starting point. Every line of it is on the published scope, all 123 deliverables.
What is a hardening sprint, and what we mean by a sprint
A sprint in software development is, in the Scrum Guide’s words, a fixed length event of one month or less, and every increment must be usable and meet the Definition of Done. A hardening sprint is an end-of-project sprint kept for the fixes and testing pushed there. The Scrum Guide names no such sprint type.
The Scrum Guide, in its November 2020 version, is strict on the second point: “Work cannot be considered part of an Increment unless it meets the Definition of Done.” TechTarget’s feature on the practice describes organizations that “push defect fixes” to the end of a project, along with regression, integration, end-to-end and third-party testing, and reports that those who favor an approach truer to the Agile Manifesto and the Scrum Guide call the hardening sprint “a Scrum anti-pattern.” That is the hardening sprint in Scrum as people search for it: something organizations add, which the guide itself never names.
The Production Hardening Sprint uses the word for something else. It runs 10 working days: days 1-2 to inspect and begin, days 3-7 to implement and test, and days 8-10 to verify and hand over, against 123 published deliverables.
Common questions about hardening an app
What does hardening mean in coding?
Hardening in coding means removing the ways code can be abused or can fail quietly: authorization checked on the server for every request, input validated on the server, an app that refuses to start when a required setting is missing, no secret in the browser bundle, and dependencies pinned to known versions. Each is a change in the code rather than a setting on a server.
What is a hardening guide?
A hardening guide is a written list of the settings and checks that close the known ways into one kind of system, each with how to confirm it. Server and operating-system guides list ports, services and patches. For a web app the list is the 13 areas on this page, and the checklist linked in the first section gives every check its proof.
Is hardening sprint allowed in Scrum?
Not within the framework the Scrum Guide describes: the guide names no hardening or stabilizing iteration, and it asks that every Increment be usable and meet the Definition of Done. In my reading, an iteration set aside only for stabilizing sits outside that model, because work that still needs stabilizing has not met the Definition of Done yet. Teams do run them; the guide gives them no place.
Is vibe coding really coding?
Yes, in what it produces: vibe coding generates real source code in real programming languages, which runs, deploys and can be read like any other code. The difference is whether anyone reads and owns that code before users depend on it.
If this checklist left you with more open items than you expected, the sprint below works through all of them in ten working days.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase