Do not rewrite it, and do not refactor anything in week 1. A legacy code takeover goes well when the first month runs in a fixed order: get the app running and deployable, pin its current behavior with tests in week 2, ship the first small changes in week 3, and only in week 4 decide what to clean up, replace or leave alone.
The legacy code takeover plan: four weeks, in order
A legacy code takeover plan runs 4 weeks in a fixed order: week 1 makes the app run, deploy and roll back in your hands, week 2 pins what it does today with tests, week 3 ships small real changes through that net, and week 4 makes the keep, fix or replace decision with evidence.
This plan is one piece of the engineering standards for AI-assisted teams I hold code to, written from the seat of the person who has just been handed the repository. “Legacy” here means code you are afraid to change, whatever its age, so an AI-built app written last quarter qualifies as much as a monolith nobody has touched in years.
| Week | Goal | What you do | What you do not do | Exit test | Artifact left behind |
|---|---|---|---|---|---|
| Day 0 | Access | Get the repository, hosting, database, error tracker, a staging or local database, and a named person who can answer business questions | Start work on borrowed passwords | You can log in to every system without asking anyone | Access list with an owner per account |
| Week 1 | Run, deploy, roll back | Run it locally, deploy it unchanged, roll it back, check backups, draw the system | Refactor, upgrade or “tidy” anything | A second machine goes from clone to running app using only your notes | Setup notes, one-page system drawing, first risk list |
| Week 2 | Pin behavior | Write characterization tests around the flows that make money or hold data, and run them in CI | Fix the wrong behavior the tests reveal | An intentional regression in one core flow fails the suite | Test list, with wrong behavior marked |
| Week 3 | Small real changes | Ship a few small tickets through the normal release path | Start cleanup projects | A real change shipped with no incident | README, change log |
| Week 4 | Decide | Weigh each module against the evidence and write the memo | Decide for the whole codebase at once | The memo is accepted or argued with | Decision memo |
The day 0 row is the incoming engineer’s half of a handover. The owner’s half, the accounts, keys and billing that have to change hands, is the subject of owning an app someone else built, and the checks an owner runs before hiring a developer to take over your project come before any of this.
Three rules hold all month. Every change goes through the normal release path, one at a time, so that when something breaks you know which change did it. Every surprise gets written down the day you meet it, with the date, while you still remember what you expected instead. And nothing gets fixed until you have reproduced it, because a fix for a bug you have only read about is a guess with a commit message.
The order is fixed for one reason: each week’s exit test is the safety net the next week stands on. Tests written before you can run the app pin the wrong thing. Changes shipped before the tests exist have no net. A keep or replace decision made in the first week is made on reading alone, which is the least reliable evidence you will ever have about this code.
How to work each week
Each week below ends on its exit test, a check that passes or fails, so you know whether to move on or stay.
Week 1: run it, deploy it, roll it back, and change nothing
Week 1 of a takeover has a single goal: you can run the app locally, deploy it and roll it back without the previous developer. Nothing is refactored. Every surprise goes into a notes file with the date, because those notes become the README the codebase never had.
When you have inherited a codebase nobody could understand, reading it front to back is the wrong first move and running it is the right one. The running app shows you which small part of the files actually carries the traffic, and that is the part worth reading first.
- 01 Clone the repository and run it locally from whatever instructions exist, writing down every step that was missing
- 02 Find every environment variable the app reads and where each value lives today
- 03 Identify the deploy path and do one deploy with no code change
- 04 Do one rollback, then confirm that normal deploys go live again afterwards
- 05 Confirm backups exist and restore one into a scratch database or a new project, never over the live database
- 06 Find the error tracker, or turn one on, and read its recent history
- 07 Draw the system on one page: entry points, data stores, third parties and scheduled jobs
- 08 List who and what has production access
The notes you write when you cannot run the project locally on the first try are the first draft of the setup guide, so keep them rough and complete rather than tidy.
Item 4 is where you learn what a rollback plan is by doing one, not by reading one. Check that deploys resume afterwards, because hosts behave differently here: on Vercel, for one, a rollback turns off auto-assignment of production domains, so new pushes to the production branch won’t go live automatically until you undo the rollback by promoting a different deployment.
A missing backup goes on the risk list this week, not into a fix, because this week changes nothing. Supabase, for one, automatically backs up Pro, Team and Enterprise Plan projects daily and recommends that free tier projects regularly export their data with the Supabase CLI db dump command. Its Restore to a New Project feature is “exclusive to users on paid plans and requires that physical backups are enabled for the source project”, so on the free tier the restore you test is that exported dump, loaded into a scratch database.
For the error tracker, my working rule is to read about the last 30 days, or whatever shorter history it keeps. If there was no tracker at all, that absence is the first line of the risk list.
The files will mislead you if you let them. In my June and July 2026 audits, a Q&A app’s code called a stored database function that none of its seven migrations defines, and nothing in the repository applied those migrations to anything. The migrations folder was not the schema’s source of truth, so nobody could rebuild or reason about the live database from the code. The lesson I take from it is that the files are a claim about the system, and the running app and the live database are the evidence.
If the app came out of an AI builder, the platform side of the first week, meaning control of the builder account, exposed keys and open database rules, belongs to first-week triage for an inherited vibe-coded app, which runs alongside this list.
Exit test: someone else’s machine reaches a running app from a fresh clone, with your notes as the only help.
Week 2: pin what it does today with characterization tests
Characterization tests record what the code does today, not what it should do. In week 2, my working rule is to write about 5 to 10 of them around the flows that make money or hold data: sign-up, login, checkout, the main write path. A wrong behavior gets pinned too, with a comment, and fixed later on purpose.
This week exists because you cannot assume the net is there. Across the 21 third-party apps from my June and July 2026 audits, at least 18 had no working test anywhere: 17 with literally none, plus a retail POS whose checkout “test suite” never executed the actual checkout code. Those 21 were third-party apps I selected and audited, a selected set and not a random sample, so treat the count as a reason to look rather than a rate for AI-built apps in general.
Michael Feathers coined the term characterization test, and his book Working Effectively with Legacy Code was published by Pearson in 2004. The idea is simple: you run the code, record what comes out, and assert exactly that, so any later change that alters the output fails loudly.
Write these end to end, through the browser or the API, before any unit test. Knowing how to write end-to-end smoke tests is the skill this week leans on most, because a flow test survives the refactors that would break a unit test written against today’s function names.
| Flow | What to assert | Why it is pinned even if wrong |
|---|---|---|
| Sign-up | A new account exists, with the role and default settings the app gives it today | A changed default changes every new customer’s experience without anyone deciding it |
| Login | A known user gets in; a wrong password and a disabled account are refused the way they are refused today | Changing who gets refused is an access decision, not a cleanup |
| Checkout or payment | The total charged for a known cart, including today’s rounding and tax handling | Customers have already paid that total, so changing it is a billing change that needs notice |
| The main write path | The record saved, its fields, and what other screens show afterwards | Other code and reports read that record in its current shape |
| Delete or export of customer data | What is removed or returned, and what is left behind | Data handling changes can have legal and support consequences the code does not show |
For a behavior you know is wrong, put a comment on the test that says so, add a line to the risk list, and leave it, because fixing it now would be a change nobody asked for, shipped by the person with the least context. The tests run in CI on every push from the day they exist.
AI assistants are useful here for drafting tests against what the app does today. My rule: every generated test is run red once, by breaking the code it covers and watching it fail, before it is trusted.
Exit test: break one core flow on purpose, and the suite goes red.
Week 3: ship small real changes through the net
My working rule is to pick about three to five small, real tickets: a copy fix, a validation rule, a dependency patch. None of them is cleanup. Each one takes the normal path of branch, review and deploy, and each one teaches you one part of the code you would not have read otherwise.
The map grows from these tickets. You learn which module owns which rule, and where state is duplicated, the classic symptom being when two screens show different values. You also notice which files nothing imports. Write those down now and remove them later, once you know how to find unused code in a repo without deleting something a scheduled job still loads.
Keep to the boy-scout limit: tidy only the function you are already changing, and do it in a separate commit so the tidy-up can be reverted without reverting the fix.
By the end of the week the notes file becomes the README. Most of what belongs in a README is already in your first-week notes: the steps that were missing, the variables nobody documented, the command that actually deploys. If the previous developer is reachable, ask for one recorded session; knowing what a codebase walkthrough is before the call lets you ask for the parts the code cannot show, and that recording is worth more than a week of reading.
Exit test: a real change shipped with no incident, and the README lets someone else do the same.
Week 4: the keep, fix or replace decision, per module
Line count says little about how hard a takeover is. My working rule: 700k lines with tests, types and one clear entry point can be safer than 7,000 lines nobody can run, so I weigh four questions instead: can you build it, can you deploy it, is anything tested, and does anyone know the business rules.
If you inherited 700k lines of bad code, my working rule is to run the same month-long plan; what changes is how much of the code you leave untouched at the end. The decision is made per module, never for the whole codebase at once, and it uses what three weeks produced: the test list, the risk list, the change log and the error tracker.
The signals below are how I read a module; the cells are my judgment, not a standard.
| Signal | Keep as is | Fix in place | Replace behind an interface |
|---|---|---|---|
| Does it change often? | Rarely touched | Changes often, and changes land cleanly | Changes often, and every change breaks something |
| Does it break often? | Quiet in the error tracker | Breaks in known, reproducible ways | Breaks in ways nobody can reproduce |
| Can it be tested? | Already pinned by characterization tests | Testable once a seam is added | Cannot be tested without rewriting it |
| Does anyone understand the rule it encodes? | Yes, and the rule is still right | Yes, but the code gets it wrong | Nobody, and the rule has to be rediscovered anyway |
| Is it on the money or data path? | Off the path, or on it and covered by tests | On the path, with characterization tests around it | On the path, untested and untestable as written |
Replacing one module behind an interface while the old path keeps serving has a name. In Martin Fowler’s description of the strangler fig pattern, the gradual approach “begins with small additions, often new features, that are built on top of, yet separate to the legacy code base”, and behavior moves across in small pieces.
Rebuilding the whole app is a different sum, with costs on both sides, and that decision belongs to whether to rebuild or fix the app. This plan’s last week stays at module level.
By my working rule, the output is a decision memo of about two pages: what stays, what gets fixed and in which order, what gets replaced, and what each costs in weeks, written for the person who pays.
Exit test: the memo is accepted or argued with, and either is fine.
A filled example: month one on a small AI-built SaaS
I constructed this example: a made-up app, no real client, and every number in it illustrative. The app I made up is a Next.js and Supabase SaaS of about 40,000 lines, with no tests and one contractor who built it and has moved on.
Constructed example, illustrative numbers: the plan table filled in for that app.
| Week | What turned up | What was produced |
|---|---|---|
| Week 1 | Setup took two days because three environment variables were undocumented and the seed script was missing. The no-change deploy showed that production deploys ran from a laptop | Setup notes, a one-page drawing, a risk list headed by the laptop deploy |
| Week 2 | Eight tests around sign-up, login, checkout and the two main write paths; two of them pinned a total known to be wrong | Test list, with the two wrong totals marked |
| Week 3 | Four tickets shipped; the same price calculation turned up in three places | README, change log, a note on the duplicated calculation |
| Week 4 | The memo: keep the UI layer, fix billing in place with tests first, replace the hand-rolled job runner behind an interface, do not rebuild | Decision memo |
At the end of the month the engineer in this example holds four artifacts: a setup guide that someone else has followed, a test list, a dated risk list and the decision memo. That is a starting set, not a full one; the complete documentation set for a vibe-coded app adds runbooks, business rules and decision records on top.
The useful part of the example is the order of the findings. The laptop deploy was found in the first week because the plan forced a deploy before anything else. The wrong totals were pinned, not fixed, so billing could be fixed in the last week with tests around it. The duplicated price calculation was found by shipping a ticket, not by reading.
What is legacy modernization, and when a takeover turns into one
Legacy modernization is replacing or restructuring an old system so it can keep changing; legacy transformation and product modernization name the same move. AWS’s migration guide lists 7 strategies, from retire and rehost to refactor or re-architect, which it calls the most complex. A takeover comes first and shows which parts need it.
The names in AWS’s seven migration strategies were written for moving workloads into its cloud, but they still describe the choices you have for one module. The first two columns below are AWS’s; the third is my reading for a small team.
| Strategy (AWS’s name) | What AWS says it is | When a small team picks it |
|---|---|---|
| Retire | ”the migration strategy for the applications that you want to decommission or archive” | A feature week 3 showed nobody uses |
| Retain | For “applications that you want to keep in your source environment or applications that you are not ready to migrate” | Every module the memo marked keep |
| Rehost | ”also known as lift and shift”: moving applications “without making any changes to the application” | Moving the app to another host as it is, when the host itself is the problem |
| Relocate | Transferring servers to “a cloud version of the platform”, or moving instances “to a different virtual private cloud (VPC), AWS Region, or AWS account” | Moving the project into an account the business owns |
| Repurchase | ”also known as drop and shop. You replace your application with a different version or product.” | Swapping a hand-rolled piece, such as a job runner, for a hosted product |
| Replatform | ”also known as lift, tinker, and shift”: moving to the cloud while introducing “some level of optimization” | Moving to a managed service with few code changes |
| Refactor or re-architect | Moving to the cloud and modifying the architecture “by taking full advantage of cloud-native features” | The module the memo marked replace, done behind an interface |
AWS’s own advice for large migrations is to rehost, relocate or replatform first and modernize after the move, because refactoring during the migration means modernizing at the same time. The same logic applies to a takeover: get control first, change shape second.
Legacy code modernization is that work at the level of one codebase, module by module, in the order the memo sets. Legacy modernization solutions and takeover products are also sold for estates far bigger than a small SaaS: one takeover product’s page names COBOL, RPG, AS400, mainframes and SAP among the systems it manages. A small SaaS rarely needs legacy systems modernization services as a program; it needs the decision memo and a quarter of steady work.
The case for keeping a way back comes from a bank. Over the weekend of 20 to 22 April 2018, TSB migrated the majority of the operations of its corporate systems, customer services and customer data to a new platform, in what the FCA’s final notice to TSB calls “a predominantly single event data migration”. The notice is plain about the consequence: “A ‘roll-back’ plan was not possible for TSB in these circumstances.” TSB is a bank, far larger than any app this plan is written for, and the point is the shape of the decision, not the scale. My opinion on the lesson: replace the module the memo marked replace behind an interface, while the old path still serves, and keep a way back until the new one has held.
How to verify the takeover worked
A takeover has worked when 5 things are true and none depends on the previous developer: a new machine goes from clone to running app by following the README, a deploy and a rollback have both been done, the core flows have tests, the risk list is written, and a real change has shipped.
Each check below can fail at day 30, and each leaves evidence you can show someone else.
- 01 Someone else, on a machine that has never run the app, gets it running using the README and nothing else. Evidence: their notes
- 02 You have done a deploy and a rollback yourself, and deploys went back to normal after the rollback. Evidence: the two dated entries in the host's deployment list, or the CI run links
- 03 An intentional regression in a core flow fails CI. Evidence: the failed run
- 04 The risk list exists, is dated, and its top three items each have an owner. Evidence: the dated file
- 05 One real change shipped without the previous developer's help. Evidence: the pull request
A sixth check is optional and harder to fake: someone else can explain the system back to you from your one-page drawing.
Two of these checks are how the Production Hardening Sprint verifies its own code health work. Deliverable 10.10, Reproducible local development setup, is verified this way: go from clone to running application on a new machine by following the guide alone. Deliverable 10.5, Critical-path smoke tests, is verified this way: run the suite in CI and demonstrate that an intentional regression fails it.
Where the sprint fits
The sprint starts from the app you already have: its current framework and hosting setup are the starting point, and components are refactored or replaced where the production work requires it. Its code health work covers 10.10 and 10.5, verified as the section above describes, and four of its other items are checked the same evidence-first way. Deliverable 10.4, Documented repository structure, is verified this way: follow the README from a clean checkout and trace a core feature through its documented modules. Next, deliverable 10.7, Auth and billing unit tests, is verified this way: run the tests and show rejection of unauthorized access and incorrect billing transitions. Deliverable 10.1, Dead code and duplicate logic cleanup, is verified this way: record removed or consolidated code and run regression checks on affected flows. Last in the area, deliverable 10.9, Recorded codebase tour, is verified this way: deliver an accessible recording with chapter markers or a short contents list. The production readiness report, 13.1, is verified this way: account for all 123 IDs; keep failures visible until resolved and explain genuine non-applicable items. Outside the sprint: building new product features or modules, completing unfinished core features or business workflows, and rebuilding core functionality that does not yet perform its intended job. After handover, we cover 14 calendar days of fixes for defects in the delivered sprint work and 30 calendar days of async questions about the handover and architecture. The full list is under the code health items in the sprint’s scope.
Common questions about legacy code
What is a legacy code?
Legacy code is code its owners are afraid to change, and age has nothing to do with it. Michael Feathers’ book Working Effectively with Legacy Code defines legacy code as code without tests. A codebase, meaning all the source files, configuration and scripts that make up one app, does not need to be old to qualify: months-old code with no tests and no author to ask fits both definitions.
Can AI understand legacy code?
Partly. An assistant such as Claude Code can give you an overview of a codebase and find the files behind a feature: Claude Code’s guide to understanding a new codebase suggests asking it to “give me an overview of this codebase” and to “find the files that handle user authentication”. On a long session, how Claude Code handles a full context window matters too: it “compacts automatically as you approach the limit”, summarizing the conversation history to fit. An assistant sees only what you give it or connect it to, and a business rule nobody wrote down is in none of that, so my rule is to check every claim it makes against the running app.
Is replacing a legacy system worth it?
Sometimes, and the call is per module: replace a part when it changes often, breaks often and cannot be tested where it stands. For the whole app, the answer is a cost sum, which my article “Should You Rebuild Your Vibe Coded App or Fix It?” works through.
Why do companies still use legacy systems?
Because they work. An old system encodes rules nobody has written down anywhere else, and replacing it means rediscovering every one of those rules while customers keep using the product, which carries its own risk.
Owning an app means being able to run it, change it and recover it without guessing. The sprint below leaves you with the runbooks and documentation to do that.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase