Every engineer eventually inherits code that works, earns money and frightens everyone. Nobody remembers why a branch exists, the one person who understood the billing module left two years ago, and a change that should take an hour takes a week because nobody can tell what else it will break. The temptation is to rewrite. The usual right answer is to make the next change safely, and leave the area a little easier to change than you found it.

This guide is about how to do that at the level of code, the functions and classes you are about to edit. It explains what actually makes code legacy, how to pin its current behaviour before you touch it, how to find places to cut in, a small set of mechanical techniques for adding new behaviour without destabilising old behaviour, and how to replace a whole subsystem gradually. A worked example runs through all of it. Moving a whole system or its data to a new platform is a different problem, covered in migration architecture, in depth.

Advertisement

What makes code legacy

Age is not the definition. Michael Feathers, in Working Effectively with Legacy Code (2004), defines legacy code as code without tests, and that definition is useful because it points at the cure. Code written last month with no tests has the same problem as code written in 2009: you cannot change it and know that you have not broken it.

In practice untested code comes with companions. Functions are long and do several things. Dependencies are created inside the code rather than passed in, so a function that calculates an invoice also opens a database connection, reads the clock and calls a tax service. Global state is mutated from many places. And the code's real specification is its current behaviour, including its bugs, because some caller somewhere depends on them.

That last point drives the whole approach. With legacy code you are not trying to make the code correct. You are trying to change one thing while keeping everything else exactly as it was, and you cannot do that unless you first know what it was.

The legacy change loop: make one change safe, make it, leave the area better tested1. Identifychange points2. Find seamswhere to cut in3. Break depsminimal, mechanical4. Characterizepin current behaviour5. Changetest-first, small6. Refactorunder green tests7. Shipbehind a flag if riskynext changeInside one modulesprout method, wrap method, extract and overrideAcross a subsystemstrangler fig: route, replace, retireEach pass leaves a few more tests behind; coverage grows where change actually happens
The legacy change loop. Each change starts by making the area testable and pinning its behaviour, and ends with the area slightly better covered than before. Small techniques work inside a module; the strangler fig works across a subsystem.

Pin the behaviour first: characterization tests

A characterization test records what the code does now, not what it should do. You call the code with some input, observe the output, and write a test asserting that output, even if it looks wrong. If it looks wrong, you note it and ask; you do not fix it in the same change.

The quickest way to write many of them is to let the test tell you the answer. Write an assertion you know will fail, run it, and paste the actual value into the test. For outputs too big to paste, such as a generated report or a rendered page, use an approval test: store the output in a file the first time, and fail whenever it differs.

import datetime, json, pathlib
from billing.legacy import build_invoice   # the code we are about to change

CASES = [
    {"customer": "acme", "items": [("widget", 3, 9.99)], "region": "EU"},
    {"customer": "acme", "items": [], "region": "EU"},                    # empty basket
    {"customer": "zed",  "items": [("widget", 1, 0.0)], "region": "US-NY"},
    {"customer": "zed",  "items": [("gadget", 1000, 1.25)], "region": "US-CA"},
]

def test_build_invoice_characterization():
    golden = pathlib.Path(__file__).with_name("invoice_golden.json")
    actual = [build_invoice(**case, today=datetime.date(2026, 10, 1),
                            tax_rate=lambda r, d: 0.20)       # fake tax service
              for case in CASES]
    text = json.dumps(actual, indent=2, sort_keys=True, default=str)
    if not golden.exists():
        golden.write_text(text)            # first run records current behaviour
    assert text == golden.read_text(), "behaviour changed; review the diff"

Choose inputs by reading the code, not by guessing. Every branch is a reason for a case: an empty list, a zero price, the region that takes a different tax path, the quantity large enough to trigger a discount. Run with coverage to see which branches your cases reach, and add cases until the lines you are about to change are all covered. Coverage of the whole module is not the goal; coverage of the change point is.

Notice the today argument. If the real function reads the clock, the test is not repeatable, and you cannot pass a date in until you have created a seam for it. That is the next step.

Advertisement

Seams: places to change behaviour without editing the code

A seam, in Feathers' sense, is a place where you can alter what a program does without editing it at that place. Every seam has an enabling point, where you choose which behaviour runs. The common kinds in modern languages:

SeamEnabling pointTypical use
ParameterThe caller passes a different object or valueInject a fake clock, repository or HTTP client
ObjectA subclass overrides one methodNeutralise a method that sends email or calls a payment API
Module or importTest configuration replaces a modulePatch a module-level function in Python or a package import in Node
Link or buildA different library is linked or a build flag setC and C++ code with hard-wired system calls
ConfigurationAn environment variable or config filePoint the code at a local database or a stub service

Prefer the parameter seam whenever you can create one safely, because it is explicit and survives refactoring. Monkeypatching a module works but couples the test to the import structure, and it is the first thing to break when someone moves a file.

Breaking dependencies, carefully

To create a seam you usually have to change the code before you have tests for it, which is exactly what you were trying to avoid. The way out is to restrict yourself to changes so mechanical that they are very unlikely to alter behaviour, do them one at a time, and lean on the compiler, the type checker or an IDE's automated refactorings wherever possible.

The most useful single move is parameterise with a default. The function gains a parameter for the dependency, and the default value is exactly what the code used to create internally, so every existing caller behaves as before:

# before: the clock and the tax client are created inside
def build_invoice(customer, items, region):
    today = datetime.date.today()
    tax = TaxServiceClient(os.environ["TAX_URL"]).rate(region, today)
    ...

# after: same behaviour for every existing caller, injectable for tests
def build_invoice(customer, items, region, today=None, tax_rate=None):
    today = today or datetime.date.today()
    if tax_rate is None:
        tax_rate = lambda r, d: TaxServiceClient(os.environ["TAX_URL"]).rate(r, d)
    tax = tax_rate(region, today)
    ...

Other standard moves include extracting a method so it can be overridden in a test subclass, wrapping a static or global call in an instance method, and introducing an interface in front of a concrete class. Whatever you use, commit the dependency-breaking change separately from the behaviour change, with a message saying it is a pure refactoring. A reviewer can then check one diff for behaviour preservation and the other for the feature, instead of untangling both.

Adding behaviour: sprout and wrap

Once the area is pinned, resist the urge to restructure the whole function. Feathers describes two small techniques for adding behaviour that keep the new code testable and the old code untouched.

Sprout method. Write the new logic as a new function, developed test-first in isolation, and call it from one line in the old code. The old function gains a line; the new logic has full tests from day one.

Wrap method. When the new behaviour must happen before or after the old behaviour, such as auditing every invoice, rename the old function, create a new function with the original name that calls the old one and adds the new step, and leave the old body unchanged.

# sprout: new rule developed and tested on its own
def bulk_discount(quantity, unit_price):
    return round(unit_price * 0.05, 2) if quantity >= 500 else 0.0

# one-line call site inside the untouched legacy loop
line_total = qty * price - qty * bulk_discount(qty, price)

# wrap: rename the old function, keep its name for callers
def _build_invoice_unaudited(*args, **kwargs):
    ...                                   # original body, unchanged

def build_invoice(*args, audit=None, **kwargs):
    invoice = _build_invoice_unaudited(*args, **kwargs)
    (audit or default_audit_log).record(invoice)
    return invoice

These techniques leave the old code no worse and the new code well tested. Over many changes, the tested islands grow and the untested sea shrinks, which is exactly where you want coverage: in the places the business keeps changing.

The strangler fig inside a codebase

Sometimes a whole module needs replacing, perhaps because its data model is wrong rather than its code merely untidy. The strangler fig pattern, named by Martin Fowler, replaces it gradually rather than in one switch. Put a single entry point in front of the old module, route a narrow slice of calls to a new implementation, compare, widen the slice, and delete the old code when nothing routes to it any more.

Inside one codebase the entry point is usually a facade function or class, and the routing is a feature flag keyed by customer, region or call type. Run both implementations in shadow mode first: call the old one, return its answer, also call the new one, and log any difference. Only switch the flag when the difference log is empty for traffic you care about. Flag delivery, evaluation and clean-up are covered in feature flag delivery architecture.

The pattern also tells you when a full rewrite might be justified: when there is no seam at which to route, the old system cannot be run alongside a new one, and the cost of keeping it alive is rising faster than the cost of replacement. Those conditions are rarer than they feel at 2 a.m. during an incident. A rewrite rediscovers every undocumented behaviour by breaking it in production, which is why the incremental path is the default.

Worked example: adding a bulk discount to the billing module

The request is simple: orders of 500 units or more get 5 percent off the unit price. The build_invoice function is 400 lines long, has no tests, reads the clock and calls a tax service.

  1. Read the function and mark the change point: the loop that computes line totals.
  2. Apply parameterise-with-default for today and tax_rate. Commit as a pure refactoring.
  3. Write the characterization test above with a fake tax rate, adding cases until coverage shows every line of the loop is reached. One case reveals that a zero price produces a negative tax line. Record it, open a ticket, and do not fix it now.
  4. Sprout bulk_discount test-first, with cases at 499, 500 and a large quantity, and add the one-line call.
  5. Run the characterization test. It fails only for the case with 1,000 gadgets, and the diff shows exactly the expected discount. Re-record the golden file in the same commit, so the reviewer sees the intended behaviour change in the diff.
  6. Ship behind a flag, compare invoice totals for a day, then remove the flag.

The change touched one line of legacy code and added two tested functions and a test file that the next engineer will be glad to find.

Using AI assistants on legacy code

Coding assistants are good at the tedious parts: reading a long function and listing its branches, proposing characterization inputs, and drafting the mechanical refactorings. They are also confident about behaviour they have not observed. Treat their explanation of what legacy code does as a hypothesis, and confirm it with a characterization test before relying on it. Never accept a suggested fix to an odd-looking branch in the same change as your feature; that branch may be load-bearing. Reviewing such changes is covered in reviewing AI-generated code, and the test-first loop with an agent in TDD with AI coding agents.

Failure modes

  • Fixing bugs while refactoring. A caller depended on the bug, and now it breaks in a way nobody connects to your change.
  • Characterizing the wrong thing. Tests that assert only that no exception was thrown pass through any behaviour change.
  • Non-deterministic goldens. Timestamps, random identifiers or dictionary order in outputs make approval tests flaky until people stop trusting them.
  • Big-bang cleanup. A refactoring branch that lives for weeks, cannot be reviewed and conflicts with every other change.
  • Mocking everything. Tests so full of fakes that they verify the fakes, not the code.
  • Rewrite fever. Starting a replacement before the old system's behaviour is written down, then running both forever.

What to do next

  1. For your next change in legacy code, mark the change point and write characterization tests that cover it before editing anything.
  2. Introduce one parameter seam for the worst hidden dependency, usually the clock, the database or an external client, and commit it as a pure refactoring.
  3. Add new behaviour with sprout or wrap, test-first, and keep the legacy edit to as few lines as possible.
  4. Strip non-deterministic values from approval outputs so golden files stay stable.
  5. Keep a list of odd behaviours your tests uncovered and decide on each one deliberately, not inside a feature change.
  6. For a module that needs replacing, put a facade in front of it and plan a flagged, shadow-compared strangler fig instead of a rewrite.
Key takeaway: Legacy code is code you cannot change safely because nothing tells you when you broke it. Pin current behaviour with characterization and approval tests, create seams with small mechanical refactorings committed on their own, add new behaviour through sprouted or wrapped functions developed test-first, and replace whole modules gradually behind a facade. Each change leaves the area better tested, so the codebase gets safer exactly where it changes most.