Technical DebtArchitectureSoftware Delivery

Legacy software: when to live with it, when to evolve it, when to rewrite it

Not all legacy software needs rewriting, and not all of it can be left alone. The decision comes down to three variables that almost nobody measures before choosing. A guide for anyone deciding what to do with a system that works but is holding things back.

QMates· Software Advisory14 April 202611 min read

The problem with the word "legacy"

In everyday usage, "legacy software" has become an almost entirely negative term: old, troublesome, something to replace as soon as possible. That emotional framing leads to expensive decisions. Once a system gets labelled legacy, it goes straight onto the rewrite list, regardless of how much it is actually slowing delivery down, what replacing it would cost, or how much value it is still generating.

The definition most often cited in the literature comes from Michael Feathers: legacy code is simply code without tests. Useful, but it captures only one symptom. In practice, we prefer a broader definition: a legacy system is one that runs in production and generates value, but shows at least one of these traits: technology the community no longer maintains, knowledge concentrated in a handful of people, or changes that keep getting harder for structural reasons. Age isn't the deciding factor; a system written three years ago can already have accumulated enough debt to be genuinely hard to change.

The question isn't "is this system legacy?" but "what does it cost to keep it as it is, what would it cost to change it, and what would it cost to replace it?" Without answers to those three questions backed by numbers, any decision is arbitrary.

The three variables we look at

In our experience, the choice between living with a system, evolving it or rewriting it mainly comes down to three variables, and we recommend measuring them rather than guessing.

The first is change frequency. A system that almost never changes has a low maintenance cost, however technically complex it might be. Technical debt gets paid every time you touch the system; leave it alone and the debt stays latent rather than active. A legacy module that does one thing, has been stable for years and sits outside the critical path of growth can stay exactly as it is, indefinitely. Acting on it only becomes urgent once new requirements disturb that stability.

The second is its weight in the dependency graph. An isolated legacy system, with few connections to the rest of the architecture, is far less of a problem than one sitting at the centre of a dense web of dependencies. If every new feature has to pass through that system, its debt compounds with every change to the wider system. Where a legacy system sits in the architecture often matters more than its internal state.

The third is the risk of losing institutional knowledge. This is the most underrated factor of the three. A system that works well but that only two people truly understand is a real operational risk. The moment those people are unavailable, whether on holiday, off sick, or because they've left, that part of the system turns opaque. Knowledge risk can turn a manageable system into an emergency very quickly.

Scenario one: living with it

Living with a legacy system is the right call when all three variables point the same way: the cost of changing or replacing it outweighs the benefit.

The typical profile is a system that's highly stable (low change frequency), light on dependencies (few connections to the rest of the architecture), and whose knowledge is either spread across several people or well documented. In this scenario, the legacy system isn't an operational problem; it's an asset that keeps generating value at an acceptable maintenance cost.

Living with it doesn't mean ignoring it. It means treating it as a stable component and committing to three things: don't add new responsibilities to it (build every new feature elsewhere), document the knowledge that already exists (not to rewrite the system, but to reduce the risk of depending on specific people), and keep an eye on the variables that could change the assessment, such as a strategic integration that forces you to touch it, the retirement of a component it depends on, or the key person leaving.

The sign that living with it has stopped being the right choice is when one of those triggers actually materialises. Until then, any energy spent refactoring or rewriting a stable system is energy taken away from more pressing problems.

Scenario two: evolving it progressively

Progressive evolution is the right choice in most cases where a legacy system sits on the critical path of growth but isn't degraded enough to warrant a full replacement. It's also the least glamorous strategy, and the most effective one.

Evolving a legacy system means systematically applying the principles of progressive refactoring we covered in the previous article, with a few considerations specific to systems carrying years of accumulated debt.

The first principle is never touch a part of the system before covering it with tests. Test coverage in legacy systems is almost always inadequate. Adding tests for current behaviour, even behaviour that isn't the behaviour you'd actually want, is the prerequisite for any safe change. This phase takes time, but it isn't optional: without that safety net, every change is an unquantified risk.

The second principle is isolate before you modify. Before restructuring any part of the legacy system, isolate it from everything else behind clear interfaces. Isolation lets you test it independently, replace the internals without affecting anything outside it, and check behaviour against specific cases. It's the Strangler Fig pattern applied at module scale rather than to the whole system.

The third principle, specific to legacy systems, is don't try to understand everything before you start. In systems with years of history and little documentation, full understanding before you intervene is often impossible. The approach that works is exploratory: narrow down an area, add tests, make the change, see what surfaces. You build up knowledge of the system incrementally, through the process of changing it safely.

Scenario three: rewriting it

A full rewrite is justified in a limited number of scenarios. Pinning them down precisely matters because the emotional instinct is to reach for a rewrite even when the problem could be solved with something more targeted and less risky.

The first scenario where a rewrite is justified is when the system rests on technology that is no longer maintained and whose unsupported status is becoming a security or continuity risk. We don't mean old technology as such, which, as we've seen, isn't debt if it still works. We mean technology that no longer receives security patches, has no active community behind it, or is incompatible with regulatory requirements the organisation has to meet. Here the risk sits outside the system and can't be managed through internal refactoring.

The second scenario is when the system's architecture is so constraining that any evolution costs more than a rewrite would. That threshold is rarely reached, but it exists: systems with coupling so pervasive that isolation becomes impossible, systems where the data model has drifted so far from current requirements that it needs a complete overhaul, systems where the cost of every single change outstrips its business value. The practical test: if the average cost of change consistently exceeds the value of each individual feature, a rewrite could pay off over the medium term.

The third scenario is strategic discontinuity: a change in business model, an acquisition, a shift to a new technology platform that forces you to rebuild from scratch. Here it isn't technical debt that justifies the rewrite but the change in requirements, which makes the existing system inadequate regardless of its technical quality.

In every other scenario, a rewrite is riskier than it looks and less necessary than the accumulated debt makes it feel.

How to decide: a three-question framework

When you need to make a call on a legacy system, these three questions, asked in sequence, produce a more reliable answer than intuition does:

How much will each change to this system cost over the next twelve months? Estimate the number of changes you'll need and the average cost of each, given the current debt. This is the cost of living with the system: not the cost of leaving it untouched, but the cost of every change the business will demand anyway.

What would it cost to get the system to a state where those changes cost half as much? This is the cost of progressive evolution. Not the cost of reaching some technical ideal, but the cost of halving the operational cost of future changes. If that investment pays for itself within six to twelve months of faster changes, evolution is justified.

What would it cost to replace the system entirely, including the hidden costs: data migration, running both systems in parallel, the risk of regressions, and the loss of domain knowledge baked into the old system? This is the cost of a rewrite. It's often significantly higher than the initial estimate, precisely because those hidden costs get underestimated. If this figure is higher than the cost of progressive evolution over a reasonable period, a rewrite doesn't stack up economically.

The special case: the legacy system nobody wants to touch

One specific case deserves separate attention: the legacy system the team systematically avoids, where changes get put off for as long as possible and anxiety builds every time a release is about to touch it. This pattern isn't always a sign that the system needs rewriting; it's almost always a sign that the team lacks the tests and the knowledge to change it with confidence.

The fix isn't a rewrite, it's stabilisation: add tests for current behaviour, document the murkiest parts, introduce monitoring that surfaces problems before they turn into emergencies. These steps can turn a system nobody wants to touch into one the team knows how to work on, even if it's still ugly under the bonnet.

Stabilisation alone isn't enough, though, unless it's backed by technical practices that keep the system under control over time. Systematic code review, frequent, reversible deployments, incremental refactoring as part of normal work, a CI pipeline that works and that people actually respect: without these mechanisms, debt tends to build back up at the same rate you remove it. Tests protect against regressions and monitoring surfaces problems, but it's the team's day-to-day practices that decide whether the system improves or degrades from one sprint to the next. That's the difference between an emergency fix and a lasting one.

What to do in practice

  1. Apply the three-question framework to your most critical legacy system: not every legacy system at once, just the one causing the most friction. Estimate the cost of living with it, the cost of evolving it and the cost of rewriting it using real numbers, not impressions. The result is often surprising: systems that looked like they needed a rewrite turn out to be manageable with targeted evolution, and systems that looked stable turn out to be expensive to live with already.
  2. Identify the trigger that would change your current decision: if you've decided to live with the system, what signal would make you change your mind? The key person leaving? A new integration the business demands? A security vulnerability you can't patch? Knowing the trigger lets you monitor the variables that matter instead of constantly second-guessing a decision you've already made.
  3. Document the knowledge before anything else: whatever strategy you choose, documenting current knowledge is always step one. Not documenting the code itself, which takes time and goes stale quickly, but documenting the decisions and the constraints: why this system is structured the way it is, which edge cases it handles in non-obvious ways, where the behaviour the business takes for granted actually lives. That knowledge is the hardest thing to rebuild once it's lost.

The QMates perspective

The question we're asked most often about a legacy system is: "Should we rewrite it?" Our answer is almost never a straight yes or no; it's "that depends on these three variables, and you need to measure them before deciding". In many cases, the analysis reveals that the legacy system isn't actually the main problem: the problem is that there are no tests to let anyone touch it safely, or that knowledge sits with one person who's thinking of leaving, or that there's no monitoring in place to understand what it's actually doing in production.

A rewrite is rarely the right answer to the right question. It's often the emotionally satisfying answer to the wrong one. A system that comes under control, with tests, distributed knowledge, monitoring and shared technical practices, stops being legacy in any operational sense: it becomes a system that generates value at a predictable cost, whatever the age of the technology. A new system without these practices won't stay new for long.

If you want to work out the right strategy for your specific system, our Embedded Teams service works inside your team to pass on the method: not to replace the team, but to build its own capability to manage debt systematically. It's the approach that lasts.

Tell us where you're stuck

A fragile prototype, a burdensome legacy codebase or unpredictable delivery: that's where we start

We use your data to respond to your request. Read our Privacy Policy.