Technical DebtArchitectureSoftware Delivery

How to reduce technical debt: strategies that work

Reducing technical debt doesn't mean rewriting everything from scratch, or stopping delivery for months. It means stepping in at the right points, in the right order, while keeping the ability to ship value throughout the whole process.

QMates· Software Advisory14 April 202611 min read

The big-bang rewrite: a high-risk choice, not a solution

Once technical debt becomes visible enough to land on the management agenda, the proposal that comes up most often is also the riskiest one: rewrite everything from scratch. The logic is intuitive — the current system is the problem, so a new system would fix it — but it ignores some operational realities that make this approach statistically risky: it hides significant costs, it requires keeping the old system running in parallel, and it tends to reproduce the old system's problems in the new one if the structural causes aren't addressed.

The first is that the current system, however messy it might be, embodies years of domain knowledge: edge cases nobody remembers handling, behaviour the business treats as a given, integrations documented only in the code itself. A rewrite starts from zero on that knowledge too, and tends to discover the hidden complexity at the worst possible moment — when the new system has to go live.

The second is the cost of running two systems at once. During a full rewrite, the team has to keep the old system in production while building the new one. The old system stays slow because of its debt; the new one eats resources in proportion to its ambition. The usual outcome is that both systems move more slowly than planned, and the timeline to cut over stretches well past the original estimate.

The third is deployment risk. A new system that's meant to fully replace an old one requires a cutover moment that is, by definition, the point of maximum exposure to risk. You can't eliminate that risk entirely; you can only postpone it until there's no postponing it any longer.

There are scenarios where a full rewrite is the right call — but you identify them with precise criteria, not accumulated frustration. The specific triggers (technology that's no longer maintained, architecture so tightly coupled that isolating anything is impossible, a strategic discontinuity) are covered in detail in the next article in this series: legacy software — when to live with it, when to evolve it and when to actually rewrite it.

The principle behind effective debt reduction

The strategy that works for reducing technical debt is always the same, whatever the technology or context: work incrementally, and keep the ability to ship value throughout the whole process.

This principle has three practical implications. The first is that every piece of work has to be small enough to finish and ship within a reasonable time — usually one or two sprints — without stalling feature delivery. The second is that the order you tackle things in should be driven by their impact on delivery, not by how "clean" the code ends up. The third is that every change has to leave the system better than it found it, not stranded halfway to some distant goal.

That last point is critical. Technical debt builds up precisely in those intermediate states: a refactor starts, gets paused halfway through for something urgent, and the halfway state becomes the new permanent one. Every debt-reduction effort has to be self-contained: if you can't finish it, don't start it.

The Strangler Fig: a strategy for legacy systems

The most effective approach for replacing a legacy system without a big-bang rewrite is the Strangler Fig pattern, named by Martin Fowler. The name comes from the strangler fig plant, which grows around a host tree, gradually envelops it, and replaces it over time without ever going through an abrupt transition.

In practice, you build the new system around the old one. Every new feature goes into the new system; existing features are migrated one at a time, from simplest to most complex, as the new system matures. The old system keeps running for whatever hasn't been migrated yet; the new one grows step by step until it has taken over completely.

The advantage of this approach is that risk is spread out over time: every migration is a small, reversible step. The new system can be tested in production before it fully replaces the old one. There's no single day when everything switches over; the transition happens gradually, reaching a point of no return once enough traffic has moved to the new system to make the old one redundant.

The downside is the cost of running two systems during the transition instead of one. That's why the length of the coexistence period needs to be capped and planned up front.

Progressive refactoring: where and how to apply it

For debt that doesn't call for a full rewrite — which is most real-world debt — progressive refactoring is the most effective approach. But the term "refactoring" gets used too loosely: any improvement to the code gets labelled refactoring, whether or not it has any measurable effect on delivery.

Refactoring that counts isn't the kind that makes the code prettier; it's the kind that lowers the cost of future changes in areas that change often. The distinction is a practical one. Before starting any refactoring effort, ask: once this change is done, will it reduce the cycle time of future changes in this area? If the answer isn't clear, the effort probably doesn't belong on the right priority list.

The starting point for progressive refactoring is always the debt map built from the metrics covered in the previous article: which areas of the system have the highest cycle time? Where is the bug rate worst? Where is the cost of change disproportionately high? Those are the areas where refactoring pays off most. Refactoring the wrong areas is one of the most common ways to waste effort: it makes the code more readable but does nothing for delivery speed.

The order that works: lowest complexity first, highest impact first. You start with the moves that have the best impact-to-effort ratio: often not the big architectural overhauls, but extracting duplicated logic, clarifying module boundaries, removing circular dependencies in areas that change frequently. These are changes you can finish in a few days: they cut operational complexity immediately and create room for the bigger efforts that follow.

How to keep delivery moving during the work

The main risk with refactoring is that it slows delivery down while it's under way. That almost always happens when the scope is too big: too much code gets touched at once, the system becomes unstable, tests start failing unpredictably, and the team spends more time stabilising than building.

The techniques for reducing that risk are well established. The first is branch by abstraction — a technique introduced by Stacy Curl in 2007 and later popularised by Paul Hammant. Instead of changing the existing code directly, you introduce an abstraction that isolates the part you're rewriting, reimplement behind that abstraction, confirm the behaviour is equivalent, and only then remove the old code. That way the system stays in a working state throughout.

The second is the feature flag: the new code exists in production but stays switched off until someone deliberately turns it on. That lets you ship the refactor dark, test it against a slice of traffic, and roll it back in seconds if something goes wrong.

The third is test coverage before you start. Refactoring without adequate coverage is like stripping down an engine without knowing how it's supposed to run. Before touching an area of the system, you add tests that describe its current behaviour — even when that behaviour isn't the one you actually want. Those tests become the safety net that lets you move forward with confidence.

Prioritisation: where to start

When a system has debt spread across several areas, prioritisation is a real problem: where do you start? The answer isn't "the most complex area" or "the area with the oldest code". It's the area where the investment pays off fastest for delivery.

Apply three criteria, in this order.

The first is change frequency. An area that changes every sprint and carries high debt is more urgent than an area that changes rarely with the same level of debt. The cost of debt is paid every time that part of the system is touched: the more often it changes, the faster the refactoring investment pays for itself.

The second is dependency weight. A module that ten other modules depend on blocks far more delivery than an isolated one. If the architectural bottleneck is a single point of coupling slowing down changes in five different areas of the system, fixing it there unblocks five workstreams at once.

The third is how reversible the change is. If you're not confident in the approach, start with whatever is easiest to undo: something you could unwind in a few hours if it doesn't work out as planned. Irreversible changes need more careful planning and should be validated in controlled environments before they reach production.

What to do in practice

  1. Pick one area, not the whole system: the first debt-reduction effort shouldn't be the most ambitious one. It should be the one with the clearest impact-to-effort ratio. Take the area with the highest cycle time and the worst bug rate. Define a specific piece of work, completable in one or two sprints, that reduces the cost of changes in that area. Measure cycle time afterwards. If the trend improves, that confirms the prioritisation was right.
  2. Add tests before touching the code: if the area you want to refactor doesn't have adequate test coverage, the first move is to add tests for its current behaviour. Not tests that validate the design — tests that describe what the system actually does. These tests are the safety net that lets you change the code with confidence. Without them, every change is an unquantified risk.
  3. Define a success metric before you start: every debt-reduction effort needs a measurable success criterion. Not "the code is cleaner" but "the cycle time for changes in this area dropped from X days to Y". That metric is how you know whether the work paid off, and how you communicate the return on investment to management in terms they understand.

The QMates perspective

Our approach to debt reduction rests on one principle: never do anything that stops the team shipping value. Technical debt has already slowed delivery down; any step that blocks it further, even temporarily, makes the problem worse instead of solving it.

In practice, this means every piece of work we propose is designed to run alongside ordinary delivery, not instead of it. We don't ask for an isolated "cleanup sprint" carved out of the rest of the work; we propose a way of folding refactoring into the normal development flow, with clear criteria for when an area is stable enough to stop needing attention.

Our Targeted Intervention service is built exactly for this: stabilising the blocked areas of the system in the shortest possible time, while keeping delivery running throughout. It isn't cosmetic refactoring; it's surgical work on whatever is blocking the flow of value, with entry and exit metrics defined before it begins.

The sixth and final article in this series looks at a specific case of debt reduction: the legacy system — when to live with it, when to evolve it and when to actually rewrite it.

Tell us where you're stuck

A fragile prototype, a burdensome legacy codebase or unpredictable delivery: that's where we start

We use your data to respond to your request. Read our Privacy Policy.