Why the wrong metric is worse than no metric at all
The most common temptation when you want to measure technical debt is to install a static analysis tool, get a code quality score, and use that number as a proxy for the debt. The trouble is that this number measures the complexity of the code, not the cost that complexity imposes on your ability to change the system.
A module with high cyclomatic complexity in an area that never changes isn't an operational problem. A module with average complexity in an area that changes every sprint, that every new feature touches, and that ten other modules depend on, is a blocker. The difference isn't in the number the tool produces; it's in how often the area changes and how much weight its dependencies carry.
Measuring technical debt means answering a specific question: how much does it cost to change this system, and where does it cost more than it should? The metrics that answer this question are different from the ones that measure static code quality. This guide focuses on the former.
The three operational metrics
There aren't dozens of equally relevant metrics for measuring debt. There are three that, read together, give a picture precise enough to base decisions on:
Cycle time is the time from when a piece of work is picked up to when it reaches production — the time it takes to move through the active workflow. It excludes analysis and estimation (that's a process issue) and covers only the time actually spent on implementation, review, testing and release. It's distinct from DORA's Lead Time for Changes (from first commit to production deployment), which also includes the time a commit spends queued before work even starts. In this article we use cycle time in the Lean sense, which makes it the most sensitive metric for detecting internal resistance in the system. A cycle time that grows quarter on quarter without a corresponding increase in feature complexity is the most reliable sign that debt is at work. The reason is that technical debt shows up exactly there: every change takes longer because hidden dependencies multiply the work required, tests fail unpredictably, and review means tracing impact into parts of the system that shouldn't be involved at all.
The normalised bug rate is the number of incidents and regressions per feature shipped, not the absolute number. A team that ships twice as many features can have twice as many bugs without that being a sign of debt. The signal is when the bug rate per feature rises: each new feature generates more problems than the last, because the system is becoming less predictable in how it reacts to change. This metric is particularly useful for pinpointing debt: the highest bug rate almost always marks the area of the system with the most hidden coupling.
Cost of change is the real cost of a specific change, calculated not just from development time but across the whole cycle: impact analysis, cross-team coordination, in-depth review, extra manual testing to make up for gaps, deployment coordination, post-release verification. This metric can't be automated: it has to be built from concrete cases. But it's the one that translates most directly into an argument management can follow. "This change should have taken three days; it took fifteen because it cut across these four dependent modules" is a specific number on a specific case. It's more persuasive than any aggregate estimate.
The Consistency Model: the dimension local metrics don't see
The three operational metrics measure the cost of debt in the present. They don't measure the structural cause that generates it. For that you need a different kind of analysis: the Consistency Model.
The Consistency Model measures the degree of alignment between three dimensions that, in a healthy system, evolve together: the software architecture, the structure of the teams that build it, and the business processes it needs to support. When these three dimensions are aligned, the system is easy to change: every team has clear ownership of its own area, the architectural boundaries reflect the boundaries of responsibility, and business processes map directly onto the software modules.
When the three dimensions drift apart — and they always do during growth — the debt becomes systemic. The telltale sign of misalignment isn't code quality; it's that changes requested by the business systematically cross architectural boundaries they shouldn't have to cross.
In practice: an organisation that has restructured its teams but not its architecture ends up with teams that have to coordinate on every change because the code doesn't reflect the new structure. An organisation that has changed its business processes but not its software modules ends up with bounded contexts modelling flows that no longer exist, and workarounds everywhere to handle the new ones. The debt isn't in the individual components; it's in the gap between how the system is structured and how the organisation works today.
How to measure the misalignment
You don't need sophisticated tooling to apply the Consistency Model. It needs answers to a structured set of questions across three levels.
At the architecture–team level: can every change requested by the business be handled by a single team on its own, or does it need coordination across several teams? If the answer is "often more than one team", where do these boundaries sit? The teams that have to coordinate most on every feature are almost always the ones whose architectural boundaries don't reflect their boundaries of responsibility. Teams that block one another don't have a communication problem; they have an architecture problem.
At the architecture–process level: do the system's bounded contexts reflect the business's current flows? Or are there areas where every new process requires touching three or four different modules because the abstractions were designed for an organisation that no longer exists? A simple test: take a business process that has changed in the last twelve months. How many software modules needed changes to support it? If the answer is more than two or three, the system no longer reflects the structure of the business.
At the team–process level: do teams' areas of responsibility reflect the value streams that reach the customer? Or are teams organised around technologies or components rather than around products or services? A team that owns the database but not the API that exposes it creates structural dependencies that slow down any change to that data.
From data to decision
Measuring debt isn't an end in itself. It's the prerequisite for deciding where and how to intervene. The operational metrics and the Consistency Model answer two different but complementary questions.
The operational metrics answer: how much debt are we paying right now, and where? Cycle time shows how much the system is slowing delivery down. Bug rate shows where the system is least predictable. Cost of change quantifies the concrete cost of specific cases. Together, they show where the pressure is highest and where intervening would pay off the most.
The Consistency Model answers: why does debt accumulate in those specific areas? Understanding the structural misalignment is necessary to address the cause, not just the effect. A refactor that reduces cyclomatic complexity without touching the architectural boundaries treats a symptom, not the cause: the debt will come back, because the conditions that generated it are still there.
The Consistency Impact Calculator is the tool we use to structure this analysis so that it produces an output management can actually read: not "we have a lot of debt" but "the cost of the next change in this area is X, because it crosses Y misaligned modules, and could be reduced to Z with a specific intervention".
The metrics that distract
Not every available metric is useful for measuring debt. Some are easier to collect but less informative; using them leads you to optimise the number instead of the underlying problem.
Test coverage is the most common one. High coverage doesn't imply low debt: you can have tests on every function and still have an architecture so coupled that every change is a risk. Coverage is a prerequisite for refactoring safely, not a measure of the debt you need to eliminate.
Absolute cyclomatic complexity is another misleading metric when used without context. As discussed, the complexity that matters isn't the static complexity of the code but the operational one: how much it costs to change that part of the system given how often it needs changing. A module with high complexity that never changes isn't an urgent problem.
The number of open tickets isn't a debt metric: it's a load metric. A long backlog can reflect prioritisation, not debt. A backlog that grows in specific areas of the system every sprint, on the other hand, is an indirect sign of debt: that area is harder to manage than the rest.
What to do in practice
- Instrument cycle time by area of the system: not the average global cycle time, which is too aggregated to be useful. Segment by codebase area or by type of change. Comparing areas almost always reveals patterns that aggregate analysis hides: an area that's four times slower than the rest is almost certainly an architectural bottleneck.
- Calculate the cost of change on three recent, complex changes: pick three changes the team felt were harder than expected. Reconstruct the real cost: impact analysis time, cross-team coordination, review, extra testing, deployment coordination. The sum of these costs, compared against the original estimate, gives you the delta that measures the cost of debt on those specific cases.
- Map the Consistency Model in an hour: get the tech lead and the leads of two or three teams together. Answer these questions for every strategic change over the last six months: how many teams had to coordinate? How many modules of the system needed changes? Where do the architectural boundaries not reflect the boundaries of responsibility? The result won't be a complete analysis, but it will be enough to identify the areas with the sharpest misalignment.
The QMates perspective
Our analysis starts from the flow, not the code: we trace how changes propagate through the system, where they get stuck, how long they take to complete. The metrics we collect aren't static quality metrics; they're dynamic friction metrics. The difference is that dynamic friction correlates directly with delivery cost, while static code quality doesn't necessarily.
What we find almost every time is that the most expensive debt isn't where the code is most complex. It's where the architecture and the team structure have stopped evolving together. It can be a bounded context that used to serve three clients and now has to serve a hundred, with completely different processes. It can be a team whose composition and responsibilities have changed but that still works on an architecture designed for the previous composition. The measure of debt that matters is always relational: not how complex the code is, but how expensive it is to change in the context of the organisation that has to change it.
The fifth article in this series covers the next step: how to reduce technical debt with strategies that actually work, without stalling delivery during the intervention.