Skip to main content
InfromatinTechnologies
Engineering8 min read

Testing a codebase nobody dares change

You do not need 90% coverage to make a system safer. Characterisation tests around the parts that matter buy more per hour.

Infromatin Technologies

Testing a codebase nobody dares change

A codebase with no tests is not untested. It has thousands of tests — every production incident, every support ticket, every "we tried that and it broke". They just are not executable.

The goal is not coverage. It is to make changing the system less frightening, in the places where changing it carries real risk.

Characterisation before refactoring

Before you restructure anything, capture what it currently does.

Write tests against the existing code that assert current behaviour, including behaviour you suspect is wrong. Those tests are not a specification of what should happen. They are a record of what does, which is what you need to know before you change it.

Then refactor. If a characterisation test fails, you have found either a bug or a behaviour you did not know about. Both are useful. Without the tests, you would have found out in production.

Test where the money is, not everywhere

Ninety percent coverage sounds rigorous and is a poor use of a quarter. Coverage is uniform by nature; risk is not.

Rank by blast radius:

  • code that moves money, issues credit or settles claims
  • code with a compliance or audit obligation
  • code that has changed most often, since that is where changes continue
  • code you do not understand at all

Test those thoroughly. For genuinely low-risk peripheral code, a smoke test that the service starts and responds is usually sufficient.

Prefer characterisation to assertion for unknown behaviour

For code nobody understands, an assertion-based test requires someone to decide what correct behaviour is — which is the thing you cannot do without reading the whole thing.

A characterisation test records what happens and asserts it stays that way. That is achievable in an afternoon and it converts an unknown into a known, which is the actual prerequisite for change.

Put a seam around the parts you cannot test

Sometimes a component cannot be tested in isolation: it talks to a mainframe, a proprietary device driver, or a vendor API with no sandbox.

Introduce an interface at the boundary. Test your side of it with a fake; keep the real integration covered by a small number of tests that run in a controlled environment.

This is worth doing even if the component is stable, because it converts an untestable component into a testable one plus an integration you trust more narrowly.

Automate the regression you already have

Every organisation has a set of failures that recur. Manual workarounds, hotfixes applied under pressure, data patches run by hand.

Each of those is a regression test that has not been written down. Convert them one at a time, starting with the ones applied most often — they are the most reliable signal you will get about where the design has failed.

Make the suite cheap enough to run

A suite that takes forty minutes does not get run before a release. It gets run before a release that mattered, which means it provides less protection exactly when protection is most valuable.

The fix is usually structural, not heroic: fewer end-to-end tests, more at the boundary, parallelise the slow ones, and split the suite so the fast subset runs on every commit and the full set runs nightly.

Target a few minutes for the pre-merge gate.

Get the whole team running it

A test suite nobody trusts produces a green tick and no information. That happens when tests fail intermittently, when a known failure is ignored, or when the suite has been failing for so long that everyone stopped looking.

Budget for it: fix the flaky ones properly rather than adding retries, remove tests for behaviour that no longer exists, and make the suite's status visible in the same place as the build.

The realistic outcome

A legacy estate does not reach high coverage without a rewrite, and a rewrite is usually worse than the disease.

What is achievable, and sufficient, is that the risky paths are characterised, the unknown boundaries are isolated, the recurring regressions are automated, and the suite runs fast enough to be a genuine gate. That makes the system changeable — which was the goal all along.

In this article

  • testing
  • legacy
  • technical debt

Working on something similar?

These articles come from real engagements. If the problem here sounds familiar, a 30-minute call is usually enough to tell you whether we can help.

Start a conversation

Related reading

Continue from here

Articles connected to the same delivery problems.

Have a related problem in front of you?

Send us the problem in whatever detail you have. A senior engineer replies within one business day, and you will get an honest read on whether we are the right partner for it.

We would like to use Google Analytics to understand how this website is used. No analytics are loaded unless you accept. Your choice is stored for six months.

See our Privacy Policy for details.