Model governance an examiner can actually read
If your ML system cannot produce a defensible explanation for a decision, the governance framework is decorative. Here is what to keep.
Model governance an examiner can actually read
Governance frameworks for machine learning tend to be read once by the people who wrote them and never again by the people operating the models.
The test of whether governance is real is narrow: when a regulator asks why a specific decision was made in a specific month, can you produce the answer in under a day?
If the honest answer is no, the framework is decorative.
What has to exist for the answer to be findable
For any automated decision — a credit approval, a claim outcome, a fraud flag — you need four things retained together:
The inputs. Exactly the values used, with a pointer to where each came from and when it was captured. Not the current value of the customer record; the value at decision time.
The model version. Which model, which version, deployed when, trained on what. A model that is overwritten in place makes historical decisions unreproducible, which is the same as indefensible.
The output. The decision, the confidence or score, and which decision path produced it — including any rules applied before or after the model.
The human outcome. What was decided subsequently, by whom, and whether it was overridden.
Linked by a single decision identifier. If these four live in different systems and cannot be joined, you do not have governance; you have logging.
The rules-before-and-rules-after problem
Most lending and insurance decisions are not a model's output. They are a pipeline: rules screen the input, a model scores what survives, and rules interpret the score.
An examiner asking why a customer was declined is asking about the pipeline. Systems that record only the model invocation cannot answer, because the decline may have come from a rule that never reached the model.
Record the whole path: rules evaluated, outcomes, the model score, and the final interpretation. It is more storage and considerably more defensibility.
Model change control is the part that decays
Initial governance is usually sound. The framework degrades as:
- models are retrained without a version bump
- thresholds are tuned in a spreadsheet and applied directly
- a feature is added for one use case and quietly used for two
- a model is validated once and never again
The controls that survive are the automated ones. A retraining pipeline that cannot deploy without a version identifier, a threshold change that requires a code review, and a drift report that goes to a named owner rather than a dashboard nobody opens.
Write for the reader who has to follow it
The most useful artefact is a one-page model card per production model, written in plain language:
- what the model decides, in one sentence
- what data it uses, and what it deliberately does not
- how it performs, measured on data held out at the time
- which populations it performs worse for
- known failure modes and what happens when they occur
- who owns it, and how to reach them
That last section is the one that determines whether governance works. Every model in production should have a named owner who will notice if its behaviour changes.
What an examiner will actually ask
Plan for these and most of what follows is answerable:
- Which model made this decision, and what version?
- What training data was used, and when?
- What were the inputs at decision time?
- Why was this applicant treated differently from a similar one?
- Has this model changed since the policy that authorises it?
- Can you reproduce this decision today?
Notice that none of these are about the model being good. They are about the organisation being able to account for it.
Where to start if you have nothing
You do not need a framework. You need one decision identifier, one immutable record per automated decision, and one owner.
Pick the highest-volume automated decision your organisation makes. Implement the identifier and the record for that one path. Prove you can answer all six questions above for it in under a day. Then extend.
That demonstration is more persuasive to a regulator, and to an internal risk committee, than any policy document.
In this article
- MLOps
- governance
- banking
- insurance
Working on something similar?
These articles come from real engagements. If the problem here sounds familiar, a 30-minute call is usually enough to tell you whether we can help.
Start a conversationRelated reading
Continue from here
Articles connected to the same delivery problems.
Have a related problem in front of you?
Send us the problem in whatever detail you have. A senior engineer replies within one business day, and you will get an honest read on whether we are the right partner for it.