Skip to main content
InfromatinTechnologies
AI & Data8 min read

Extracting fields from lending documents without losing the audit trail

Accuracy is the easy part. Being able to explain where each extracted value came from is what makes it defensible.

Infromatin Technologies

Extracting fields from lending documents without losing the audit trail

Getting an accurate extraction out of a payslip or a bank statement is mostly solved. The accuracy number on a vendor slide is achievable and repeatable.

The hard part for a regulated lender is not accuracy. It is being able to demonstrate, eighteen months later, where a specific figure came from.

Accuracy is table stakes; provenance is the requirement

A lending decision that uses a figure from a document needs to answer a question that will genuinely be asked: where did this number come from?

That requires more than a confidence score. It requires, for every extracted field:

  • the source document and its hash at the time of extraction
  • the page it was found on
  • the position on the page, or the bounding region
  • the raw extracted text, before any normalisation
  • the model version and the extraction date
  • any confidence threshold that was applied downstream

Store the raw text. Not the parsed number. The number is derivable from the text; the text cannot be reconstructed from the number.

Design for verification, not just extraction

If the workflow is extraction with no verification step, you have built something nobody trusts and therefore nobody uses.

Two patterns that work:

Confidence bands with human review. Above a high threshold, auto-accept and log. Below it, route to a human who sees the document region and the extracted text side by side. The reviewer's correction becomes training data and a quality signal.

Sample-based verification. Accept everything, and have a second reviewer check a random sample per batch. Far cheaper, and it detects drift in the model rather than individual errors.

The threshold is a business decision, not a technical one, because it trades straight-through-processing rate against error exposure. Set it explicitly.

Post-processing is where errors are introduced

The extraction is rarely the problem. The problem is the logic between the raw value and the field you store.

  • A payslip figure that represents a monthly total may be stored as an annual one
  • Currency conversion applied without recording the rate and date
  • A statement balance read from the wrong column
  • An annual figure divided by twelve to produce a monthly, which is wrong for bonuses
  • A negative sign lost in parsing

Every one of these produces a confident, clean number that is wrong. Record the transformation as part of the lineage, not just the extraction.

Multi-document consistency is the real test

Real applications contain contradictions: two payslips for the same month, a bank statement that does not reconcile with the loan application, a salary letter that disagrees with the payslip.

A single-document extractor will not notice. A lender cares about this more than about per-field accuracy.

Build explicit reconciliation rules and surface conflicts rather than silently choosing a value. Which document wins is a policy question — usually the most recent, sometimes the employer letter — and it should be visible when it is applied.

Watch for distribution drift

Document extraction degrades quietly. New lenders, new form layouts, a change in scan quality, or a shift to mobile photograph uploads all change the input distribution without changing a line of your code.

Monitor the input population: document sources, page counts, image resolutions, and the confidence distribution. A shift in confidence is the earliest signal that something has changed upstream.

Schedule periodic re-validation on a fixed human-labelled set. Not a full annotation project — a few hundred documents reviewed quarterly, scored against the same criteria each time.

The defensible position

Regulators are not asking whether your extraction model is accurate. They are asking whether the decision that used it can be explained and reproduced.

That is achievable, and it is entirely a design decision made at the start:

  • immutable source retention with document hashing
  • field-level provenance down to the page and region
  • raw text preserved alongside every parsed value
  • explicit, logged post-processing
  • a reconciliation layer that surfaces rather than resolves
  • monitoring on input distribution, not just model confidence

Built this way, the accuracy problem stays tractable and the audit question never arises. Built the other way, accuracy is fine and the audit question ends the project.

In this article

  • document AI
  • OCR
  • banking

Working on something similar?

These articles come from real engagements. If the problem here sounds familiar, a 30-minute call is usually enough to tell you whether we can help.

Start a conversation

Related reading

Continue from here

Articles connected to the same delivery problems.

Have a related problem in front of you?

Send us the problem in whatever detail you have. A senior engineer replies within one business day, and you will get an honest read on whether we are the right partner for it.

We would like to use Google Analytics to understand how this website is used. No analytics are loaded unless you accept. Your choice is stored for six months.

See our Privacy Policy for details.