KYB document check automation starts from a very ordinary moment: a business customer applying for an account sends a folder of files called `articles.pdf`, `IMG_4821.pdf` and `scan_final2.pdf`. For a European payment institution, a compliance officer then opens each file, works out what it is, checks it against the onboarding rulebook, and writes back to the customer about what is missing. Often more than once. We built a document pre-flight that does the first pass in minutes: it identifies every document, checks the whole pack against the institution’s own rules, and produces one complete list of questions for the customer.
The problem: the rules are clear, the documents are not
The institution’s onboarding guidelines are specific. A certificate of incumbency must be no older than six months. A power of attorney without an expiry date is treated as valid for one year. A residence card or a driving licence is not an acceptable identity document. Beneficial ownership must be documented.
The hard part is everything before the rules can be applied:
- File names say nothing. Each file has to be opened to find out whether it is articles of association, a registry extract or a passport.
- Company structures cross borders. A local company owned by a holding in another EU country means documents in two languages and two registry formats.
- Dates hide in the text. Whether a document is expired depends on the issue date, the rule and the day of the review.
- Personal data is everywhere. Birth dates and document numbers should not travel further than they have to.
- Questions come in rounds. Missing documents are found one at a time, so the customer is asked again and again.
What we built
A pre-flight check that reads the pack and prepares the decision, without making it:
- Mask personal data first. Birth dates and document numbers are removed from the files before any text is extracted or any page is rendered.
- Identify every document. Each file is classified and placed into the right slot of the institution’s onboarding API, whatever its name.
- Extract the key fields with a source. Company details, directors, authorised persons, parent company and activities, each with the document and page it came from. Every quote is verified word for word on that page; in the test pack all 52 of 52 quotes matched.
- Apply the rulebook in code. The 24 onboarding rules, taken word for word from the institution’s public guidelines, are applied by ordinary code, not by the model. Change the review date and the results recalculate.
- Produce one list for the customer. Expired documents, unacceptable ID types and missing documents are gathered into a single set of questions. In the test pack that was 16 questions, asked once.
The same approach extends to other regulatory checks, for example reviewing supplier contracts against the contract requirements of the EU Digital Operational Resilience Act (DORA).

What changes for the compliance team
| Before | With KYB document check automation | |
|---|---|---|
| Sorting the pack | Open every file to see what it is | Documents identified and placed in the right slot |
| Finding the data | Read each document for names, dates and numbers | Key fields extracted with document and page |
| Applying the rules | From memory and a checklist | 24 rules applied in code, each quoted from the guidelines |
| Expiry checks | Manual date arithmetic | Calculated for the review date, shown line by line |
| Questions to the customer | Several rounds as gaps are discovered | One complete list |
| Audit trail | Notes in the case file | Every finding linked to its source and rule |
Where people stay in control
The model never decides whether a customer passes. It reads and quotes; the rules are applied by deterministic code that the compliance team can read, and the final onboarding decision is always a person’s. Every field can be clicked to open the exact page it came from. This is also the answer to the question a regulator will ask: how was this decision made, and on what evidence.
The limit is honest too: masking lists are defined per document set, so a new document type needs its personal data fields added before it is processed.
Where else this works
The same pattern of reading, quoting and applying rules in code fits any regulated review:
- Insurance: policy terms turned into structured parameters with a reference to each clause.
- Banking: a clear map of where AI is allowed to help under the EU AI Act.
- Wholesale and construction: documents reconciled against orders and tender scopes.
Related use cases: Supplier invoice reconciliation · EU AI Act for banks · Insurance policy document extraction. All Document checks & reconciliation use cases · How we deliver this: Custom AI solutions
FAQ
Can AI make KYB decisions?
In this setup, no. AI identifies documents and extracts fields with verified quotes. The rules are applied by deterministic code, and a compliance officer makes the decision.
How is personal data protected during the check?
Sensitive fields such as birth dates and document numbers are removed from the files before text extraction and page rendering, so they never reach the later steps.
What happens when a document is missing or expired?
It is listed with the rule it breaks, the rule’s exact wording and, for expiry, the date calculation. All gaps go into one list of questions for the customer.
Want to see how your onboarding rulebook would work on a real document pack? Talk to us