Skip to content

Home » Proiecte » AI document processing: data extracted, checked and exported

Typical project

AI document processing: data extracted, checked and exported

How we build a workflow that reads scanned invoices, contracts and forms, extracts the fields, sends uncertain cases for review and exports the data to your company’s software.

An operator opens a PDF, reads the supplier, the number, the date and the total, then types them into the accounting software. They repeat this hundreds of times a month, and at some point a transposed figure turns up. An AI document processing system takes over the reading and the typing and leaves the person what people do better: checking the unclear cases and deciding.

It is worth saying at the outset what there is no point putting through AI. Invoices received from Romanian suppliers through e-Factura, Romania’s national e-invoicing system, already arrive in a structured format, so they are imported directly, with no reading at all. AI extraction helps with the rest: invoices from abroad, till receipts, contracts, delivery notes, forms filled in by hand, the scanned archive.

Which kinds of document are suitable and which are not

Documents with a repetitive structure and printed text work best: invoices, delivery notes, statements, orders, insurance policies, standard forms. Even though every supplier lays out its invoice differently, the fields we are looking for are the same, and current models find them without a separate template for each issuer.

Contracts are a different case. The factual elements can be extracted: the parties, the dates, the term, the value, the payment deadlines. Interpreting a clause remains the lawyer’s job; the system can only show them where it is in the text.

What works poorly: hurried handwriting, photos taken at an angle with a phone, copies of copies and stamps placed over figures. Here, human review will be the rule, not the exception.

A document’s journey, from scan to your company’s software

Say you receive, by email every day, invoices from foreign suppliers and delivery notes scanned at the warehouse. The workflow picks them up from there, with no manual upload.

Of the steps below, it is the review screen that decides whether operators adopt the system: the document sits on the left, the extracted fields on the right, and selecting a field highlights the place on the page it comes from.

  • Intake: from email, a shared folder or the scanner; files are classified by document type.
  • Reading: skewed pages are straightened, then OCR (optical character recognition) turns the image into text.
  • Extraction: a language model returns the requested fields in a fixed structure, including table lines.
  • Checks: rules that verify the result, independently of the model.
  • Human review: only for documents that fail the checks or have missing fields.
  • Export: the confirmed data goes into the accounting software or the ERP, through an API (its data exchange interface) or an import file.

Why we do not rely on the model’s “confidence”

A language model does not provide a measure of confidence you can depend on: it can return a wrong figure just as confidently as a right one. That is why the decision “goes through automatically or goes to a person” is taken by the rules, not by the model.

The rules use what can be calculated. The line amounts must add up to the total. The VAT must correspond to the taxable amount. A Romanian tax identification code (CUI) and an IBAN have check digits that can be verified mathematically. The same pair of supplier plus invoice number must not come in twice.

For important fields we can also run a second, independent reading with a different tool and send any mismatch to a person. It costs more per document, so we apply it only where the stakes call for it.

Where sensitive documents are processed

Contracts, identity documents, medical documents and payroll records contain personal data and trade secrets. Before choosing the technology, we settle with your data protection officer which documents enter the workflow and on what conditions.

There are three options. The first: cloud services that process in a European Union region under a data processing agreement, where we check whether documents are used for training. The second: models run on your own infrastructure, with an open-source OCR engine such as Tesseract and a local language model; the data goes nowhere, but accuracy and the cost of the hardware have to be weighed up. The third: a combination, in which only ordinary documents go to the cloud.

Whichever option is chosen, we keep the original document, the extracted data and the name of the person who confirmed it, with role-based access rights.

How we measure an AI document processing workflow before go-live

We do not promise an accuracy figure before seeing your documents. We ask for a representative sample, the ugly ones included, run it through the workflow and compare the result, field by field, with the data entered correctly by your people.

After go-live, the operators’ corrections are analysed periodically: they show which rules need adding and which issuers consistently send hard-to-read documents. Before go-live, the trial answers a few questions:

  • How many documents go through with no intervention at all.
  • Which fields attract the most corrections and why.
  • How long it takes to review a document compared with entering it from scratch.
  • Which document types should be left out of the workflow for now.

Frequently asked questions

Are people still needed if documents are read by AI?

Yes, for the unclear cases and for spot checks on the ones passed automatically. The work shifts from typing to checking.

Does it work with documents in other languages?

Current models read the main European languages well, including Romanian with diacritics. For less common alphabets or formats, we check on your sample.

How is the cost per document made up?

Of the monthly page volume, the variety of document types, where the processing happens (in the cloud or on your own infrastructure) and the complexity of the export. Cloud reading services are generally charged per page or by usage.

This is a typical project description: it shows how we usually approach this kind of work and does not present a project carried out for a particular client. Every real project starts from your company’s situation, and the stages, timescales and price are agreed after the initial discussion.

Want a project like this?

Write us a few lines about what you want to build or what no longer works. We will reply with concrete steps and a quote with a price for each stage.

WhatsApp