Skip to content
Balkan Tecnologia

Data from PDFs and images, with human review

We read PDFs, photos and scans, pull out the fields your system needs and send anything uncertain to a person to check.

A mint-green thread magnified under a brass lens, on a stack of blank sheets.

Who it's for

For companies that retype data from boletos, receipts, contracts or forms that arrive as PDFs or photos.

What it solves

  • The team spends the day copying data from PDFs into the system, field by field.
  • Every supplier sends documents in a different layout, and fixed rules break with each new one.
  • Typing errors only show up later, at payment or reconciliation.
  • An automated tool was tried before, but it got things wrong silently and nobody knew what to check.

What we do

Field schema
Which fields to take from each document type, in what format and under which rules.
Reading with OCR and models
OCR and language models fill in the schema; each field comes with its page and source passage.
Rule-based validation
CPF and CNPJ check digits, sums, dates and amounts checked against your records.
Human review queue
Anything that fails a rule or comes back uncertain goes to a person, with the document alongside.
Into your system
Approved fields go into the ERP or internal workflow through an API, recording who approved them.
Corrections become tests
Every correction made in review joins the test set that measures the next version.

What you get

  • An extraction workflow in production for the agreed document types
  • A review queue showing the document and where each field came from
  • Validation rules for each document type
  • A record for every document: what was extracted, corrected and approved
  • A set of test documents to measure every change

How we do it

  1. Sample

    We gather real documents of each type, including the hardest to read.

  2. Schema and rules

    Fields, formats and validations defined with the people who do the typing today.

  3. Measurement

    We run the sample and compare the result with what your team typed.

  4. Live, with review

    At first a person checks everything; review is reduced only where the measurements allow.

  5. Follow-up

    Corrections and new documents join the tests before each new version.

Technology examples

  • OCR
  • Python
  • JSON Schema
  • XML
  • PostgreSQL
  • RabbitMQ

Related services

FAQ

What is the accuracy rate?

We do not give a number before measuring with your own documents: it depends on image quality and how varied the layouts are. We measure on the sample and decide together what can go through without review.

Can human checking be removed altogether?

For some fields and documents it can be reduced. Where a mistake is costly, as with amounts and bank details, it makes sense to keep a person in the loop. The level of review is your team's decision.

Does it work for tax invoices?

For an NF-e, the right path is the XML, which already carries structured data; we cover that under tax document integration. Extraction makes sense for what has no XML: receipts, contracts, forms and older documents.

What about handwritten documents?

Handwriting and poor photos are the hardest cases. They go into the sample test; if the result is not good enough, those documents go to human review from the start.