Data from PDFs and images, with human review
We read PDFs, photos and scans, pull out the fields your system needs and send anything uncertain to a person to check.

Who it's for
For companies that retype data from boletos, receipts, contracts or forms that arrive as PDFs or photos.
What it solves
- The team spends the day copying data from PDFs into the system, field by field.
- Every supplier sends documents in a different layout, and fixed rules break with each new one.
- Typing errors only show up later, at payment or reconciliation.
- An automated tool was tried before, but it got things wrong silently and nobody knew what to check.
What we do
- Field schema
- Which fields to take from each document type, in what format and under which rules.
- Reading with OCR and models
- OCR and language models fill in the schema; each field comes with its page and source passage.
- Rule-based validation
- CPF and CNPJ check digits, sums, dates and amounts checked against your records.
- Human review queue
- Anything that fails a rule or comes back uncertain goes to a person, with the document alongside.
- Into your system
- Approved fields go into the ERP or internal workflow through an API, recording who approved them.
- Corrections become tests
- Every correction made in review joins the test set that measures the next version.
What you get
- An extraction workflow in production for the agreed document types
- A review queue showing the document and where each field came from
- Validation rules for each document type
- A record for every document: what was extracted, corrected and approved
- A set of test documents to measure every change
How we do it
Sample
We gather real documents of each type, including the hardest to read.
Schema and rules
Fields, formats and validations defined with the people who do the typing today.
Measurement
We run the sample and compare the result with what your team typed.
Live, with review
At first a person checks everything; review is reduced only where the measurements allow.
Follow-up
Corrections and new documents join the tests before each new version.
Technology examples
- OCR
- Python
- JSON Schema
- XML
- PostgreSQL
- RabbitMQ
Related services
FAQ
What is the accuracy rate?
We do not give a number before measuring with your own documents: it depends on image quality and how varied the layouts are. We measure on the sample and decide together what can go through without review.
Can human checking be removed altogether?
For some fields and documents it can be reduced. Where a mistake is costly, as with amounts and bank details, it makes sense to keep a person in the loop. The level of review is your team's decision.
Does it work for tax invoices?
For an NF-e, the right path is the XML, which already carries structured data; we cover that under tax document integration. Extraction makes sense for what has no XML: receipts, contracts, forms and older documents.
What about handwritten documents?
Handwriting and poor photos are the hardest cases. They go into the sample test; if the result is not good enough, those documents go to human review from the start.
Write to Balkan
Talk to us about your project