Skip to content
Balkan Tecnologia

Answers from your own documents, with the source alongside

We build search that answers questions from your company's documents, shows where each answer came from and uses only what each person is allowed to see.

A mint-green thread magnified under a brass lens, on a stack of blank sheets.

Who it's for

For companies where the answer sits in a manual, contract or policy, but nobody finds it in time.

What it solves

  • The same question reaches the legal or support team several times a week.
  • Information is spread across folders, the wiki, emails and scanned PDFs.
  • The current search only finds the exact word, not the subject.
  • A generic assistant answers confidently but never says where the answer came from.
  • Not everyone may see every document, and the search does not know that.

What we do

Collecting and refreshing documents
Folders, wiki and systems read on a schedule; changes are updated and deleted documents leave the index.
Preparing the text
Documents split into passages, with tables kept intact and OCR for scanned files.
Search by meaning and by term
Semantic search combined with keyword search, to find both the subject and the exact code.
Answers with sources
Each answer names the document and passage used; without enough source material, the system says it does not know.
Permissions respected
The search considers only documents the person asking is already allowed to open.
Personal data in documents
With the LGPD in mind: what goes into the index, who can query it and how long it stays.
Test questions
Real questions with checked answers, used to measure every change before release.

What you get

  • A search API ready for your product, intranet or internal chat
  • A collection routine that keeps the index up to date
  • Answers linked to the source document and passage
  • A set of test questions, with a report on answer quality
  • A query log with restricted access

How we do it

  1. Document set

    Which documents go in, where they are, who may see each and which hold personal data.

  2. Test questions

    We gather real questions and the right answers with your team.

  3. First version

    Collection, index and sourced answers for part of the documents, measured against the questions.

  4. Expansion

    The rest of the documents come in, with permissions checked source by source.

  5. Follow-up

    Unanswered questions and disputed answers become new test cases.

Technology examples

  • PostgreSQL
  • pgvector
  • OpenSearch
  • Python
  • OCR
  • OpenTelemetry

In Brazil, we handle

LGPD
Brazil's General Personal Data Protection Law (Law 13,709/2018).

Related services

FAQ

What if the documents do not contain the answer?

The system should say it found nothing rather than make something up. This is set up and tested with questions whose answers are not in the documents; it can still get things wrong, which is why the source is always there to check.

Does it work with scanned PDFs?

Yes, with OCR. Quality depends on the scan: skewed pages, stamps and handwriting make results worse. We assess a sample at the start; forms and documents with fixed fields are better served by data extraction.

Do our documents leave the company?

It depends on the model chosen. With an external provider, passages are sent with each question; with a model in your own cloud, they stay in your environment. The decision is your company's, with the LGPD in mind.

Do we need to organise the documents first?

They do not have to be perfect, but old and duplicate versions muddle the answers. While surveying the documents, we point out what is worth cleaning up first.