Answers from your own documents, with the source alongside
We build search that answers questions from your company's documents, shows where each answer came from and uses only what each person is allowed to see.

Who it's for
For companies where the answer sits in a manual, contract or policy, but nobody finds it in time.
What it solves
- The same question reaches the legal or support team several times a week.
- Information is spread across folders, the wiki, emails and scanned PDFs.
- The current search only finds the exact word, not the subject.
- A generic assistant answers confidently but never says where the answer came from.
- Not everyone may see every document, and the search does not know that.
What we do
- Collecting and refreshing documents
- Folders, wiki and systems read on a schedule; changes are updated and deleted documents leave the index.
- Preparing the text
- Documents split into passages, with tables kept intact and OCR for scanned files.
- Search by meaning and by term
- Semantic search combined with keyword search, to find both the subject and the exact code.
- Answers with sources
- Each answer names the document and passage used; without enough source material, the system says it does not know.
- Permissions respected
- The search considers only documents the person asking is already allowed to open.
- Personal data in documents
- With the LGPD in mind: what goes into the index, who can query it and how long it stays.
- Test questions
- Real questions with checked answers, used to measure every change before release.
What you get
- A search API ready for your product, intranet or internal chat
- A collection routine that keeps the index up to date
- Answers linked to the source document and passage
- A set of test questions, with a report on answer quality
- A query log with restricted access
How we do it
Document set
Which documents go in, where they are, who may see each and which hold personal data.
Test questions
We gather real questions and the right answers with your team.
First version
Collection, index and sourced answers for part of the documents, measured against the questions.
Expansion
The rest of the documents come in, with permissions checked source by source.
Follow-up
Unanswered questions and disputed answers become new test cases.
Technology examples
- PostgreSQL
- pgvector
- OpenSearch
- Python
- OCR
- OpenTelemetry
In Brazil, we handle
- LGPD
- Brazil's General Personal Data Protection Law (Law 13,709/2018).
Related services
FAQ
What if the documents do not contain the answer?
The system should say it found nothing rather than make something up. This is set up and tested with questions whose answers are not in the documents; it can still get things wrong, which is why the source is always there to check.
Does it work with scanned PDFs?
Yes, with OCR. Quality depends on the scan: skewed pages, stamps and handwriting make results worse. We assess a sample at the start; forms and documents with fixed fields are better served by data extraction.
Do our documents leave the company?
It depends on the model chosen. With an external provider, passages are sent with each question; with a model in your own cloud, they stay in your environment. The decision is your company's, with the LGPD in mind.
Do we need to organise the documents first?
They do not have to be perfect, but old and duplicate versions muddle the answers. While surveying the documents, we point out what is worth cleaning up first.
Write to Balkan
Talk to us about your project