Skip to content
Balkan Tecnologia

Language models inside your product, with limits

We connect language models to your product with versioned instructions, validated response formats, cost and usage limits, and the option to change provider without a rewrite.

A mint-green thread magnified under a brass lens, on a stack of blank sheets.

Who it's for

For companies that want an AI feature in their product without losing control of cost or data.

What it solves

  • The prototype worked in the demo, but nobody knows what it will cost with every user on board.
  • The model's instructions are scattered through the code, and nobody knows which version is live.
  • Sometimes the model answers in a format that breaks the screen or the next step.
  • When the provider slows down or goes offline, the whole feature stops.
  • It is unclear what customer data leaves the company with each call.

What we do

Model access layer
A single point in the back-end talks to the provider; the rest of the product does not depend on it.
Versioned instructions
Prompts kept as code, with history, review and a way back to the previous version.
Structured, validated responses
Responses follow a defined format and are checked before reaching a screen or another system.
Cost and usage limits
Caps per user, per customer and per period, with an alert before the budget runs out.
A fallback when the provider fails
Timeouts, retries and, if agreed, an alternative model or a default response.
Control over the data sent
Only what is needed goes to the model; personal data is removed or masked where the case allows.
A record of every call
Cost, response time and instruction version are recorded for ongoing evaluation.

What you get

  • An AI feature in production, integrated with your back-end
  • A model access layer, with the provider switchable by configuration
  • Versioned instructions in your repository
  • Cost limits and usage alerts
  • A document on what the feature does and where it can get things wrong

How we do it

  1. Use case

    What task the model does, with what data, and what happens when it gets it wrong.

  2. Measured prototype

    We test with real examples and measure quality, response time and cost.

  3. Integration

    Access layer, limits and call records wired into your back-end.

  4. Gradual release

    First to a small group of users, then to everyone.

  5. Follow-up

    Cost, errors and complaints tracked, with adjustments to the instructions.

Technology examples

  • Python
  • TypeScript
  • JSON Schema
  • OpenAPI
  • Redis
  • OpenTelemetry

Related services

FAQ

Which model do you use?

It is chosen case by case, depending on the task, the cost and where the data is allowed to go. It can be an external provider or a model running in your own cloud.

Can the model get things wrong?

Yes, and the product has to allow for it. That is why responses are validated, sensitive tasks go through a person and quality is measured against test cases before every change.

Will our data be used to train someone else's model?

That depends on the provider's terms and the plan you have. We check this before starting and send only what is needed; if the terms do not suit you, a model in your own cloud is an option.

How much does it cost to run each month?

It depends on usage volume, message size and the chosen model. During the prototype we measure the cost per use and project the monthly figure; the limits keep consumption within the agreed budget.