Skip to content
Balkan Tecnologia

AI measured on every release: quality, latency and cost

We measure the answer quality, latency and cost of your AI feature against test cases on every release, and keep tracking all of it in production.

A mint-green thread magnified under a brass lens, on a stack of blank sheets.

Who it's for

For companies that already run AI in production and change instructions or models without knowing if results got worse.

What it solves

  • An instruction tweaked for one case broke others, and only the customer noticed.
  • A new version of the model came out, and nobody knows whether switching helps or hurts.
  • The AI bill went up, and nobody can say which feature or customer used it.
  • Quality is judged by feel, glancing at a few conversations now and then.
  • Slow answers make users give up, but nobody measures response time.

What we do

Test case set
Real questions and tasks with the expected answer, including hard cases and requests that should be refused.
Quality criteria
What counts as a good answer, agreed with your team: correctness, source, format and tone.
Evaluation on every release
A model, instruction or data change goes live only once it has been tested and compared with the current version.
Sampled human review
A share of the answers is judged by people, because automated evaluation gets things wrong too.
Production dashboard
Quality, latency, errors and cost per feature and per customer, with alerts.
Budget and limits
A spending cap per feature, with an alert before the limit and a cut-off once it is reached.
Rolling back a version
If a new version does worse in production, the previous one returns by configuration, with no new release.

What you get

  • A versioned set of test cases in your repository
  • Automated evaluation built into the release process
  • A comparison report for every version: quality, latency and cost
  • A production dashboard with alerts on cost, errors and slowness
  • A rollback procedure to the previous version

How we do it

  1. Inventory

    Which AI features exist, which models they use and what they cost today.

  2. Cases and criteria

    We build the test cases and the definition of a good answer with your team.

  3. Baseline

    We measure the current version so there is something to compare against.

  4. Automation

    The tests start running on every change, before release.

  5. Operation

    Dashboard, alerts and periodic review of the test cases with your team.

Technology examples

  • OpenTelemetry
  • Prometheus
  • Grafana
  • Python
  • PostgreSQL

Related services

FAQ

Can text quality be measured objectively?

Partly. Format, whether a source is cited and forbidden answers can be checked by rule; correctness and tone need comparison with expected answers and sampled human review. We use both together.

Does it work with AI another team built?

Yes. We need access to the calls to the model, the instructions and examples of real use. There is no need to rewrite the feature to start measuring.

Is it worth switching models to save money?

Sometimes. With the test cases, a cheaper model can be compared with the current one to see what is lost in quality before deciding.

Are users' conversations stored?

Only what is needed for measurement, with restricted access and a retention period set by your company. Personal data can be masked before it enters the logs.