AI measured on every release: quality, latency and cost
We measure the answer quality, latency and cost of your AI feature against test cases on every release, and keep tracking all of it in production.

Who it's for
For companies that already run AI in production and change instructions or models without knowing if results got worse.
What it solves
- An instruction tweaked for one case broke others, and only the customer noticed.
- A new version of the model came out, and nobody knows whether switching helps or hurts.
- The AI bill went up, and nobody can say which feature or customer used it.
- Quality is judged by feel, glancing at a few conversations now and then.
- Slow answers make users give up, but nobody measures response time.
What we do
- Test case set
- Real questions and tasks with the expected answer, including hard cases and requests that should be refused.
- Quality criteria
- What counts as a good answer, agreed with your team: correctness, source, format and tone.
- Evaluation on every release
- A model, instruction or data change goes live only once it has been tested and compared with the current version.
- Sampled human review
- A share of the answers is judged by people, because automated evaluation gets things wrong too.
- Production dashboard
- Quality, latency, errors and cost per feature and per customer, with alerts.
- Budget and limits
- A spending cap per feature, with an alert before the limit and a cut-off once it is reached.
- Rolling back a version
- If a new version does worse in production, the previous one returns by configuration, with no new release.
What you get
- A versioned set of test cases in your repository
- Automated evaluation built into the release process
- A comparison report for every version: quality, latency and cost
- A production dashboard with alerts on cost, errors and slowness
- A rollback procedure to the previous version
How we do it
Inventory
Which AI features exist, which models they use and what they cost today.
Cases and criteria
We build the test cases and the definition of a good answer with your team.
Baseline
We measure the current version so there is something to compare against.
Automation
The tests start running on every change, before release.
Operation
Dashboard, alerts and periodic review of the test cases with your team.
Technology examples
- OpenTelemetry
- Prometheus
- Grafana
- Python
- PostgreSQL
Related services
FAQ
Can text quality be measured objectively?
Partly. Format, whether a source is cited and forbidden answers can be checked by rule; correctness and tone need comparison with expected answers and sampled human review. We use both together.
Does it work with AI another team built?
Yes. We need access to the calls to the model, the instructions and examples of real use. There is no need to rewrite the feature to start measuring.
Is it worth switching models to save money?
Sometimes. With the test cases, a cheaper model can be compared with the current one to see what is lost in quality before deciding.
Are users' conversations stored?
Only what is needed for measurement, with restricted access and a retention period set by your company. Personal data can be masked before it enters the logs.
Write to Balkan
Talk to us about your project