Skip to content
Balkan Tecnologia

Know something broke before your customer tells you

We instrument your system with logs, metrics, tracing and alerts, test that backups actually restore and leave a plan for when something fails.

A mint-green thread stretched over a brass pulley, unwinding from a wooden spool.

Who it's for

For companies that find out about failures from a customer complaint.

What it solves

  • The system slows down and nobody can say which part is to blame.
  • The first sign that something is down is a message from a customer.
  • Alerts fire all the time, and the team has learnt to ignore them.
  • There are backups, but nobody has ever tried restoring one.
  • After each incident the same problem comes back, because nobody recorded the cause.

What we do

Logs and tracing
Searchable records and the path of each request across services, without needless personal data.
Metrics and dashboards
Traffic, errors, latency and queues for each service, on dashboards the team understands.
Service level objectives (SLOs)
Availability and response-time targets agreed with you for flows such as confirming a Pix payment.
Alerts that call for action
Every alert has an owner, a severity and a runbook; the noise gets cut.
Backup and restore
Automatic backups and a real restore rehearsal in a separate environment, with the time it took recorded.
Post-incident review
After each significant failure, the cause, timeline and actions are recorded, without hunting for culprits.

What you get

  • Logs, metrics and tracing instrumented in the main services
  • Dashboards per service and per business flow
  • Alert rules, each with its own response runbook
  • A report on the backup restore rehearsal
  • A post-incident review template, ready for your team to use

How we do it

  1. System map

    Services, dependencies and the flows that cannot stop.

  2. Instrumentation

    Logs, metrics and tracing added with as little code change as possible.

  3. Targets and alerts

    Service objectives agreed with you, and alerts tied to them.

  4. Restore rehearsal

    A backup actually restored, with the outcome and the time taken recorded.

  5. Routine

    Regular review of alerts and incident reviews with your team.

Technology examples

  • OpenTelemetry
  • Prometheus
  • Grafana
  • Loki
  • Sentry
  • pgBackRest

Related services

FAQ

Who gets the alerts in the middle of the night?

It depends on what is agreed. Alerts can go to your team, to us or to both. Any out-of-hours cover is set out in the maintenance contract.

Do we need a paid monitoring tool?

Not necessarily. Everything can be built with open-source tools or on a paid service; the choice weighs cost, data volume and who will run it.

What is an SLO, in practice?

A target agreed for a flow and measured continuously: for example, how many payment confirmations must complete within a set time. When the target is at risk, stability comes before new features.

What if the system still goes down?

Observability does not prevent every failure. It shortens the time until someone notices and finds the cause, and the review afterwards makes a repeat less likely.