Skip to content

What we build

We build the part that has to work.

Four things, each measured in production: the agents, the evaluation that keeps them honest, the integrations they run on and the people who operate them.

01

Agents in production

A demonstration works on ten examples and fails on the eleventh. In an operation the eleventh case arrives every hour, and someone has to own what happens next.

What we deliver
A task-specific agent inside the tools your team already uses, with the data it is allowed to see, a path for the cases it must not handle alone, a log of every action and a named owner for the outcome.
How we measure it
Against the evaluation harness built in the second step: accuracy, escalation rate, cost and latency per case, reported on every change and watched in production.
What you own
The agent, its prompts and configuration, the evaluation set, the logs and a runbook your team can follow without us.
Where it connects
Ticketing and service desks, CRM, ERP, document repositories, email and chat platforms.
Agents in productionAPI

02

Evaluation harnesses

A model or prompt change looks better on a handful of examples, ships, and quietly degrades. Without a harness, improvement and luck are indistinguishable.

What we deliver
A test set drawn from your real cases, the metrics that match the task, baselines, confidence intervals and a report that runs on every change, from the first prototype to the version in production.
How we measure it
By the harness itself: a change ships only when it beats the baseline by more than chance, at a sample size that can tell the difference.
What you own
The evaluation set, the scoring code, the baselines and the history of every run, in your own repository.
Where it connects
Model APIs and open models, your data warehouse, continuous integration pipelines, monitoring and alerting.
Evaluation harnesses708090100v1v2

03

Integration and automation

The model is the easy part. The work is the plumbing between it and the systems of record, and that plumbing has to stand up when finance, legal or a regulator asks what happened.

What we deliver
Connectors and automations between the model and ERP, CRM, ticketing, finance and document systems, with authentication, permissions, retries, fallbacks and a log of every read and write.
How we measure it
Throughput, error rate, time saved per case and reconciliation against the source system, reported with sources and dates.
What you own
The integration code, the access configuration, the audit trail and the documentation needed to change any of it.
Where it connects
ERP, CRM, ticketing, finance and accounting, document management, spreadsheets and databases.
Model
ERP
CRM
Ticketing
Finance
Documents

04

AI training for organisations

Teams are asked to adopt AI without the vocabulary to judge a proposal or the practice to use a tool well. Slides do not change that.

What we deliver
Programmes for companies, schools, foundations and individuals: executive briefings, hands-on workshops for engineering teams and teacher training, in Portuguese or English, built on the systems we ship.
How we measure it
By what participants can do afterwards: an agent built and evaluated in the workshop, a policy drafted, a lesson designed and tested.
What you own
The materials and the exercises, and a team that can carry on the practice without us.
Where it connects
Your own tools and data where permitted; otherwise generic examples drawn from the same kind of process.
AI training for organisations01020304

Bring us the process that keeps breaking.

One conversation is usually enough to tell whether an agent belongs there, what it should be measured against, and what it would take to keep it running.