What we build
We build the part that has to work.
Four things, each measured in production: the agents, the evaluation that keeps them honest, the integrations they run on and the people who operate them.
01
Agents in production
A demonstration works on ten examples and fails on the eleventh. In an operation the eleventh case arrives every hour, and someone has to own what happens next.
- What we deliver
- A task-specific agent inside the tools your team already uses, with the data it is allowed to see, a path for the cases it must not handle alone, a log of every action and a named owner for the outcome.
- How we measure it
- Against the evaluation harness built in the second step: accuracy, escalation rate, cost and latency per case, reported on every change and watched in production.
- What you own
- The agent, its prompts and configuration, the evaluation set, the logs and a runbook your team can follow without us.
- Where it connects
- Ticketing and service desks, CRM, ERP, document repositories, email and chat platforms.
02
Evaluation harnesses
A model or prompt change looks better on a handful of examples, ships, and quietly degrades. Without a harness, improvement and luck are indistinguishable.
- What we deliver
- A test set drawn from your real cases, the metrics that match the task, baselines, confidence intervals and a report that runs on every change, from the first prototype to the version in production.
- How we measure it
- By the harness itself: a change ships only when it beats the baseline by more than chance, at a sample size that can tell the difference.
- What you own
- The evaluation set, the scoring code, the baselines and the history of every run, in your own repository.
- Where it connects
- Model APIs and open models, your data warehouse, continuous integration pipelines, monitoring and alerting.
03
Integration and automation
The model is the easy part. The work is the plumbing between it and the systems of record, and that plumbing has to stand up when finance, legal or a regulator asks what happened.
- What we deliver
- Connectors and automations between the model and ERP, CRM, ticketing, finance and document systems, with authentication, permissions, retries, fallbacks and a log of every read and write.
- How we measure it
- Throughput, error rate, time saved per case and reconciliation against the source system, reported with sources and dates.
- What you own
- The integration code, the access configuration, the audit trail and the documentation needed to change any of it.
- Where it connects
- ERP, CRM, ticketing, finance and accounting, document management, spreadsheets and databases.
04
AI training for organisations
Teams are asked to adopt AI without the vocabulary to judge a proposal or the practice to use a tool well. Slides do not change that.
- What we deliver
- Programmes for companies, schools, foundations and individuals: executive briefings, hands-on workshops for engineering teams and teacher training, in Portuguese or English, built on the systems we ship.
- How we measure it
- By what participants can do afterwards: an agent built and evaluated in the workshop, a policy drafted, a lesson designed and tested.
- What you own
- The materials and the exercises, and a team that can carry on the practice without us.
- Where it connects
- Your own tools and data where permitted; otherwise generic examples drawn from the same kind of process.
Bring us the process that keeps breaking.
One conversation is usually enough to tell whether an agent belongs there, what it should be measured against, and what it would take to keep it running.