Main content

Custom AI agents that run one real process

We build custom AI agents for back-office work: meeting notes, the monthly finance pack, invoice intake. Each agent runs the routine end to end; money and mail still wait for a person.

From our own back office

Admin that runs itself

Of admin a month, now done by agents

Meeting notes filed minutes after the call

From call to filed meeting note

Company finance down to a few approvals

Approvals a month, for company finance

How an AI agent runs one process

Plain code decides when the agent runs and checks its work. A person signs off once.

FIG. 1One process, one approval
Agent pipeline: a trigger such as a new mail, a finished call or the clock starts a quick check with no AI; only new work reaches the agent; rules in plain code check what it returns; a person approves; the result is written where the work lives.A case that fails a rule goes back to a person, flagged with the reason. Every step is logged.TriggerMail, call, clockCheckNo AIAgentClaudeRulesPlain codePerson approvesMoney and mailSystem of recordERP, bank, mailAutomaticA personClient

What a custom AI agent includes

  • A working agent for one process, in the client's environment
  • Access where the team already works: email, Slack, Teams or a web screen
  • Tests built from the client's real cases, run before launch and on every change
  • An action log the client can read, plus a runbook and a handover
Scope
One process per agent
Models
Chosen per step, Claude by default
Cost
Build and monthly run, quoted separately

Discover, build, deploy

Discover
The process is mapped with the people who run it, and the success metric is written down first.
Build
Progress shows on a staging copy, tested against real past cases.
Deploy
The agent goes live in the client's environment, then gets monitored, tuned and handed over.

Rules first: which step goes to code, a model or a person

Some steps are cheaper and safer as plain rules. The first meeting sorts every step.

StepDecided byWhy
The supplier's IBAN checksum is validRuleArithmetic gives the same answer every time.
The amount in words matches the figuresRuleChecked by code, never by a model.
Which mail attachment is a company receiptModelNo two suppliers send the same layout.
A supplier writes that last month's invoice was wrongModelFree text, with the meaning in the wording.
Paying a supplier whose IBAN just changedPersonIrreversible, and a known fraud pattern.

AI agents already at work

LifeOS in one line: a finished call, a new mail and the clock trigger one agent, which files a meeting note, prepares the finance pack for a person to approve and writes the daily log.CallMailClockAgentClaude CodeNoteFinanceLog

When the agent is unsure, a person decides

Payments, customer emails and deletions wait for a yes. Every step is logged.

Approval gateDemo data
Filed
1,184
Held
25

Last held: Total does not match

Six guardrails, and where each one runs

Money and mail wait for a personLifeOS finance
The agent writes the bank's payment file and a mail draft. It has no way to pay or to send: a person imports the file and presses Send. LifeOS case study
Admin changes wait for a yesMozar admin
An agent can change mission types and prices only after a person sees the change and confirms it. Mozar case study
Exact checks stay in codeLifeOS finance
Plain code does every exact check, from bank details to totals. The model only sorts receipts.
Names are never guessedLifeOS meeting notes
A speaker the agent cannot confirm is flagged for a person to check, never filled in.
Access starts read-onlyMozar admin
The agent can look before it can change anything, and every change is shown to a person first.
Every change is on recordMozar
Every change lands in an audit trail that nobody can rewrite.

Prompt injection heads the OWASP Top 10 for LLM applications, and excessive agency, meaning permissions wider than the task, is on the same list. How the agents we build are secured: cybersecurity and NIS2.

Where the data goes

Model input
Only the text a step needs. Personal data becomes placeholders wherever the process allows it.
Training
Anthropic does not train models on API inputs and outputs without explicit permission. Any other provider gets the same check before use.
Region
The first-party Claude API processes requests in the US or globally. When data must stay in the EU, the model runs in an EU region on Amazon Bedrock or Google Vertex AI. Last checked 28 September 2026.
Contract and access
A GDPR data processing agreement (Article 28) is signed before real data reaches the agent, and every account and API key it needs is approved up front.
Billing
Model usage can run on the client's own provider account, so the bill goes straight to the client.

Custom AI agents: common questions

Difference from ChatGPT with a custom prompt

A custom agent runs one defined process with code around the model: integrations with the client's systems, rules that check every value, an approval step and a log. ChatGPT with a prompt answers one person at a time and leaves no record in the client's systems.

In LifeOS, most of the engineering sits outside the model call: the schedule, the checks and the approval points.

AI agent vs chatbot

A chatbot answers questions. An agent takes actions: it reads a document, updates a record or drafts a payment, within permissions set in advance. Many projects need both: a RAG chatbot for company documents answers, and an agent acts on what it reads.

When the agent is wrong or unsure

An agent that is unsure hands the case to a person, with the reason and the source. Irreversible steps (payments, emails to customers, deletions) always wait for approval, and every decision lands in the log.

Wrong answers found after launch become test cases, so the same mistake fails before the next release.

Data location, model training and EU hosting

Data goes only to the systems listed before the build, and the model receives only the text a step needs. Anthropic does not train models on API data without explicit permission. When data must stay in the EU, the model runs in an EU region on Amazon Bedrock or Google Vertex AI. The full list is under where the data goes.

Cost to build and cost to run

The build is quoted after the process is mapped, because the integrations drive most of the effort. The run cost is estimated per month from the expected volume and the model chosen for each step, and it appears as a separate line.

The AI workshop produces both estimates for one process.

Time to production

It depends on how many systems the agent touches and how clean the inputs are. The first release covers one process and one metric, so it reaches production before a second process starts. The workshop brief gives the estimate for a specific process.

AI models used, and switching later

Claude is the default, and each step gets the smallest model that passes its tests. The model call sits behind one interface, so switching provider means a configuration change and a rerun of the tests.

EU AI Act and the agents built here

Most back-office agents fall outside the Act's high-risk categories, which cover areas such as hiring, credit scoring and critical infrastructure. The transparency duties in Article 50 still apply when people interact with an AI system. Each project's classification is checked during scoping and written into the brief.

One process on the table: the AI workshop

Where an agent pays off, and what it costs to build and run.

Studio
Cluj-Napoca, Romania, EU