Skip to content
Mirai Minds

AI agent development for business workflows

In short

Mirai Minds builds AI agents that work inside a company's own systems. They call internal APIs through MCP, operate browsers, remember users across sessions and pass risky steps to a person for approval. Every tool call is traced, and each release is checked against replayed history and an LLM judge. Most builds ship in two to eight weeks.

Problems we take on, and the systems we ship for them

Operations assistant over your APIs

ProblemStaff answer routine questions by clicking through several screens of an internal platform.

SystemA chat agent reaches the platform's APIs as MCP tools. One model picks the tool, a second writes the arguments against the tool's schema, and validation errors go back to the model before anything runs.

ResultAnswers come from live records, and malformed tool calls are caught before they reach the API.

Browser and computer-use agents

ProblemUseful work sits in apps that have no API.

SystemA desktop app runs an agent that reads screenshots, clicks and types in a browser, calls MCP tools, runs scheduled tasks and keeps long-term memory with Mem0.

ResultA working prototype that completes multi-step browser tasks on request.

Research pipelines with an approval gate

ProblemA database has to grow every day, but nothing unverified can reach customers.

SystemA crawler reads news feeds every 6 hours, AI gates extract and verify entities, and a reviewer approves each record from a side-by-side diff.

ResultThe database grows without unreviewed records going live.

Agents that are graded every night

ProblemAgent quality drifts after launch, and nobody notices until a customer does.

SystemA nightly LLM judge scores the previous day's outputs from 0 to 10 against a fixed list of issues. Reviewer feedback overrides the judge.

ResultQuality problems show up on a dashboard the operations team checks daily.

Prompt changes tested on real outcomes

ProblemPrompts get edited by feel, and regressions ship.

SystemAn optimizer built on GEPA replays past conversations. Each candidate prompt is scored 60% on matching the real outcome and 40% on an LLM judge's quality score.

ResultPrompt changes are compared on the same history before they reach production.

How the pieces fit together

Reference architecture for AI agents and automationINPUTRequest or scheduleMODELPlanner model picksa toolSYSTEMArguments checkedagainst the toolschemaHUMAN REVIEWPerson approvesrisky actionsSYSTEMYour APIs via MCPOUTPUTAnswer and audittrail

How it flows

  1. 01 Request or schedule → Planner model picks a tool
  2. 02 Planner model picks a tool → Arguments checked against the tool schema
  3. 03 Arguments checked against the tool schema → Person approves risky actions
  4. 04 Arguments checked against the tool schema → Your APIs via MCP
  5. 05 Person approves risky actions → Your APIs via MCP
  6. 06 Your APIs via MCP → Answer and audit trail

Published

What is an AI agent, in practice?

An AI agent is a loop. A language model reads the request, picks a tool, reads what the tool returns and decides what to do next, until the task is done or it has to ask. The model is the easy part. The work is in the tools, the permissions around them and the tests that tell you when the loop goes wrong.

Mirai Minds builds agents that live inside existing systems. For a US logistics software company we exposed the client's platform APIs as MCP tools and put a chat agent in front of them. For Rouh we built a desktop agent that operates a browser from screenshots, for tasks that have no API at all. For iKoMatch the agent is a pipeline that discovers, verifies and matches investors every day.

How do you keep a person in the loop?

We decide with you which actions an agent may take alone. Reading data usually needs no approval. Writing records, spending money or messaging a customer does. Those steps pause and show a person what will happen, with the evidence the agent used. In the iKoMatch data pipeline, no investor record goes live until a reviewer approves it from a side-by-side diff.

Every tool call is logged with its inputs, outputs and timing. When something goes wrong, you can see which step failed and why, instead of rereading a chat transcript.

How do you know it works?

Before we build, we collect real examples of the task and agree what a correct result looks like. That becomes the test set. Each release runs against it, an LLM judge scores the results, and a person reviews the cases the judge flags. For iKoMatch, a judge re-scores the previous day's recommendations every night and files concrete fixes. On our Voice Agents platform, a GEPA-based optimizer tests prompt changes against past calls before they ship.

Judges are models too, and they drift. We check them against human labels and let reviewer feedback override them.

What does a typical build look like?

Week 1 is access, examples and the test set. Weeks 2 to 6 cover the tools, the agent loop, approval flows and tracing, with a working version running on real data early. Then comes a launch with a person reviewing outputs, and a period of tuning prompts and tools against what production shows. Most builds ship in two to eight weeks.

If the agent needs to talk on the phone or on WhatsApp, it pairs with our voice AI agents and WhatsApp AI agents work. If it needs to read documents reliably, see custom LLM and document AI.

Where we've built this

What we usually build it with

Chosen per project. We'll tell you when something simpler will do.

Agent frameworks
OpenAI Agents SDKModel Context Protocol (FastMCP)Custom tool loops
Models
GPT-4.1ClaudeGeminiLlama 4 via OpenRouter
Memory
Mem0Redis session summaries
Tracing and evals
LangSmithLLM judgesGEPA
Automation
PlaywrightScheduled jobs
Backend
FastAPIPostgreSQLMongoDBRedis

How the engagement runs

  1. Week 01

    Discovery call

    Thirty minutes with an engineer. You describe the job; we say whether AI is the right tool and what it would take.

  2. Week 12

    Scope and evals

    We agree what done looks like and build a test set from your real data before writing the system.

  3. Weeks 2–63

    Build

    Working software every week, measured against the test set, with your team trying it early.

  4. Launch4

    With human review

    The system goes live with a person checking the risky steps and a clear way to reach a human.

  5. Ongoing5

    Run and improve

    We watch real traffic, fix what breaks and tune on real cases, or hand over with documentation.

Asked on the first call

When is an AI agent the wrong choice?

When the steps never change. A fixed script, a form or a scheduled job is cheaper and easier to test. Agents earn their cost when inputs vary, the right tool depends on the request, and a person would otherwise read, decide and click through several systems.

How do you stop an agent from doing something it shouldn't?

Three layers. The agent only sees the tools we expose, scoped to the account it acts for. Every tool call is checked against the tool's schema before it runs. Actions that write data, spend money or message a customer wait for a person to approve them, and every call is logged.

What do you need from us to start?

Access to the systems the agent will use (a sandbox is ideal), 30 to 50 real examples of the task with the right outcome, and one person who owns the process and can say what correct looks like. We turn the examples into the first test set.

How do you know the agent is good enough to launch?

We agree a pass bar before building, on a test set drawn from your real history. Each release runs against it, an LLM judge scores the results and a person checks the cases the judge flags. After launch, a judge can re-score a sample of live runs every night, as it does for iKoMatch.

Which models do you use?

Whichever fits each step. We have shipped agents on GPT-4.1, Claude, Gemini and Llama 4, sometimes two in one agent: a stronger model to plan and another to fill in tool arguments. The model sits behind an interface, so it can be swapped when prices or quality change.

Where does our data go?

The agent runs in your cloud or in ours and calls model providers through their business APIs. Credentials stay in your secret store. We log tool calls for debugging and audit, and agree retention with you before launch.

What drives the cost of an agent project?

The number of systems the agent touches, how many actions need approval flows, how much test data already exists, and model spend at your volume. A read-only assistant over one API is the smallest build; agents that write to several systems take longer.

Other services

Have a system in mind? Let's scope it.

A 30-minute call with an engineer who has shipped this before. You leave with a plan, a rough timeline and what it would take — whether or not we build it.