# AI agent development for business workflows

> Mirai Minds builds AI agents that work inside a company's own systems. They call internal APIs through MCP, operate browsers, remember users across sessions and pass risky steps to a person for approval. Every tool call is traced, and each release is checked against replayed history and an LLM judge. Most builds ship in two to eight weeks.

## Key facts

- **Typical build:** 2–8 weeks
- **Tool access:** MCP servers, REST APIs, browsers
- **Human in the loop:** Approval before writes, payments and outbound messages
- **Evaluation:** Replayed history and an LLM judge before each release
- **Hosting:** Your cloud or ours


## What we build

### Operations assistant over your APIs

- Problem: Staff answer routine questions by clicking through several screens of an internal platform.
- System: A chat agent reaches the platform's APIs as MCP tools. One model picks the tool, a second writes the arguments against the tool's schema, and validation errors go back to the model before anything runs.
- Result: Answers come from live records, and malformed tool calls are caught before they reach the API.

### Browser and computer-use agents

- Problem: Useful work sits in apps that have no API.
- System: A desktop app runs an agent that reads screenshots, clicks and types in a browser, calls MCP tools, runs scheduled tasks and keeps long-term memory with Mem0.
- Result: A working prototype that completes multi-step browser tasks on request.

### Research pipelines with an approval gate

- Problem: A database has to grow every day, but nothing unverified can reach customers.
- System: A crawler reads news feeds every 6 hours, AI gates extract and verify entities, and a reviewer approves each record from a side-by-side diff.
- Result: The database grows without unreviewed records going live.

### Agents that are graded every night

- Problem: Agent quality drifts after launch, and nobody notices until a customer does.
- System: A nightly LLM judge scores the previous day's outputs from 0 to 10 against a fixed list of issues. Reviewer feedback overrides the judge.
- Result: Quality problems show up on a dashboard the operations team checks daily.

### Prompt changes tested on real outcomes

- Problem: Prompts get edited by feel, and regressions ship.
- System: An optimizer built on GEPA replays past conversations. Each candidate prompt is scored 60% on matching the real outcome and 40% on an LLM judge's quality score.
- Result: Prompt changes are compared on the same history before they reach production.

## What is an AI agent, in practice?

An AI agent is a loop. A language model reads the request, picks a tool, reads what the tool returns and decides what to do next, until the task is done or it has to ask. The model is the easy part. The work is in the tools, the permissions around them and the tests that tell you when the loop goes wrong.

Mirai Minds builds agents that live inside existing systems. For a [US logistics software company](/work/logistics-ai-agents) we exposed the client's platform APIs as MCP tools and put a chat agent in front of them. For [Rouh](/work/rouh-computer-use-agent) we built a desktop agent that operates a browser from screenshots, for tasks that have no API at all. For [iKoMatch](/work/ikomatch-ai-fundraising-assistant) the agent is a pipeline that discovers, verifies and matches investors every day.

## How do you keep a person in the loop?

We decide with you which actions an agent may take alone. Reading data usually needs no approval. Writing records, spending money or messaging a customer does. Those steps pause and show a person what will happen, with the evidence the agent used. In the iKoMatch data pipeline, no investor record goes live until a reviewer approves it from a side-by-side diff.

Every tool call is logged with its inputs, outputs and timing. When something goes wrong, you can see which step failed and why, instead of rereading a chat transcript.

## How do you know it works?

Before we build, we collect real examples of the task and agree what a correct result looks like. That becomes the test set. Each release runs against it, an LLM judge scores the results, and a person reviews the cases the judge flags. For iKoMatch, a judge re-scores the previous day's recommendations every night and files concrete fixes. On our [Voice Agents platform](/work/voice-agents-platform), a GEPA-based optimizer tests prompt changes against past calls before they ship.

Judges are models too, and they drift. We check them against human labels and let reviewer feedback override them.

## What does a typical build look like?

Week 1 is access, examples and the test set. Weeks 2 to 6 cover the tools, the agent loop, approval flows and tracing, with a working version running on real data early. Then comes a launch with a person reviewing outputs, and a period of tuning prompts and tools against what production shows. Most builds ship in two to eight weeks.

If the agent needs to talk on the phone or on WhatsApp, it pairs with our [voice AI agents](/services/voice-ai-agents) and [WhatsApp AI agents](/services/whatsapp-ai-agents) work. If it needs to read documents reliably, see [custom LLM and document AI](/services/custom-llm-development).

## Typical stack

- **Agent frameworks:** OpenAI Agents SDK, Model Context Protocol (FastMCP), Custom tool loops
- **Models:** GPT-4.1, Claude, Gemini, Llama 4 via OpenRouter
- **Memory:** Mem0, Redis session summaries
- **Tracing and evals:** LangSmith, LLM judges, GEPA
- **Automation:** Playwright, Scheduled jobs
- **Backend:** FastAPI, PostgreSQL, MongoDB, Redis

## Related work

- [iKoMatch: an AI fundraising assistant on WhatsApp and voice](https://www.miraiminds.co/work/ikomatch-ai-fundraising-assistant.md) — 6,717 investor records; each day's matches graded by an LLM judge
- [Voice Agents: our platform for AI phone agents](https://www.miraiminds.co/work/voice-agents-platform.md) — Live at voice-agents.miraiminds.co
- [Operations chat agent for a US logistics software company](https://www.miraiminds.co/work/logistics-ai-agents.md) — Delivered; every tool call's arguments are checked before it runs
- [Rouh: a computer-use agent for the desktop](https://www.miraiminds.co/work/rouh-computer-use-agent.md) — Working desktop prototype, reviewed by the founder

## Frequently asked questions

### When is an AI agent the wrong choice?

When the steps never change. A fixed script, a form or a scheduled job is cheaper and easier to test. Agents earn their cost when inputs vary, the right tool depends on the request, and a person would otherwise read, decide and click through several systems.

### How do you stop an agent from doing something it shouldn't?

Three layers. The agent only sees the tools we expose, scoped to the account it acts for. Every tool call is checked against the tool's schema before it runs. Actions that write data, spend money or message a customer wait for a person to approve them, and every call is logged.

### What do you need from us to start?

Access to the systems the agent will use (a sandbox is ideal), 30 to 50 real examples of the task with the right outcome, and one person who owns the process and can say what correct looks like. We turn the examples into the first test set.

### How do you know the agent is good enough to launch?

We agree a pass bar before building, on a test set drawn from your real history. Each release runs against it, an LLM judge scores the results and a person checks the cases the judge flags. After launch, a judge can re-score a sample of live runs every night, as it does for iKoMatch.

### Which models do you use?

Whichever fits each step. We have shipped agents on GPT-4.1, Claude, Gemini and Llama 4, sometimes two in one agent: a stronger model to plan and another to fill in tool arguments. The model sits behind an interface, so it can be swapped when prices or quality change.

### Where does our data go?

The agent runs in your cloud or in ours and calls model providers through their business APIs. Credentials stay in your secret store. We log tool calls for debugging and audit, and agree retention with you before launch.

### What drives the cost of an agent project?

The number of systems the agent touches, how many actions need approval flows, how much test data already exists, and model spend at your volume. A read-only assistant over one API is the smallest build; agents that write to several systems take longer.


---

Canonical: https://www.miraiminds.co/services/ai-agents
Last updated: 2026-09-23
Publisher: Mirai Minds LLP, 906 Sarthana Business Hub, Nana Varachha, Surat, Gujarat 395013, India. hello@miraiminds.co
