Skip to content
Mirai Minds

Case studyAI agents and automation

Operations chat agent for a US logistics software company

By Sagar DavaraPublished

Delivered

In short

Mirai Minds built an operations chat agent for a US logistics software company. Users ask about loads and customers in plain language, and the agent answers from live records by calling the client's APIs through MCP. A second model writes and checks every tool call's arguments, and each conversation keeps per-user memory and the user's time zone.

The problem

The client's users manage freight operations in the client's software. Simple questions, such as which loads a customer has this week, meant several screens and filters. The client wanted a chat assistant that answers from live data without giving a model free rein over their APIs.

What we built

  • An MCP tool layer. The client's APIs are exposed as MCP servers registered in a database. The agent loads the active ones at start-up and caches tool descriptions in Redis.
  • A two-stage tool loop. GPT-4.1 reasons privately, chooses one tool per step and decides between single-record and list tools. A second model, Llama 4 Maverick through OpenRouter, fills the arguments against the tool's JSON schema. Validation errors go back for up to three retries before the agent gives up cleanly.
  • Memory. Conversations are stored in MongoDB. A rolling session summary and per-user memory live in Redis and go into the prompt with the user's local time, so "this week" means their week.
  • Visible progress. Short progress notes appear in the interface while the agent works, followed by a final answer in plain language.

How it works in production

Each message becomes a task. The loop runs until the model produces a final answer or hits its retry limit, and tool schemas are fetched once and cached rather than on every call. The service sits behind JWT authentication and reports to New Relic. We built it between May and July 2025, with engineers from our team working through prompt, memory and tool-validation changes in short cycles with the client. The general pattern is described under AI agents and automation.

What we'd change

  • Use native structured tool calls. We used an XML tool protocol, which is why argument validation needed a second model. JSON-schema tool calling in current models would remove most of that stage and one source of latency.
  • Start with an evaluation set of real operator questions. We tuned the prompt through many rounds of manual testing. A fixed set of questions with expected answers, re-run on every change, would have made each round measurable instead of anecdotal.

What it runs on

Agent
Custom tool loopModel Context Protocol (FastMCP client)OpenRouter
Models
GPT-4.1Llama 4 Maverick
Backend
FastAPIMongoDBRedisJWT authentication
Monitoring
New Relic

Asked on the first call

Why use two models for one tool call?

They do different jobs. GPT-4.1 reads the conversation and picks the tool. A second model then writes the arguments against that tool's input schema, and if validation fails the error goes back to it, up to three times. Splitting the steps meant arguments were checked before anything reached the client's API.

Can users see what the agent is doing?

Yes. Alongside its private reasoning, the agent writes a short progress note for the user, such as 'I have the customer list, now checking loads for each customer', which the interface shows while it works. Internal IDs and tool names never appear in the final answer.

What happens when the agent can't complete a request?

It asks a clarifying question when a request is vague, and after repeated failed tool calls it stops and says so rather than guessing. It never calls a tool on a disconnected server; it tells the user that server is unavailable.

Have a system in mind? Let's scope it.

A 30-minute call with an engineer who has shipped this before. You leave with a plan, a rough timeline and what it would take — whether or not we build it.