Case studyAI agents and automation
Operations chat agent for a US logistics software company
By Sagar DavaraPublished
DeliveredIn short
Mirai Minds built an operations chat agent for a US logistics software company. Users ask about loads and customers in plain language, and the agent answers from live records by calling the client's APIs through MCP. A second model writes and checks every tool call's arguments, and each conversation keeps per-user memory and the user's time zone.
The story
The problem
The client's users manage freight operations in the client's software. Simple questions, such as which loads a customer has this week, meant several screens and filters. The client wanted a chat assistant that answers from live data without giving a model free rein over their APIs.
What we built
- An MCP tool layer. The client's APIs are exposed as MCP servers registered in a database. The agent loads the active ones at start-up and caches tool descriptions in Redis.
- A two-stage tool loop. GPT-4.1 reasons privately, chooses one tool per step and decides between single-record and list tools. A second model, Llama 4 Maverick through OpenRouter, fills the arguments against the tool's JSON schema. Validation errors go back for up to three retries before the agent gives up cleanly.
- Memory. Conversations are stored in MongoDB. A rolling session summary and per-user memory live in Redis and go into the prompt with the user's local time, so "this week" means their week.
- Visible progress. Short progress notes appear in the interface while the agent works, followed by a final answer in plain language.
How it works in production
Each message becomes a task. The loop runs until the model produces a final answer or hits its retry limit, and tool schemas are fetched once and cached rather than on every call. The service sits behind JWT authentication and reports to New Relic. We built it between May and July 2025, with engineers from our team working through prompt, memory and tool-validation changes in short cycles with the client. The general pattern is described under AI agents and automation.
What we'd change
- Use native structured tool calls. We used an XML tool protocol, which is why argument validation needed a second model. JSON-schema tool calling in current models would remove most of that stage and one source of latency.
- Start with an evaluation set of real operator questions. We tuned the prompt through many rounds of manual testing. A fixed set of questions with expected answers, re-run on every change, would have made each round measurable instead of anecdotal.
Stack
What it runs on
- Agent
- Custom tool loopModel Context Protocol (FastMCP client)OpenRouter
- Models
- GPT-4.1Llama 4 Maverick
- Backend
- FastAPIMongoDBRedisJWT authentication
- Monitoring
- New Relic