Skip to content
Mirai Minds

In-houseWhatsApp AI agents

WhatsApp Agents: our service for AI WhatsApp conversations

By Sagar DavaraPublished

In production

In short

WhatsApp Agents is Mirai Minds' own service for running AI conversations on WhatsApp. It receives messages from Meta's webhook, waits for the customer to finish typing, builds context from a rolling summary, recent messages and Mem0 memories, and answers with Gemini. One deployment serves several business numbers, and it runs iKoMatch's founder conversations.

The problem

Our WhatsApp builds kept needing the same core: receive Meta's webhook, route the message to the right business account, remember the customer across weeks, answer from documents, and follow WhatsApp's rules on templates and the 24-hour window. Rebuilding that per client wasted time, so we built it once as a service.

What we built

  • Message handling. A FastAPI webhook receives messages, looks up the account by phone number ID, saves the message and starts a 3-second debounce, so a burst of short messages gets one reply.
  • Context building. Every 20 messages are folded into a rolling summary of at most 150 words. Each reply sees that summary, the last 20 unsummarized messages and the Mem0 memories relevant to the current message, which keeps prompts small in long conversations.
  • Gemini with tools. Replies come from Gemini, with tools and document retrieval through a retrieval service shared with our Voice Agents platform. Account owners can set their own Gemini key and model.
  • Media and documents. Images and files go to S3 and are described for the agent; decks are converted before reading.
  • Outbound messaging. Templates, broadcasts, follow-ups and conversation stages, so campaigns stay inside WhatsApp's rules.

How it works in production

The service runs iKoMatch's founder conversations: signup, deck intake and the daily investor slates of a 45-day campaign. For every known number, iKoMatch's backend returns the founder's status with a matching prompt, which replaces the account's default prompt. Replies are split into short bubbles with natural pauses, so they read like texting. The patterns are described under WhatsApp AI agents.

What we'd change

  • Put human takeover in the core service. Today the team inbox for staff replies lives in the client's dashboard, as it does for iKoMatch. Building it into the service would give every deployment the same takeover flow.
  • Grade conversations every night. iKoMatch's recommendations get a nightly LLM judge; the conversations themselves deserve one too, scoring a sample of each day's chats against a fixed list of issues.

What it runs on

AI
Gemini (google-genai)Mem0
Backend
FastAPIPostgreSQLSQLAlchemy (async)Alembic
Messaging
WhatsApp Business Platform (Cloud API)
Storage and documents
Amazon S3Shared retrieval serviceunoserver

Asked on the first call

Why wait 3 seconds before replying?

People send several short messages in a row. Waiting 3 seconds after the last one lets the agent answer everything at once instead of replying to each fragment.

How is one deployment shared across numbers?

Each WhatsApp Business number is an account with its own credentials and prompt. Every incoming webhook carries the phone number ID, which selects the account. For iKoMatch, a role-specific prompt from their backend replaces the default one for founders who are already registered.

What happens to images and documents customers send?

Media is stored in S3 and described by Gemini, so the agent can use it in the conversation. Pitch decks and other documents are converted to a readable format first.

Have a system in mind? Let's scope it.

A 30-minute call with an engineer who has shipped this before. You leave with a plan, a rough timeline and what it would take — whether or not we build it.