Mirai Minds builds voice AI agents that answer and place phone calls for support, sales and follow-ups. We build on Pipecat and LiveKit with Deepgram speech recognition, including Hindi, Cartesia or ElevenLabs voices, and Plivo or Twilio numbers. A caller can reach a live person mid-call, and every call is recorded, transcribed and turned into structured data.
Problems we take on, and the systems we ship for them
Live handoff to a person
ProblemCallers who need a person get stuck in a loop with a bot.
SystemThe agent raises a handoff that appears on the team's dashboard. The first person to accept joins the same LiveKit room; the claim is atomic, so two people can't take one caller.
ResultIf nobody accepts within 60 seconds, the agent says so and offers a callback instead of leaving the caller waiting.
Outbound campaigns within set limits
ProblemDialing a contact list with a script that ignores how many calls are live, or what time it is.
SystemA queue-driven campaign worker with per-campaign and global caps on live calls, daily dial windows, international number validation and a voicemail status per contact.
ResultCampaigns run unattended inside the limits the operator sets.
Not talking to voicemail
ProblemOutbound agents start their script into ringback, call screening or a voicemail greeting.
SystemThe agent stays silent until a person speaks. A detector recognizes voicemail and carrier messages in English and Spanish and ends those calls with a short goodbye.
ResultThe greeting reaches a person, and voicemail outcomes are recorded per contact.
Structured data after every call
ProblemCall outcomes sit in recordings nobody listens to.
SystemAfter each call, a model reads the recording and fills fields: did a person answer, did they engage, was a callback requested and for when, was a meeting booked, and the exact words of any conflict.
ResultCallbacks, meetings and WhatsApp follow-ups are driven by data instead of memory.
Calls in Hindi
ProblemMany Indian customers would rather speak Hindi than English on the phone.
SystemDeepgram Nova-3 recognition configured for Hindi, Cartesia voices and Indian phone numbers.
ResultAgents that hold Hindi calls for Indian businesses on our own platform.
Reference architecture
How the pieces fit together
How it flows
01 Caller on a phone line → Plivo or Twilio into LiveKit SIP
02 Plivo or Twilio into LiveKit SIP → Deepgram speech-to-text and turn detection
03 Deepgram speech-to-text and turn detection → Language model with tools
04 Language model with tools → Cartesia or ElevenLabs voice
05 Language model with tools → Live agent on handoff
06 Cartesia or ElevenLabs voice → Recording, transcript, structured fields
07 Live agent on handoff → Recording, transcript, structured fields
The long version
Published
How does a call move through the system?
A call arrives on a Plivo or Twilio number and joins a LiveKit room over SIP. A Pipecat pipeline takes it from there: Deepgram turns speech into text, a language model decides what to say and which tools to call, and Cartesia or ElevenLabs speaks the reply while the caller can still interrupt. When the call ends, the recording and transcript are stored, and a model fills structured fields such as whether a callback was requested and when.
How does the agent know the caller has finished?
Turn-taking is where most voice agents feel wrong. Cut in too early and the agent talks over people; wait too long and it feels slow. In the outbound calling platform we use three layers: voice activity detection so short answers like "yes" register, end-of-turn detection from the speech model (Deepgram Flux where it fits), and a watchdog for turns that end with nothing usable. On outbound calls the agent also waits for a real person before its first line, instead of greeting a voicemail box.
Off-the-shelf platforms (Vapi, Retell) or a custom Pipecat/LiveKit build?
Consideration
Hosted platform (Vapi, Retell)
Custom build (Pipecat + LiveKit)
Time to a first call
Fastest: configure in a dashboard
Slower: pipeline, hosting and telephony to set up
Cost per minute
Provider costs plus a platform fee
Provider costs plus your hosting
Turn-taking and interruptions
The platform's settings
Full control, tuned per use case
Speech and voice providers
The platform's supported list
Any, including self-hosted models
Where audio and transcripts live
The vendor's cloud
Your cloud or ours
Who runs it
The vendor
You, or us on your behalf
Good fit
Pilots, standard flows, lower volume
High volume, Indian languages, strict data rules, custom handoff
Our own Voice Agents platform supports both, per assistant: a hosted runtime such as Vapi, or our Pipecat pipeline. A pilot on a hosted runtime is often the right first step. We move to a custom build when per-minute cost, language quality or data rules make the case, and we say so when they don't.
What happens after launch?
We watch real calls. Each call has a recording, a transcript and structured fields, and CallAuditAI can score recorded calls against a rubric. The fixes are usually small: a prompt line, a turn-taking threshold, a pronunciation. We ship them like code, tested on past calls first, and a person stays reachable through handoff the whole time.
Thirty minutes with an engineer. You describe the job; we say whether AI is the right tool and what it would take.
Week 12
Scope and evals
We agree what done looks like and build a test set from your real data before writing the system.
Weeks 2–63
Build
Working software every week, measured against the test set, with your team trying it early.
Launch4
With human review
The system goes live with a person checking the risky steps and a clear way to reach a human.
Ongoing5
Run and improve
We watch real traffic, fix what breaks and tune on real cases, or hand over with documentation.
FAQ
Asked on the first call
How does a voice agent hand off to a human?
On our platform the agent raises a handoff request, which appears at once on the team's dashboard. The first person to accept joins the same live call; the claim is atomic, so two people can't take one caller. If nobody accepts within 60 seconds, the agent tells the caller no one is free and offers a callback.
How quickly does the agent reply after the caller stops talking?
We measure it on every turn: the gap between the end of the caller's speech and the agent's first audio. The levers are how the end of a turn is detected, which model answers, how fast the voice starts streaming and how far the servers sit from the phone network. We tune those on your real calls rather than quote a lab figure.
Can the agent speak Hindi?
Yes. We run Hindi calls on Deepgram Nova-3 recognition with Cartesia voices. For other languages we first check speech-recognition and voice support, then test on recordings from your own callers before anyone relies on it.
Do we need Vapi or Retell, or a custom build?
Either can be right. A hosted platform is the fastest way to a pilot. A custom Pipecat and LiveKit build costs more up front but gives control over turn-taking, providers, data location and per-minute cost at volume. We have shipped both, and a pilot on one can move to the other.
What do you need from us?
Your call scripts or a few recorded calls, access to the systems the agent should read or update (CRM, calendar, store), phone numbers or permission to provision them, and a person or team to take handoffs. For outbound work, the contact lists and the hours you are allowed to call.
Where are recordings and transcripts stored?
Where you decide. On our platform they are scoped to your workspace. For one client we wrote each call's recording straight to their own Azure Blob storage using workload identity, so the app held no storage keys. Retention is agreed per deployment.
What drives the cost of running a voice agent?
Each minute pays for telephony, speech-to-text, the language model and text-to-speech, plus hosting; hosted platforms add their own fee. Call volume, call length and model choice move the total most. We model it on your expected minutes before you commit.
A 30-minute call with an engineer who has shipped this before. You leave with a plan, a rough timeline and what it would take — whether or not we build it.