Skip to content
Mirai Minds

Voice AI agents for inbound and outbound calls

In short

Mirai Minds builds voice AI agents that answer and place phone calls for support, sales and follow-ups. We build on Pipecat and LiveKit with Deepgram speech recognition, including Hindi, Cartesia or ElevenLabs voices, and Plivo or Twilio numbers. A caller can reach a live person mid-call, and every call is recorded, transcribed and turned into structured data.

Problems we take on, and the systems we ship for them

Live handoff to a person

ProblemCallers who need a person get stuck in a loop with a bot.

SystemThe agent raises a handoff that appears on the team's dashboard. The first person to accept joins the same LiveKit room; the claim is atomic, so two people can't take one caller.

ResultIf nobody accepts within 60 seconds, the agent says so and offers a callback instead of leaving the caller waiting.

Outbound campaigns within set limits

ProblemDialing a contact list with a script that ignores how many calls are live, or what time it is.

SystemA queue-driven campaign worker with per-campaign and global caps on live calls, daily dial windows, international number validation and a voicemail status per contact.

ResultCampaigns run unattended inside the limits the operator sets.

Not talking to voicemail

ProblemOutbound agents start their script into ringback, call screening or a voicemail greeting.

SystemThe agent stays silent until a person speaks. A detector recognizes voicemail and carrier messages in English and Spanish and ends those calls with a short goodbye.

ResultThe greeting reaches a person, and voicemail outcomes are recorded per contact.

Structured data after every call

ProblemCall outcomes sit in recordings nobody listens to.

SystemAfter each call, a model reads the recording and fills fields: did a person answer, did they engage, was a callback requested and for when, was a meeting booked, and the exact words of any conflict.

ResultCallbacks, meetings and WhatsApp follow-ups are driven by data instead of memory.

Calls in Hindi

ProblemMany Indian customers would rather speak Hindi than English on the phone.

SystemDeepgram Nova-3 recognition configured for Hindi, Cartesia voices and Indian phone numbers.

ResultAgents that hold Hindi calls for Indian businesses on our own platform.

How the pieces fit together

Reference architecture for Voice AI agentsINPUTCaller on a phonelineSYSTEMPlivo or Twiliointo LiveKit SIPMODELDeepgramspeech-to-text andturn detectionMODELLanguage model withtoolsMODELCartesia orElevenLabs voiceHUMAN REVIEWLive agent onhandoffOUTPUTRecording,transcript,structured fields

How it flows

  1. 01 Caller on a phone line → Plivo or Twilio into LiveKit SIP
  2. 02 Plivo or Twilio into LiveKit SIP → Deepgram speech-to-text and turn detection
  3. 03 Deepgram speech-to-text and turn detection → Language model with tools
  4. 04 Language model with tools → Cartesia or ElevenLabs voice
  5. 05 Language model with tools → Live agent on handoff
  6. 06 Cartesia or ElevenLabs voice → Recording, transcript, structured fields
  7. 07 Live agent on handoff → Recording, transcript, structured fields

Published

How does a call move through the system?

A call arrives on a Plivo or Twilio number and joins a LiveKit room over SIP. A Pipecat pipeline takes it from there: Deepgram turns speech into text, a language model decides what to say and which tools to call, and Cartesia or ElevenLabs speaks the reply while the caller can still interrupt. When the call ends, the recording and transcript are stored, and a model fills structured fields such as whether a callback was requested and when.

How does the agent know the caller has finished?

Turn-taking is where most voice agents feel wrong. Cut in too early and the agent talks over people; wait too long and it feels slow. In the outbound calling platform we use three layers: voice activity detection so short answers like "yes" register, end-of-turn detection from the speech model (Deepgram Flux where it fits), and a watchdog for turns that end with nothing usable. On outbound calls the agent also waits for a real person before its first line, instead of greeting a voicemail box.

Off-the-shelf platforms (Vapi, Retell) or a custom Pipecat/LiveKit build?

Consideration Hosted platform (Vapi, Retell) Custom build (Pipecat + LiveKit)
Time to a first call Fastest: configure in a dashboard Slower: pipeline, hosting and telephony to set up
Cost per minute Provider costs plus a platform fee Provider costs plus your hosting
Turn-taking and interruptions The platform's settings Full control, tuned per use case
Speech and voice providers The platform's supported list Any, including self-hosted models
Where audio and transcripts live The vendor's cloud Your cloud or ours
Who runs it The vendor You, or us on your behalf
Good fit Pilots, standard flows, lower volume High volume, Indian languages, strict data rules, custom handoff

Our own Voice Agents platform supports both, per assistant: a hosted runtime such as Vapi, or our Pipecat pipeline. A pilot on a hosted runtime is often the right first step. We move to a custom build when per-minute cost, language quality or data rules make the case, and we say so when they don't.

What happens after launch?

We watch real calls. Each call has a recording, a transcript and structured fields, and CallAuditAI can score recorded calls against a rubric. The fixes are usually small: a prompt line, a turn-taking threshold, a pronunciation. We ship them like code, tested on past calls first, and a person stays reachable through handoff the whole time.

Where we've built this

What we usually build it with

Chosen per project. We'll tell you when something simpler will do.

Orchestration
PipecatLiveKit rooms, SIP and Egress
Telephony
PlivoTwilioIndian DID numbers
Speech-to-text
Deepgram Nova-3Deepgram FluxSarvam
Text-to-speech
CartesiaElevenLabsRime
Turn-taking
Silero VADSpeech-model end-of-turn detectionSilence watchdog
Language models
GPT-4o-miniGemini
Hosted runtimes
Vapi

How the engagement runs

  1. Week 01

    Discovery call

    Thirty minutes with an engineer. You describe the job; we say whether AI is the right tool and what it would take.

  2. Week 12

    Scope and evals

    We agree what done looks like and build a test set from your real data before writing the system.

  3. Weeks 2–63

    Build

    Working software every week, measured against the test set, with your team trying it early.

  4. Launch4

    With human review

    The system goes live with a person checking the risky steps and a clear way to reach a human.

  5. Ongoing5

    Run and improve

    We watch real traffic, fix what breaks and tune on real cases, or hand over with documentation.

Asked on the first call

How does a voice agent hand off to a human?

On our platform the agent raises a handoff request, which appears at once on the team's dashboard. The first person to accept joins the same live call; the claim is atomic, so two people can't take one caller. If nobody accepts within 60 seconds, the agent tells the caller no one is free and offers a callback.

How quickly does the agent reply after the caller stops talking?

We measure it on every turn: the gap between the end of the caller's speech and the agent's first audio. The levers are how the end of a turn is detected, which model answers, how fast the voice starts streaming and how far the servers sit from the phone network. We tune those on your real calls rather than quote a lab figure.

Can the agent speak Hindi?

Yes. We run Hindi calls on Deepgram Nova-3 recognition with Cartesia voices. For other languages we first check speech-recognition and voice support, then test on recordings from your own callers before anyone relies on it.

Do we need Vapi or Retell, or a custom build?

Either can be right. A hosted platform is the fastest way to a pilot. A custom Pipecat and LiveKit build costs more up front but gives control over turn-taking, providers, data location and per-minute cost at volume. We have shipped both, and a pilot on one can move to the other.

What do you need from us?

Your call scripts or a few recorded calls, access to the systems the agent should read or update (CRM, calendar, store), phone numbers or permission to provision them, and a person or team to take handoffs. For outbound work, the contact lists and the hours you are allowed to call.

Where are recordings and transcripts stored?

Where you decide. On our platform they are scoped to your workspace. For one client we wrote each call's recording straight to their own Azure Blob storage using workload identity, so the app held no storage keys. Retention is agreed per deployment.

What drives the cost of running a voice agent?

Each minute pays for telephony, speech-to-text, the language model and text-to-speech, plus hosting; hosted platforms add their own fee. Call volume, call length and model choice move the total most. We model it on your expected minutes before you commit.

Other services

Have a system in mind? Let's scope it.

A 30-minute call with an engineer who has shipped this before. You leave with a plan, a rough timeline and what it would take — whether or not we build it.