Case studyVoice AI agents
Outbound AI calling: a Pipecat voice bot and campaign dialer
By Sneh MehtaPublished
In productionIn short
Mirai Minds engineers worked inside an AI outbound-calling startup's platform team to build the parts that make calls work: a Pipecat voice bot over LiveKit SIP, turn detection that doesn't talk over people or into voicemail, a campaign dialer that caps live calls, and one recording per call. The platform runs outbound campaigns in production.
The story
The problem
The startup's platform lets an operator define an agent (a language model, speech-to-text, text-to-speech and functions), load a contact list into a campaign and have the system call every contact. The first version ran its voice logic in an older Node service. The client needed calls that sounded natural, a dialer that respected concurrency and calling-hour limits, and recordings kept in their own cloud.
What we built
Our engineers worked inside the client's monorepo and deployment pipeline, alongside their team.
- A voice bot on Pipecat. In July 2025 we replaced the older voice service with a Python Pipecat workflow that joins each call through LiveKit SIP. Later we added Deepgram Flux end-of-turn detection, outbound calling, handoff handling and local audio recording.
- Turn-taking in three layers. Silero voice activity detection so short answers like "yes" register, end-of-turn detection from Deepgram Nova or Flux, and a watchdog for turns that end with nothing usable.
- Human presence detection on outbound calls, in English and Spanish, so the greeting goes to a person and voicemail gets a polite goodbye.
- A campaign dialer. Campaign and call-status APIs, bulk contact import, a headless worker on BullMQ, per-campaign and global live-call caps, daily dial windows, international number validation and a voicemail status for each contact.
- Recording and operations. Per-call recording through LiveKit Egress to Azure Blob Storage with workload identity, a recordings browser and a pod-logs viewer in the admin dashboard, and Kubernetes deployments for the API and the worker.
How it works in production
Creating a campaign, or any call ending, puts a "fill" job on the queue. The worker reserves a slot against both caps, dials through LiveKit SIP and hands the audio to the Pipecat bot, which runs speech-to-text, the language model and text-to-speech and posts results back to the API. The API and the worker ship in one image but run as separate deployments, and only the API migrates the database, so replicas never race on schema changes. A repeatable backstop job reaps stuck calls, reconciles counters and marks finished campaigns complete.
For the patterns behind this, see voice AI agents; the same ideas run on our own Voice Agents platform.
What we'd change
- Build human presence detection before the first outbound campaign. The bot originally greeted as soon as the line went active, which often meant talking into ringback or a voicemail greeting.
- Make turn-taking settings per agent from the start. Thresholds that suit a quick yes-or-no call are wrong for a long, detailed conversation, and environment-wide settings turn every adjustment into a compromise.
Stack
What it runs on
- Voice
- PipecatLiveKit SIP and EgressDeepgram Nova and FluxElevenLabsOpenAISilero VAD
- Backend
- NestJS on FastifyMongoDBBullMQRedis
- Admin
- ReactViteMUI
- Infrastructure
- Kubernetes on AzureAzure Blob StorageAzure DevOps pipelines