Skip to content
Mirai Minds

Case studyVoice AI agents

Outbound AI calling: a Pipecat voice bot and campaign dialer

By Sneh MehtaPublished

In production

In short

Mirai Minds engineers worked inside an AI outbound-calling startup's platform team to build the parts that make calls work: a Pipecat voice bot over LiveKit SIP, turn detection that doesn't talk over people or into voicemail, a campaign dialer that caps live calls, and one recording per call. The platform runs outbound campaigns in production.

The problem

The startup's platform lets an operator define an agent (a language model, speech-to-text, text-to-speech and functions), load a contact list into a campaign and have the system call every contact. The first version ran its voice logic in an older Node service. The client needed calls that sounded natural, a dialer that respected concurrency and calling-hour limits, and recordings kept in their own cloud.

What we built

Our engineers worked inside the client's monorepo and deployment pipeline, alongside their team.

  • A voice bot on Pipecat. In July 2025 we replaced the older voice service with a Python Pipecat workflow that joins each call through LiveKit SIP. Later we added Deepgram Flux end-of-turn detection, outbound calling, handoff handling and local audio recording.
  • Turn-taking in three layers. Silero voice activity detection so short answers like "yes" register, end-of-turn detection from Deepgram Nova or Flux, and a watchdog for turns that end with nothing usable.
  • Human presence detection on outbound calls, in English and Spanish, so the greeting goes to a person and voicemail gets a polite goodbye.
  • A campaign dialer. Campaign and call-status APIs, bulk contact import, a headless worker on BullMQ, per-campaign and global live-call caps, daily dial windows, international number validation and a voicemail status for each contact.
  • Recording and operations. Per-call recording through LiveKit Egress to Azure Blob Storage with workload identity, a recordings browser and a pod-logs viewer in the admin dashboard, and Kubernetes deployments for the API and the worker.

How it works in production

Creating a campaign, or any call ending, puts a "fill" job on the queue. The worker reserves a slot against both caps, dials through LiveKit SIP and hands the audio to the Pipecat bot, which runs speech-to-text, the language model and text-to-speech and posts results back to the API. The API and the worker ship in one image but run as separate deployments, and only the API migrates the database, so replicas never race on schema changes. A repeatable backstop job reaps stuck calls, reconciles counters and marks finished campaigns complete.

For the patterns behind this, see voice AI agents; the same ideas run on our own Voice Agents platform.

What we'd change

  • Build human presence detection before the first outbound campaign. The bot originally greeted as soon as the line went active, which often meant talking into ringback or a voicemail greeting.
  • Make turn-taking settings per agent from the start. Thresholds that suit a quick yes-or-no call are wrong for a long, detailed conversation, and environment-wide settings turn every adjustment into a compromise.

What it runs on

Voice
PipecatLiveKit SIP and EgressDeepgram Nova and FluxElevenLabsOpenAISilero VAD
Backend
NestJS on FastifyMongoDBBullMQRedis
Admin
ReactViteMUI
Infrastructure
Kubernetes on AzureAzure Blob StorageAzure DevOps pipelines

Asked on the first call

How does the dialer stop too many calls running at once?

A call lasts minutes but a queue job finishes in milliseconds, so the queue's own concurrency can't limit live calls. We use two atomic counters in Redis, one per campaign and one global, reserved in a single Lua step before each dial. When a call ends its slot is released and the next dial is queued at once; a backstop job reconciles the counters and reaps stuck calls.

How does the bot avoid greeting a voicemail box?

On outbound calls, a line going 'active' doesn't mean a person is there. The bot stays silent through ringback, call screening and silence, starts as soon as someone says hello, and ends voicemail calls with a short goodbye. The detector handles English and Spanish.

How are recordings handled?

LiveKit Egress writes one audio-only file per call. In production those files go to the client's Azure Blob storage using workload identity, so the app holds no storage keys. Operators can search, play and delete recordings from the admin dashboard.

Have a system in mind? Let's scope it.

A 30-minute call with an engineer who has shipped this before. You leave with a plan, a rough timeline and what it would take — whether or not we build it.