Skip to content
Mirai Minds

In-houseCustom LLM and document AI

CallAuditAI: rubric scoring for counseling and sales calls

By Sneh MehtaPublished

In production

In short

CallAuditAI is Mirai Minds' service for auditing recorded sales and counseling calls. Gemini listens to each recording and scores it against a weighted rubric covering opening and rapport, needs discovery, value and objection handling, and closing, then returns an Excel scorecard with zero scores flagged in red. An admissions counseling company uses it to review calls.

How it works

System diagram: CallAuditAI: rubric scoring for counseling and sales callsINPUTRecorded callMODELGemini scores eachcriterionSYSTEMWeighted totals:30/30/20/20OUTPUTExcel scorecardwith red flagsHUMAN REVIEWManager reviewsflagged calls

How it flows

  1. 01 Recorded call → Gemini scores each criterion
  2. 02 Gemini scores each criterion → Weighted totals: 30/30/20/20
  3. 03 Weighted totals: 30/30/20/20 → Excel scorecard with red flags
  4. 04 Excel scorecard with red flags → Manager reviews flagged calls

The problem

Sales and counseling teams record their calls, but managers can listen to only a few. The admissions counseling company wanted every recorded call scored the same way, against a fixed rubric, so coaching could start from the calls that most needed it.

What we built

A small service with one job: take a recording and the counselor's name, and return a scorecard.

  • Scoring. Gemini listens to the audio and scores each criterion of the rubric. The rubric has four categories: opening and rapport building (30%), needs discovery and information gathering (30%), value proposition and objection handling (20%), and closing (20%). Each criterion carries its own weight inside its category.
  • The scorecard. An Excel sheet with section headers, the weight of every criterion, category totals and an overall score. Any criterion scored zero is highlighted as a red flag. The same result is available as JSON for teams that want to load it elsewhere.
  • Transcripts and costs on request. A diarized transcript, labeling the company representative and the customer, comes from a separate, lighter Gemini call. Model cost per call is returned only when the request carries a secret key, so cost details stay out of everyday responses.
  • A log and an upload page. Each processed call and its duration are logged, and a simple web page lets staff upload recordings without touching the API.

How it works in production

The client's team uploads recordings, one per request, and downloads the scorecards. Scoring and transcription are separate calls, so the transcript can be skipped when only the score is needed. The service was deployed for the client in January 2026; in September 2026 the product gained JSON output, transcripts and cost reporting. For the broader pattern, see custom LLM and document AI; for the calls themselves, voice AI agents.

What we'd change

  • Calibrate against human scores. Before trusting totals, we'd have managers score a sample of calls blind and compare with the model criterion by criterion. Subjective criteria, such as warmth and tone, need that check most.
  • Batch uploads and trends. One file per request keeps the service simple. Teams with higher call volumes want batch uploads and per-counselor trends over time.

What it runs on

AI
Google Gemini for scoringA lighter Gemini model for transcripts
API
FastAPIopenpyxlSQLite call log
Interface
Upload page in HTML, CSS and JavaScript

Asked on the first call

Does CallAuditAI need a transcript first?

No. Gemini scores the audio file directly. A transcript with the company representative and the customer as separate speakers is optional and comes from a separate call to a lighter model.

Can the rubric be changed?

Yes. The rubric is four weighted categories, each with its own weighted criteria, such as active listening or warmth and tone. We set the categories, criteria and weights with each client; the current deployment uses 30/30/20/20 across opening, discovery, value and closing.

How do managers use the scores?

They start with the red flags: any criterion scored zero is highlighted in the sheet. Category totals and an overall score make calls comparable, and the recording is always there to check the model's judgment.

Have a system in mind? Let's scope it.

A 30-minute call with an engineer who has shipped this before. You leave with a plan, a rough timeline and what it would take — whether or not we build it.