Skip to content
Mirai Minds

Case studyCustom LLM and document AI

Explanation and pronunciation tools for a Gurbani learning app

By Sagar DavaraPublished

Delivered

In short

Mirai Minds built the AI back end for a Gurbani learning app. Given a line of Gurbani, the API transliterates it into Gurmukhi, applies pronunciation rules, writes a plain explanation in Hindi or English with the meaning of key words, and streams that explanation as speech from a self-hosted text-to-speech model.

The problem

Learners reading Gurbani often need three things for each line: how to say it, what the words mean, and what the line says as a whole, in a language they are comfortable with. The app team wanted all three on demand, as text and as audio, from one call.

What we built

A small FastAPI service with token authentication and one pipeline per request:

  1. Transliteration of the input line into Gurmukhi, for example from Devanagari.
  2. Pronunciation rules, a grapheme-to-phoneme step applied to the Gurmukhi text.
  3. An explanation in Hindi or English: the meaning of key words in a fixed pattern, then the line as a whole, in a friendly, simple register.
  4. Speech: the explanation streamed as MP3 audio from a self-hosted Kokoro model, with a blended Hindi voice or an English voice.

The app can pass an optional instruction with each request, so editors can adjust tone for a particular screen without changing the service.

How it works in production

The app calls the API with a line and a language. Each model step runs at temperature 0.1 with a capped output length, so wording stays consistent from one request to the next. Audio streams while it is being generated, so playback can start before the whole explanation is synthesized. The service ships as a Docker image, and requests without the right token are refused. The general approach, chaining model steps with checks between them, is described under custom LLM and document AI.

What we'd change

  • A reviewed test set. Explanations of scripture should be checked by people who know it well. A set of reviewed lines would make every prompt change measurable instead of a matter of taste.
  • Cache audio per line and language. Popular lines are requested again and again; caching would cut both latency and model calls.
  • Fewer model calls. For lines that already arrive in Gurmukhi, transliteration can be skipped, and some pronunciation rules could run as plain code rather than as a model step.

What it runs on

API
FastAPIToken authenticationStreaming responses
Models
OpenAI chat modelsKokoro text-to-speech
Packaging
Dockeruv

Asked on the first call

Why split the work into several model calls?

Each step has a different job and a different way to fail. Transliteration and pronunciation run first at low temperature, so the text the explanation is based on is stable. The explanation step then works from that clean text instead of fixing the script and explaining it at once.

Why a self-hosted text-to-speech model?

The app streams an explanation as audio for every line a learner opens. A self-hosted Kokoro model behind an OpenAI-compatible endpoint keeps the voices consistent, a blended Hindi voice and an English voice, and keeps them under the app team's control.

How are explanations checked?

The prompts fix the format, so every explained word follows the same pattern, and the app team reviews the output. Explanations of scripture need people who know it; a reviewed test set of lines is the first thing we'd add in a next phase.

Have a system in mind? Let's scope it.

A 30-minute call with an engineer who has shipped this before. You leave with a plan, a rough timeline and what it would take — whether or not we build it.