Case studyCustom LLM and document AI
Explanation and pronunciation tools for a Gurbani learning app
By Sagar DavaraPublished
DeliveredIn short
Mirai Minds built the AI back end for a Gurbani learning app. Given a line of Gurbani, the API transliterates it into Gurmukhi, applies pronunciation rules, writes a plain explanation in Hindi or English with the meaning of key words, and streams that explanation as speech from a self-hosted text-to-speech model.
The story
The problem
Learners reading Gurbani often need three things for each line: how to say it, what the words mean, and what the line says as a whole, in a language they are comfortable with. The app team wanted all three on demand, as text and as audio, from one call.
What we built
A small FastAPI service with token authentication and one pipeline per request:
- Transliteration of the input line into Gurmukhi, for example from Devanagari.
- Pronunciation rules, a grapheme-to-phoneme step applied to the Gurmukhi text.
- An explanation in Hindi or English: the meaning of key words in a fixed pattern, then the line as a whole, in a friendly, simple register.
- Speech: the explanation streamed as MP3 audio from a self-hosted Kokoro model, with a blended Hindi voice or an English voice.
The app can pass an optional instruction with each request, so editors can adjust tone for a particular screen without changing the service.
How it works in production
The app calls the API with a line and a language. Each model step runs at temperature 0.1 with a capped output length, so wording stays consistent from one request to the next. Audio streams while it is being generated, so playback can start before the whole explanation is synthesized. The service ships as a Docker image, and requests without the right token are refused. The general approach, chaining model steps with checks between them, is described under custom LLM and document AI.
What we'd change
- A reviewed test set. Explanations of scripture should be checked by people who know it well. A set of reviewed lines would make every prompt change measurable instead of a matter of taste.
- Cache audio per line and language. Popular lines are requested again and again; caching would cut both latency and model calls.
- Fewer model calls. For lines that already arrive in Gurmukhi, transliteration can be skipped, and some pronunciation rules could run as plain code rather than as a model step.
Stack
What it runs on
- API
- FastAPIToken authenticationStreaming responses
- Models
- OpenAI chat modelsKokoro text-to-speech
- Packaging
- Dockeruv