Skip to content
Mirai Minds

Custom LLM development: RAG, document AI and evals

In short

Mirai Minds builds custom LLM systems for work that depends on a company's own documents and data: answers grounded in retrieval, document AI that turns files into structured records, and fine-tuned models when the numbers justify them. Every system ships with an evaluation set and a human review step, such as engineers checking each extracted row before it reaches SAP.

Problems we take on, and the systems we ship for them

Engineering drawings to SAP-ready data

ProblemEngineers hand-keyed hundreds of line items per drawing into SAP material-master fields.

SystemGemini finds the bill-of-materials table on PDF, scanned or CAD drawings and extracts every row. A reviewer checks each row beside the drawing, can re-run extraction with a plain-English hint on a stronger model, and exports a 16-column SAP upload sheet.

ResultNothing reaches SAP without an engineer's sign-off, and illegal SAP value combinations can't be entered.

Call recordings scored against a rubric

ProblemManagers can't listen to every recorded sales or counseling call.

SystemGemini listens to each recording and scores it against a weighted rubric for opening, discovery, value and closing, producing an Excel scorecard that flags zero scores.

ResultEvery uploaded call gets a scorecard, and managers start with the red flags.

Answers grounded in a catalog

ProblemA general chatbot invents products and prices.

SystemAn agent calls a retrieval tool that ranks catalog entries with hybrid semantic and keyword search, and shows product cards only from real results.

ResultProduct details come from catalog records, not from the model's memory.

Language tools for a specialist text

ProblemLearners need the pronunciation and meaning of each line of scripture, in their own language.

SystemA pipeline transliterates the line into Gurmukhi, applies pronunciation rules, writes a plain explanation in Hindi or English and streams it as speech from a self-hosted text-to-speech model.

ResultThe app team calls one API and gets a spoken explanation back.

How the pieces fit together

Reference architecture for Custom LLM and document AIINPUTYour documents anddataSYSTEMRetrieval ordocument parsingMODELLLM: API orfine-tunedSYSTEMSchema andbusiness-rulechecksHUMAN REVIEWReviewer correctsand signs offOUTPUTStructured outputto your system

How it flows

  1. 01 Your documents and data → Retrieval or document parsing
  2. 02 Retrieval or document parsing → LLM: API or fine-tuned
  3. 03 LLM: API or fine-tuned → Schema and business-rule checks
  4. 04 Schema and business-rule checks → Reviewer corrects and signs off
  5. 05 Reviewer corrects and signs off → Structured output to your system

Published

What counts as a custom LLM system?

Anything where a general chatbot isn't enough because the answer depends on your documents, your data or your rules. In practice that means three kinds of build. Retrieval systems answer from your material and cite it. Document AI turns PDFs, scans, drawings or recordings into structured records your systems can load. Specialist pipelines chain several model steps with checks between them. Fine-tuning is a tool inside those builds, not the starting point.

Our largest example is the engineering-drawings-to-SAP system we built for an Indian public-sector refinery. CallAuditAI applies the same idea to call recordings, and the catalog assistant to a product catalog.

How do you choose a model?

We start from the task and a test set, not from a model. The first run uses a strong hosted model to learn what good output looks like on your documents. Then we try cheaper and faster options against the same test set and keep the cheapest one that clears the bar.

Often the answer is two models. In the refinery system a fast Gemini model does the first pass on every drawing, and the premium model runs only when a reviewer asks for a re-evaluation, with a plain-English hint such as "the quantity column is on the right". Cost is tracked per document, so the trade-off stays visible.

Fine-tuning comes in when the test set says a smaller model can't reach the bar with instructions and retrieval alone, when volume makes per-call cost dominant, or when data must stay on your servers. Licenses matter too: we check that open-weight models allow commercial use before we benchmark them.

What drives the cost?

Two things: building the system and running it. Building cost depends on how messy the inputs are, how many fields and rules the output needs, how much review tooling people need, and which systems it must load into. Running cost is mostly model usage: pages, images or audio minutes in, structured text out, times the number of retries and re-runs, plus hosting if you run open-weight models on your own GPUs. We measure both on your samples before launch rather than guess.

How do people stay in the loop?

By making review fast instead of optional. The refinery reviewer sees the drawing beside the extracted table, edits only the fields that are wrong, and can't enter an illegal SAP combination. CallAuditAI flags zero scores in red so managers start there. When the reviewers' corrections show a field is reliably right, review can shrink for that field, and the numbers, not our confidence, make that call.

Where we've built this

What we usually build it with

Chosen per project. We'll tell you when something simpler will do.

Models
Gemini (fast and premium tiers)OpenAI GPT-4o-mini and GPT-4.1Llama 4
Retrieval
Hybrid semantic and keyword searchPineconeLangChain
Documents
PDF, TIFF, JPEG, DWG and DXF intakeAmazon S3 staging for large filesServer-Sent Events progress
Evaluation
Test sets from real documentsLLM judgesOffline replay of past outputs
Speech
Kokoro text-to-speech (self-hosted)
Application
FastAPIMySQLPostgreSQLVue 3

How the engagement runs

  1. Week 01

    Discovery call

    Thirty minutes with an engineer. You describe the job; we say whether AI is the right tool and what it would take.

  2. Week 12

    Scope and evals

    We agree what done looks like and build a test set from your real data before writing the system.

  3. Weeks 2–63

    Build

    Working software every week, measured against the test set, with your team trying it early.

  4. Launch4

    With human review

    The system goes live with a person checking the risky steps and a clear way to reach a human.

  5. Ongoing5

    Run and improve

    We watch real traffic, fix what breaks and tune on real cases, or hand over with documentation.

Asked on the first call

Do we need to fine-tune a model?

Usually not at the start. Most projects reach their quality bar with a strong API model, good retrieval and clear instructions. Fine-tuning pays off when the task is narrow and high-volume, when a smaller model must match a larger one to cut cost or latency, or when data can't leave your servers. We decide on your eval numbers, not by default.

How accurate is document extraction?

It depends on the documents, so we measure it on yours. We build a test set from real files, report how many rows and fields come out right, and keep a person reviewing output until the numbers support less review. In the refinery build, every extracted row is reviewed before export by design.

What do you need from us to start?

Twenty to fifty representative documents or examples, including the messy ones, the target format (for example your SAP upload sheet), and a domain expert who can answer 'is this right?' questions for about an hour a week.

Can it run on our own servers?

Yes, with trade-offs. The application and open-weight models can run in your cloud or data center; the strongest hosted models can't. We compare quality and cost on your test set before you choose.

How do you stop the model from making things up?

Ground it and check it. The model works from retrieved passages or the document itself, its output follows a schema that is validated before saving, and fields that must match a fixed list, like SAP picklists, are checked against that list. In iKoMatch, each extracted fact must quote the source text it came from. Anything uncertain goes to a reviewer.

What drives the cost of a custom LLM system?

Document volume and size, the number of fields you need, how often a stronger model has to re-run, and how much review tooling and integration the workflow needs. We estimate model spend per document on your samples before launch.

Other services

Have a system in mind? Let's scope it.

A 30-minute call with an engineer who has shipped this before. You leave with a plan, a rough timeline and what it would take — whether or not we build it.