# Custom LLM development: RAG, document AI and evals

> Mirai Minds builds custom LLM systems for work that depends on a company's own documents and data: answers grounded in retrieval, document AI that turns files into structured records, and fine-tuned models when the numbers justify them. Every system ships with an evaluation set and a human review step, such as engineers checking each extracted row before it reaches SAP.

## Key facts

- **Inputs handled:** PDF, scans, TIFF, DWG and DXF drawings, call audio
- **Review:** Human sign-off before data reaches your system of record
- **Model choice:** API models first; fine-tuning only when evals justify it
- **Evaluation:** A test set from your real documents, re-run on every change
- **Typical build:** 2–8 weeks; large document AI takes longer


## What we build

### Engineering drawings to SAP-ready data

- Problem: Engineers hand-keyed hundreds of line items per drawing into SAP material-master fields.
- System: Gemini finds the bill-of-materials table on PDF, scanned or CAD drawings and extracts every row. A reviewer checks each row beside the drawing, can re-run extraction with a plain-English hint on a stronger model, and exports a 16-column SAP upload sheet.
- Result: Nothing reaches SAP without an engineer's sign-off, and illegal SAP value combinations can't be entered.

### Call recordings scored against a rubric

- Problem: Managers can't listen to every recorded sales or counseling call.
- System: Gemini listens to each recording and scores it against a weighted rubric for opening, discovery, value and closing, producing an Excel scorecard that flags zero scores.
- Result: Every uploaded call gets a scorecard, and managers start with the red flags.

### Answers grounded in a catalog

- Problem: A general chatbot invents products and prices.
- System: An agent calls a retrieval tool that ranks catalog entries with hybrid semantic and keyword search, and shows product cards only from real results.
- Result: Product details come from catalog records, not from the model's memory.

### Language tools for a specialist text

- Problem: Learners need the pronunciation and meaning of each line of scripture, in their own language.
- System: A pipeline transliterates the line into Gurmukhi, applies pronunciation rules, writes a plain explanation in Hindi or English and streams it as speech from a self-hosted text-to-speech model.
- Result: The app team calls one API and gets a spoken explanation back.

## What counts as a custom LLM system?

Anything where a general chatbot isn't enough because the answer depends on your documents, your data or your rules. In practice that means three kinds of build. Retrieval systems answer from your material and cite it. Document AI turns PDFs, scans, drawings or recordings into structured records your systems can load. Specialist pipelines chain several model steps with checks between them. Fine-tuning is a tool inside those builds, not the starting point.

Our largest example is the [engineering-drawings-to-SAP system](/work/refinery-drawings-to-sap-bom) we built for an Indian public-sector refinery. [CallAuditAI](/work/callauditai) applies the same idea to call recordings, and the [catalog assistant](/work/resin-catalogue-assistant) to a product catalog.

## How do you choose a model?

We start from the task and a test set, not from a model. The first run uses a strong hosted model to learn what good output looks like on your documents. Then we try cheaper and faster options against the same test set and keep the cheapest one that clears the bar.

Often the answer is two models. In the refinery system a fast Gemini model does the first pass on every drawing, and the premium model runs only when a reviewer asks for a re-evaluation, with a plain-English hint such as "the quantity column is on the right". Cost is tracked per document, so the trade-off stays visible.

Fine-tuning comes in when the test set says a smaller model can't reach the bar with instructions and retrieval alone, when volume makes per-call cost dominant, or when data must stay on your servers. Licenses matter too: we check that open-weight models allow commercial use before we benchmark them.

## What drives the cost?

Two things: building the system and running it. Building cost depends on how messy the inputs are, how many fields and rules the output needs, how much review tooling people need, and which systems it must load into. Running cost is mostly model usage: pages, images or audio minutes in, structured text out, times the number of retries and re-runs, plus hosting if you run open-weight models on your own GPUs. We measure both on your samples before launch rather than guess.

## How do people stay in the loop?

By making review fast instead of optional. The refinery reviewer sees the drawing beside the extracted table, edits only the fields that are wrong, and can't enter an illegal SAP combination. CallAuditAI flags zero scores in red so managers start there. When the reviewers' corrections show a field is reliably right, review can shrink for that field, and the numbers, not our confidence, make that call.

## Typical stack

- **Models:** Gemini (fast and premium tiers), OpenAI GPT-4o-mini and GPT-4.1, Llama 4
- **Retrieval:** Hybrid semantic and keyword search, Pinecone, LangChain
- **Documents:** PDF, TIFF, JPEG, DWG and DXF intake, Amazon S3 staging for large files, Server-Sent Events progress
- **Evaluation:** Test sets from real documents, LLM judges, Offline replay of past outputs
- **Speech:** Kokoro text-to-speech (self-hosted)
- **Application:** FastAPI, MySQL, PostgreSQL, Vue 3

## Related work

- [Catalog assistant with photo search for a resin manufacturer](https://www.miraiminds.co/work/resin-catalogue-assistant.md) — Pilot on a 15-product catalog sample
- [CallAuditAI: rubric scoring for counseling and sales calls](https://www.miraiminds.co/work/callauditai.md) — Scorecards for an admissions counseling team's recorded calls
- [Engineering drawings to SAP-ready bills of material](https://www.miraiminds.co/work/refinery-drawings-to-sap-bom.md) — Delivered and handed over; every row reviewed before it reaches SAP
- [Explanation and pronunciation tools for a Gurbani learning app](https://www.miraiminds.co/work/gurbani-learning-ai-tools.md) — Delivered to the app team as an authenticated API
- [Myva.ai and CallPaaS: document chatbots and call-center AI](https://www.miraiminds.co/work/myva-callpaas-support-and-call-centre.md) — Both products built and extended across 2023–2025

## Frequently asked questions

### Do we need to fine-tune a model?

Usually not at the start. Most projects reach their quality bar with a strong API model, good retrieval and clear instructions. Fine-tuning pays off when the task is narrow and high-volume, when a smaller model must match a larger one to cut cost or latency, or when data can't leave your servers. We decide on your eval numbers, not by default.

### How accurate is document extraction?

It depends on the documents, so we measure it on yours. We build a test set from real files, report how many rows and fields come out right, and keep a person reviewing output until the numbers support less review. In the refinery build, every extracted row is reviewed before export by design.

### What do you need from us to start?

Twenty to fifty representative documents or examples, including the messy ones, the target format (for example your SAP upload sheet), and a domain expert who can answer 'is this right?' questions for about an hour a week.

### Can it run on our own servers?

Yes, with trade-offs. The application and open-weight models can run in your cloud or data center; the strongest hosted models can't. We compare quality and cost on your test set before you choose.

### How do you stop the model from making things up?

Ground it and check it. The model works from retrieved passages or the document itself, its output follows a schema that is validated before saving, and fields that must match a fixed list, like SAP picklists, are checked against that list. In iKoMatch, each extracted fact must quote the source text it came from. Anything uncertain goes to a reviewer.

### What drives the cost of a custom LLM system?

Document volume and size, the number of fields you need, how often a stronger model has to re-run, and how much review tooling and integration the workflow needs. We estimate model spend per document on your samples before launch.


---

Canonical: https://www.miraiminds.co/services/custom-llm-development
Last updated: 2026-09-23
Publisher: Mirai Minds LLP, 906 Sarthana Business Hub, Nana Varachha, Surat, Gujarat 395013, India. hello@miraiminds.co
