What counts as a custom LLM system?
Anything where a general chatbot isn't enough because the answer depends on your documents, your data or your rules. In practice that means three kinds of build. Retrieval systems answer from your material and cite it. Document AI turns PDFs, scans, drawings or recordings into structured records your systems can load. Specialist pipelines chain several model steps with checks between them. Fine-tuning is a tool inside those builds, not the starting point.
Our largest example is the engineering-drawings-to-SAP system we built for an Indian public-sector refinery. CallAuditAI applies the same idea to call recordings, and the catalog assistant to a product catalog.
How do you choose a model?
We start from the task and a test set, not from a model. The first run uses a strong hosted model to learn what good output looks like on your documents. Then we try cheaper and faster options against the same test set and keep the cheapest one that clears the bar.
Often the answer is two models. In the refinery system a fast Gemini model does the first pass on every drawing, and the premium model runs only when a reviewer asks for a re-evaluation, with a plain-English hint such as "the quantity column is on the right". Cost is tracked per document, so the trade-off stays visible.
Fine-tuning comes in when the test set says a smaller model can't reach the bar with instructions and retrieval alone, when volume makes per-call cost dominant, or when data must stay on your servers. Licenses matter too: we check that open-weight models allow commercial use before we benchmark them.
What drives the cost?
Two things: building the system and running it. Building cost depends on how messy the inputs are, how many fields and rules the output needs, how much review tooling people need, and which systems it must load into. Running cost is mostly model usage: pages, images or audio minutes in, structured text out, times the number of retries and re-runs, plus hosting if you run open-weight models on your own GPUs. We measure both on your samples before launch rather than guess.
How do people stay in the loop?
By making review fast instead of optional. The refinery reviewer sees the drawing beside the extracted table, edits only the fields that are wrong, and can't enter an illegal SAP combination. CallAuditAI flags zero scores in red so managers start there. When the reviewers' corrections show a field is reliably right, review can shrink for that field, and the numbers, not our confidence, make that call.