From AI Model to AI Product

Evaluation-driven development of LLM-based products

Author

Vitaly Rubinovich

An engineer adjusts a luminous model-like instrument on a workbench surrounded by reference books, selected pages, a magnifying lens, and precision tools.

Building dependable software systems around large language models: prompt interfaces, evaluation-driven development (EDD), retrieval-augmented generation (RAG), deterministic workflows, controlled agents, and evidence-grounded release.

Contents

Front matter

  • How an AI model becomes an AI product
    A language model becomes a product when surrounding software supplies information, tests answers, controls actions, and records evidence for release decisions.

Part I - The language-model interface

How language models receive text, generate responses, and interact with the software that uses them.

Part II - Evaluation-based system improvement

Testing answer quality, comparing changes, and deciding whether the evidence supports a release.

Part III - RAG: Supplying external evidence

Finding relevant source material, supplying it to the model, and checking whether the answer uses it correctly.

Part IV - Controlled AI workflows and agents

Controlling multi-step work, tool calls, agent coordination, and information saved across sessions.

Part V - System cases, diagnosis, and release decisions

Complete system cases, failure diagnosis, and release decisions based on quality, cost, latency, and safety.

Appendices