Introduction: The Operating Problem

A model can be impressive in an isolated experiment and still fail as a product. Production introduces traffic, failures, shared resources, changing data, delayed feedback about correctness, cost limits, and deployment risk. When a result changes or a job fails, the team also needs to know which code, data, model, and settings produced it. Together, these constraints turn a one-time result into a system whose behavior and changes the team can explain.

The supporting systems store model versions, allocate compute, handle requests, and record what happened. Their design depends on the workload. A service can use a pretrained model without running its own training job. Training or adaptation becomes relevant when the existing model does not meet the task’s needs.

The running case: A grounded-answer service

Suppose a company offers its employees a question-answering service over internal policy documents. A typical question is how many remote-work days a contractor may take. The application checks which documents this employee may read, retrieves relevant passages from the current policies, and supplies the question and passages to a large language model (LLM), a model that generates text from the supplied input. The application returns an answer with citations to the passages it used. Policies change weekly, so the service must answer from documents that the model never saw during training.

  • Retrieval-augmented generation (RAG): A design in which the application retrieves passages from a document collection at request time and supplies them to the model as context for generating the answer.1

  • Groundedness: The degree to which every claim in an answer is supported by the passages supplied to the model. A grounded answer can still be wrong if the supplied passage is outdated or incorrect.2

  • Hallucination: Generated content that the supplied evidence does not support or that contradicts it, such as an answer stating a renewal date that appears in no retrieved passage.3

The model has different roles in three operating modes. In inference, the serving system turns a question and passages into an answer using the model’s learned numerical settings, called weights. In training, a job compares model predictions with target outputs and updates those weights. In evaluation, a job scores answers against reference labels or evidence and records a verdict without changing any weights.

The records of Chapter 1 identify which model, data, and settings produced an answer. When training is needed, Chapters 2 and 3 explain how to distribute the work and keep its data and saved state recoverable. Chapter 4 serves the model within a latency target, Chapter 5 retrieves authorized evidence, and Chapter 6 measures whether answers remain grounded. Chapter 7 automates rebuilds, and Chapter 8 follows one groundedness incident from alert to a verified release.

Scope boundaries

The starting knowledge is basic programming, familiarity with code and data files, and arithmetic with rates and percentages. The reader can recognize a trained model but need not know how training clusters, serving engines, vector search, or workflow systems operate. Those ideas are explained where they are used.

The book focuses on the systems around a model: allocating compute, moving and storing data, scheduling requests, finding relevant passages, measuring performance and answer quality, and checking releases. Low-level GPU kernel authoring, proprietary hardware compilation, and model architecture design are outside its scope. Software products illustrate component roles. They are not required purchases or a prescribed deployment stack.

Eight operating stages and handoffs

The book follows eight stages grouped into three parts: Reproducible training at cluster scale (Part I, Chapters 1–3), Reliable model services with retrieval and measurement (Part II, Chapters 4–6), and Automated workflows and reliable system operation (Part III, Chapters 7–8). Each stage solves one operating problem and passes a defined result to the next stage.

Table 1: The roadmap uses the same three part names and eight stages as the chapters that follow.
Part / stage Problem solved Result passed forward
Part I: Reproducible training at cluster scale Turn changing code, data, and compute into repeatable, recoverable training runs. Reproducible inputs, a recoverable checkpoint, and a versioned training bundle
1. Reproducible runs Identify the exact code, data, configuration, environment, and outputs of one run. Tracked run with versioned inputs and artifacts
2. Distributed training Use many devices while controlling memory, communication, scheduling, and recovery. Checkpoint contents, recovery objectives, and efficiency evidence
3. Storage for training Keep GPUs fed with data and publish only complete checkpoints on the right storage. Committed checkpoint generation and measured storage path
Part II: Reliable model services with retrieval and measurement Serve the trained bundle, supply it with authorized evidence, and measure answer quality. Measured service behavior with grounded, traceable answers
4. Predictable serving Meet latency and throughput targets while managing batching, KV-cache memory, and releases. Versioned serving bundle with performance gates
5. Retrieval for grounded answers Build rebuildable embeddings and indexes and retrieve authorized passages with source identity. Versioned retrieval collection and cited passages
6. Measurement and evaluation Distinguish service health from answer quality and preserve traceable evaluation evidence. Traces, metrics, experiment records, and quality verdicts
Part III: Automated workflows and reliable system operation Automate repeatable work and turn production failures into verified improvements. Run evidence, release decisions, and regression prevention
7. Repeatable workflows Schedule dependencies, retry safely, and commit outputs as one reproducible run. Immutable run manifest with task and artifact lineage
8. Safe system operation Trace a groundedness incident to its cause, repair it, and add a gate that catches the same failure before another release. Regression case and evidence for the next pipeline run

System map

The book map groups the stages into the three operating parts. Its arrows show what passes between them: versioned model files for serving, authorized source passages for an answer, request traces and release identities, and records of the workflow that built a release. The return arrows show how serving measurements inform a change and how an incident becomes an evaluation case and a check on the next workflow run.

Three Part panels contain Chapters 1 to 3, 4 to 6, and 7 to 8. The training branch passes versioned model files to serving, while a separate pretrained-model input reaches serving directly. Retrieval sends authorized passages to the grounded-answer application. Request records and source passages reach measurement, which separates health evidence from quality verdicts. Workflows pass run and release evidence to system operation. Orange dashed arrows return measurement to serving, incident cases to evaluation, and regression checks to the next workflow run.
Figure 1: The three Parts connect reproducible runs, model services, and repeatable workflows. Labeled arrows show the model files, source passages, and records they exchange; dashed arrows return serving evidence and incident cases to later checks. A service may use a pretrained model without running its own training.

Learning path

The learning path presents the same eight stages as a reading sequence. Each chapter adds knowledge used by the connected stages, from identifying a run to investigating and testing a repair. The feedback links match the system map: measurement informs serving changes, and incident evidence changes evaluation and checks on the next workflow run.

Three Part panels list eight numbered chapters in reading order. Part I develops run identity, distributed training, and recoverable storage. Part II develops serving, authorized retrieval, and separate health and quality measurement. Part III develops repeatable workflows and incident repair. Each chapter has a short capability statement. Orange dashed links run from Chapter 6 to 4, from Chapter 8 to 6, and from Chapter 8 to 7. A note distinguishes the reading sequence from a required deployment pipeline.
Figure 2: The reading path follows the same eight stages as the roadmap, grouped into three Parts. Each chapter develops a capability for understanding the connected system. Feedback links revisit serving, evaluation, and workflow checks when production evidence reveals a problem.

Reference material and appendices

The Formula reference collects the quantitative relationships and links to their explanations. Questions and concise answers provide conceptual and calculation practice. The Glossary collects definitions, and References gives complete details for the cited research, standards, and documentation. Footnotes identify the claim each source supports and its limits.

Table 2: The back matter collects formulas, practice questions, definitions, and cited sources.
Reference section Title Focus and content Calling chapters
Appendix A Formula reference Quantitative relationships with links to their teaching locations Chapters 1–8
Appendix B Questions and concise answers Conceptual and calculation questions with answers Chapters 1–8
Glossary Glossary Definitions of terms introduced in the book Chapters 1–8
References References Complete details for cited research, standards, and documentation Chapters 1–8

  1. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474. Curran Associates. https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html. The original RAG paper combines a retriever with a generator for knowledge-intensive tasks. Its model trains both parts jointly, while the service in this book retrieves with a separately built index.↩︎

  2. Rashkin, H., Nikolaev, V., Lamm, M., Aroyo, L., Collins, M., Das, D., Petrov, S., Tomar, G. S., Turc, I., & Reitter, D. (2023). Measuring attribution in natural language generation models. Computational Linguistics, 49(4), 777–840. DOI: 10.1162/coli_a_00486. https://doi.org/10.1162/coli_a_00486. The Attributable to Identified Sources framework judges whether generated statements are supported by an identified source. Support is distinct from the source being true or current.↩︎

  3. Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248. DOI: 10.1145/3571730. https://doi.org/10.1145/3571730. The survey defines hallucination as generated content that is nonsensical or unfaithful to the provided source. This book uses the source-unfaithful sense for answers built from retrieved passages.↩︎