Part II: Reliable model services with retrieval and measurement
A serving system must keep model requests within its time and memory budgets. For the policy-answering service, the application also needs current passages the user may read. Operators then need evidence of whether the model used those passages correctly.
Chapter 4 establishes the request path and release controls. Chapter 5 builds the retrieval path that supplies authorized evidence. Chapter 6 separates service health from answer quality and checks the evaluator itself. The same request and release identifiers connect these measurements, so a wrong answer can be investigated across components.
Chapters in this part
- Serving LLMs with predictable performance: Users of the grounded-answer service expect the first words of an answer within about a second and a steady stream after that. Part I produced a recoverable model checkpoint, but a checkpoint answers no one. Serving replaces the synchronized steps of training with independent requests of different lengths, latency objectives, per-request attention memory, and rolling changes across a fleet of replicas.
- Retrieving evidence for grounded answers: Suppose the remote-work policy changes after the model served in Chapter 4 was trained. Its weights cannot contain that update, but it can still produce a fluent answer. The grounded-answer service finds current passages the user is allowed to read and supplies them with the question. This gives the model evidence for an answer. It does not guarantee that retrieval finds the right passage or that the model uses it faithfully.
- Measuring service health and model quality: The grounded-answer service can return every request in under a second with HTTP status 200 and still tell an employee the wrong remote-work rule. Chapter 4 made the service fast and reversible, and Chapter 5 made it retrieve authorized evidence. Neither shows whether an answer used that evidence faithfully, whether the questions users ask have changed, or which component caused a bad answer.