Part II - Evaluation-based system improvement
An LLM response can follow the requested format and still contain incorrect information or fail to answer the question. Evaluating response quality requires representative test questions and clear criteria for judging the answers.
A useful comparison starts with a definition of success and fixed test cases, including important edge cases. The proposed change is compared with a baseline for the complete set and for meaningful groups of questions. When exact matching cannot assess answer quality, human reviewers or a separately prompted model can apply a scoring rubric. The release decision must also account for uncertainty in these measurements.
Chapter 4 starts with the product decision, builds a golden evaluation set, compares one change with a baseline, and carries uncertainty into a release decision. Chapter 5 then matches measures to different response types and asks whether benchmark scores, judge-model results, and human ratings really support the quality claim.