Part II: From Scores to Loss

Prepared features need a rule that turns their values into a prediction or score. Chapter 7 uses that score with fixed parameters to compare generation choices. Chapter 8 then compares predictions with learning targets to form a loss.

Chapter 5 establishes affine scores and, in §5.4, local slopes, partial derivatives, gradients, and the chain rule. Chapter 6 uses that foundation when relating scores to probabilities. Because generation with fixed parameters needs only those probabilities, Chapter 7 can compare generation choices before Chapter 8 introduces training loss. Chapter 9 then assesses generalization using held-out evidence.

The shared score calculation therefore supports different uses. Loss and evaluation show how predictions compare with references. Part III extends the derivative foundation through model computations and uses gradients to update parameters.

Part II opener showing the chapters in From scores to loss, with each chapter's purpose, input, operation, result, and dependency arrows.
Figure II.1: Linear scores become probabilities, losses, and evaluation results that can be compared.

Chapters in this part