Part III: From Error to Learning

A loss measures prediction error, but its value alone does not specify a parameter change. Gradients describe local sensitivity to parameters. An update rule uses those sensitivities to change the model. The available representations determine which patterns those changes can express.

Section 5.4 established slopes, partial derivatives, gradients, and the chain rule. Chapter 10 can compare update rules using supplied gradients before Chapter 11 follows their calculation through graphs and completes the classifier updates in §11.6. Readers who want the gradient derivation first can read Chapter 11 before returning to Chapter 10.

Chapter 12 shows how nonlinear hidden features separate XOR. Chapter 13 connects token IDs to trainable rows and masked document vectors. Chapter 14 supplies context-prediction objectives and examines the limits of static vectors. Those limits motivate the sequence models in Part IV.

Five chapter panels combine gradient arrows, activation sketches, embedding rows, and word-vector plots.
Figure III.1: The panels associate optimizer updates, graph derivatives, nonlinear features, token lookup, and vector relationships.

Chapters in this part