Part IV: Modeling Sequences
A static token vector does not identify which surrounding words matter to a particular occurrence. A sequence model needs both an input representation and a way to use ordered context. Fixed histories, recurrent states, and attention provide different access to earlier information.
Chapter 15 begins with output and target layouts, then recalls the count model’s fixed-history limit before introducing learned recurrent state. Chapter 16 replaces long state paths with direct attention over permitted positions. Chapter 17 composes attention, feature transformations, normalization, positions, and optional expert routing into blocks.
Chapter 18 compares how model families arrange visible inputs and prediction targets. Chapter 19 then connects those training objectives to selected data, packed storage, and compute allocation. Data preparation supports training rather than forming another inference stage. Part V considers later task adaptation and generation state.