Roadmap

The book follows how data becomes a model input, how a calculation produces a prediction, and how training can change that calculation. The distinctions between inference, training, and evaluation introduced in Foundations determine which calculations each task requires.

The six Parts organize learning dependencies. Two representation routes share the task definitions: document features can feed a classifier, while ordered token IDs can select vectors for a sequence model. Neither representation must pass through the other.

Six-panel book map showing inputs, operations, and results for the book's six parts, with a training feedback arrow from loss to parameter learning.
Figure 1: The six parts connect model inputs to implementation checks.

The arrows relate prerequisites. Later chapters use fixed parameters during decoding, compare predictions with learning targets during training, or measure fixed-model behavior against task criteria during evaluation, as the local calculation requires.

The second map connects representation, prediction, learning, sequence context, serving, and implementation. These are related study routes, with the task determining which operations are required.

Learning path across six parts, showing token representations, probabilities and loss, gradients, attention, model adaptation and serving, and implementation checks.
Figure 2: Representation, prediction, learning, context, serving, and implementation connect the study routes.

Foundations distinguishes tasks, model families, and learning settings. Four choices remain separate throughout: the required result, the model structure, the training objective, and the inference procedure. For example, next-token targets differ from selected masked-token targets, and sequence-to-sequence tasks specify an output sequence for an input sequence. Representation learning instead shapes vectors for uses such as similarity and retrieval. A task’s output form does not select its model family automatically.

The family comparison below is a lookup aid to the later mechanisms. A recurrent model uses the preceding hidden state when processing the next sequence position. Attention combines information from permitted positions. Vision and multimodal models first convert images, audio, or other inputs into suitable vectors.

Table 1: Model families differ in the context they can use and the outputs they can produce.
Family Context used Typical output Chapters
Linear and logistic models A fixed feature vector A number or one class 5-6, 9
N-gram models A fixed token history A next-token probability 3.4, 15.1
RNNs and LSTMs A recurrent hidden state Sequence labels or generated tokens 15
Transformers Attention over permitted positions Classes, tokens, or a target sequence 16-21
Vision and multimodal models Patches, frames, fields, or text vectors A label, matched pair, or generated text 22

The inputs, model parameters, intermediate results, and training state have different roles and storage costs. Chapter 23 brings these objects together after their calculations have been introduced.

Starting knowledge and tensor notation

This book is for software engineers and technically trained readers who want to understand the calculations behind language-model training and inference. It assumes school-level algebra, functions, averages, summation notation, and familiarity with scalars, vectors, and matrices. Basic probability and the ability to follow Python lists, indexing, functions, and imports are also assumed.

Machine-learning concepts are introduced as they become necessary. The book explains prediction tasks, loss, empirical risk, gradients, attention, and their practical consequences. It uses familiar mathematics directly while explaining what each quantity and operation means for the model. The derivative, gradient, and chain-rule foundation is taught in §5.4. Chapter 11 extends it to automatic differentiation through a graph, including the operations exposed by PyTorch. Readers who need a reminder of matrix multiplication can use the dimension calculation in §5.3 before continuing to embeddings and attention.

A tensor is an array with explicit axes, a numerical type, and a storage device. An array with dimensions \([B,T,D]\) contains \(B\) examples, \(T\) positions per example, and \(D\) coordinates per position. For \([2,3,4]\), there are \(2\cdot3\cdot4=24\) values. Integer token IDs select rows, floating-point arrays hold learned vectors, and Boolean arrays can mark permitted positions. The device identifies where the array is stored and computed, such as central processing unit (CPU) memory or a graphics processing unit (GPU). CUDA is NVIDIA’s software and device platform through which the PyTorch examples can select a supported GPU, such as device cuda:0.

Matrix multiplication sums over a shared axis. For example, multiplying dimensions \([B,F]\) by \([F,D_{\mathrm{out}}]\) removes the shared feature axis \(F\) and produces dimensions \([B,D_{\mathrm{out}}]\). §5.3 presents the calculation and the orientation used by nn.Linear. Chapter 23 later develops storage, strides, copies, and device costs.

Reading paths and scope

Read Chapters 1–6 for targets, representations, and probability outputs. From there, Chapter 7 follows inference with fixed parameters. Chapter 8 develops losses, Chapter 9 evaluation, Chapter 10 update rules, and Chapter 11 gradient propagation. The default path already has the calculus foundation from §5.4 before probability and loss derivatives rely on it. Readers who prefer derivation before optimizer comparisons can read Chapter 11 before Chapter 10.

For text classification, read the loss and evaluation foundations in Chapters 8–9 and the update and gradient calculations in Chapters 10–11. Then continue through nonlinear classifiers, embedding lookup, and Word2Vec in Chapters 12–14 before the classification lab. For the count-model route, begin with n-gram counts and smoothing, read logarithmic loss and perplexity in §§8.1 and 8.5, then continue to the count-model lab. These establish the baseline before recurrent state, attention, and transformers. For adaptation and serving, complete Chapters 17–21 before the QLoRA lab. The multimodal lab examines image representations, captioning, and visual question answering.

The book develops the mathematics behind the supplied LLM Architecture lectures and notebooks, with tensor examples from Python for Data Processing, Lecture 3. References to the course mean these source materials unless another course is identified. It explains the objects and calculations needed to inspect training and inference: tensor dimensions, probabilities, gradients, attention, adapters, and cache storage. Kernel programming, Tensor Core scheduling, memory coalescing, production benchmarking, and complete distributed training configurations are outside this book’s scope. Mixed precision and serving engines appear only where they change a mathematical assumption, memory estimate, or experimental choice.

References and execution guidance

Use the notation reference to check symbols, dimensions, and mask conventions. The library map connects each operation to its first introduction and practical exercise. The Glossary expands abbreviations and returns to the section that explains them.

Each code example states its scope. A runnable calculation checks the listed inputs and outputs. An abbreviated fragment requires the surrounding model or data objects. A course experiment additionally needs its dataset split, checkpoint, training configuration, and held-out results. Passing a tensor assertion establishes neither model quality nor a speed improvement. The Environment and source notebooks explains the environment and evidence to retain.

The References collect the sources cited beside the explanations.