Mathematical Foundations of LLM Training and Inference

A language model’s behavior depends on calculations you can inspect: how text becomes numbers, how errors change weights, and how a trained model produces an answer. This book develops the mathematical foundations of large language models through small examples and computational labs. It helps you connect those calculations to practical decisions about model quality, adaptation, memory, and inference.
Contents
Front matter
Part I: From Text to Mathematical Objects
Part II: From Scores to Loss
- Chapter 5: Linear models: Predictions, boundaries, and slopes
- Chapter 6: Sigmoid and softmax: Scores as probabilities
- Chapter 7: Decoding Strategies: Generating Text Token by Token
- Chapter 8: Cross-Entropy & Perplexity: Measuring Prediction Error and Uncertainty
- Chapter 9: Validation & Testing: Catching Overfitting and Measuring Generalization
Part III: From Error to Learning
- Chapter 10: Optimization and regularization
- Chapter 11: Backpropagation: Tracing Gradients through a Computation
- Chapter 12: Activation Functions: Nonlinear Features for Classification
- Chapter 13: Embedding Tables: Token Lookup, Gradients, and Pooling
- Chapter 14: Word2Vec: Learning Vector Relationships from Context