Models, tasks, and learning
A model output can be judged only against a task: a required result and a criterion for success. A task may require a topic category for a text, a duration estimate for a program run, or a continuation of a prompt, the text supplied to start generation. The same input can support different tasks, so producing an output and knowing whether it is useful are separate steps. Using a model with fixed parameters, changing those parameters through training, and evaluating its behavior also serve different purposes.
Models and tasks
A topic category and a measured duration require different forms of answer. A task defines the required relationship between an input and an output. Classification produces a category, while regression produces a numerical value. Sequence generation produces an ordered output such as text. These task types specify the form of the required result without specifying how it will be calculated.
Artificial intelligence, or AI, is the broad field of building systems for tasks such as perception, reasoning, prediction, and generation. Machine learning, or ML, is the part of AI that fits computational behavior from data. A machine learning model is the fitted numerical system used to produce outputs. A neural network is a machine learning model that processes inputs through layers of numerical calculations. Neural networks underlie familiar chat assistants. A parameter is a stored model coefficient, such as a weight or bias. These coefficients control how the model transforms its inputs. Training changes the coefficients selected for learning, while inference uses their stored values.
The output produced by a model for an input is a prediction. The word does not imply a forecast of the future. Assigning a category to an existing text is a prediction because the model supplies an estimated answer, and that answer may be wrong.
Judging a prediction requires a reference value or another success criterion. A supplied category can serve as the reference for classification, while a measured quantity can serve as the reference for regression. Such a reference answer or measured value is called a target. The prediction is produced by the model, while the target is used to judge that prediction.
Inference and training
Producing a prediction and improving future predictions require different operations. Inference applies a model with its current parameters to an input and returns a prediction. The correct answer is not required as part of the input, and the calculation does not change the parameters. An incorrect prediction alone does not correct the model.
Training adjusts a model’s parameters using information in the training data. In prediction tasks with reference answers, the model first predicts an output. A loss assigns a numerical penalty to the error when that prediction is compared with its reference answer. The average loss can serve as the training objective, the quantity training tries to reduce. An update procedure changes the parameters to pursue that reduction. Later chapters derive the calculation, and the learning settings below explain other sources of training information.
One update may improve some predictions while making others worse. A separate process is needed to measure the resulting behavior without continuing to change the model.
Evaluation measures predictions against task criteria while keeping the tested parameters fixed. It usually uses held-out examples that did not determine those parameters. Evaluation produces evidence about model quality, such as classification accuracy or average prediction error. Repeatedly choosing settings from the same evaluation results can bias that evidence.
Inference, training, and evaluation all include a prediction calculation. The surrounding information and resulting state change distinguish them. Inference needs an input and returns an output. Training also needs information that guides parameter updates, called a learning signal, and changes parameters. Evaluation needs reference information but leaves the parameters unchanged.
Learning signals
Parameter updates require information about what the model should produce or which behavior should improve. Different learning settings are distinguished by the source and form of that information.
Supervised learning uses supplied input-target pairs. A category attached to a text can supervise classification, while a measured value can supervise regression. Human judgment or an instrument may supply the target. An existing record can also supply it.
Text can also provide a reference answer from within itself. In self-supervised learning, the training procedure constructs the target from the data. Given the recorded sentence “The team won,” it can present “The team” and use the following word, “won,” as the reference answer. The model predicts from the preceding text. The recorded continuation supplies the training reference even though other continuations might also be valid. Chapter 1 makes this input-target separation precise using the model’s text units, called tokens.
Unsupervised learning seeks structure without supplied target answers for each example. A clustering method can group texts by similarities in their contents. The resulting groups are properties found by the method and need not match categories required for a later task. Self-supervised learning is often placed within this broader family. This book distinguishes it because its constructed prediction targets have a specific role in training.
Reinforcement learning uses consequences of actions rather than a supplied target for every decision. A decision-making program, called an agent, selects an action. The environment responds with a new observation and possibly a reward. The reward evaluates a consequence of the action but need not identify a uniquely correct action. A poorly chosen reward can encourage unwanted behavior.
These learning settings describe the available training signal. Classification, regression, and sequence generation instead describe the required output. The main calculations in this book begin with supervised and self-supervised prediction. The source of the learning signal identifies what guides fitting. The loss rule and the examples encountered determine the penalty of the resulting predictions.
Loss and risk
A training procedure needs a numerical comparison between a prediction and its target. The loss introduced above supplies that comparison under a chosen rule. Squared error can compare a numerical prediction with a measured value, while cross-entropy can compare predicted class probabilities with a target category. Loss values from different rules do not necessarily have the same meaning or units.
A model can make different prediction errors on different inputs. Before a new input arrives, its error is unknown. The expected penalty across the inputs and reference outcomes the model will encounter is called risk. It depends on which examples occur, how frequently they occur, and how strongly the loss penalizes their errors. For squared error, occasional large mistakes can contribute substantially to risk.
The average loss on an available dataset is called empirical risk. It estimates risk when the dataset represents the intended use. Training tries to reduce this average on the training examples. Because those examples also guide the parameter choices, their average loss can be optimistic about new inputs. A separate evaluation set helps assess performance beyond the training data. §1.3 connects this purpose to the calculation.
Generalization is the ability to perform the task on examples beyond those used for fitting. Evaluation on suitable separate data provides evidence about it. Even that evidence depends on how well the evaluation collection represents the intended use.
Training and evaluation may use different measures. A classifier can be trained by penalizing a low probability for the target category, while evaluation may report the fraction of predictions whose selected category is correct. Chapters 1 and 8 develop loss calculations, and Chapter 9 examines what evaluation supports. Comparing models requires both a clear measure and data appropriate to the intended use.
Model families
The required output does not determine the calculation that produces it. The task specifies the required input-output relationship, while the model family specifies the structure used to calculate it. Models with different structures can perform the same task. A linear model forms scores from weighted numerical features, possibly with an added constant. A feature is a numerical property of an input, such as a word count. The model’s parameters determine each feature’s contribution to the resulting score.
A fixed weighted sum cannot express every relationship between input features and a useful output. A neural network composes layers of numerical transformations. It usually includes nonlinear operations, so the overall calculation cannot be reduced to a fixed weighted sum of the original features. Its parameters control those transformations. Layers construct intermediate representations, collections of values passed to later calculations. Learning through many such layers is called deep learning. Larger networks require more storage and computation, and the value of that cost must be measured for the task.
A language model represents probabilities for language sequences or their components given surrounding context. Some language models use simple counts, while others use neural networks. A large language model, or LLM, is a neural language model with a large parameter set. There is no universal parameter-count cutoff for “large.” Its text predictions can be used for sequence generation and can also be arranged to produce categories or numerical text.
Another distinction concerns what a probabilistic model represents. A discriminative model directly models a decision or a conditional relationship, such as a category given an input. A generative model represents a distribution over possible data, which can support producing new content. Systems described as generative AI use models to produce content such as text, images, or audio. Producing content does not guarantee that it differs from every training example. A generative language model can still perform a classification task when the requested output is a category.
Model family, task, and learning setting answer separate questions. A neural network can perform regression, and a language model can use self-supervised training before further supervised training. The model is also one component of an application. Surrounding software prepares the inputs and applies rules governing the use of outputs. Choosing a model therefore depends on the required result, measured quality, storage and computation costs, and the work of maintaining that application. For text generation, the remaining question is how the model receives and extends a sequence.
Tokens and context
Language-model software receives text, but the neural network operates on numerical representations. Tokenization divides text into supported units and assigns each unit an integer identifier. Each unit is a token. Tokens may correspond to whole words, parts of words, punctuation, or control markers. Chapters 2 and 3 explain how token IDs are assigned and how the token choices affect sequence length.
The context window is the bounded span of tokens available to the model for a computation. Information outside that span cannot directly affect the current prediction. When an input exceeds the limit, the surrounding software must decide what to retain. It may truncate the text or divide it into smaller portions. Longer context also increases memory use and computation, with the cost depending on the model and its implementation.
In autoregressive generation, each next token is predicted from the tokens already available in the sequence. Starting from “The team,” the model might select “won,” append it, and predict another token from the extended prefix. Generation repeats this process until a stopping condition is reached. Each choice changes the available text, but ordinary inference does not change the model’s parameters.
Context and learned parameters play different roles. Context supplies information for the current prediction. Parameters encode the numerical behavior produced by training. Adding text to a prompt can affect the next output without training the model on that text.
Part I now turns these roles into precise computational objects. Chapter 1 constructs inputs and targets, and Chapter 2 prepares token identifiers and batches. The Roadmap shows the paths through the rest of the book. Notation provides a separate reference for symbols when the calculations begin. The Glossary collects short definitions for later lookup. References lists the publications and documentation cited in the book.