Glossary
This glossary defines recurring concepts and expands abbreviations. Each entry links to the first explanation or identifies a later full mechanism. Symbol shapes and local notation belong in Notation.
- ablation: A controlled comparison that changes or removes one component to examine its effect. Explanation.
- absolute error: The magnitude of the difference between a numerical prediction and its target, in the target’s units. Explanation.
- absolute positional encoding: A vector determined by an absolute position index and added to the corresponding token representation. Explanation.
- accuracy: The fraction of all evaluated classification decisions that match their reference labels. Explanation.
- activation: An intermediate value produced by a model’s forward calculation, potentially retained for backward. The hidden-layer calculation appears in §12.2, and storage lifetime in §23.4. Explanation.
- activation function: A function transforming a hidden score, commonly with a nonlinear rule between affine layers. Explanation.
- activation memory: Intermediate forward outputs retained until their gradients can be calculated during training. Explanation.
- AdaGrad: An optimizer that scales each gradient coordinate using accumulated squared gradients. Explanation.
- Adam: An optimizer combining exponential averages of gradients and squared gradients, with zero-initialization corrections. Explanation.
- AdamW: Adam with parameter shrinkage applied separately from its adaptive gradient statistics. Explanation.
- add-alpha smoothing: Adding the same positive pseudocount to every vocabulary continuation, then normalizing by the augmented total. Explanation.
- advantage estimate: An estimate of an action or response outcome relative to the expected-reward baseline for its context. Explanation.
- affine map: A linear combination followed by addition of a constant bias. Explanation.
- aliasing: A sampling distortion in which distinct input frequencies become indistinguishable in the samples. Filtering before sampling limits the frequencies that can cause it. Explanation.
- ALiBi: Attention with Linear Biases, adding head-specific distance penalties to permitted causal attention scores before normalization. Explanation.
- alignment: Training and evaluation intended to guide behavior toward specified instructions, preferences, or constraints. Explanation.
- all-reduce: An aggregation across participating devices that returns the aggregate array to each replica. Explanation.
- application programming interface (API): The functions, classes, and input-output conventions exposed by a software library or service. Explanation.
- area under the ROC curve (AUC): Area summarizing the ROC curve, a measure of score ranking that does not itself select a threshold or measure calibration. Explanation.
- artificial intelligence: The field of building systems for tasks such as perception, reasoning, prediction, and generation. The abbreviation is AI. Explanation.
- attention: A calculation of coefficients over permitted content vectors followed by their weighted combination. Explanation.
- attention head: One query-key-value attention calculation with its own learned projections. Explanation.
- attention matrix: An array with query rows and key columns containing the normalized attention coefficients. Explanation.
- attention output: The weighted combination of value vectors calculated for a query position. Explanation.
- attention weight: A normalized coefficient multiplying one permitted value vector in an attention output. Explanation.
- attribution: A contribution assigned to part of an input under a specified comparison, such as the change in a fixed classifier score when a token is removed. Explanation.
- AudioLM: An audio model using semantic and acoustic token representations, staged conditional prediction, and codec decoding to produce a waveform. Explanation.
- autoencoder: An encoder and decoder trained to reconstruct inputs through a constrained intermediate representation. Explanation.
- automatic differentiation: Computation of derivatives by applying local derivative rules to recorded or transformed operations. Explanation.
- autoregressive generation: Producing a sequence by repeatedly selecting a next token and appending it to the available prefix, ordinarily with fixed parameters. Explanation.
- backpropagation: Application of the chain rule in reverse dependency order to compute gradients through a recorded computation. Explanation.
- backpropagation through time: Applying reverse differentiation to the repeated shared operations of an unrolled recurrent computation, abbreviated BPTT. Section 15.5 develops repeated derivative factors. Explanation.
- backward pass: Reverse propagation of sensitivities through a computation to obtain requested derivatives. Explanation.
- bag-of-words representation: Document features that count terms without retaining their full order. Explanation.
- batch: A group of examples processed together by numerical software. Explanation.
- batch size: The number of examples contributing to one update in the stated training procedure. Explanation.
- bayesian optimization: Sequential selection of configurations using a model of previous trial outcomes and a criterion involving predicted performance and uncertainty. Explanation.
- beam search: A sequence search that expands retained partial candidates, scores the expansions, and prunes them to a bounded set. Explanation.
- beam width: The configured number of candidates retained at a beam pruning step, under the declared completion policy. Explanation.
- benchmark: A dataset and associated input, target, and scoring protocol used to compare task performance. Explanation.
- bernoulli distribution: A distribution over outcomes 1 and 0 with probabilities \(p\) and \(1-p\). Explanation.
- BERT: Bidirectional Encoder Representations from Transformers, an encoder model whose original pretraining combined selected-token reconstruction and next sentence prediction. Explanation.
- bias: An adjustable value added independently of input features. A multiclass model has one per output score. Explanation.
- bias correction: Division by initialization factors to correct moving averages that started at zero under the stated moment assumptions. Explanation.
- bidirectional attention: Attention permitting context on both sides of a query position when allowed by the task and mask. Explanation.
- bidirectional recurrence: Separate forward and reverse scans of an available sequence whose outputs are combined at each position. Explanation.
- big-O notation: Notation describing a cost’s growth while ignoring fixed multiplicative factors in the stated regime. Explanation.
- bilingual alignment: Fitting a mapping between separately trained language-vector coordinates using seed translation pairs, then evaluating retrieval on held-out pairs. Explanation.
- binary cross-entropy: The negative log probability of a binary target under predicted probabilities p and 1-p. Explanation.
- BLEU: A translation evaluation measure based on matched short token sequences against references, with a length penalty. Explanation.
- block quantization: Quantization that uses scale information for separate groups of values, trading local range adaptation against metadata cost. Explanation.
- block table: A request’s mapping from logical sequence-block indices to physical cache-block locations. Explanation.
- [BPE:] See Byte Pair Encoding (BPE). Explanation.
- [BPTT:] See backpropagation through time. Explanation.
- broadcasting: Reusing array values across compatible axes, such as adding the same class-bias vector to every example row. Explanation.
- byte fallback: A tokenizer policy that represents otherwise unsupported text through byte units. Explanation.
- Byte Pair Encoding (BPE): A subword method that repeatedly merges frequent adjacent pieces during vocabulary construction and applies the saved merge order during encoding. Explanation.
- byte-level BPE: BPE that starts from UTF-8 byte units, so ordinary text can be represented before learned merges even when a character was absent from the training text. Explanation.
- cache key: An identifier for compatible saved computation, including the relevant preceding context and configuration. Explanation.
- calibration: Agreement between an event’s predicted probability and its observed frequency among cases assigned that probability. Explanation.
- catastrophic forgetting: Loss of previously useful behavior during further fitting on new data. Explanation.
- categorical distribution: Probabilities summing to one over a finite set of mutually exclusive outcomes. Explanation.
- causal attention mask: A position-to-position restriction that prevents a prediction from using future tokens. §16.3 develops the attention calculation. Explanation.
- causal language-model objective: Negative log probability of included recorded tokens conditioned on earlier context, under a declared sum or mean reduction. Explanation.
- CBOW: Continuous bag-of-words, a task that predicts a center word from aggregated surrounding-word vectors. Explanation.
- cell state: The LSTM’s additive state vector, combining retained previous coordinates with gated candidate content. Explanation.
- channel: One measurement coordinate at each image position, such as red intensity. Explanation.
- checkpoint: A saved version of model parameters. A resumable training checkpoint may additionally contain update state and configuration. Explanation.
- Chinchilla rule of thumb: A rough planning ratio of about twenty training tokens per parameter, limited by the empirical setup and chosen cost objective. Explanation.
- chunked prefill: Processing a prompt in portions while preserving its earlier state, allowing other work between portions. Explanation.
- class token: An added learned input vector whose final representation can supply a classifier after interacting with input positions. Explanation.
- class-weight matrix: A matrix with one weight per feature and output class. Explanation.
- classification: A task that assigns an input to one or more defined categories according to the task’s label rules. Explanation.
- classifier head: The part of a model that maps a supplied representation to class scores. Explanation.
- CLIP: Contrastive Language-Image Pretraining, fitting separate image and text encoders using paired comparisons in both retrieval directions. Explanation.
- CLS pooling: Using a designated classification token’s representation as a sequence vector under a suitable model and training objective. Explanation.
- [CNN:] See convolutional neural network. Explanation.
- codebook: A finite table of learned reconstruction vectors selected by discrete codes. Explanation.
- computation graph: A representation of operations and the dependencies connecting their inputs and outputs. Explanation.
- compute dtype: The numerical type used by an operation, which may differ from the stored weight format. Explanation.
- compute-optimal training: Selecting parameter and training-token counts to minimize a fitted loss under a specified training-compute budget. Explanation.
- concatenation: Joining arrays along a specified axis while preserving their entries and order. Explanation.
- conditional generation: Generation of an output conditioned on an input such as a prompt, document, or image. Explanation.
- confidence interval: An interval calculated by a procedure with a stated long-run coverage rate under its sampling assumptions. Explanation.
- confusion matrix: An array counting predicted classes against reference classes under a declared axis and label order. Explanation.
- context window: The bounded span of token positions available directly to the model for a computation. Explanation.
- contextual representation: A representation whose value also depends on the available surrounding sequence. Explanation.
- contiguous tensor: A tensor meeting the layout of a specified memory format, such as ordinary row-major storage. Explanation.
- continuous batching: Changing request membership between execution steps as requests finish and waiting work is admitted. Explanation.
- convex combination: A weighted sum with nonnegative coefficients whose sum is one. Explanation.
- convex objective: A function whose graph lies below or on each segment joining two graph points. Its local minima are global. Explanation.
- convolution: A shared local weighted sum over spatial neighborhoods and input channels. The book uses the deep-learning cross-correlation convention. Explanation.
- convolutional neural network: A network composing shared local filters and nonlinear transformations, abbreviated CNN. Explanation.
- coordinate: One fixed position in a feature vector, with the same assigned meaning across examples. Explanation.
- copy: A tensor with separately allocated storage containing copied values. Explanation.
- copy-on-write: Creating a private copy of shared storage before changing it, preserving the views of other users of that storage. Explanation.
- corpus: A collection of texts used for a language-processing task. Explanation.
- corpus mixture: A distribution combining retained source distributions using nonnegative weights summing to one. Explanation.
- cosine similarity: The dot product of two nonzero vectors divided by their norm product, measuring directional similarity. Explanation.
- count vector: A fixed-order vector whose entries count occurrences of the corresponding features. Explanation.
- cross-attention: Attention deriving queries from the sequence being updated and keys and values from a separate source sequence. Explanation.
- cross-entropy: The expected negative log probability assigned by a prediction distribution, weighted by a target distribution. Explanation.
- CUDA: NVIDIA’s software and device platform for computation on supported graphics processing units. A framework can identify one target as
cuda:0. Explanation. - data leakage: Use of information reserved for evaluation in fitting or development that invalidates the intended evaluation. Explanation.
- data parallelism: Processing different data portions on model replicas, then combining gradients under a declared loss/count convention before consistent updates. Explanation.
- dead ReLU unit: A ReLU unit whose relevant inputs all produce negative preactivations and therefore supply no incoming-weight gradient through that activation. Explanation.
- decision boundary: The boundary separating input regions assigned different outputs by a decision rule. Explanation.
- decode: Repeated processing of a selected next token with earlier cached state to produce logits for the following selection. Explanation.
- decoder step: One forward calculation producing next-token logits for the available prefix. Explanation.
- decoder-only transformer: A transformer commonly using causal self-attention and a vocabulary head for continuation from an available prefix. Explanation.
- deduplication: Identifying repeated or sufficiently similar records and applying a retention policy under a stated matching rule. Explanation.
- deep learning: Machine learning through multilayer neural networks that construct intermediate representations. Explanation.
- dense tensor: An array representation that stores a value at every coordinate, including coordinates whose value is zero. Explanation.
- depth: The number of repeated model blocks in the stated architecture. Explanation.
- derivative: The limiting local slope as an input change tends to zero, when that limit exists. Explanation.
- derivative chain rule: The rule that combines local derivatives through a composition, multiplying along a path and adding contributions from different dependent paths. It is distinct from the probability chain rule. Explanation.
- device: The memory and execution location associated with a tensor, such as CPU or a CUDA GPU. Explanation.
- diffusion language model: A language model using a corruption-and-reconstruction formulation with iterative generation or refinement of token positions. Explanation.
- discriminative model: A model that directly represents a decision or conditional relationship, such as a class given an input. Explanation.
- distillation: Training a student to match teacher-supplied outputs or compatible representations, possibly combined with other target losses. Explanation.
- distributional hypothesis: The idea that words appearing in similar contexts often have related meanings or roles. Context similarity does not by itself establish synonymy. Explanation.
- document frequency: The number of fitted documents containing a term at least once. Explanation.
- document-term matrix: A matrix with document rows and shared term-feature columns. Explanation.
- domain token: A vocabulary piece retained or added for text from a particular field. Explanation.
- dot product: The sum of products of corresponding coordinates in two equal-length vectors. Explanation.
- double quantization: Quantization of scale metadata in addition to weight values, requiring nested reconstruction steps. Explanation.
- DPO: Direct Preference Optimization, fitting a policy on preferred-rejected response pairs using reference-relative log probability differences. Explanation.
- dropout: Random removal and appropriate rescaling of activation coordinates during training, disabled for ordinary evaluation. Explanation.
- dtype: The numerical type used to represent a tensor element. Explanation.
- early stopping: Selecting a saved model version using validation behavior and ending further fitting under a declared stopping rule. Explanation.
- elementwise product: Multiplication of corresponding coordinates of compatible arrays, written \(\odot\). Explanation.
- embedding: A numerical vector associated with an object such as a token. A learned embedding is adjusted to support a training objective. Explanation.
- embedding lookup: Selection of a vector row using a discrete identifier. Explanation.
- embedding matrix: A parameter matrix containing one vector row per vocabulary entry. Explanation.
- empirical risk: Average loss over an observed dataset. Its interpretation as an estimate of risk depends on the data and fitting procedure. The arithmetic is in §1.3. Explanation.
- encoder: A component that computes a representation from an input. Later chapters distinguish recurrent, attention-based, and multimodal encoders. Explanation.
- encoder-decoder transformer: A model encoding a source sequence and using a decoder with causal target self-attention and cross-attention to that encoded source. Explanation.
- encoder-only transformer: A transformer producing input representations through bidirectional permitted attention, with a task head determining its output and loss. Explanation.
- entailment: A relation in which a hypothesis follows from a premise. Section 18.1 explains classification of that relationship. Explanation.
- entropy: The expected surprisal of outcomes when their occurrence probabilities and the probabilities used in the logarithm come from the same distribution. Explanation.
- epoch: One pass through the training examples. Explanation.
- evaluation: Measuring predictions against task criteria while keeping the tested parameters fixed. Explanation.
- exact match: An output score that checks equality with an accepted reference under a declared text-normalization rule. Explanation.
- expert: A subnetwork selected to process some token representations and return an output of the required width. Explanation.
- expert capacity: The maximum number of token assignments admitted to an expert within a specified routing batch. Explanation.
- exploding gradient: A derivative signal that grows enough along a computation path to disrupt useful updates or numerical calculation. Explanation.
- exposure bias: The mismatch between recorded prefixes used during teacher-forced training and self-generated prefixes encountered during generation. Explanation.
- F1 score: The harmonic mean of precision and recall, with a declared policy for undefined component values. Explanation.
- false negative: An actual positive example predicted negative, abbreviated FN. Explanation.
- false positive: An actual negative example predicted positive, abbreviated FP. Explanation.
- false-positive rate (FPR): The fraction of actual negatives predicted positive. Explanation.
- fast Fourier transform: An algorithm for calculating a discrete Fourier transform efficiently, abbreviated FFT. Window selection is a separate operation. Explanation.
- feature: A numerical property of an input used by a model, such as a word count. Explanation.
- feed-forward network: A network transforming each token’s feature coordinates independently using the same parameters, abbreviated FFN. Section 17.4 develops its mechanism. Explanation.
- [FFN:] See feed-forward network. Explanation.
- [FFT:] See fast Fourier transform. Explanation.
- filtering: Removing or downweighting data records according to declared inclusion criteria. Explanation.
- fine-tuning: Continuing training from a pretrained checkpoint on selected data, updating all or some parameters or added components. Explanation.
- finite difference: A change in a function value divided by a nonzero change in its input, used to measure an average slope over that interval. Explanation.
- first-moment estimate: Adam’s exponential average of signed gradient coordinates. Explanation.
- FlashAttention: An exact attention algorithm that processes position comparisons in blocks while combining normalized contributions without storing the full score matrix. Explanation.
- forget gate: An LSTM sigmoid vector scaling the previous cell state’s contribution. Explanation.
- frozen base weight: A pretrained weight retained for computation but excluded from parameter updates. Explanation.
- full fine-tuning: Further fitting with all base-model parameters eligible for updates. Explanation.
- full-batch gradient descent: An update procedure that computes each gradient using the entire training set. Explanation.
- gated linear unit: A two-branch activation multiplying a linear projection by a sigmoid of another projection, abbreviated GLU. Explanation.
- gated recurrent unit: A recurrent cell using reset and update gates with one carried state rather than separate LSTM cell and hidden states, abbreviated GRU. Explanation.
- GELU: Gaussian Error Linear Unit, multiplying an input by the standard normal cumulative probability at that input. Explanation.
- generalization: Ability to perform the task on examples beyond those used for fitting. Explanation.
- generative AI: Systems using models to produce content such as text, images, or audio. Novelty relative to all training data is not guaranteed. Explanation.
- generative model: A model of a distribution over possible data that can support producing content. Explanation.
- global minimum: The lowest objective value across all allowed parameter choices. Explanation.
- gradient: The collection of partial derivatives of a scalar function in the order of its input or parameter coordinates. Explanation.
- gradient accumulation: Adding gradients from multiple backward calls into persistent parameter-gradient buffers before an optimizer update. This differs from summing graph-path contributions inside one backward call. Explanation.
- gradient clipping: Rescaling or restricting gradient values before an optimizer consumes them. Norm clipping bounds the combined gradient norm. Explanation.
- gradient descent: Updating parameters by subtracting a positive multiple of the current loss gradient. Explanation.
- gradient-boosted decision trees: An additive tree model whose successive trees fit negative loss derivatives. Residual fitting is the squared-error special case. Explanation.
- greedy decoding: Selecting a token with maximum current probability, with a specified rule for ties. Explanation.
- grid search: Evaluation of combinations from declared lists of hyperparameter values. Explanation.
- grouped-query attention: Attention assigning groups of query heads to shared key-value heads, abbreviated GQA. Explanation.
- hallucination: Generated content presented as established fact without support in the available evidence. A claim can lack support or contradict the evidence. A formatting failure is separate. Explanation.
- head dimension: The coordinate width used by an individual attention head in the specified layout. Explanation.
- head-of-line blocking: Delay of later work behind earlier long work sharing a queue or execution schedule. Explanation.
- heatmap: A display that encodes numerical entries as colored cells using a stated quantity and color scale. Explanation.
- hidden representation: An intermediate feature vector computed inside a model before its final output. Explanation.
- hidden state: An internal model representation carrying information through a computation. Explanation.
- hierarchical softmax: A normalized word-prediction model that multiplies conditional branch probabilities along a tree path to a word leaf. Explanation.
- hop size: The number of waveform samples between successive analysis-window starts. Explanation.
- hyperparameter: A training or architecture setting selected outside the fitting of model weights, such as learning rate or width. Explanation.
- hyperplane: The set of points in a \(D\)-dimensional space satisfying one affine equation with a nonzero coefficient vector. Explanation.
- ignore value: A reserved target value instructing a loss function to exclude a position. It is not an input vocabulary ID. Explanation.
- implementation mapping: The relationship between mathematical objects and code objects, including axes, numerical types, and devices. Explanation.
- in-context learning: Supplying task examples in the prompt while keeping the model parameters fixed. Explanation.
- inference: Applying a model with fixed parameters to produce an output for an input. Explanation.
- input gate: An LSTM sigmoid vector scaling candidate content added to the cell state. Explanation.
- instruction tuning: Supervised fitting on instructions and corresponding responses, possibly with additional context. Explanation.
- inter-token latency: The measured gap between consecutive received output events, with event-versus-token boundaries declared. Explanation.
- jacobian: A matrix whose entry in output row i and input column j is the partial derivative of output i with respect to input j. Explanation.
- k-fold cross-validation: Repeatedly fitting on all but one of several development-data folds and evaluating on the held-out fold, while keeping a final test set separate. Explanation.
- key: A projected source-position vector used in query-key compatibility scores. Explanation.
- key projection: The learned map from a source representation to key coordinates. Explanation.
- KL penalty: A term penalizing divergence of a trained distribution from a specified reference distribution. Explanation.
- Kullback-Leibler (KL) divergence: The extra expected log loss from using a predicted distribution in place of the target distribution. Finite values require positive predictions wherever target mass is positive. Explanation.
- KV cache: Per-layer attention keys and values from processed positions, retained for reuse in later compatible causal calculations. Explanation.
- KV-cache offloading: Moving retained attention state between memory tiers to trade accelerator capacity for transfer and scheduling work. Explanation.
- l1 regularization: Addition of an absolute-parameter penalty to a data-fitting objective. Explanation.
- l2 penalty: A multiple of the parameters’ squared magnitudes added to a data loss. Explanation.
- l2 regularization: See L2 penalty. Explanation.
- label: A supplied category that serves as the target for a classification example. Explanation.
- labeled dataset: A collection that pairs each input with its reference label or labels. Explanation.
- language model: A model of probabilities for language sequences or components conditioned on available context. Explanation.
- large language model: A neural language model with a large parameter set, with no universal numerical cutoff. The abbreviation is LLM. Explanation.
- LayerNorm: Normalization using the mean and variance across a token’s feature coordinates, followed by learned gain and bias. Section 17.3 develops the calculation. Explanation.
- leaf tensor: A tensor without a recorded operation producing it. Directly created trainable parameters commonly act as leaves in autograd. Explanation.
- learning rate: A multiplier controlling the size of an update step. Explanation.
- learning-rate decay schedule: A declared rule for lowering the update multiplier over a training horizon, such as step, linear, or cosine decay. Explanation.
- length penalty: A declared adjustment to sequence scores that changes their preference for output length. Explanation.
- linear model: A model that combines numerical features through weights, possibly adding a constant bias. Explanation.
- linear regression: Prediction of a numerical target through an affine combination of input features. Explanation.
- [LLM:] See large language model. Explanation.
- load balancing: Training or scheduling pressure encouraging routed work to be distributed across experts, without guaranteeing identical use. Explanation.
- local minimum: An objective value no greater than values at nearby allowed parameter choices. Explanation.
- log-mel features: Logarithmically scaled mel-band values with the floor and scaling convention expected by the model. Explanation.
- log-odds: The logarithm of odds, mapped from probability by the logit function. Explanation.
- log-sum-exp: The logarithm of a sum of exponentials, evaluated stably by factoring out the maximum finite score. Explanation.
- logistic regression: In the binary case, a model that makes log-odds an affine function of the input features, with sigmoid giving the event probability. Explanation.
- logit: An unnormalized classification score. In a binary logistic model it represents log-odds. Multiple softmax scores specify probabilities through their differences. Explanation.
- long short-term memory: A gated recurrent cell with separate cell and hidden states, abbreviated LSTM. Explanation.
- LoRA: Low-Rank Adaptation, training two factors whose scaled product adds a change to a frozen base matrix. Explanation.
- LoRA rank: The inner dimension \(r\) of the two LoRA factors. It bounds the rank of their product \(BA\) without guaranteeing that the product has rank \(r\). Explanation.
- loss: A numerical penalty assigned to a prediction under a chosen comparison rule. Explanation.
- loss mask: A rule or array identifying which target comparisons contribute to the loss and its reduction. Explanation.
- low-rank approximation: Reconstruction using a limited number of independent matrix directions. Explanation.
- [LSTM:] See long short-term memory. Explanation.
- machine learning: The part of artificial intelligence that fits computational behavior from data. The abbreviation is ML. Explanation.
- machine learning model: A computational system whose numerical behavior is adjusted using data. Explanation.
- macro average: An average of per-class metrics giving equal weight to each class, with an explicit undefined-value policy. Explanation.
- markov assumption: In the n-gram setting, restricting prediction to a fixed suffix of the preceding tokens. Explanation.
- MASK: A configured vocabulary placeholder for text hidden by a masked-token training procedure. Explanation.
- masked language modeling: Predicting original tokens at selected positions from a changed input and its visible context. The fuller objective is in §18.2. Explanation.
- masked pooling: Aggregation that excludes invalid positions, with the denominator also restricted to valid positions when taking a mean. Explanation.
- max pooling: Taking the largest retained value separately in each vector coordinate. Explanation.
- maximum-likelihood estimation: Fitting parameters to maximize the probability assigned to recorded answers. In the unsmoothed count model, this yields relative continuation frequencies. Explanation.
- mean pooling: Averaging retained vectors across a sequence axis to produce one vector of the same coordinate width. Explanation.
- mean squared error: The mean of squared residuals over a nonempty set of examples, in squared target units. Explanation.
- mel spectrogram: A representation grouping spectral magnitude or power into bands on a perceptual frequency scale under a stated filterbank convention. Explanation.
- micro average: A metric calculated after pooling the relevant outcome counts across classes. Explanation.
- mini-batch gradient descent: Descent using the mean gradient of a selected subset of examples processed together. Explanation.
- mixed precision: Using different floating-point formats for different computations or stored state, with explicit numerical checks. Explanation.
- mixture of experts: A layer storing several expert networks and routing each token through a selected subset, abbreviated MoE. Explanation.
- MLE: See maximum-likelihood estimation. Explanation.
- [MLM:] See masked language modeling. Explanation.
- [MLP:] See multilayer perceptron. Explanation.
- [MoE:] See mixture of experts. Explanation.
- momentum: An update method that combines the current gradient with a decayed stored direction. The book uses an unnormalized decayed sum. Explanation.
- multi-head attention: Combining several separately projected attention calculations on the same input, then joining and projecting their outputs. Explanation.
- multi-query attention: Attention in which multiple separate query heads share one key-value head, abbreviated MQA. Explanation.
- multilayer perceptron: A feed-forward network composing affine layers with nonlinear activations, abbreviated MLP. Explanation.
- n-gram: A contiguous sequence of \(n\) tokens. A bigram has two and a trigram three. Explanation.
- n-gram language model: A model that predicts a token using at most the preceding \(n-1\) tokens, often from conditional frequency counts. Explanation.
- nat: The unit of information or logarithmic loss based on natural logarithms. Explanation.
- nearest neighbor: The candidate with highest selected similarity or lowest selected distance to a query. Explanation.
- negative log likelihood: The negative logarithm of the model probability assigned to an observed event or target, abbreviated NLL. Explanation.
- negative predictive value: The fraction of negative predictions that are actually negative, abbreviated NPV and defined when negative predictions exist. Explanation.
- negative sample: A word drawn from a noise distribution and assigned a negative label in a sampled comparison, even if linguistically plausible. Explanation.
- negative sampling: Training an observed-pair versus sampled-noise classifier using a limited number of noise draws. Explanation.
- neural network: A model that composes layers of parameter-controlled numerical transformations, usually including nonlinear operations. Explanation.
- next sentence prediction: The original BERT classification task distinguishing a continuing second segment from a randomly sampled one, abbreviated NSP. Explanation.
- NormalFloat 4 (NF4): Four-bit codes selecting 16 nonuniform reconstruction levels designed around a normalized normal distribution, combined with block scales. Explanation.
- normalization: In text preparation, a transformation making selected text forms consistent, potentially discarding distinctions. Numerical normalization has a different locally defined meaning. Explanation.
- object detection: Predicting categories and spatial locations for individual objects in an image, commonly as labeled bounding boxes. Explanation.
- odds: The ratio \(p/(1-p)\) of an event’s probability to its complement, for \(0<p<1\). Explanation.
- one-hot target: A class-target vector with one at the reference class and zero at other classes. Explanation.
- optimizer: A rule that converts gradients and any stored update state into parameter changes. Explanation.
- optimizer state: Update information retained between steps separately from model parameters and the current gradient. Explanation.
- optimizer step: An operation that changes parameters using their gradients and any stored optimizer state. Explanation.
- output gate: An LSTM sigmoid vector scaling the transformed new cell state to produce the exposed hidden state. Explanation.
- output projection: The learned map from joined attention-head outputs to the model’s required output width. Explanation.
- output throughput: Aggregate output tokens divided by elapsed observation time across the stated requests. Explanation.
- overfitting: Fitting that adapts to training-sample details at the expense of performance on relevant unseen data. Explanation.
- overflow: A numerical result exceeding the finite range of its floating-point type. Explanation.
- padding: Filling added batch positions with a chosen value so unequal sequences share a rectangular stored shape. Explanation.
- padding mask: An array recording retained and added positions. The illustrated convention uses one for retained positions and zero for padding. Explanation.
- PagedAttention: Attention over blocked KV storage through a mapping from logical sequence blocks to physical memory blocks. Explanation.
- parameter: A stored model coefficient, such as a weight or bias. Training changes selected coefficients, while frozen coefficients remain parameters. Explanation.
- partial derivative: A derivative with respect to one input or parameter while holding the others fixed. Explanation.
- pass@k: The fraction of programming tasks for which at least one of \(k\) generated solutions passes the declared tests. Explanation.
- patch token: A vector representing one image region, usually produced by flattening and a learned projection. Explanation.
- path length: The number of computational transitions connecting two positions or values in the specified graph. Explanation.
- PCA: Principal component analysis, a linear projection onto orthogonal directions of greatest sample variance. Explanation.
- PCM: Pulse-code modulation, storing quantized amplitudes at successive sample times. Explanation.
- PEFT: Parameter-efficient fine-tuning, fitting selected parameters or smaller added components rather than the whole base model. Explanation.
- per-token transformation: Applying the same learned transformation independently to each position’s vector. Explanation.
- permutation equivariant: Having outputs permute in the same way as permuted inputs, under the specified operation and visibility assumptions. Explanation.
- perplexity: The reciprocal geometric mean probability assigned to included target tokens, equal to the exponential of their mean natural-log loss. Explanation.
- policy: A conditional distribution over output actions, such as tokens forming a response to a prompt. Explanation.
- pooling: Combining a sequence’s vectors into one fixed-width vector. Explanation.
- post-norm transformer: A block arrangement applying normalization after a sublayer output is added to its input. Explanation.
- PPO: Proximal Policy Optimization, a policy-update method whose clipped surrogate objective discourages large changes from a rollout policy. Explanation.
- pre-norm transformer: A block arrangement normalizing the input of each sublayer before adding that sublayer’s output to the residual stream. Explanation.
- preactivation: The affine score before an activation function is applied. Explanation.
- precision: The fraction of predicted positives that are actually positive, defined when at least one positive prediction exists. Explanation.
- precision-recall curve: Precision plotted against recall as a binary classification threshold changes. Explanation.
- prediction: The model’s estimated output for an input. It need not concern a future event. Explanation.
- prefill: Processing a supplied prompt to produce initial key-value state and logits, whose final position supports the first output selection. Explanation.
- prefix: The sequence preceding the token being predicted. Explanation.
- prefix cache: Stored state reusable across requests sharing a compatible already processed prompt prefix. Explanation.
- pretraining: Initial broad model fitting before later task use or adaptation. Explanation.
- probability chain rule: The factorization of a joint sequence probability into successive probabilities, each conditioned on the preceding sequence elements. Explanation.
- prompt: Input text supplied to start generation, potentially including a task instruction or examples. Explanation.
- proximal step: A separate minimization balancing closeness to a proposed value against a penalty. Explanation.
- pruning: Removing selected weights or structures under a declared criterion, potentially followed by further fitting. Explanation.
- QLoRA: Quantized Low-Rank Adaptation, fitting higher-precision adapters while retaining the frozen base in a quantized format. Explanation.
- quantization: Representing values with a limited set of compact codes. Reconstructing a value from its code and metadata generally introduces approximation error. Section 20.5 gives the encoding and reconstruction arithmetic. Explanation.
- query: A projected vector at the position whose attention output is being computed, compared with source keys. Explanation.
- query projection: The learned map from an input representation to query coordinates. Explanation.
- RAG: Retrieval-augmented generation, selecting external material and supplying it as context for a generator. Explanation.
- random search: Evaluation of hyperparameter configurations drawn from declared distributions. Explanation.
- rank: The number of independent directions in a matrix. Retaining fewer singular directions limits the rank of its reconstruction. Explanation.
- recall: The fraction of actual positives predicted positive, defined when actual positives exist. Explanation.
- receiver operating characteristic (ROC) curve: A plot of true-positive rate against false-positive rate as a binary classification threshold changes. Explanation.
- receptive field: The region of an input that can affect a particular output value. Composing local layers can enlarge it. Explanation.
- recurrent neural network: A network that repeatedly combines a current input vector with a carried hidden state using shared parameters, abbreviated RNN. Explanation.
- reduction: An operation that combines array entries, such as summing or averaging per-example losses into a scalar. Explanation.
- reference policy: A fixed response distribution used as the comparison for a policy-training objective. Explanation.
- regression: A task that predicts a numerical value to be compared with a numerical target. Explanation.
- regularization: Changes to fitting intended to discourage reliance on patterns that do not extend beyond training data, with benefits assessed on held-out examples. Explanation.
- reinforcement learning: Learning behavior from the consequences of actions, including rewards supplied by an environment. Explanation.
- relative position bias: An adjustment to an attention score based on the displacement between query and key positions. Explanation.
- ReLU: The rectified linear unit, which returns the maximum of zero and its input. Explanation.
- request batch: A group of requests sharing an execution step while retaining separate sequence and cache state. Explanation.
- residual: The target minus its numerical prediction under this book’s sign convention. Explanation.
- residual connection: Addition of a transformation’s output to its compatible input representation. Explanation.
- ResNet: A residual network with bypass paths around learned transformations, using compatible shapes or projections at additions. Explanation.
- reward model: A learned scorer of prompt-response pairs fitted from judgments or other declared reward information. Explanation.
- risk: Expected prediction loss under the distribution of inputs and outcomes relevant to the intended use. Explanation.
- RLAIF: Reinforcement learning from AI feedback, using AI-generated judgments for some feedback records. Explanation.
- RLHF: Reinforcement learning from human feedback, using human judgments to supply or train a reward signal for policy learning. Explanation.
- RMSNorm: Root mean square normalization, which divides coordinates by their root-mean-square scale without centering and applies learned coordinate gains. Explanation.
- RMSProp: An optimizer that scales gradient coordinates using an exponential average of squared gradients. Explanation.
- [RNN:] See recurrent neural network. Explanation.
- RNN cell: The shared recurrent function that produces a new hidden state from the current input and preceding state. Explanation.
- role token: A configured marker for a chat span’s role, such as user or assistant. The marker itself grants no application permissions. Explanation.
- rollout: In the preference-training example, a response sampled under a saved language-policy version. Explanation.
- [RoPE:] See rotary positional embedding. Explanation.
- rotary positional embedding: Position-dependent rotation of paired query and key coordinates, producing relative-angle dependence in their dot products, abbreviated RoPE. Explanation.
- ROUGE: A family of summary evaluation measures based on text overlap with references, including n-gram and longest-common-subsequence measures. Explanation.
- router: A learned component assigning scores or coefficients to candidate expert networks for each input token. Explanation.
- sampling rate: The number of amplitude samples per second in each audio channel. Explanation.
- scaled dot-product attention: Attention using query-key dot products divided by the square root of their width before masking and row normalization. Explanation.
- scaling law: An empirical fitted relationship between resource quantities and a measured outcome. Explanation.
- score matrix: An array with one row per example and one column per output class score. Explanation.
- second-moment estimate: Adam’s exponential average of squared gradient coordinates, distinct from variance around the first moment. Explanation.
- self-attention: Attention deriving queries, keys, and values from the same representation sequence. Explanation.
- self-supervised learning: Learning from prediction targets constructed from the data itself, such as recorded following tokens. Explanation.
- semantic segmentation: Assigning a category to each image pixel, with spatial reference annotations used in supervised training. Explanation.
- SentencePiece: A toolkit that supports BPE and unigram tokenization on raw sentences, including an explicit whitespace representation. Explanation.
- sequence acceptor: A model that reads a sequence and returns one result, such as a sequence class label. Explanation.
- sequence generation: A task that produces an ordered output, such as a text sequence. Explanation.
- sequence packing: Placing several shorter records in shared sequence storage with explicit boundary, attention, position, and target policies. Explanation.
- sequence score: The accumulated conditional log probability of a candidate continuation. Explanation.
- sequence transducer: A model that maps an input sequence to an output sequence, whose length may differ from the input length. Explanation.
- sequence-to-sequence modeling: A task arrangement that produces an output sequence conditioned on an input sequence, as in translation. Explanation.
- sequential dependency: A dependency requiring an earlier result before the next operation can be computed. Explanation.
- [SFT:] See supervised fine-tuning. Explanation.
- short-time Fourier transform: Sine and cosine comparisons over successive weighted waveform windows, producing time-indexed complex frequency coefficients, abbreviated STFT. Explanation.
- sigmoid: The function 1/(1+exp(−z)), mapping a finite real logit to a probability strictly between zero and one. Explanation.
- SiLU: Sigmoid linear unit, \(u\sigma(u)\), also called Swish with parameter one. Explanation.
- singular value decomposition: Factorization into left directions, nonnegative singular values, and right directions, abbreviated SVD. Explanation.
- skip-gram: A task that predicts surrounding context words from a center word. Explanation.
- small language model: A relatively low-parameter language model considered within a stated resource and task budget, with no universal numerical cutoff. Explanation.
- smoothing: Adjusting count-based estimates to allocate probability to unseen events. The seen and unseen calculations are in §3.4. Explanation.
- soft thresholding: Coordinate shrinkage toward zero by a threshold, setting values within the threshold exactly to zero. Explanation.
- softmax: Exponentiating competing scores and dividing each by their shared exponential sum to form a categorical distribution. Explanation.
- source faithfulness: Agreement of an output with its supplied evidence, distinct from external factual correctness when that evidence is wrong or incomplete. Explanation.
- source-mixture weight: The probability of choosing a source in a sampler whose item type, such as document or token block, is specified. Explanation.
- sparse: Having many zero-valued entries. This describes the values, while a sparse storage format describes how they are stored. Explanation.
- sparse activation: Evaluating only a selected subset of stored model parameters for an input. Explanation.
- sparse matrix format: A storage representation recording nonzero values and their locations instead of storing every matrix coordinate. Explanation.
- special token: A vocabulary entry that marks structure or serves a processing role, such as an end marker or padding value. Explanation.
- specificity: The fraction of actual negatives predicted negative, defined when actual negatives exist. Explanation.
- SpeechVerse: An audio-language system combining continuous audio features and text instructions to produce text, with staged training of connections and adapters. Explanation.
- standard error: An estimate of how much a sample statistic would vary across repeated samples under the stated sampling assumptions. Explanation.
- static batching: A scheduling policy that keeps a fixed request group until it finishes before admitting a replacement group. Explanation.
- static embedding: A token-type vector that stays the same across occurrences when the fitted table is fixed. Explanation.
- stationary point: A point where a differentiable function’s gradient is zero, without necessarily being a minimum. Explanation.
- statistical bias: At a fixed input, the average fitted prediction across possible training samples minus the population mean target. This differs from an affine bias parameter. Explanation.
- [STFT:] See short-time Fourier transform. Explanation.
- stochastic gradient descent: In the strict sampling sense, descent using one sampled example per update. Software also uses SGD for the corresponding rule applied to mini-batches. Explanation.
- storage offset: The element index of a tensor view’s first logical value relative to the underlying storage start. Explanation.
- storage precision: The numerical format retained in memory, distinct from the format used by a calculation. Explanation.
- storage stride: The number of stored elements traversed when one logical index increases by one. Explanation.
- straight-through approximation: A backward rule passing a gradient through a discrete selection as though that selection were an identity map. Explanation.
- stride: The spacing between successive image-window starts. Tensor storage stride is a separate element-offset quantity in §23.3. Explanation.
- subgradient: In one dimension, the slope of a line touching a convex graph at the chosen point and staying below or on it everywhere. Explanation.
- supervised fine-tuning: Further fitting on input-output examples of desired behavior, abbreviated SFT. Explanation.
- supervised learning: Learning from supplied input-target pairs. Explanation.
- support: The candidates assigned nonzero probability by a distribution. Explanation.
- surprisal: The negative logarithm of the probability assigned to an event. Explanation.
- SwiGLU: A gated activation multiplying one linear projection by the SiLU-transformed second projection before output projection. Explanation.
- t-SNE: T-distributed stochastic neighbor embedding, a nonlinear display method that tries to preserve local similarities in fewer dimensions. Explanation.
- TabTransformer: A model applying attention to categorical field embeddings, then joining their outputs with continuous features for an MLP. Explanation.
- tanh: The hyperbolic tangent, mapping a real input to a signed value between negative one and one. Explanation.
- target: A reference answer or measured value used to judge a prediction. Explanation.
- task: The required relationship between an input and an output, with a criterion for judging success. Explanation.
- teacher forcing: Supplying recorded previous tokens as decoder inputs during training rather than feeding back its sampled predictions. Explanation.
- temperature: A positive divisor on logits before softmax that changes their relative probability gaps. Explanation.
- tensor: A numerical array with explicit axes. A software tensor also has a numerical type and storage device. Explanation.
- term frequency and inverse document frequency (TF-IDF): Weighting a term using its frequency within a document and its rarity across the fitted corpus. Formula, smoothing, and normalization conventions must be specified. Explanation.
- test set: Examples reserved for assessment after development choices, rather than used to make those choices. Explanation.
- [TF-IDF:] See term frequency and inverse document frequency (TF-IDF). Explanation.
- tied weight: A parameter matrix reused by different operations, with derivatives contributed by all its tracked uses. Explanation.
- token: One supported unit in a model’s text sequence, such as a word piece, punctuation mark, or control marker. Explanation.
- token budget: The number of token occurrences processed during training, including repetitions when present. Explanation.
- token ID: An integer identifying one entry in a tokenizer’s vocabulary. Its numerical distance from another ID is not semantic similarity. Explanation.
- token loss: The negative log probability assigned to one recorded token at a supervised prediction position. Explanation.
- token mixing: Combining information across sequence positions, as attention does. Explanation.
- tokenization: Dividing text into supported units and assigning their vocabulary identifiers. The mapping is developed in §2.1. Explanation.
- tokenizer: Software that prepares text units and assigns IDs using a configured vocabulary, normalization, and control-token rules. Explanation.
- top-k sampling: Sampling after retaining and renormalizing the k highest-probability candidates. Explanation.
- top-p sampling: Sampling from the shortest probability-sorted prefix whose cumulative mass reaches or exceeds a specified threshold. Explanation.
- TPOT: Time per output token for a request, using its post-first-token interval divided by output-token count minus one. Explanation.
- trace: An ordered account of intermediate values and state changes in a computation. Explanation.
- training: Adjusting model parameters using information supplied or constructed from training data or interactions. Explanation.
- training example: An input and associated learning information used to adjust a model. In supervised prediction this is an input-target pair. Explanation.
- training objective: The quantity a training procedure tries to reduce, such as average prediction loss. Explanation.
- training set: Examples used to adjust model parameters. Explanation.
- transformer: A neural-network architecture that combines attention with transformations at each position. Visibility rules determine which positions can supply information. Explanation.
- transformer block: A repeated sequence-model layer combining attention, per-position feature transformation, normalization, and residual additions. Explanation.
- true negative: An actual negative example predicted negative, abbreviated TN. Explanation.
- true positive: An actual positive example predicted positive, abbreviated TP. Explanation.
- true-positive rate (TPR): The fraction of actual positives predicted positive, identical to recall. Explanation.
- truncated BPTT: Recurrent training that stops derivative propagation at selected boundaries while optionally carrying state values onward. Explanation.
- truncation: Removing tokens outside a retained span, potentially discarding useful input information. Explanation.
- TTFT: Time to first token, from a declared request-start event to the first visible output at the stated measurement boundary. Explanation.
- underfitting: Failure of the fitted rule to capture useful structure even in the training task. Explanation.
- unigram tokenizer: A tokenizer whose piece probabilities score segmentations by products and whose vocabulary learning uses likelihood summed over valid segmentations. Explanation.
- [UNK:] See unknown token. Explanation.
- unknown token: A fallback vocabulary entry used when a tokenizer cannot represent an input piece, often written UNK. Explanation.
- unsupervised learning: Finding structure in data without supplied target answers for each example. Explanation.
- upstream gradient: Loss sensitivities received from later operations when differentiating a local operation. Explanation.
- validation set: Separate examples used to guide development decisions such as stopping or selecting settings. Explanation.
- value: A source-position content vector combined using normalized attention coefficients. Explanation.
- value projection: The learned map from a source representation to the content coordinates combined by attention. Explanation.
- vanishing gradient: A derivative signal that shrinks along a long chain of operations, limiting learning from distant dependencies. Explanation.
- variance: Variation in predictions caused by changes in the training sample. Explanation.
- vector-Jacobian product: A row upstream gradient multiplied by an output-by-input Jacobian to obtain input sensitivities. Explanation.
- view: A tensor sharing underlying storage with its own shape, strides, and storage offset. Explanation.
- ViLT: Vision-and-Language Transformer, jointly processing projected patches and word embeddings in one encoder. Explanation.
- vision transformer: A model applying transformer blocks to projected image patches with position information, abbreviated ViT. Explanation.
- visual grounding: Locating the image region described by text. A similarity score alone does not supply the required spatial location. Explanation.
- visual question answering: Producing an answer to a supplied question using an image, abbreviated VQA. Explanation.
- [ViT:] See vision transformer. Explanation.
- vocabulary: The finite collection of supported tokens with their assigned IDs. Explanation.
- vocabulary distribution: A probability distribution over possible next-token vocabulary entries. Explanation.
- VQ-VAE: Vector-quantized variational autoencoder, reconstructing inputs from nearest learned codebook vectors with a straight-through encoder-gradient approximation. Explanation.
- [VQA:] See visual question answering. Explanation.
- warmup: An optional schedule that gradually increases the learning rate during an initial interval. Explanation.
- waveform: A sequence of sampled signal amplitudes indexed by time. Explanation.
- weight decay: Direct subtraction of a fraction of the current parameter value during an update. Explanation.
- weight vector: One adjustable coefficient per input feature in an affine calculation. Explanation.
- Whisper: An encoder-decoder speech model using log-mel inputs and text outputs with task, language, and optional timestamp markers. Explanation.
- window: A local selection and weighting of waveform samples for an analysis frame. Explanation.
- word error rate: The number of word substitutions, deletions, and insertions in a transcript divided by its reference word count under declared normalization and alignment rules. WER is its abbreviation. Explanation.
- Word2Vec: Methods that learn static word vectors through center-to-context or context-to-center prediction tasks. Explanation.
- WordPiece: A subword method that encodes words through longest matching pieces in its saved vocabulary, using configured continuation and unknown policies. Explanation.
- zero-one loss: A penalty of zero for a correct classification decision and one for an incorrect decision. Explanation.
The reading paths locate a chapter’s prerequisites. A short glossary entry restores a term’s meaning. Its linked explanation supplies the calculation and conditions.