Notation

Several equations reuse symbols for examples, positions, dimensions, and parameters. This reference collects recurring meanings so a reader can resolve a symbol or shape and return to its calculation. The linked sections explain the mechanisms and provide examples. A chapter’s local definitions take precedence over shorthand in this page.

For a corpus, \(C\) collects \(N\) documents \(d_i\). Corpus construction and counting appear in §3.1:

\[ C = \{d_1,\ldots,d_N\} \tag{1}\]

The tokenization process in §2.1 maps text to an ordered sequence of pieces. This reference writes that map as \(\operatorname{tok}\), with \(T\) pieces \(s_i\). Probability equations reserve \(\tau\) for temperature and identify that meaning beside the formula:

\[ \operatorname{tok}(\text{text}) = (s_1,\ldots,s_T) \tag{2}\]

Vocabulary lookup assigns each piece an integer ID \(x_i\). Here \(V\) is the vocabulary set and \(|V|\) its size:

\[ x_i = \operatorname{id}(s_i),\quad x \in \{0,\ldots,\lvert V\rvert-1\}^{T} \tag{3}\]

A rectangular batch of \(B\) token sequences has \(T\) stored positions in every row. The integer-array notation below gives its shape. Actual entries must also be valid vocabulary IDs. §2.2 explains how padding creates this layout and records which positions are real:

\[ X \in \mathbb{Z}^{B\times T} \tag{4}\]

Symbols and dimensions

The same letter can denote different objects in different equations. Its local dimensions identify which meaning applies. Each row below links to the relevant explanation rather than substituting a formula list for that teaching.

Table 1: Recurring quantities and their dimensions.
Symbol Meaning and dimensions Explanation
\(B,T,V,D\) Batch size, stored sequence length, active vocabulary size, carried representation width Token batches
\(\mathcal D,N_{\mathrm{ex}}\) Dataset and its number of examples Training pairs
\(F,C,D_{\mathrm{in}},D_{\mathrm{out}},D_{\mathrm{ff}}\) Feature count, class count, layer input and output widths, and feed-forward intermediate width Linear shapes, transformer blocks
\(X\) Feature rows \([B,F]\) or integer token IDs \([B,T]\), as stated locally Batch scores
\(E\) Embedding table \([V,D]\) Embedding lookup
\(\theta,W,b\) All model parameters, a weight matrix, and a bias Linear models
\(z,p,y\) Logits, probabilities, and observed targets Probability outputs
\(L,\widehat R\) Per-example loss and its average over an observed sample Loss reduction
\(\sigma,\phi\) Sigmoid and a locally specified activation function Sigmoid, activations
\(\nabla_\theta L,\eta\) Parameter gradient, with the dimensions of \(\theta\), and learning rate Gradients, updates
\(Q,K,V\) Query, key, and value arrays, commonly \([B,H,T,d_h]\) Attention
\(H,d_h,d_k\) Number of heads, per-head width, query/key width used for score scaling Multiple heads
\(m,v\) First and second optimizer moments, with parameter dimensions Adam
\(r,\alpha_{\mathrm{LoRA}}\) LoRA rank and scale numerator, giving scale \(\alpha_{\mathrm{LoRA}}/r\) LoRA
\(N_{\mathrm{param}},N_{\mathrm{tok}}\) Parameter count and training-token count in scaling calculations Compute allocation

In Chapter 2, \(V\) is the vocabulary set and \(|V|\) its size. Some later chapters use \(V\) directly for that size. In attention, \(V\) is an array of value vectors. Likewise, \(L\) can denote scalar loss or a layer count in a cache estimate. These are local uses of a letter, not interchangeable quantities. Natural logarithms give losses in nats and pair with ordinary exponentiation for perplexity. Base-two logarithms give bits and require base-two exponentiation. §8.5 relates that base choice to the calculation.

The notation \(\mathbb{E}_{z\sim P}[g(z)]\) means the average value of \(g(z)\) when \(z\) is drawn from distribution \(P\). For discrete outcomes, multiply each value by its probability and add the products. If \(g\) is 2 with probability 0.25 and 6 with probability 0.75, its expectation is \(0.25(2)+0.75(6)=5\). §14.1 uses this notation for sampled training examples, and §20.3 uses it for responses drawn from a policy.

Masks and matrix orientation

A token value cannot identify an added padding position when the same value also marks a real end of sequence. Separate control arrays specify which positions each operation may use. The table distinguishes the operations that use them, their values, and links to their explanations.

Table 2: Masks exclude different positions from attention, pooling, or loss calculations.
Object Values and dimensions Consumer and effect
Tokenizer attention_mask \(M^{\mathrm{valid}}\in\{0,1\}^{B\times T}\), with entry \(m_{b,t}=1\) for a retained token and \(0\) for inserted padding Supplies padding visibility to the model. Padding validity
Additive attention mask \(M^{\mathrm{attn}}\), with \(0\) for permitted positions and \(-\infty\) for blocked positions, broadcastable to scores Added before softmax, making blocked attention weights zero. Attention visibility
Loss labels with ignore_index Target IDs with a sentinel such as \(-100\), \([B,T]\); \(N_{\mathrm{tgt}}\) counts included targets Excludes positions from both summed loss and the supervised-position count. Loss exclusion
Pooling mask \(1\) for included vectors and \(0\) otherwise, \([B,T]\) Excludes padding vectors and sets the mean’s denominator. Pooling

Boolean conventions depend on the receiving API: True may mean permitted or blocked. A tokenizer’s validity array therefore cannot be passed to every mask argument without checking its contract. Loss exclusion and pooling also require their own included-position counts. Attention requires a defined policy when no position remains visible. The linked chapters explain these different operations.

For a row-vector batch, mathematical weights \(W\) have shape \([D_{\mathrm{in}},D_{\mathrm{out}}]\). PyTorch stores a linear weight \(A\) with shape \([D_{\mathrm{out}},D_{\mathrm{in}}]\). Thus \(W=A^{\mathsf T}\) and the same output can be written \(XW+b\) or \(XA^{\mathsf T}+b\). One input written as a column vector instead uses \(y=A x+b\), with the stored matrix’s output-by-input shape. Chapters 12, 15, 17, and 20 use that column-vector convention in some equations. Read each equation’s declared shapes before combining it with a row-batch formula. §5.3 gives the row-batch multiplication and Chapter 23 develops storage details.

The word decoding also depends on the object being decoded. Converting token IDs back to text is tokenizer decoding, whereas generation decoding selects the next ID from model scores. A transformer decoder computes representations under a specified visibility rule, after which generation chooses a token from its scores. For audio, a decoder can map stored sound features toward a waveform. A model checkpoint is a saved model state, while a chapter checkpoint is a question for the reader.