Library Map

The examples in this book use Python libraries to prepare data, calculate model outputs, train models, and evaluate predictions. A library’s application programming interface (API) exposes operations through functions, classes, and their expected inputs and outputs. The links below connect each operation with its mathematical explanation and code. The original course notebooks provide broader exercises, with their dataset and model requirements described in Chapter 24.

The table provides the shortest route from a mathematical role to the operation used in the examples. The paragraphs that follow add the surrounding data, evaluation, and notebook context.

Table 1: API roles link to complete explanations and local code.
Role Operation Explanations and code
Token row selection nn.Embedding Lookup and gradients
Affine class scores nn.Linear Supplied-vector head
Aligned target loss F.cross_entropy Ignored-target mean
Derivative accumulation loss.backward() Fresh graphs
Parameter updates and gradient clearing step(), zero_grad() Running-batch SGD

Tabular data and tokenization. pandas.DataFrame keeps text-label pairs together in §1.1, with runnable preparation code. transformers.AutoTokenizer turns text into IDs and a padding mask in §2.2, with batch code requiring cached or downloaded tokenizer files. Chapter 3 develops tokenizer algorithms. Its WordPiece score calculation is not a complete tokenizer trainer. The original LLM_Architectures_hometask_2_Tel_Aviv.ipynb supplies BPE and WordPiece practice associated with §24.2.

Fixed document features. scikit-learn vectorizers fit vocabulary and document-frequency information on training text, then reuse it for new rows. §4.2 distinguishes count and TF-IDF conventions and provides the executable vectorizer example. Its returned matrices use SciPy sparse storage. The original LLM_Architectures,_hometask_1.ipynb uses bag-of-words features for the classification exercise.

Arrays and affine layers. PyTorch tensors hold the token batch, class-score matrix, and three-row affine calculation. The last two use explicit multiplication. torch.nn.Linear is explained as an affine layer in §5.3. Its first code use is in the runnable classifier head in §12.3. That calculation receives an existing representation. Chapter 23 develops array storage, numerical types, strides, and devices. The NumPy library provides numerical arrays for CPU calculations.

Probabilities, loss, and updates. §6.2 contains sigmoid and softmax code. torch.optim supplies the update procedures explained in Chapter 10. The running-batch optimizer example applies SGD in §11.6. §11.6 provides the complete hand calculation and the random-batch forward, backward, and update code. The masked target check in §8.5 verifies alignment and the included-position mean. The original logistic-regression homework adds repeated batches, validation, and optimizer comparisons to these local checks.

Evaluation. §9.2 explains confusion counts and classification metrics. The metric code uses sklearn.metrics in §9.3, after ranking and calibration are explained. The text-classification exercise in §24.2 requires held-out results and error inspection in addition to these calculations.

Embeddings and count models. Gensim trains the Word2Vec example in §14.1. The original second homework compares learned vectors and downstream classifiers. nltk connects the count estimator in §3.4 to the tutorial Ngram_Language_Model_with_NLTK_(1).ipynb and §24.6, which covers counting and generation. The course tutorial continues through generation, saving, and reloading the corpus model. Smoothing and held-out perplexity are book extensions in §24.6.

Recurrence and attention plots. torch.nn.RNN and torch.nn.LSTM are discussed in Chapter 15, with an LSTM dimension example in §15.4 and a character-model exercise. Matplotlib and Seaborn render the attention heatmap code in §16.2. The original From_Finetuning_to_Attention_Inside_LLMs.ipynb supplies the controlled attention comparison in §24.3.

Generation, training interfaces, and adapters. GenerationConfig and model.generate are interface references in Chapter 7, whose code example selects one token rather than running a full generation loop. datasets supplies data-loading utilities. §20.4 explains where Trainer can coordinate a supervised training loop, although its adapter setup fragment does not run that loop. peft and bitsandbytes appear in the adapter configuration fragment in §20.4. The original fine-tuning notebook supplies the wider data and model experiment associated with §24.4.

Other modalities and serving. Chapter 22 introduces image, audio, and table representations, including patch-construction code. The original LLM_Architecture_HW4.ipynb underlies the captioning and visual-question-answering work in §24.7. vLLM, SGLang, and TensorRT-LLM appear in §21.5 as serving implementations.

The Environment and source notebooks provides notebook access, dataset and checkpoint requirements, environment versions, and evidence from experiments. The linked snippets isolate particular calculations. Reproducing a result requires the library version and relevant tokenizer or model revision as well as the listed inputs. A fragment also requires its surrounding objects, and a complete experiment needs the data and evaluation procedure specified in the lab. The Glossary supplies short definitions, and the reading paths locate the mathematical prerequisites.