References
These sources support the claims cited beside the explanations. Entries follow their first use in the book. The footnotes identify the relevant sections and any limits of the cited evidence.
- Sennrich, R., Haddow, B., & Birch, A. (2016). Neural machine translation of rare words with subword units. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1715–1725.
OpenAI. (2019). GPT-2 byte pair encoding implementation. GPT-2 source code.
Hugging Face. (n.d.). Tokenization algorithms. Transformers documentation. Retrieved September 25, 2026.
Hugging Face. (n.d.). WordPiece tokenization. LLM Course, Chapter 6. Retrieved September 24, 2026.
- Kudo, T. (2018). Subword regularization: Improving neural network translation models with multiple subword candidates. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 66–75.
Google. (n.d.). SentencePiece: README. Retrieved September 24, 2026.
Scikit-learn developers. (n.d.). Feature extraction: TF-IDF term weighting. Scikit-learn documentation. Retrieved September 24, 2026.
Holtzman, A., Buys, J., Du, L., Forbes, M., & Choi, Y. (2020). The Curious Case of Neural Text Degeneration. International Conference on Learning Representations.
Hugging Face. (n.d.). Generation. Transformers documentation. Retrieved September 24, 2026.
Jurafsky, D., & Martin, J. H. (2026). Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models. Third edition draft, released August 19, 2026.
PyTorch Contributors. (2026). CrossEntropyLoss. PyTorch 2.12 documentation.
Hugging Face. (2026). Transformers 5.10.2: Causal language-model loss. Version 5.10.2.
Hugging Face. (2026). Transformers 5.10.2: GPT-2 forward. Version 5.10.2.
Hastie, T., Tibshirani, R., & Friedman, J. (2009). The elements of statistical learning: Data mining, inference, and prediction. Second edition, Springer.
Scikit-learn developers. (2026). Cross-validation: evaluating estimator performance. Scikit-learn 1.9.1 documentation.
Scikit-learn developers. (2026). roc_auc_score. Scikit-learn 1.9.1 documentation.
Scikit-learn developers. (2026). Probability calibration. Scikit-learn 1.9.1 documentation.
- Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002). BLEU: a Method for Automatic Evaluation of Machine Translation. Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 311–318.
Lin, C.-Y. (2004). ROUGE: A Package for Automatic Evaluation of Summaries. Text Summarization Branches Out, 74–81.
Chen, M., et al. (2021). Evaluating Large Language Models Trained on Code. arXiv:2107.03374.
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., & Steinhardt, J. (2021). Measuring Massive Multitask Language Understanding. International Conference on Learning Representations.
- Lin, S., Hilton, J., & Evans, O. (2022). TruthfulQA: Measuring How Models Mimic Human Falsehoods. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3214–3252.
- Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., & Choi, Y. (2019). HellaSwag: Can a Machine Really Finish Your Sentence? Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 4791–4800.
Phan, L., et al. (2025). Humanity’s Last Exam. arXiv:2501.14249, version 1.
ARC Prize. (n.d.). ARC-AGI-1 & ARC-AGI-2 Guide. Retrieved September 24, 2026.
PyTorch Contributors. (2026). SGD. PyTorch 2.12 documentation.
Kingma, D. P., & Ba, J. (2015). Adam: A Method for Stochastic Optimization. International Conference on Learning Representations.
Loshchilov, I., & Hutter, F. (2019). Decoupled Weight Decay Regularization. International Conference on Learning Representations.
PyTorch Contributors. (2026). AdamW. PyTorch 2.12 documentation.
Loshchilov, I., & Hutter, F. (2017). SGDR: Stochastic Gradient Descent with Warm Restarts. International Conference on Learning Representations.
Parikh, N., & Boyd, S. (2014). Proximal algorithms. Foundations and Trends in Optimization, 1(3), 127–239.
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. (2014). Dropout: a simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(56), 1929–1958.
PyTorch Contributors. (2026). Dropout. PyTorch 2.12 documentation.
PyTorch Contributors. (2026). clip_grad_norm_. PyTorch 2.12 documentation.
Bergstra, J., & Bengio, Y. (2012). Random search for hyper-parameter optimization. Journal of Machine Learning Research, 13, 281–305.
Snoek, J., Larochelle, H., & Adams, R. P. (2012). Practical Bayesian optimization of machine learning algorithms. Advances in Neural Information Processing Systems, 25.
PyTorch Contributors. (2026). Autograd mechanics. PyTorch 2.12 documentation.
Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323, 533–536.
Hendrycks, D., & Gimpel, K. (2016). Gaussian Error Linear Units (GELUs). arXiv:1606.08415.
Elfwing, S., Uchibe, E., & Doya, K. (2018). Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning. Neural Networks, 107, 3–11.
Glorot, X., & Bengio, Y. (2010). Understanding the difficulty of training deep feedforward neural networks. Proceedings of Machine Learning Research, 9, 249–256.
PyTorch Contributors. (2026). Reproducibility. PyTorch 2.12 documentation.
PyTorch Contributors. (2026). Embedding. PyTorch 2.12 documentation.
Harris, Z. S. (1954). Distributional Structure. WORD, 10(2–3), 146–162.
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., & Dean, J. (2013). Distributed representations of words and phrases and their compositionality. Advances in Neural Information Processing Systems, 26.
Gensim Contributors. (2025). Word2Vec implementation. Gensim 4.4.0 source.
Gensim Contributors. (2025). Keyed vector queries. Gensim 4.4.0 source,
get_mean_vectorandmost_similar.Mikolov, T., Le, Q. V., & Sutskever, I. (2013). Exploiting similarities among languages for machine translation. arXiv:1309.4168.
van der Maaten, L., & Hinton, G. (2008). Visualizing data using t-SNE. Journal of Machine Learning Research, 9, 2579–2605.
Eckart, C., & Young, G. (1936). The approximation of one matrix by another of lower rank. Psychometrika, 1(3), 211–218.
Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735–1780.
Cho, K., van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., & Bengio, Y. (2014). Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), 1724–1734.
PyTorch Contributors. (2026). LSTM. PyTorch 2.12 documentation.
Vaswani, A., et al. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.
PyTorch Contributors. (2026). Scaled dot-product attention. PyTorch 2.12 documentation.
PyTorch Contributors. (2026). MultiheadAttention. PyTorch 2.12 documentation.
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., & Ré, C. (2022). FlashAttention: Fast and memory-efficient exact attention with IO-awareness. Advances in Neural Information Processing Systems, 35.
Ba, J. L., Kiros, J. R., & Hinton, G. E. (2016). Layer normalization. arXiv:1607.06450.
Zhang, B., & Sennrich, R. (2019). Root mean square layer normalization. Advances in Neural Information Processing Systems, 32.
Shazeer, N. (2020). GLU variants improve transformer. arXiv:2002.05202, equations 5–6.
Press, O., Smith, N. A., & Lewis, M. (2022). Train short, test long: Attention with linear biases enables input length extrapolation. International Conference on Learning Representations.
Su, J., et al. (2024). RoFormer: Enhanced transformer with rotary position embedding. Neurocomputing, 568, 127063.
- Shazeer, N., et al. (2017). Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. International Conference on Learning Representations.
Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training. OpenAI.
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT, 4171–4186.
Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the knowledge in a neural network. arXiv:1503.02531.
Nie, S., et al. (2025). Large language diffusion models. arXiv:2502.09992, version 1.
Kaplan, J., et al. (2020). Scaling laws for neural language models. arXiv:2001.08361, version 1. The separate resource-limited fits and their allocation assumptions differ from the later Chinchilla formulation.
Hoffmann, J., et al. (2022). Training compute-optimal large language models. Advances in Neural Information Processing Systems, 35, 30016–30030.
Brown, T. B., et al. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33.
Hugging Face. (2025). Soft prompts. PEFT 0.17.0 documentation.
Lewis, P., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33.
Bradley, R. A., & Terry, M. E. (1952). Rank analysis of incomplete block designs. I. The method of paired comparisons. Biometrika, 39(3–4), 324–345.
Ouyang, L., et al. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744.
Williams, R. J. (1992). Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8, 229–256.
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv:1707.06347.
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., & Finn, C. (2023). Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36.
Hu, E. J., et al. (2022). LoRA: Low-rank adaptation of large language models. International Conference on Learning Representations.
Hugging Face. (2025). PEFT checkpoint format. PEFT 0.17.0 documentation.
Hugging Face. (2025). Quantization. PEFT 0.17.0 documentation.
PyTorch Contributors. (2026). Tensor attributes. PyTorch 2.12 documentation.
PyTorch Contributors. (2026). Type Info. PyTorch 2.12 documentation.
Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient finetuning of quantized LLMs. Advances in Neural Information Processing Systems, 36.
Ainslie, J., et al. (2023). GQA: Training generalized multi-query transformer models from multi-head checkpoints. Proceedings of EMNLP, 4895–4901.
vLLM Contributors. (n.d.). Benchmark CLI: Understanding the latency metrics. Retrieved September 24, 2026.
vLLM Contributors. (2025). Automatic prefix caching. vLLM 0.10.2 design documentation.
- Kwon, W., et al. (2023). Efficient memory management for large language model serving with PagedAttention. Proceedings of the 29th Symposium on Operating Systems Principles, 611–626.
vLLM Contributors. (2025). Engine arguments. vLLM 0.10.2 documentation.
PyTorch Contributors. (2026). Conv2d. PyTorch 2.12 documentation.
He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 770–778.
- Dosovitskiy, A., et al. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. International Conference on Learning Representations.
Radford, A., et al. (2021). Learning transferable visual models from natural language supervision. Proceedings of Machine Learning Research, 139, 8748–8763.
Kim, W., Son, B., & Kim, I. (2021). ViLT: Vision-and-language transformer without convolution or region supervision. Proceedings of Machine Learning Research, 139, 5583–5594.
van den Oord, A., Vinyals, O., & Kavukcuoglu, K. (2017). Neural discrete representation learning. Advances in Neural Information Processing Systems, 30.
PyTorch Contributors. (2026). STFT. PyTorch 2.12 documentation.
Radford, A., et al. (2023). Robust speech recognition via large-scale weak supervision. Proceedings of Machine Learning Research, 202, 28492–28518.
Das, N., et al. (2024). SpeechVerse: A large-scale generalizable audio language model. arXiv:2405.08295, version 2.
Borsos, Z., et al. (2023). AudioLM: A language modeling approach to audio generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31, 2523–2533.
Huang, X., Khetan, A., Cvitkovic, M., & Karnin, Z. (2020). TabTransformer: Tabular data modeling using contextual embeddings. arXiv:2012.06678.
Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29(5), 1189–1232.
PyTorch Contributors. (2026). Broadcasting semantics. PyTorch 2.12 documentation.
PyTorch Contributors. (2026). Tensor views. PyTorch 2.12 documentation.
PyTorch Contributors. (2026). Tensor.detach. PyTorch 2.12 documentation.
PyTorch Contributors. (2026). Autograd mechanics: Evaluation mode. PyTorch 2.12 documentation.
PyTorch Contributors. (2026). DistributedDataParallel. PyTorch 2.12 documentation.
PyTorch Contributors. (2026). Automatic mixed precision examples. PyTorch 2.12 documentation.
fancyzhx. (n.d.). AG News dataset card. Repository
fancyzhx/ag_news, revisioneb185aade064a813bc0b7f42de02595523103ca4.NLTK Project. (2025). Language-model preprocessing. NLTK 3.9.2 source,
padded_everygram_pipeline.NLTK Project. (2025). Language-model API. NLTK 3.9.2 source,
logscore,entropy,perplexity, andgenerate.NLTK Project. (2025). Language-model vocabulary. NLTK 3.9.2 source,
Vocabularyandunk_cutoff.NLTK Project. (2025). Treebank tokenization and detokenization. NLTK 3.9.2 source,
TreebankWordDetokenizer.