Appendix C — References
Identify the original papers, established reports, and official software references supporting the guide’s claims.
The sources below are numbered by their first citation in the book. The footnotes identify the supported claim and any important limits. Versioned software documentation describes an interface, not experimental evidence that a training method is better.
- Ouyang, L., et al. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744. DOI: 10.52202/068431-2011. Official proceedings. Link checked 2026-09-16.
- Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., & Finn, C. (2023). Direct Preference Optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36. Source. Link checked 2026-09-16.
- Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal Policy Optimization Algorithms. arXiv:1707.06347. Source. Link checked 2026-09-16.
- Shao, Z., et al. (2024). DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (arXiv:2402.03300). Source. Link checked 2026-09-16.
- Lambert, N., Morrison, J., Pyatkin, V., et al. (2024). Tülu 3: Pushing Frontiers in Open Language Model Post-Training (arXiv:2411.15124, v5). Source. Link checked 2026-09-16.
- Nakano, R., Hilton, J., Balaji, S., et al. (2021). WebGPT: Browser-assisted question-answering with human feedback (arXiv:2112.09332). Source. Link checked 2026-09-16.
- Lhoest, Q., et al. (2021). Datasets: A community library for natural language processing. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (pp. 175–184). Association for Computational Linguistics. DOI: 10.18653/v1/2021.emnlp-demo.21. Paper. Link checked 2026-09-16.
- Wolf, T., et al. (2020). Transformers: State-of-the-art natural language processing. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (pp. 38–45). Association for Computational Linguistics. DOI: 10.18653/v1/2020.emnlp-demos.6. Paper. Link checked 2026-09-16.
- Paszke, A., et al. (2019). PyTorch: An imperative style, high-performance deep learning library. Advances in Neural Information Processing Systems, 32. Original paper. Link checked 2026-09-16.
- Hugging Face. (2025). TRL: Transformer Reinforcement Learning (v0.22.2 documentation). Official documentation. Link checked 2026-09-16.
- Wei, J., et al. (2022). Finetuned language models are zero-shot learners. International Conference on Learning Representations. OpenReview. Link checked 2026-09-16.
- Bengio, Y., Ducharme, R., Vincent, P., & Jauvin, C. (2003). A neural probabilistic language model. Journal of Machine Learning Research, 3, 1137–1155. Source. Link checked 2026-09-16.
- Hugging Face. (2025). SFT Trainer (TRL v0.22.2 documentation). Official documentation. Link checked 2026-09-16.
- Hugging Face. (2025). Chat templates (Transformers v4.56.2 documentation). Official documentation. Link checked 2026-09-16.
- PyTorch Contributors. (2025). CrossEntropyLoss (PyTorch 2.8 documentation). Official documentation. Link checked 2026-09-16.
- Hu, E. J., et al. (2022). LoRA: Low-rank adaptation of large language models. International Conference on Learning Representations. arXiv. Link checked 2026-09-16.
- Stiennon, N., et al. (2020). Learning to summarize from human feedback. Advances in Neural Information Processing Systems, 33. Official paper. Link checked 2026-09-16.
- Bai, Y., et al. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv:2212.08073. arXiv. DOI: 10.48550/arXiv.2212.08073. Link checked 2026-09-16.
- Bradley, R. A., & Terry, M. E. (1952). Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika, 39(3–4), 324–345. Oxford Academic. https://doi.org/10.1093/biomet/39.3-4.324. Link checked 2026-09-16.
- Hugging Face. (2025). Reward Trainer (TRL v0.22.2 documentation). Official documentation. Link checked 2026-09-16.
- Gao, L., Schulman, J., & Hilton, J. (2023). Scaling laws for reward model overoptimization. Proceedings of the 40th International Conference on Machine Learning, PMLR 202, 10835-10866. Source. Link checked 2026-09-16.
- Park, R., Rafailov, R., Ermon, S., & Finn, C. (2024). Disentangling length from quality in direct preference optimization. Findings of the Association for Computational Linguistics: ACL 2024, 4998–5017. ACL Anthology. DOI: 10.18653/v1/2024.findings-acl.297. Link checked 2026-09-16.
- Hugging Face. (2025). DPO Trainer (TRL v0.22.2 documentation). Official documentation. Link checked 2026-09-16.
- Hugging Face. (2025). Accelerate (Accelerate v1.10.1 documentation). Official documentation. Link checked 2026-09-16.
- Hugging Face. (2025). Online DPO Trainer (TRL v0.22.2 documentation). Official documentation. Link checked 2026-09-16.
- Gheshlaghi Azar, M., et al. (2024). A general theoretical paradigm to understand learning from human preferences. Proceedings of the 27th International Conference on Artificial Intelligence and Statistics, PMLR 238, 4447–4455. PMLR. Link checked 2026-09-16.
- Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., & Kiela, D. (2024). Model alignment as prospect theoretic optimization. Proceedings of the 41st International Conference on Machine Learning, PMLR 235, 12634–12651. PMLR. Link checked 2026-09-16.
- Meng, Y., Xia, M., & Chen, D. (2024). SimPO: Simple preference optimization with a reference-free reward. Advances in Neural Information Processing Systems, 37. Official proceedings. https://doi.org/10.52202/079017-3946. Link checked 2026-09-16.
- Hong, J., Lee, N., & Thorne, J. (2024). ORPO: Monolithic preference optimization without reference model. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 11170–11189. ACL Anthology. DOI: 10.18653/v1/2024.emnlp-main.626. Link checked 2026-09-16.
- Yu, Q., et al. (2025). DAPO: An open-source LLM reinforcement learning system at scale. Advances in Neural Information Processing Systems, 38. DOI: 10.52202/085713-3775. Source. Link checked 2026-09-16.
- Sutton, R. S., McAllester, D. A., Singh, S. P., & Mansour, Y. (2000). Policy gradient methods for reinforcement learning with function approximation. Advances in Neural Information Processing Systems, 12, 1057–1063. Source. Conference held in 1999. Proceedings volume published in 2000. Link checked 2026-09-16.
- Hugging Face. (n.d.). Reducing Memory Usage, TRL v0.22.2 documentation. Source. Link checked 2026-09-16.
- Sheng, G., et al. (2025). HybridFlow: A flexible and efficient RLHF framework. Proceedings of the Twentieth European Conference on Computer Systems, 1279–1297. DOI: 10.1145/3689031.3696075. Source. Link checked 2026-09-16.
- Schulman, J., Moritz, P., Levine, S., Jordan, M. I., & Abbeel, P. (2016). High-dimensional continuous control using generalized advantage estimation. International Conference on Learning Representations. Source. Link checked 2026-09-16.
- Williams, R. J. (1992). Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8(3–4), 229–256. Source. Link checked 2026-09-16.
- Schulman, J., Levine, S., Abbeel, P., Jordan, M. I., & Moritz, P. (2015). Trust Region Policy Optimization. Proceedings of the 32nd International Conference on Machine Learning, PMLR 37, 1889–1897. Source. Link checked 2026-09-16.
- Hugging Face. (n.d.). PPO Trainer, TRL v0.22.2 documentation. Source. Link checked 2026-09-16.
- Hugging Face. (n.d.). GRPO Trainer, TRL v0.22.2 documentation. Source. Link checked 2026-09-16.
- Ahmadian, A., et al. (2024). Back to basics: Revisiting REINFORCE-style optimization for learning from human feedback in LLMs. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12248–12267. DOI: 10.18653/v1/2024.acl-long.662. Source. Link checked 2026-09-16.
- Uesato, J., et al. (2022). Solving Math Word Problems With Process- and Outcome-Based Feedback. arXiv:2211.14275. Source. Link checked 2026-09-16.
- DeepSeek-AI. (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (version 1). arXiv:2501.12948. Source. Link checked 2026-09-16.
- Zelikman, E., Wu, Y., Mu, J., & Goodman, N. D. (2022). STaR: Self-Taught Reasoner, bootstrapping reasoning with reasoning. Advances in Neural Information Processing Systems, 35, 15476–15488. Source. Link checked 2026-09-16.
- NVIDIA. (n.d.). Algorithms, NeMo RL documentation. Source. Link checked 2026-09-16.
- Kwon, W., et al. (2023). Efficient memory management for large language model serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles (pp. 611–626). Source. Link checked 2026-09-16.
- Souppaya, M., Morello, J., & Scarfone, K. (2017). Application Container Security Guide (NIST SP 800-190). National Institute of Standards and Technology. Source. DOI: 10.6028/NIST.SP.800-190. Link checked 2026-09-16.
- Lightman, H., et al. (2024). Let’s verify step by step. International Conference on Learning Representations. Source. Link checked 2026-09-16.
- Yang, S., Chiang, W.-L., Zheng, L., Gonzalez, J. E., & Stoica, I. (2023). Rethinking Benchmark and Contamination for Language Models with Rephrased Samples (arXiv:2311.04850). Source. Link checked 2026-09-16.
- Zheng, L., et al. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems, 36, Datasets and Benchmarks Track. DOI: 10.52202/075280-2020. Source. Link checked 2026-09-16.
- Dubois, Y., Galambosi, B., Liang, P., & Hashimoto, T. B. (2024). Length-Controlled AlpacaEval: A simple way to debias automatic evaluators. Conference on Language Modeling. Source. Link checked 2026-09-16.
- Chen, M., Tworek, J., Jun, H., et al. (2021). Evaluating Large Language Models Trained on Code (arXiv:2107.03374). Source. Link checked 2026-09-16.
- Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., & Meger, D. (2018). Deep reinforcement learning that matters. Proceedings of the AAAI Conference on Artificial Intelligence, 32(1), 3207-3214. Source. DOI: 10.1609/aaai.v32i1.11694. Link checked 2026-09-16.
- Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press. ISBN 9780262039246. Source. Link checked 2026-09-16.
- Konda, V. R., & Tsitsiklis, J. N. (2000). Actor-critic algorithms. Advances in Neural Information Processing Systems, 12 (pp. 1008–1014). MIT Press. Source. Conference held in 1999. Proceedings volume published in 2000. Link checked 2026-09-16.
- Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems, 30. Source. Link checked 2026-09-16.
- Ziegler, D. M., et al. (2019). Fine-tuning language models from human preferences. OpenAI research report, arXiv:1909.08593. Report. Link checked 2026-09-16.
- Cobbe, K., et al. (2021). Training verifiers to solve math word problems. OpenAI research report, arXiv:2110.14168. Source. Link checked 2026-09-16.