Appendix A — Course sources and further reading
The lecture decks introduce the training methods, while the summaries and notebooks connect those ideas to calculations and software. The source map below identifies the strongest course references for each course topic. The research papers and versioned software documentation provide further detail beyond the classroom examples.
The source inventory covers 11 files and 662 individual units: 368 PDF pages, 176 DOCX paragraphs, and 118 notebook cells. These include 34 cells from the updated PPO notebook. Each unit has a recorded decision: 29 are directly covered, 95 repeat other units, 519 are incorporated into broader explanations, and 19 are intentionally omitted, such as empty cells and material outside the guide’s scope. These are source-accounting categories, not measures of how much a topic is taught.
The accompanying source records allow a reader or editor to check these connections.
Equations:
ledgers/equation-registry.yamlrecords the 25 principal equations and their source locations. Worked arithmetic is explained next to the relevant equation.Visuals:
ledgers/visual-ledger.yamlrecords each figure’s role, source or generation history, and review decision.Individual source units:
ledgers/source-coverage.csvrecords the pages, paragraphs, and notebook cells, with the reason for each coverage decision.
The course-source table connects six topics to their main course sources. Notebook cell references are zero-based and refer to the indicated notebook version. The individual source record disambiguates overlapping files.
| Guide material | Source units used | How used |
|---|---|---|
| Foundations and SFT | Session 1 pp. 16, 21. Week 1 summary paras. 10-25 | RL mapping, SFT limitation, state trace, and objective |
| Software responsibilities | SFT/DPO notebook cells 19, 21, 23, 33, 37, 39, 43, 45, 50; PPO notebook cells 4, 6, 10, 12, 14, 16, 17, 19 | Data, model, adapter, trainer, tensor, evaluation, and verifier responsibilities with explicit project checks |
| Reward models and DPO | Session 2 pp. 29, 31; Week 1 summary; SFT/DPO cells 33, 37, 39, 43, 45, 50 | Pairwise loss, full RLHF objective, DPO workflow, LoRA, and evaluation |
| PPO | Session 2 pp. 47, 61, 66-67; Week 2 summary; PPO cells 10, 12, 14, 16, 17, 19, 30; updated PPO notebook cells 2, 4, 16 | Policy ratio, clipping, GAE, tensor trace, diagnostics, and exercises |
| GRPO, R1, and DAPO | Session 3 v1 pp. 17, 19-22, 30-35. Week 3 summary paras. 1-36 | Group baseline, reasoning tokens, staged pipeline, DAPO clipping, sampling, token weighting, overlong shaping, and the separate no-KL trade-off |
| RLVR and agents | Session 3 v1 pp. 36-50, especially p. 41. Week 3 summary paras. 37-48 | Verifier rules, partial rewards, tool loop, and evaluation |
The following primary sources develop the methods further.
- DPO: Direct Preference Optimization derives the policy-to-reference preference objective without a separately fitted reward model.1
- PPO and GAE: Proximal Policy Optimization Algorithms explains the clipped surrogate. Generalized Advantage Estimation develops the weighted combination of temporal-difference residuals.23
- Group-based reasoning training: DeepSeekMath introduces GRPO. The course-era DeepSeek-R1 report distinguishes R1-Zero from the staged R1 training process.45
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale explains separate clipping bounds, dynamic sampling, token-level weighting, and overlong-answer reward handling.6
- Feedback on reasoning steps: Let’s Verify Step by Step studies feedback on intermediate steps rather than only final outcomes. Its results concern the reported experiments, not a guarantee that any verifier prevents reward gaming.7
- Software: TRL v0.22.2 documentation provides a fixed reference for trainer interfaces. Installed versions may expose different arguments and defaults.8
- Preference-training extensions: Disentangling Length from Quality in DPO investigates length as a shortcut. Online DPO in TRL v0.22.2 shows how freshly generated and judged pairs enter training.910
- Training resources: TRL’s memory guide explains storage-saving options. HybridFlow connects model roles to distributed scheduling and generation.1112
- Evaluation limits: HumanEval develops multiple-sample code evaluation. MT-Bench and Chatbot Arena examine LLM judges. Rephrased benchmark contamination explains why exact text matching can miss prior test exposure.131415
A.1 Reading routes for the methods
The original papers answer different questions from the practical notebooks. Some establish the mathematical update, some test a training recipe, and others investigate failure modes. The groups below connect those questions to the relevant method sections, so a reader can follow the derivation or experiment that needs more detail.
A.1.1 Foundations and early language-model cases
The methods in Chapters 1 through 3 use older results about learning from examples, preferences, and sampled rewards. These readings connect those mathematical ideas to early language-model applications.
- A broad RL reference: Sutton and Barto, Reinforcement Learning: An Introduction, second edition develops states, actions, returns, value functions, temporal-difference learning, and policy gradients. It supplies the wider mathematical context for the token and action examples in Chapter 3.16
- Sampled policy gradients: Williams (1992) introduces the REINFORCE family used as background for the score-function explanation in Section 3.4. Sutton, McAllester, Singh, and Mansour give the function-approximation policy-gradient result. The proceedings are normally cited as 2000, although the archive labels the conference NIPS 1999.1718
- Pairwise preferences: Bradley and Terry (1952) develop the Bradley-Terry comparison model used in Chapter 2. This citation identifies the model within the wider history of paired comparisons.19
- Actor-critic learning: Konda and Tsitsiklis analyze actor-critic methods in which the policy and value estimates evolve on different time scales. The paper’s two-time-scale analysis connects directly to the separate policy and value roles in Section 3.1 and the credit estimates in Section 3.3.20
- Trust regions and adapters: TRPO supplies the trust-region motivation before PPO clipping. It does not give clipping a hard trust-region guarantee. LoRA is the source for the low-rank adapter method used in Section 2.3.2122
- Human-feedback lineage: Christiano et al. learn rewards from human trajectory comparisons. Ziegler et al. apply that direction to language tasks, and Stiennon et al. provide a reward-model-plus-RL summarization case. InstructGPT provides a concrete recipe combining demonstrations, rankings, and RL. That sequence is not compulsory for every project.23242526
- Instruction and reasoning supervision: FLAN is a useful instruction-tuning case for Section 1.3. Uesato et al. compare process and outcome feedback, while Cobbe et al. show a verifier-ranking case. Read these with the guide’s verifier limits in Chapter 4 and Chapter 5.272829
- Overoptimization: Gao, Schulman, and Hilton measure proxy-reward overoptimization in a synthetic setup. It motivates the cautions in Section 2.1 and Section 3.6, but its fitted behavior is not a universal law.30
A.1.2 Modern extensions and engineering cases
The basic updates leave practical questions about feedback quality, sampling cost, and reliable evaluation. These readings investigate those questions in the settings discussed in Chapters 2 through 6.
- Reasoning data and RLVR terminology: STaR is a rationale-bootstrapping method that retains successful generated rationales. It is not the origin of generic rejection sampling. Tulu 3 describes its RLVR stage as a novel method in that open recipe. It is a useful terminology and recipe source for Section 5.1, not proof that no earlier similar method existed.3132
- Critic-free online RL: Ahmadian et al. revisit REINFORCE-style learning from feedback, including RLOO. Its leave-one-out baseline still depends on multiple completions for a prompt, so it does not eliminate grouped sampling.33
- Offline preference options: IPO, KTO, SimPO, and ORPO illustrate different choices about paired versus binary feedback, reference models, and objectives. Their paper results are setup-specific rather than a ranking of methods for every dataset. See Section 2.4 before choosing among them.34353637
- AI feedback and systems: Constitutional AI is a specific route from human-written principles to AI critiques, preferences, and RL. It does not remove judge bias or reward hacking. For engineering choices, PagedAttention/vLLM explains KV-cache sharing. HybridFlow describes a distributed RLHF framework. Cache savings, throughput, and framework behavior depend on the stated implementation and version.383940
- Evaluation discipline: Length-Controlled AlpacaEval adjusts an LLM-judge comparison for output-length differences. It complements, but does not replace, task accuracy and human review in Section 6.2. Henderson et al. is a general RL reminder to report uncertainty and conditions, not a rule that every post-training result needs one fixed number of seeds.4142
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., & Finn, C. (2023). Direct Preference Optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36. Source. Link checked 2026-09-16. Abstract. Section 4. The paper’s experiments use selected tasks and models up to 6B parameters.↩︎
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal Policy Optimization Algorithms. arXiv:1707.06347. Source. Link checked 2026-09-16. Abstract. Sections 2-3. Algorithm 1. The original experiments are not LLM post-training experiments.↩︎
Schulman, J., Moritz, P., Levine, S., Jordan, M. I., & Abbeel, P. (2016). High-dimensional continuous control using generalized advantage estimation. International Conference on Learning Representations. Source. Link checked 2026-09-16. Sections 3-4. Equations 11 and 16. The experiments are continuous-control tasks, not token-level LLM training.↩︎
Shao, Z., et al. (2024). DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (arXiv:2402.03300). Source. Link checked 2026-09-16. Abstract. Section 4. The empirical evidence is for the reported 7B mathematics models.↩︎
DeepSeek-AI. (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (version 1). arXiv:2501.12948. Source. Link checked 2026-09-16. Section 2.2. Producing-organization report, not an independent replication.↩︎
Yu, Q., et al. (2025). DAPO: An open-source LLM reinforcement learning system at scale. Advances in Neural Information Processing Systems, 38. DOI: 10.52202/085713-3775. Source. Link checked 2026-09-16. Sections 3.1-3.4. Reported evidence is tied to the paper’s Qwen2.5-32B mathematics setup.↩︎
Lightman, H., et al. (2024). Let’s verify step by step. International Conference on Learning Representations. Source. Link checked 2026-09-16. Sections 2.5-2.6 and 4.1. Step-level human labels and best-of-N selection on MATH, not agent-tool checkpoints or universal effectiveness.↩︎
Hugging Face. (n.d.). GRPO Trainer, TRL v0.22.2 documentation. Source. Link checked 2026-09-16. GRPOConfig and vLLM sections. Arguments and defaults are release-specific.↩︎
Park, R., Rafailov, R., Ermon, S., & Finn, C. (2024). Disentangling length from quality in direct preference optimization. Findings of the Association for Computational Linguistics: ACL 2024, 4998–5017. ACL Anthology. DOI: 10.18653/v1/2024.findings-acl.297. Link checked 2026-09-16. Sections 3–5. Observed on selected summarization and dialogue datasets. Length is not mechanically rewarded by every log-ratio.↩︎
Hugging Face. (2025). Online DPO Trainer (TRL v0.22.2 documentation). Official documentation. Link checked 2026-09-16. Overview. Quick start. The documented model-judge setup is one form of online preference feedback.↩︎
Hugging Face. (n.d.). Reducing Memory Usage, TRL v0.22.2 documentation. Source. Link checked 2026-09-16. Truncation, padding-free, activation-offloading, and online-generation sections. Documentation does not define universal PPO memory use or measured performance.↩︎
Sheng, G., et al. (2025). HybridFlow: A flexible and efficient RLHF framework. Proceedings of the Twentieth European Conference on Computer Systems, 1279–1297. DOI: 10.1145/3689031.3696075. Source. Link checked 2026-09-16. Sections 3-5. Reported speedups are hardware-, workload-, revision-, and baseline-specific.↩︎
Chen, M., Tworek, J., Jun, H., et al. (2021). Evaluating Large Language Models Trained on Code (arXiv:2107.03374). Source. Link checked 2026-09-16. Section 2.1, Equation 1. HumanEval program synthesis with unit-test correctness and the paper’s sampling procedure. Not model confidence or automatic selection without a test oracle.↩︎
Zheng, L., et al. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems, 36, Datasets and Benchmarks Track. DOI: 10.52202/075280-2020. Source. Link checked 2026-09-16. Sections 3.3-3.4 and 4. The measured effects and agreement rates are tied to the tested 2023 models, prompts, and answer sets.↩︎
Yang, S., Chiang, W.-L., Zheng, L., Gonzalez, J. E., & Stoica, I. (2023). Rethinking Benchmark and Contamination for Language Models with Rephrased Samples (arXiv:2311.04850). Source. Link checked 2026-09-16. Abstract and Sections 2-4. Selected models, datasets, transformations, and the authors’ detector. Not proof of a clean corpus when no overlap is found.↩︎
Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press. ISBN 9780262039246. Source. Link checked 2026-09-16. Chapters 3, 6 and 13, as listed in the publisher Table of Contents. General RL background, not evidence for a particular LLM training recipe.↩︎
Williams, R. J. (1992). Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8(3–4), 229–256. Source. Link checked 2026-09-16. Sections 2-3. Historical algorithm source, not evidence about modern LLM training systems.↩︎
Sutton, R. S., McAllester, D. A., Singh, S. P., & Mansour, Y. (2000). Policy gradient methods for reinforcement learning with function approximation. Advances in Neural Information Processing Systems, 12, 1057–1063. Source. Conference held in 1999. Proceedings volume published in 2000. Link checked 2026-09-16. Abstract. Actor-critic section. General RL theory, not an LLM implementation specification.↩︎
Bradley, R. A., & Terry, M. E. (1952). Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika, 39(3–4), 324–345. Oxford Academic. https://doi.org/10.1093/biomet/39.3-4.324. Link checked 2026-09-16. Biometrika 39, pp. 324–345. The original model does not establish reward calibration, truth, or safety for language models.↩︎
Konda, V. R., & Tsitsiklis, J. N. (2000). Actor-critic algorithms. Advances in Neural Information Processing Systems, 12 (pp. 1008–1014). MIT Press. Source. Conference held in 1999. Proceedings volume published in 2000. Link checked 2026-09-16. Abstract and actor-critic formulation. The analysis uses its stated approximation and time-scale assumptions, not modern LLM PPO.↩︎
Schulman, J., Levine, S., Abbeel, P., Jordan, M. I., & Moritz, P. (2015). Trust Region Policy Optimization. Proceedings of the 32nd International Conference on Machine Learning, PMLR 37, 1889–1897. Source. Link checked 2026-09-16. Sections 3-4. The practical TRPO algorithm itself uses approximations to the theoretical scheme.↩︎
Hu, E. J., et al. (2022). LoRA: Low-rank adaptation of large language models. International Conference on Learning Representations. arXiv. Link checked 2026-09-16. Abstract. Section 4. LoRA specifies trainable parameters, not the training evidence or objective.↩︎
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems, 30. Source. Link checked 2026-09-16. Abstract and reported human trajectory-comparison experiments. The tasks are Atari and simulated locomotion, not language modeling.↩︎
Ziegler, D. M., et al. (2019). Fine-tuning language models from human preferences. OpenAI research report, arXiv:1909.08593. Report. Link checked 2026-09-16. Abstract and reported continuation and summarization cases. The results do not establish general conversational-model reliability.↩︎
Stiennon, N., et al. (2020). Learning to summarize from human feedback. Advances in Neural Information Processing Systems, 33. Official paper. Link checked 2026-09-16. Section 3.4. The experiments concern TL;DR summarization with 1.3B and 6.7B models.↩︎
Ouyang, L., et al. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744. DOI: 10.52202/068431-2011. Official proceedings. Link checked 2026-09-16. Figure 2. Sections 3.1 and 3.4. One InstructGPT pipeline does not make the stages compulsory for every project.↩︎
Wei, J., et al. (2022). Finetuned language models are zero-shot learners. International Conference on Learning Representations. OpenReview. Link checked 2026-09-16. Abstract. Sections 2–3. Experiments use a 137B model and more than 60 selected NLP datasets.↩︎
Uesato, J., et al. (2022). Solving Math Word Problems With Process- and Outcome-Based Feedback. arXiv:2211.14275. Source. Link checked 2026-09-16. Sections 2-3. The empirical comparison is task-specific and does not prove process supervision is always better.↩︎
Cobbe, K., et al. (2021). Training verifiers to solve math word problems. OpenAI research report, arXiv:2110.14168. Source. Link checked 2026-09-16. Abstract and GSM8K verifier evaluation. A learned verifier ranks candidate solutions. This is not verifier-based policy optimization.↩︎
Gao, L., Schulman, J., & Hilton, J. (2023). Scaling laws for reward model overoptimization. Proceedings of the 40th International Conference on Machine Learning, PMLR 202, 10835-10866. Source. Link checked 2026-09-16. Abstract. Sections 1–3. Uses a synthetic gold reward model rather than direct human preference judgments.↩︎
Zelikman, E., Wu, Y., Mu, J., & Goodman, N. D. (2022). STaR: Self-Taught Reasoner, bootstrapping reasoning with reasoning. Advances in Neural Information Processing Systems, 35, 15476–15488. Source. Link checked 2026-09-16. Abstract. Section 2. Related bootstrapping method, not the origin of rejection sampling generally.↩︎
Lambert, N., Morrison, J., Pyatkin, V., et al. (2024). Tülu 3: Pushing Frontiers in Open Language Model Post-Training (arXiv:2411.15124, v5). Source. Link checked 2026-09-16. Sections 6 and 6.2, especially Figure 18 and the opening of Section 6. One model family, recipe, and task suite. The report does not prove that arbitrary checks represent deployment goals.↩︎
Ahmadian, A., et al. (2024). Back to basics: Revisiting REINFORCE-style optimization for learning from human feedback in LLMs. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12248–12267. DOI: 10.18653/v1/2024.acl-long.662. Source. Link checked 2026-09-16. Section 3. Task, model, and reward comparisons are paper-specific. RLOO is not GRPO with only a changed denominator.↩︎
Gheshlaghi Azar, M., et al. (2024). A general theoretical paradigm to understand learning from human preferences. Proceedings of the 27th International Conference on Artificial Intelligence and Statistics, PMLR 238, 4447–4455. PMLR. Link checked 2026-09-16. Sections 4–5. Empirical demonstrations are illustrative and do not establish universal superiority.↩︎
Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., & Kiela, D. (2024). Model alignment as prospect theoretic optimization. Proceedings of the 41st International Conference on Machine Learning, PMLR 235, 12634–12651. PMLR. Link checked 2026-09-16. Sections 3–4. Results cover selected model families and benchmarks from 1B to 30B parameters.↩︎
Meng, Y., Xia, M., & Chen, D. (2024). SimPO: Simple preference optimization with a reference-free reward. Advances in Neural Information Processing Systems, 37. Official proceedings. https://doi.org/10.52202/079017-3946. Link checked 2026-09-16. Sections 3–4. Removing a reference pass does not determine total system cost across all configurations.↩︎
Hong, J., Lee, N., & Thorne, J. (2024). ORPO: Monolithic preference optimization without reference model. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 11170–11189. ACL Anthology. DOI: 10.18653/v1/2024.emnlp-main.626. Link checked 2026-09-16. Sections 3–4. Experiments cover selected 125M–7B models and still rely on paired evidence.↩︎
Bai, Y., et al. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv:2212.08073. arXiv. DOI: 10.48550/arXiv.2212.08073. Link checked 2026-09-16. Abstract. Sections 2–4. A preprint about one harmlessness setup. AI labels are not independent human or factual ground truth.↩︎
Kwon, W., et al. (2023). Efficient memory management for large language model serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles (pp. 611–626). Source. Link checked 2026-09-16. Sections 3-4. Serving evidence does not establish exact training-rollout savings or policy correctness.↩︎
Sheng, G., et al. (2025). HybridFlow: A flexible and efficient RLHF framework. Proceedings of the Twentieth European Conference on Computer Systems, 1279–1297. DOI: 10.1145/3689031.3696075. Source. Link checked 2026-09-16. Sections 3-5. Reported speedups are hardware-, workload-, revision-, and baseline-specific.↩︎
Dubois, Y., Galambosi, B., Liang, P., & Hashimoto, T. B. (2024). Length-Controlled AlpacaEval: A simple way to debias automatic evaluators. Conference on Language Modeling. Source. Link checked 2026-09-16. Abstract and Sections 2-3. A statistical adjustment for AlpacaEval-style comparisons, not a factual-correctness test or complete debiasing method.↩︎
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., & Meger, D. (2018). Deep reinforcement learning that matters. Proceedings of the AAAI Conference on Artificial Intelligence, 32(1), 3207-3214. Source. DOI: 10.1609/aaai.v32i1.11694. Link checked 2026-09-16. Sections 3-5. Deep-RL control benchmarks, not LLM fine-tuning. No universal rule for seed count or uncertainty method.↩︎