Part II: Learning from new attempts
Saved preference pairs need not contain the answers a changing model now produces. When a scorer can evaluate new attempts, training can generate answers, score them, and use the results for another update. This creates two practical questions: how much credit each generated token should receive, and how to avoid overreacting to one small or noisy batch.1
Chapter 3 follows one sampled answer through scoring and a PPO update, comparing the result with what the model was expected to achieve. Chapter 4 instead compares several answers to the same prompt, then considers how to use those comparisons when answers contain long reasoning steps. The comparison determines which answers training favors. It cannot make an unreliable scorer trustworthy.23
PPO and GRPO compare scored attempts in different ways. PPO uses a prediction of the result expected from an unfinished answer. GRPO uses other answers to the same prompt. Either selected comparison supports a model update and another round of attempts. The prediction and the group comparison can each be wrong, so both routes need checks on the resulting model.
Chapters in this part
- 3. PPO for scored attempts: Scored answers can improve a policy without being organized into preference pairs. Reusing those answers saves generation work, but training must account for the policy changes made since they were sampled.
- 4. GRPO and DAPO for reasoning: Training several answers for each prompt trades extra generation work for a comparison between their scores.
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal Policy Optimization Algorithms. arXiv:1707.06347. Source. Link checked 2026-09-16. Sections 2–3 and Algorithm 1. Clipping does not validate rewards or impose a hard policy-distance limit.↩︎
Shao, Z., et al. (2024). DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (arXiv:2402.03300). Source. Link checked 2026-09-16. Section 4. Grouped sampling still costs memory and computation.↩︎
Yu, Q., et al. (2025). DAPO: An open-source LLM reinforcement learning system at scale. Advances in Neural Information Processing Systems, 38. DOI: 10.52202/085713-3775. Source. Link checked 2026-09-16. Sections 3.1–3.4. The recipe is not a general guarantee of correct reasoning.↩︎