Part III: Verifying results and choosing models

The online methods depend on the scores they receive. Some results can be checked directly: a program passes tests, a mathematical answer matches a known result, or a tool interaction reaches a required state. Such checks can supply training rewards, but an incomplete test can still reward a solution that fails the real task. Tool use also requires restrictions that act before or during execution, independently of the training score.12

Chapter 5 explains how verifiers and tool environments produce rewards and observations, including the shortcuts that their checks may miss. Chapter 6 brings the methods together as alternatives chosen according to available feedback and cost. It then compares saved models on independent tasks so that higher training reward is not mistaken for success in use.

An enforced tool environment returns observations for the next model input. Recorded trajectories and separate verifier rewards feed the trainer. Saved checkpoints are compared on held-out tasks, with acceptance or retention of the previous model determined by the results.
Figure III.1: Checked interactions train models. Held-out tasks select them.

The blue observation path informs the model’s next action. Recorded interactions and the orange reward path supply training data. Tool limits are enforced during execution. The later held-out comparison tests saved models independently, so a higher training reward alone does not authorize replacing the previous model.3

Chapters in this part

  • 5. RL from tests and tool results: Automated checks make reward design part of the software that runs a task.
  • 6. Choosing and evaluating training methods: The earlier chapters connected each training method to the feedback it can use: demonstrated answers, preference pairs, scores for fresh attempts, and checks on results or tool actions. A project may have only some of that feedback. The training method must fit both the available data and the behavior the model needs to improve. The final choice is tested by comparing saved models under the same evaluation procedure.

  1. Lambert, N., Morrison, J., Pyatkin, V., et al. (2024). Tülu 3: Pushing Frontiers in Open Language Model Post-Training (arXiv:2411.15124, v5). Source. Link checked 2026-09-16. RLVR and unseen-evaluation sections. The result concerns the reported tasks and training recipe.↩︎

  2. Souppaya, M., Morello, J., & Scarfone, K. (2017). Application Container Security Guide (NIST SP 800-190). National Institute of Standards and Technology. Source. DOI: 10.6028/NIST.SP.800-190. Link checked 2026-09-16. Sections 3–4. Security guidance, not certification of a generated-code runner.↩︎

  3. Nakano, R., Hilton, J., Balaji, S., et al. (2021). WebGPT: Browser-assisted question-answering with human feedback (arXiv:2112.09332). Source. Link checked 2026-09-16. Environment and model interface. Browser-based research setup. The guide generalizes the observation/action pattern to code tools.↩︎