Part I: Learning from saved examples
A pre-trained language model can produce several plausible answers to the same request. Additional examples can make a useful response format or skill more likely. A saved comparison provides different information by identifying which of two answers is preferred. Both kinds of training can use stored records without generating new training answers during each update.1
Chapter 1 starts with demonstrated answers and explains how text becomes numerical probabilities and a supervised training loss. Chapter 2 uses that scoring procedure to compare answers and train from preference pairs. The choice between the two depends on the available records and the behavior that needs improvement. Held-out examples test whether the resulting model actually improved.
The two branches are alternatives selected by the available records. SFT uses the demonstrated tokens. DPO uses relative answer preference and a fixed comparison policy. In both cases, an optimizer updates trainable weights or adapters, and held-out examples check the result.2
Chapters in this part
- 1. SFT for task responses: A pre-trained model can produce plausible text without consistently following a task’s format or constraints. Supervised fine-tuning (SFT) uses trusted prompt-response examples to make those demonstrated responses more likely. Demonstrated text becomes recorded token probabilities and a supervised update. SFT directly rewards reproducing its supplied answers, so later methods add comparisons or outcome evidence when the task needs more.
- 2. DPO for preferred answers: SFT can imitate one demonstrated answer but cannot rank it against another plausible answer. A chosen/rejected pair supplies that missing comparison. A reward model turns many comparisons into a reusable score for fresh attempts. DPO instead learns directly from the same kind of saved pair. DPO and online RLHF are alternative uses of preference evidence, not required consecutive stages.
Ouyang, L., et al. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744. DOI: 10.52202/068431-2011. Official proceedings. Link checked 2026-09-16. Figure 2 and methods. One documented training recipe, not a required sequence.↩︎
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., & Finn, C. (2023). Direct Preference Optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36. Source. Link checked 2026-09-16. DPO objective and algorithm. The basic offline method does not sample new training responses.↩︎