2 DPO for preferred answers
SFT can imitate one demonstrated answer but cannot rank it against another plausible answer. A chosen/rejected pair supplies that missing comparison. A reward model turns many comparisons into a reusable score for fresh attempts. DPO instead learns directly from the same kind of saved pair. DPO and online RLHF are alternative uses of preference evidence, not required consecutive stages.12
Direct Preference Optimization (DPO): The original fixed-data preference-training method used in this chapter. It increases the probability of a chosen completion relative to a rejected completion, while measuring that change against a frozen reference policy. It uses recorded pairs rather than a learned reward model or fresh online rollouts. Rafailov et al. derive this objective.3
Reinforcement Learning from Human Feedback (RLHF): The classic route from preference labels to online policy optimization: train a reward model from the labels, generate fresh completions, score them with that model, and update the policy.4
A preference record has the form \((x,y^+,y^-)\): one prompt, a chosen completion, and a rejected completion. Classic RLHF first converts many such judgments into a reward model. DPO fits the comparison directly.
The same comparison data can train a separate scorer or change the policy directly. The two routes below stay separate. The DPO route scores saved texts, updates the policy, and checks the resulting behavior.
Sections 2.2 and 2.3 form the DPO calculation. Each saved text is scored under both the current policy and the frozen reference before its log-ratio is formed. Section 2.1 supplies the contrasting scorer route, and Sections 2.4 and 2.5 compare the training requirements and test the result.
2.1 Learning a scorer from comparisons
A preference record identifies the better of two saved answers, but it does not automatically assign a score to a new answer. Classical RLHF therefore needs a reusable scorer before it can evaluate fresh model attempts. Score both saved answers, turn their score difference into training feedback, and then test the scorer on unseen pairs and fresh answers.
- Reward model: A learned scalar scoring model. Given a prompt and a completion, it returns a score and is trained to rank the human-chosen completion above the rejected completion.5
Classical RLHF uses that learned score to optimize the policy on fresh completions.
Code 2.1 uses runnable Python data construction. It takes no external input and creates a dictionary named preference. It prints nothing and changes no model state. The expected result has exactly the keys prompt, chosen, and rejected, each mapped to a string.
Code example 2.1: A preference record pairs chosen and rejected text answers.
preference = {
"prompt": "Explain why a test failed.",
"chosen": "Read the first failing assertion and fix that cause.",
"rejected": "The test is probably flaky; rerun it.",
}The three fields are text strings, not precomputed policy numbers. During reward-model training, transformers tokenizes the prompt plus each completion and the reward head returns two scalar scores. During DPO, the same strings are teacher-forced through the current and reference policies to produce the four full-answer log probabilities used by its pair loss. These text fields provide the preference evidence. The model forward passes produce the numerical values.
Reward-model training fits a scalar scoring function to chosen/rejected pairs so classical RLHF can score fresh answers during training. The function predicts the preference pattern in its training data. It does not guarantee truth, safety, or task success. Evaluate the scorer before letting an online optimizer reinforce its outputs.6
- AI-generated preference label: A chosen/rejected label produced by a model judge from an explicit rubric or set of principles, rather than directly by a human rater. For example, Constitutional AI samples responses, asks a model to critique and revise them against human-written principles, then uses a model comparison to create AI-preference data for later training. Record whether a human or model produced the label. An AI judge can label many candidate pairs without asking a human to judge each one. Its preference is evidence about the rubric and judge behavior, not independent proof of factual correctness or human preference.7
A reward model makes preference data reusable for fresh on-policy rollouts, but its score is still a proxy that a policy can exploit.
Calculate the reward-model pair probability: For one preference pair, the reward-model loss penalizes a chosen answer that does not receive a higher scalar score than the rejected answer. Under the logistic Bradley–Terry assumption, the sigmoid turns that score difference into the ranking model’s predicted probability of the observed preference ordering.8 The value need not be calibrated and does not measure factual truth or safety.9
Sigmoid \(\sigma\): A function that converts any real-valued margin \(u\) into a number from zero to one: \(\sigma(u)=1/(1+e^{-u})\). A zero margin becomes 0.5. A larger positive margin moves the result toward one.
The two numeric inputs come from evaluating both answers with the same trained reward model: tokenize \(x\) with \(y^+\), run the model, and read its scalar reward-head output \(r_\psi(x,y^+)\). Then repeat with \(x\) and \(y^-\) to obtain \(r_\psi(x,y^-)\). An implementation may batch the two evaluations. These are unbounded learned scores, not probabilities. The symbol \(\psi\) identifies reward-model parameters. \(\theta\) will identify the trainable policy below. The sigmoid is applied only after the score difference is formed.
\[ \mathrm{ℓ}_{\mathrm{RM}} = -\log \sigma\left(r_{\psi}(x,y^{+}) - r_{\psi}(x,y^{-})\right) \tag{2.1}\]
Variables
- \(\ell_{\mathrm{RM}}\): Loss for one recorded preference pair.
- \(r_\psi\): Scalar reward-model function with parameters \(\psi\).
- \(x\): Prompt shared by both answers.
- \(y^+\) and \(y^-\): Chosen and rejected completions.
Mechanism: Subtract the rejected score from the chosen score, apply the sigmoid to obtain preference probability, and take the negative logarithm. Batch training averages the one-pair losses. Interpretation: Fit a ranking model that predicts which completion a rater preferred. This does not establish factual truth or safety. Scale: Reward-head outputs lie on an arbitrary learned scale, whereas the sigmoid output is a probability. The natural-log loss is reported in nats per preference pair.
The RewardTrainer example in Code 2.2 is an abbreviated, version-sensitive TRL excerpt, not a standalone runnable program. It requires RM_DIR, reward_model, tokenizer, and prepared rm_train_ds and rm_eval_ds datasets. Calling train() updates the reward-model parameters and writes trainer outputs under RM_DIR. The exact metrics depend on the data and configuration, so no numerical run is claimed here. The class and argument names reflect the course example and must be checked against the installed TRL release.10
Code example 2.2: RewardTrainer fits scalar answer scores from preference pairs.
from trl import RewardConfig, RewardTrainer
reward_cfg = RewardConfig(output_dir=RM_DIR, learning_rate=1e-5)
reward_trainer = RewardTrainer(
model = reward_model,
args = reward_cfg,
train_dataset = rm_train_ds,
eval_dataset = rm_eval_ds,
processing_class = tokenizer,
)
reward_trainer.train()rm_train_ds and rm_eval_ds contain prompt/chosen/rejected records, reward_model returns one scalar per answer, and TRL batches the pairwise loss. A successful call returns control after training. It does not by itself establish that the scorer learned the intended quality signal. The project must still check held-out ranking accuracy, deliberately misleading answers, and whether the scorer follows task quality rather than superficial style.
- Application Programming Interface (API): The names and calling rules a software package exposes to code. TRL class and argument names can change between releases, so verify them against the installed version.
Example: reward-model pair probability
For one fixed prompt, a trained reward model receives a saved chosen completion and a saved rejected completion.
Inputs and measurement
Run the reward model once on each prompt-completion pair and read the scalar reward-head outputs: 2.0 for the chosen answer and 0.5 for the rejected answer.
Calculation
- Pair probability and loss: The pair probability is \(\sigma(2.0-0.5)=\sigma(1.5)=1/(1+e^{-1.5})=0.818\). The loss is \(\ell_{\mathrm{RM}}=-\log(0.818)=0.201\) nats.
Interpretation: The logistic ranking model assigns probability 0.818 to the recorded ordering and incurs pair loss 0.201 nats. The value is not automatically calibrated and does not measure factual correctness or universal answer quality.
- Reward-model exploitation: A policy can learn patterns that raise a learned reward score without delivering the quality the score was intended to represent. For example, it may become unnecessarily verbose, imitate phrases that raters tended to favor, or exploit a gap in the scorer’s coverage rather than solve the user’s task.11
Test the scorer before using it for online optimization.
Held-out ranking: Measure pairwise accuracy on prompts and responses excluded from training.
Adversarial answers: Check fluent style, length, refusal language, and out-of-distribution text that may be mistaken for quality.
Fresh generations: Compare reward-model rankings of current policy outputs with human or trusted task-evaluator judgments.
A high score on training pairs is not sufficient evidence that fresh behavior will receive the intended reward.
This limitation motivates the fixed-data DPO path used here. A learned reward model can score fresh rollout behavior, which enables online RLHF, but it adds a proxy that must be evaluated against the behavior it represents. Fixed-data DPO removes that intermediate scorer during its offline update and instead compares the recorded chosen and rejected text answers directly.12
2.2 Turning answer probabilities into the DPO loss
Teacher-forced token probabilities already provide one score for a fixed written answer. DPO now needs four such scores from one chosen/rejected pair and a stable comparison point for interpreting their changes. The unchanged reference supplies that comparison point. Scoring each saved answer under both current and reference policies produces the four log-probabilities whose differences form the DPO margin.
- Reference policy \(\pi_{\mathrm{ref}}\): A frozen copy of the starting language model used to measure how far DPO moves the trainable policy. It is often an SFT checkpoint, but DPO only requires a suitable fixed reference with the same tokenizer and answer support. Like \(\pi_\theta\), it maps the same prompt and prefix to a complete next-token distribution. Its token probabilities and answer log-probabilities come from the forward-pass, normalization, token-selection, and aggregation steps introduced in Chapter 1.
DPO applies fixed-answer scoring to one prompt \(x\) and its saved chosen and rejected texts \(y^+\) and \(y^-\). It does not generate replacement answers while calculating the pair loss. These are four logical scores, not necessarily four serial forward calls. Software may batch the chosen and rejected sequences. Because the saved pairs and reference policy are fixed, the two reference scores may be precomputed and cached. Those cached values remain valid only when the reference checkpoint revision, tokenizer, prompt-and-answer formatting, truncation and response mask, and sequence-score reduction exactly match the training batch. Any change to those inputs invalidates the cache. Current-policy scores must be recomputed after weight updates and must not come from a live cache.
Current policy, chosen answer: Score \(x\) with \(y^+\) under the trainable policy to obtain \(\log\pi_\theta(y^+\mid x)\).
Current policy, rejected answer: Score \(x\) with \(y^-\) under the trainable policy to obtain \(\log\pi_\theta(y^-\mid x)\).
Reference policy, chosen answer: Score the same \(x\) and \(y^+\) under the frozen reference to obtain \(\log\pi_{\mathrm{ref}}(y^+\mid x)\).
Reference policy, rejected answer: Score the same \(x\) and \(y^-\) under the frozen reference to obtain \(\log\pi_{\mathrm{ref}}(y^-\mid x)\).
The four resulting scalar full-answer log-probabilities are model scores, not human ratings and not newly generated answers. Only the current-policy passes receive gradients. The reference passes use the unchanged checkpoint without gradients.
Natural logarithms turn the answer-probability product from Chapter 1 into an equivalent sum that is stable for long text.13
\[ \log \pi_{\theta}(y \mid x) = \sum_{t=1}^{T}\log \pi_{\theta}(y_t \mid x, y_{<t}) \tag{2.2}\]
Variables
- \(x\) and \(y\): Fixed prompt and complete written answer.
- \(y_t\) and \(y_{<t}\): Recorded answer token at \(t\) and the prefix before it.
- \(T\): Number of answer tokens.
- \(\pi_\theta\): Trainable next-token policy.
Mechanism: Take the log probability of each recorded next token and sum across the answer. Logarithms turn the probability product into a stable sum. Interpretation: Produce one log-probability score for a fixed text answer under the trainable policy.
Example: three token scores form one answer score
A fixed answer has three recorded tokens whose gathered log-probabilities are −0.20, −0.50, and −0.10.
Inputs and measurement
Teacher forcing supplies the three token positions. Log-softmax and gather read the current policy’s log-probability of the recorded ID at each position.
Calculation
- Answer log-probability: −0.20 + (−0.50) + (−0.10) = −0.80
Interpretation: The answer receives full-sequence log-probability −0.80 nats. DPO repeats the same sum for chosen and rejected text under both current and reference policies.
For one answer, DPO subtracts the reference policy’s full-answer log probability from the trainable policy’s full-answer log probability. This log-ratio measures how much more or less likely the current policy makes that same written answer than the reference does. DPO calculates a log-ratio for the chosen text and another for the rejected text. Subtracting the rejected-answer log-ratio from the chosen-answer log-ratio measures the relative shift toward the chosen answer, compared with the reference policy.
Example: answer length does not determine a DPO log-ratio
For one fixed prompt, compare the contribution of two added answer tokens under the current and reference policies. This isolates the length question. A complete answer comparison must also score the remaining tokens, including the termination token.
Inputs and measurement
Teacher forcing supplies the two recorded tokens. Suppose the reference scoring pass returns log-probabilities \(-0.20\) and \(-0.30\). Compare two hypothetical current checkpoints that assign different probabilities to these same tokens. At each position, subtract the reference log-probability from the current-policy log-probability, then sum the signed changes.
Calculation
A less likely extension: \([-0.30-(-0.20)]+[-0.40-(-0.30)]=-0.20\) nats.
A more likely extension: \([-0.10-(-0.20)]+[-0.20-(-0.30)]=+0.20\) nats.
Interpretation: The same two-token increase in length can lower or raise the policy/reference log-ratio. These signed contributions do not provide an automatic bonus for a longer answer.
If raters systematically choose longer answers even when a shorter answer meets the task equally well, the pair labels associate verbosity with being chosen. DPO can learn that shortcut, with the outcome depending on the data and objective. Summing in log space does not remove this risk. Park et al. document length exploitation in DPO on summarization and dialogue datasets.14
The standard DPO objective uses summed sequence log-probabilities. Replacing them with per-token averages defines a different objective and requires an explicit design choice. Section 2.5 compares held-out results in answer-length groups and checks whether added text improves the task result.
- Kullback-Leibler (KL) divergence: A numerical measure of how one probability distribution differs from another. In preference and online objectives it discourages the trainable policy from moving too far from a frozen reference, but it does not measure factual correctness.
A normalized optimal-policy expression explains why reference-policy ratios appear. For each fixed prompt, the standard derivation maximizes expected preference score minus \(\beta\,\mathrm{KL}(\pi\Vert\pi_{\mathrm{ref}})\) over a normalized answer distribution. It assumes \(\beta>0\) and that the reference assigns nonzero probability wherever the optimized policy may place probability. Under those assumptions, higher-scoring answers receive more relative weight while the reference anchors the change. A prompt-specific normalization makes the answer probabilities sum to one.
\[ \pi^{*}(y \mid x) = \frac{\pi_{\mathrm{ref}}(y \mid x)\exp\left(r(x,y)/\beta\right)}{Z(x)} \tag{2.3}\]
Variables
- Optimal policy: The probability distribution that maximizes the regularized objective.
- \(\pi^*(y\mid x)\): Ideal optimal-policy probability of answer \(y\) after prompt \(x\). It is used only in the derivation, not stored as a separate model during DPO training.
- \(\pi_{\mathrm{ref}}\): Frozen reference probability for answer \(y\) after prompt \(x\).
- \(r(x,y)\): Hypothetical preference score, not a separately trained scorer used by DPO.
- \(\beta\): Positive coefficient relating log-probability change to the preference-score scale.
- \(Z(x)\): Prompt-specific normalization that makes answer probabilities sum to one.
Mechanism: The exponential increases relative weight for higher preference scores, the reference ties that change to known behavior, and \(Z(x)\) normalizes the result. Interpretation: Express an ideal preference-weighted policy. \(Z(x)\) cancels when chosen and rejected answers for the same prompt are compared.
Solve the normalized expression for the hypothetical score before comparing the two answers. Taking the logarithm turns the policy-to-reference ratio into an additive term, while the prompt normalization remains an additive constant shared by every answer to that prompt.
\[ r(x,y) = \beta\log\left(\frac{\pi^{*}(y \mid x)}{\pi_{\mathrm{ref}}(y \mid x)}\right) + \beta\log Z(x) \tag{2.4}\]
Variables
- \(r(x,y)\): Hypothetical preference score for answer \(y\) after prompt \(x\).
- \(\pi^*\): Ideal preference-weighted policy from the previous equation.
- \(\pi_{\mathrm{ref}}\): Frozen reference policy.
- \(\beta\): Positive scale relating policy log-ratios to the preference-score scale.
- \(Z(x)\): Normalization shared by every answer to prompt \(x\).
Mechanism: Divide ideal-policy probability by reference probability, take the natural logarithm, scale it by \(\beta\), and add the prompt-specific normalization term. Interpretation: Recover the score represented by one policy-to-reference probability shift. When chosen and rejected scores for the same prompt are subtracted, their shared normalization terms cancel.
For the same prompt, substitute Equation 2.4 for both answers before applying the Bradley-Terry preference likelihood:
\[ \begin{aligned} r(x,y^+)-r(x,y^-) &=\beta\left[\log\frac{\pi^*(y^+\mid x)}{\pi_{\mathrm{ref}}(y^+\mid x)}+\log Z(x)\right]\\ &\quad-\beta\left[\log\frac{\pi^*(y^-\mid x)}{\pi_{\mathrm{ref}}(y^-\mid x)}+\log Z(x)\right]\\ &=\beta\left[\log\frac{\pi^*(y^+\mid x)}{\pi_{\mathrm{ref}}(y^+\mid x)}-\log\frac{\pi^*(y^-\mid x)}{\pi_{\mathrm{ref}}(y^-\mid x)}\right]. \end{aligned} \]
The \(\beta\log Z(x)\) terms cancel because both answers use the same prompt. Bradley-Terry then gives \(P(y^+\succ y^-\mid x)=\sigma(r(x,y^+)-r(x,y^-))\). Substitute the recovered difference, take \(-\log\), and average pairs. DPO cannot train the unavailable ideal distribution \(\pi^*\), so it parameterizes the policy as the current \(\pi_\theta\) and differentiates only those current-policy scores. The frozen reference scores and saved answer tokens are held fixed.
The coefficient \(\beta\) first appears as the positive KL regularization scale in the ideal objective. After substitution, the same value scales the logistic preference margin. Changing \(\beta\) therefore changes both the reference-relative regularization in the derivation and the margin scale in the implemented loss.15
\[ L_{\mathrm{DPO}}(\theta) = -\mathbb{E}\left[\log \sigma\left(\beta\left(\log\frac{\pi_{\theta}(y^{+}\mid x)}{\pi_{\mathrm{ref}}(y^{+}\mid x)}-\log\frac{\pi_{\theta}(y^{-}\mid x)}{\pi_{\mathrm{ref}}(y^{-}\mid x)}\right)\right)\right] \tag{2.5}\]
Variables
- \(L_{\mathrm{DPO}}\): Loss minimized by the trainable policy.
- \(x\), \(y^+\), and \(y^-\): Prompt, chosen answer, and rejected answer from one preference record.
- \(\pi_\theta\) and \(\pi_{\mathrm{ref}}\): Trainable and frozen reference policies.
- \(\beta\): Positive coefficient converting a log-ratio change to the preference-score scale.
- \(\sigma\): Sigmoid function. The expectation averages preference pairs.
Mechanism: Each log-ratio measures current change from the reference. Subtracting rejected from chosen forms the preference-margin improvement. \(\beta\) scales it. The negative log sigmoid penalizes a margin that favors the rejected answer. Interpretation: Raise the chosen-versus-rejected policy margin while measuring that change relative to reference behavior. Scale: The natural-log loss is reported in nats per preference pair.
Example: DPO turns four policy scores into one preference margin
Keep one prompt and its saved chosen and rejected answers fixed. The example compares what the current policy and frozen reference assign to those exact texts. It does not generate four new answers.
Inputs and measurement
Teacher-force the chosen and rejected token sequences through the current policy, apply log-softmax, gather each recorded token, and sum the gathered values. Repeat those two scoring passes with the frozen reference. The four sums are illustrative forward-pass outputs. Use the same DPO coefficient as the trainer configuration, \(\beta=0.10\).
Calculation
Current-policy scores: \(\log\pi_\theta(y^+\mid x)=-0.90\) and \(\log\pi_\theta(y^-\mid x)=-2.00\).
Reference-policy scores: \(\log\pi_{\mathrm{ref}}(y^+\mid x)=-1.50\) and \(\log\pi_{\mathrm{ref}}(y^-\mid x)=-1.70\).
Preference-margin calculation: The chosen-minus-rejected margin is \([-0.90-(-1.50)]-[-2.00-(-1.70)]=0.60-(-0.30)=0.90\). Scaling by \(\beta=0.10\) gives \(0.10\times0.90=0.09\).
Sigmoid and pair loss: The sigmoid is \(\sigma(0.09)=1/(1+e^{-0.09})=0.522\). The pair loss is \(L_{\mathrm{DPO}}=-\log(0.522)=0.649\) nats.
Interpretation: The current policy improves the chosen-minus-rejected margin beyond the reference by 0.90 nats. Scaling gives 0.09, the sigmoid gives 0.522, and the resulting DPO loss is 0.649 nats. Gradient flows only through the current policy or its LoRA adapters.
Counterexample: a better relative margin can lower chosen probability
For one prompt, suppose the reference assigns complete-answer probabilities 0.50 to \(y^+\) and 0.10 to \(y^-\). A current policy assigns 0.40 to \(y^+\) and 0.02 to \(y^-\). Remaining probability can belong to other answers.
Calculation
Chosen change: \(\log(0.40/0.50)=\log(0.8)=-0.223\).
Rejected change: \(\log(0.02/0.10)=\log(0.2)=-1.609\).
Relative margin: \(-0.223-(-1.609)=+1.386\) nats.
Interpretation: The relative chosen-minus-rejected margin improves by 1.386 nats even though the chosen answer’s absolute probability fell from 0.50 to 0.40. DPO’s pair loss compares the current policy’s chosen-versus-rejected margin with the reference policy’s margin. Check absolute response behavior separately when that property matters.
2.3 Preparing DPO from an SFT checkpoint with adapters
DPO needs a comparison record, current and reference policy scores, and one loss. The preference record, token scores, adapters, and loss come from different software components. Their interfaces must preserve the same prompt-and-answer alignment used in the derivation.
TRL provides the objective-and-trainer step. DPOTrainer consumes the preference record and applies the DPO loss. datasets, transformers, and peft supply the adjacent data, model, and adapter operations.
The trl-lib/ultrafeedback_binarized dataset provides both a chosen completion for SFT and a chosen/rejected pair for DPO. The course workflow fine-tunes Qwen2.5-0.5B-Instruct with LoRA and then runs DPO from that checkpoint. The abbreviated excerpt below prepares that DPO stage. Evaluation includes held-out pair ranking and matched generations. Section 2.5 defines the implicit-reward accuracy and margin used for the pair check.16
TRL handles the objective and trainer loop, while adjacent libraries perform four other jobs.
datasets: Store prompt, chosen-answer, and rejected-answer fields.transformers: Tokenize the text and run current and reference model passes.peft: Attach and save the trainable LoRA weights.Hugging Face Accelerate (
accelerate): Coordinate supported processes and place model work on the selected devices.17
Project checks still validate the preference records and generated behavior.
DPOConfig stores the objective settings, and DPOTrainer binds the comparison records to the SFT checkpoint, frozen reference, and LoRA adapters. The explicit reference_model below is a frozen copy of the exact SFT checkpoint.
The setup in Code 2.3 is an abbreviated, version-sensitive TRL/PEFT setup, not a standalone runnable program. It requires loaded policy and reference models, a compatible tokenizer, prepared dpo_train and dpo_eval datasets, output and length constants, and installed library versions whose arguments match the excerpt. The shown code constructs datasets, configurations, and a trainer. It does not call train() and therefore does not update weights or produce metrics.
Code example 2.3: DPOTrainer compares answer pairs against a reference policy.
from peft import LoraConfig
from trl import DPOConfig, DPOTrainer
# Keep only the fields that define a DPO comparison.
keep = ["prompt", "chosen", "rejected"]
dpo_train_ds = dpo_train.remove_columns(
[column for column in dpo_train.column_names if column not in keep]
)
dpo_eval_ds = dpo_eval.remove_columns(
[column for column in dpo_eval.column_names if column not in keep]
)
# Configure the DPO objective and the trainable LoRA adapter.
dpo_cfg = DPOConfig(
output_dir = DPO_DIR,
beta = 0.1,
learning_rate = 5e-5,
max_length = MAX_LEN,
)
dpo_lora_cfg = LoraConfig(
r = 16,
lora_alpha = 32,
target_modules = ["q_proj", "k_proj", "v_proj", "o_proj"],
)
dpo_trainer = DPOTrainer(
model = policy_model,
ref_model = reference_model,
args = dpo_cfg,
train_dataset = dpo_train_ds,
eval_dataset = dpo_eval_ds,
peft_config = dpo_lora_cfg,
processing_class = tokenizer,
)Comparison record: The line beginning with
keepnames the prompt, chosen answer, and rejected answer. The two dataset operations remove every other column so the trainer receives only that record.DPO controls:
dpo_cfgsets the DPO scale \(\beta\) and training limits, whiledpo_lora_cfglimits which policy parameters can change.Policy and reference:
policy_modelis trainable, whilereference_modelis the frozen SFT checkpoint. An implicit disabled-adapter reference is valid only when disabling the new DPO adapter exposes exactly that SFT checkpoint, for example after the SFT adapter has been merged or when a named frozen SFT adapter is selected. Otherwise use an explicit reference model or reference adapter.18Trainer input: The trainer constructs batches and calculates the loss.
transformerssupplies the logits whose gathered log probabilities DPO compares.
Low-Rank Adaptation (LoRA) keeps selected base projections frozen and trains small adapter matrices instead. It can reduce the trainable and stored parameter set, but it supplies neither preference data nor a loss. DPO defines the chosen-versus-rejected comparison. LoRA only limits which policy values can change.19 The adapter reference gives the matrix equation, shapes, and the 4096-by-4096 parameter-count example used by this configuration.
The fixed-data DPO form is useful when saved chosen/rejected pairs cover the behavior being trained. It increases the chosen-versus-rejected policy margin relative to a frozen reference without generating fresh rollouts or training a separate reward model. It does not explore new answers or certify factual correctness, and it need not be followed by online RL.
2.4 What DPO gains by using saved answers
The same preference pairs can train either a reusable reward model or DPO directly. The choice depends on whether training must score fresh behavior or only learn from the saved comparisons. Compare evidence coverage, required models, rollout cost, and the risk of relying on a learned score.
DPO and PPO-based RLHF begin with preference evidence but turn it into policy updates differently. The original fixed-data DPO form uses a fixed-pair loss. PPO-based RLHF keeps an online policy, reference policy, reward model, and value model active. The value model, also called a critic, estimates the expected remaining rollout score for the PPO update. Chapter 3 develops that role. Their different feedback paths create different costs, distribution shift, and reward-exploitation risks.2021
Fresh answers can also supply new DPO pairs. In iterative DPO, the current policy generates candidates, a judge labels their relative quality, and another round of DPO learns from those comparisons. Online DPO brings generation and judging into the training loop. Both differ from repeatedly scoring one fixed dataset. They add generation and judging costs while retaining a pair-based update. TRL v0.22.2 documents training from prompts with newly generated, judged completions.22
Start the choice with the evidence, not the algorithm name. Fixed chosen/rejected answers support repeatable offline DPO. When fresh behavior must be explored, online DPO needs new pair judgments, while PPO can use a scalar reward for each attempt. The available feedback and acceptable training cost determine which comparison is useful.
The comparison has four practical differences. Each shows a cost or model role that DPO removes, retains, or shifts into evaluation.
Optimization data: Fixed-data DPO learns from saved preference pairs. PPO-based RLHF generates on-policy completions and asks a reward model to score them.
Required model roles: DPO needs a trainable policy and reference. PPO-based RLHF also retains a reward model and a value model.
Strength: Fixed-data DPO has a simple stable offline objective. PPO-based RLHF can optimize a reward on fresh generated behavior.
Risk: Fixed-data DPO is limited by the preference-dataset distribution. PPO-based RLHF adds costly rollouts and opportunities to exploit reward-model errors.
Choose fixed-data DPO when the pair labels are trusted preference evidence and repeatable offline training is the main goal. In the implementation, datasets supplies the fixed triples, transformers scores their tokens, peft limits trainable weights, and TRL’s DPOTrainer calculates the loss. This avoids training and serving a separate reward model during the update.
Choose PPO-based RLHF only when a reward model or environment can score fresh rollouts and that additional coverage is worth the rollout, critic, and stability machinery. In either route, hold out prompts, not merely individual answer rows, so a near-duplicate pair cannot make an offline metric look stronger than the behavior on a new user request.
Other preference objectives change the evidence or scoring rule, rather than removing the need to evaluate held-out behavior.
| Method | Evidence and objective | Reference and cost caveat |
|---|---|---|
| Identity-PO (IPO)23 | Chosen/rejected pairs. A squared loss pulls the reference-relative chosen-minus-rejected log-ratio toward a finite target instead of rewarding an ever-larger margin. | Uses a frozen reference in its standard sampled form, so it retains paired labels and current/reference scoring. |
| Kahneman-Tversky Optimization (KTO)24 | Individual desirable or undesirable labels, so it can use unpaired binary feedback rather than one chosen/rejected pair. | Its published objective uses a reference-relative implicit reward. Binary labels still need a clear criterion and quality check. |
| Simple Preference Optimization (SimPO)25 | Chosen/rejected pairs. The loss uses average sequence log probability plus a target margin in a Bradley-Terry-style comparison. | It removes the reference model from its stated objective. That can remove a reference pass and its memory, but total training cost still depends on batching, sequence lengths, activations, and hardware. |
| Odds Ratio Preference Optimization (ORPO)26 | Chosen/rejected pairs. The objective combines chosen-answer supervised likelihood with a preference term that increases the favored answer’s odds relative to the rejected answer. | It is reference-free and combines SFT with preference alignment in one objective. It still needs paired evidence. |
These names are not interchangeable recipes. Choose the objective from the available evidence, desired behavior, reference-model budget, and held-out evaluation, then verify the exact implementation against the paper and installed library version.
2.5 Checking whether preferences improved
Choosing DPO only identifies a training path. A lower loss still does not prove that generated answers improved. The checkpoint must be tested on comparisons and tasks that were not used for its update. Separate held-out pair ranking, fixed-prompt generation, and regression checks so each check measures a distinct property.
DPO training has two distinct outcomes to verify. A held-out preference set tests whether the chosen answer has higher implicit DPO reward than the rejected answer relative to the same reference policy. A generation set tests whether those changes produce better behavior for unseen prompts. rewards/accuracies and rewards/margins measure the implicit-reward ranking, while matched outputs compare the base, SFT, and SFT-plus-DPO checkpoints.27
Use three complementary checks rather than treating one trainer metric as proof of alignment.
Held-out pair ranking: Measure the fraction of pairs with positive chosen-minus-rejected implicit-reward margin and report that margin’s distribution.
Generated behavior: Compare matched unseen prompts under fixed decoding to see whether the probability shift changes actual answers usefully.
Independent task and safety tests: Measure the correctness, helpfulness, safety, and formatting properties that motivated the labels, including regressions outside the preference dataset.
For each held-out pair, score both fixed answers under the final policy and frozen reference. The implicit reward for one answer is \(\beta\) times current-policy log-probability minus reference-policy log-probability. Subtracting rejected from chosen gives the evaluation margin.
\[ \Delta r_{\theta} = \beta\left[\left(\log\pi_{\theta}(y^{+}\mid x)-\log\pi_{\mathrm{ref}}(y^{+}\mid x)\right)-\left(\log\pi_{\theta}(y^{-}\mid x)-\log\pi_{\mathrm{ref}}(y^{-}\mid x)\right)\right] \tag{2.6}\]
Variables
- \(\Delta r_\theta\): Chosen-minus-rejected implicit DPO reward margin.
- \(\beta\): DPO margin scale.
- \(\pi_\theta\) and \(\pi_{\mathrm{ref}}\): Final trainable policy and the unchanged reference used during DPO.
- \(y^+\) and \(y^-\): Held-out chosen and rejected answers for prompt \(x\).
Mechanism: Calculate the current-minus-reference log-probability change for each answer, subtract the rejected change from the chosen change, and multiply by \(\beta\). Interpretation: Count a pair as correct when \(\Delta r_\theta\) is positive. Raw current-policy sequence log-probabilities may be reported separately, but they are length-sensitive and are not the DPO reward-accuracy metric.
Report the fraction of held-out pairs with positive implicit-reward margin and the margin distribution. A rising training margin with flat held-out accuracy warns that the policy is fitting training pairs rather than generalizing their preference rule.
For the behavior check, keep prompts, generation settings, and evaluators fixed while comparing the base, SFT, and DPO checkpoints. Ask blinded reviewers or a trusted task evaluator to judge correctness, helpfulness, safety, and formatting for the target use case. Then inspect failures by category: a better aggregate preference metric is not sufficient if DPO made refusals too broad, shortened correct reasoning, or degraded a capability absent from the pair dataset. Keep the checkpoint only when held-out pair ranking remains sound, target-task behavior improves, and regression checks show no unacceptable loss in critical capabilities. Trainer curves support diagnosis but do not replace those acceptance checks.
Example: better format can hide a worse answer
An application needs exactly one JSON object with an integer count field. Its evaluator first parses the answer, then checks the count against the requested operation. The following outputs are constructed examples of a base, SFT, and SFT-plus-DPO comparison, not saved results from trained checkpoints. For an actual comparison, all three checkpoints would receive these same prompts with greedy decoding and the same 32-new-token limit.
- Prompt:
Return only JSON with the number of items in [2, 4, 4], using the key count.The expected count is 3.- Illustrative base answer:
There are three items.The meaning is correct, but the format check fails. - Illustrative SFT answer:
{"count": 3}. Both checks pass. - Illustrative SFT-plus-DPO answer:
{"count": 3}. Both checks pass, with no further improvement on this task.
- Illustrative base answer:
- Prompt:
Return only JSON with the number of distinct values in [2, 4, 4], using the key count.The expected count is 2.- Illustrative base answer:
There are two distinct values.The meaning is correct, but the format check fails. - Illustrative SFT answer:
{"count": 2}. Both checks pass. - Illustrative SFT-plus-DPO answer:
{"count": 3}. The format passes, but correctness regresses.
- Illustrative base answer:
Conclusion: A preference for clean JSON cannot substitute for checking the operation requested by each prompt. A useful comparison records the unchanged result and the regression, not only the task that became better formatted. These invented outputs illustrate the evaluation decision, not a typical effect of SFT or DPO. The course notebook’s comparison cells contain no saved SFT/DPO outputs from which an observed improvement could be reported.
The common run record in Section 6.2 keeps prompt splits, checkpoint identities, decoding settings, evaluator versions, failure categories, and the acceptance decision. The DPO record also needs the sequence log probability from the current policy and from the reference policy, \(\beta\), answer lengths, and chosen-minus-rejected margin for each held-out pair. Store those alongside matched generations and length-stratified pair results so a better pair metric can be compared with actual output behavior.
These records distinguish a policy that fits saved preference pairs from one whose generated behavior improves. When the task must score fresh attempts rather than fixed pairs, Chapter 3 replaces the pair loss with online rewards and controlled rollout reuse.
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., & Finn, C. (2023). Direct Preference Optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36. Source. Link checked 2026-09-16. Relevant: Sections 3–4 and Equations 3–7 derive the reference-relative DPO loss and explain that fixed-data DPO omits a separate reward model and policy sampling during fine-tuning. Limit: experiments cover sentiment, summarization, and single-turn dialogue with models up to 6B parameters.↩︎
Ouyang, L., et al. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744. DOI: 10.52202/068431-2011. Official proceedings. Link checked 2026-09-16. Relevant: Figure 2 and Sections 3.1 and 3.4 document the SFT, reward-model, and PPO stages of InstructGPT. Limit: this is one RLHF implementation and does not make those stages compulsory for every post-training project.↩︎
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., & Finn, C. (2023). Direct Preference Optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36. Source. Link checked 2026-09-16. Relevant: Sections 3–4 and Equations 3–7 derive the reference-relative DPO loss and explain that fixed-data DPO omits a separate reward model and policy sampling during fine-tuning. Limit: experiments cover sentiment, summarization, and single-turn dialogue with models up to 6B parameters.↩︎
Ouyang, L., et al. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744. DOI: 10.52202/068431-2011. Official proceedings. Link checked 2026-09-16. Relevant: Figure 2 and Sections 3.1 and 3.4 document the SFT, reward-model, and PPO stages of InstructGPT. Limit: this is one RLHF implementation and does not make those stages compulsory for every post-training project.↩︎
Stiennon, N., et al. (2020). Learning to summarize from human feedback. Advances in Neural Information Processing Systems, 33. Official paper. Link checked 2026-09-16. Relevant: Sections 3.4 and 4.3 define a scalar reward model trained from pair comparisons and use its output for PPO. Section 4.3 reports a length bias in that reward model. Limit: evidence comes from Reddit TL;DR summarization with 1.3B and 6.7B models, not all tasks or reward models.↩︎
Stiennon, N., et al. (2020). Learning to summarize from human feedback. Advances in Neural Information Processing Systems, 33. Official paper. Link checked 2026-09-16. Relevant: Sections 3.4 and 4.3 define a scalar reward model trained from pair comparisons and use its output for PPO. Section 4.3 reports a length bias in that reward model. Limit: evidence comes from Reddit TL;DR summarization with 1.3B and 6.7B models, not all tasks or reward models.↩︎
Bai, Y., et al. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv:2212.08073. arXiv. DOI: 10.48550/arXiv.2212.08073. Link checked 2026-09-16. Relevant: abstract and Sections 2–4 describe self-critique/revision followed by model-generated preference comparisons and reinforcement learning from AI feedback. Limit: this is a preprint about one constitutional-harmlessness setup. An AI label is not independent proof of human preference or factual correctness.↩︎
Bradley, R. A., & Terry, M. E. (1952). Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika, 39(3–4), 324–345. Oxford Academic. https://doi.org/10.1093/biomet/39.3-4.324. Link checked 2026-09-16. Relevant: the paired-comparison probability model supplies the logistic score-difference form later used for preference likelihoods. Limit: the original statistical model concerns paired comparisons generally, not language-model reward calibration, truth, or safety.↩︎
Stiennon, N., et al. (2020). Learning to summarize from human feedback. Advances in Neural Information Processing Systems, 33. Official paper. Link checked 2026-09-16. Relevant: Sections 3.4 and 4.3 define a scalar reward model trained from pair comparisons and use its output for PPO. Section 4.3 reports a length bias in that reward model. Limit: evidence comes from Reddit TL;DR summarization with 1.3B and 6.7B models, not all tasks or reward models.↩︎
Hugging Face. (2025). Reward Trainer (TRL v0.22.2 documentation). Official documentation. Link checked 2026-09-16. Relevant: the
RewardTrainerAPI and expected preference-record fields document the class and argument behavior used by the excerpt. Limit: this is versioned implementation documentation, not evidence that a trained scorer captures the intended quality.↩︎Gao, L., Schulman, J., & Hilton, J. (2023). Scaling laws for reward model overoptimization. Proceedings of the 40th International Conference on Machine Learning, PMLR 202, 10835-10866. Source. Link checked 2026-09-16. Relevant: abstract and Sections 1–3 show that optimizing an imperfect proxy reward can eventually reduce a separate gold-standard score. Limit: the study uses a synthetic gold reward model rather than direct human judgments and does not establish every listed exploitation pattern.↩︎
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., & Finn, C. (2023). Direct Preference Optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36. Source. Link checked 2026-09-16. Relevant: Sections 3–4 and Equations 3–7 derive the reference-relative DPO loss and explain that fixed-data DPO omits a separate reward model and policy sampling during fine-tuning. Limit: experiments cover sentiment, summarization, and single-turn dialogue with models up to 6B parameters.↩︎
Bengio, Y., Ducharme, R., Vincent, P., & Jauvin, C. (2003). A neural probabilistic language model. Journal of Machine Learning Research, 3, 1137–1155. Source. Link checked 2026-09-16. Pages 1138 and 1141–1142: conditional probability product, log-likelihood and vocabulary softmax. This paper uses a fixed-context word model, not a Transformer. The probability calculation is the shared basis. The guide’s toy logits are illustrative inputs.↩︎
Park, R., Rafailov, R., Ermon, S., & Finn, C. (2024). Disentangling length from quality in direct preference optimization. Findings of the Association for Computational Linguistics: ACL 2024, 4998–5017. ACL Anthology. DOI: 10.18653/v1/2024.findings-acl.297. Link checked 2026-09-16. Relevant: Sections 3–5 report length exploitation in DPO on summarization and dialogue datasets and evaluate a length regularizer. Limit: the result is empirical and dataset-dependent. Added tokens do not mechanically improve a DPO log-ratio.↩︎
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., & Finn, C. (2023). Direct Preference Optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36. Source. Link checked 2026-09-16. Relevant: Sections 3–4 and Equations 3–7 derive the reference-relative DPO loss and explain that fixed-data DPO omits a separate reward model and policy sampling during fine-tuning. Limit: experiments cover sentiment, summarization, and single-turn dialogue with models up to 6B parameters.↩︎
Hugging Face. (2025). DPO Trainer (TRL v0.22.2 documentation). Official documentation. Link checked 2026-09-16. Relevant: “Quick start,” “Expected dataset type,” “Logged metrics,” and “Reference model considerations with PEFT” document
DPOTrainer,trl-lib/ultrafeedback_binarized, implicit-reward metrics, and reference-adapter options. Limit: this is versioned API guidance, not experimental proof that a given dataset, metric, or checkpoint improves behavior.↩︎Hugging Face. (2025). Accelerate (Accelerate v1.10.1 documentation). Official documentation. Link checked 2026-09-16. Relevant: the overview and quick tour document distributed launch support,
Accelerator.prepare, and automatic device placement for supported PyTorch objects. Limit: this is versioned implementation documentation. Exact distributed behavior depends on the selected backend and configuration.↩︎Hugging Face. (2025). DPO Trainer (TRL v0.22.2 documentation). Official documentation. Link checked 2026-09-16. Relevant: “Quick start,” “Expected dataset type,” “Logged metrics,” and “Reference model considerations with PEFT” document
DPOTrainer,trl-lib/ultrafeedback_binarized, implicit-reward metrics, and reference-adapter options. Limit: this is versioned API guidance, not experimental proof that a given dataset, metric, or checkpoint improves behavior.↩︎Hu, E. J., et al. (2022). LoRA: Low-rank adaptation of large language models. International Conference on Learning Representations. arXiv. Link checked 2026-09-16. Relevant: abstract and Section 4 freeze pretrained weights and inject trainable low-rank matrices into selected layers. Limit: LoRA reduces trainable parameters but does not provide preference records or define the DPO loss.↩︎
Stiennon, N., et al. (2020). Learning to summarize from human feedback. Advances in Neural Information Processing Systems, 33. Official paper. Link checked 2026-09-16. Relevant: Sections 3.4 and 4.3 define a scalar reward model trained from pair comparisons and use its output for PPO. Section 4.3 reports a length bias in that reward model. Limit: evidence comes from Reddit TL;DR summarization with 1.3B and 6.7B models, not all tasks or reward models.↩︎
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., & Finn, C. (2023). Direct Preference Optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36. Source. Link checked 2026-09-16. Relevant: Sections 3–4 and Equations 3–7 derive the reference-relative DPO loss and explain that fixed-data DPO omits a separate reward model and policy sampling during fine-tuning. Limit: experiments cover sentiment, summarization, and single-turn dialogue with models up to 6B parameters.↩︎
Hugging Face. (2025). Online DPO Trainer (TRL v0.22.2 documentation). Official documentation. Link checked 2026-09-16. Relevant: overview and quick start show a current model generating candidates from prompts and a judge choosing between them inside the online training workflow. Limit: this is versioned implementation guidance and its model-judge setup is one form of online preference feedback.↩︎
Gheshlaghi Azar, M., et al. (2024). A general theoretical paradigm to understand learning from human preferences. Proceedings of the 27th International Conference on Artificial Intelligence and Statistics, PMLR 238, 4447–4455. PMLR. Link checked 2026-09-16. Relevant: Sections 4–5 introduce Identity-PO as a special case of ΨPO and derive its squared finite-target loss. Limit: empirical comparisons are illustrative and do not establish universal superiority over DPO.↩︎
Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., & Kiela, D. (2024). Model alignment as prospect theoretic optimization. Proceedings of the 41st International Conference on Machine Learning, PMLR 235, 12634–12651. PMLR. Link checked 2026-09-16. Relevant: Sections 3–4 define KTO from desirable/undesirable binary feedback and a reference-relative implicit reward. Limit: reported results cover selected model families and benchmarks from 1B to 30B parameters.↩︎
Meng, Y., Xia, M., & Chen, D. (2024). SimPO: Simple preference optimization with a reference-free reward. Advances in Neural Information Processing Systems, 37. Official proceedings. https://doi.org/10.52202/079017-3946. Link checked 2026-09-16. Relevant: Sections 3–4 define average sequence log probability, a target margin, and the reference-free SimPO objective. Limit: compute and memory savings from removing a reference pass do not determine total training cost across hardware and batching choices.↩︎
Hong, J., Lee, N., & Thorne, J. (2024). ORPO: Monolithic preference optimization without reference model. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 11170–11189. ACL Anthology. DOI: 10.18653/v1/2024.emnlp-main.626. Link checked 2026-09-16. Relevant: Sections 3–4 combine chosen-answer negative log likelihood with an odds-ratio preference term in a reference-free objective. Limit: experiments cover selected 125M–7B models and UltraFeedback-derived setups. Paired evidence is still required.↩︎
Hugging Face. (2025). DPO Trainer (TRL v0.22.2 documentation). Official documentation. Link checked 2026-09-16. Relevant: “Quick start,” “Expected dataset type,” “Logged metrics,” and “Reference model considerations with PEFT” document
DPOTrainer,trl-lib/ultrafeedback_binarized, implicit-reward metrics, and reference-adapter options. Limit: this is versioned API guidance, not experimental proof that a given dataset, metric, or checkpoint improves behavior.↩︎