Reinforcement Learning for LLM Post-Training

An engineer adjusts a brass-and-glass model instrument beside example cards, an answer-comparison balance, and a test lens in a warm wooden workshop.

Contents

Front matter

  • Choosing better LLM answers
    Connect the need for useful answers to demonstrations, preferences, scored attempts, and independent evaluation.

Part I: Learning from saved examples

Demonstrations and saved answer comparisons provide two kinds of training feedback.

  • 1. SFT for task responses
    A pre-trained model can produce plausible text without consistently following a task’s format or constraints. Supervised fine-tuning (SFT) uses trusted prompt-response examples to make those demonstrated responses more likely. Demonstrated text becomes recorded token probabilities and a supervised update. SFT directly rewards reproducing its supplied answers, so later methods add comparisons or outcome evidence when the task needs more.
  • 2. DPO for preferred answers
    SFT can imitate one demonstrated answer but cannot rank it against another plausible answer. A chosen/rejected pair supplies that missing comparison. A reward model turns many comparisons into a reusable score for fresh attempts. DPO instead learns directly from the same kind of saved pair. DPO and online RLHF are alternative uses of preference evidence, not required consecutive stages.

Part II: Learning from new attempts

Scores for newly generated answers provide training feedback. Comparing those scores and limiting update incentives can reduce sensitivity to a batch, but cannot correct an unreliable scorer.

  • 3. PPO for scored attempts
    Scored answers can improve a policy without being organized into preference pairs. Reusing those answers saves generation work, but training must account for the policy changes made since they were sampled.
  • 4. GRPO and DAPO for reasoning
    Training several answers for each prompt trades extra generation work for a comparison between their scores.

Part III: Verifying results and choosing models

Task checks and tool results can supply rewards. Independent evaluation tests whether the trained model meets the task requirements.

  • 5. RL from tests and tool results
    Automated checks make reward design part of the software that runs a task.
  • 6. Choosing and evaluating training methods
    The earlier chapters connected each training method to the feedback it can use: demonstrated answers, preference pairs, scores for fresh attempts, and checks on results or tool actions. A project may have only some of that feedback. The training method must fit both the available data and the behavior the model needs to improve. The final choice is tested by comparing saved models under the same evaluation procedure.

Appendices

  • Course sources and further reading
    The lecture decks introduce the training methods, while the summaries and notebooks connect those ideas to calculations and software. The source map below identifies the strongest course references for each course topic. The research papers and versioned software documentation provide further detail beyond the classroom examples.
  • Glossary and software reference
    Look up training terms, software responsibilities, and the adapter parameter calculation after their method-level introductions.
  • References
    Identify the original papers, established reports, and official software references supporting the guide’s claims.