All Roads Lead to Likelihood: The Value of Reinforcement Learning in Fine-Tuning
Gokul Swamy, Sanjiban Choudhury, Wen Sun, Steven Wu, Drew Bagnell
Abstract
From a first-principles perspective, it may seem odd that the strongest results in foundation model fine-tuning (FT) are achieved via a relatively complex, two-stage training procedure. Specifically, one first trains a reward model (RM) on some dataset (e.g., human preferences) before using it to provide online feedback as part of a downstream reinforcement learning (RL) procedure, rather than directly optimizing the policy parameters on said dataset via offline maximum likelihood estimation. In fact, from an information-theoretic perspective, we can only lose information via passing through a reward model and cannot create any new information via on-policy sampling. To explain this discrepancy, we scrutinize several hypotheses on the value of RL in FT through both theoretical and empirical lenses. Of the hypotheses considered, we find the most support for the explanation that on problems with a generation-verification gap, (1) it is relatively easy to learn the relatively simple RM (verifier) from the preference data. Then, (2) the downstream RL procedure only returns policies (generators) that are optimal for such relatively simple verifiers. Thus, end-to-end, two-stage online FT only has to search over a reduced subset of the full space of policies, requiring less data than offline FT. INTRODUCTION Whether one refers to it as reinforcement learning from human feedback (RLHF, Christiano et al. ( 2017 )), preference fine-tuning (PFT), or even "alignment," the last step in the training pipeline of a wide variety of foundation models (FMs) is fundamentally concerned with raising the generation likelihood of preferred completions of a prompt relative to those of dis-preferred completions. From this perspective, a natural question may be why anything other than maximum likelihood estimation (MLE) -i.e., standard supervised learning -is needed for the PFT problem. Indeed, a plethora of offline approaches to PFT that directly optimize policy parameters via solving a (regularized) classification problem on preference data have been proposed in the literature (e.g., DPO (Rafailov et al., 2023), IPO (Azar et al., 2023), SLiC-HF (Zhao et al., 2023)). However, when one looks at the training procedure of today's most capable models (Achiam et al., 2023; Team et al., 2024; Dubey et al., 2024) , one almost always sees a relatively complex twostage procedure adopted instead. First, one learns a reward model (RM) -i.e., a classifier -on the preference data, before using it to provide labels for a downstream online reinforcement learning (RL) procedure that ultimately optimizes the policy's parameters (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9e6ca0a-31ed-46c9-aab7-0e61b94f722cCited by top-tier papers36
- WebDancer: Towards Autonomous Information Seeking AgencyJialong Wu, Baixuan Li, Runnan Fang, Wenbiao Yin et al.NeurIPS 2025 · 194 citations
- On the Generalization of SFT: A Reinforcement Learning Perspective with Reward RectificationYongliang Wu, Yizhou Zhou, Ziheng Zhou, Yingzhe Peng et al.ICLR 2026 · 130 citations
- Checklists Are Better Than Reward Models For Aligning Language ModelsVijay Viswanathan, Yanchao Sun, Xiang Kong, Meng Cao et al.NeurIPS 2025 · 127 citations
- e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMsAmrith Setlur, Matthew Y. R. Yang, Charlie Victor Snell, Jeremiah Greer et al.ICLR 2026 · 66 citations
- Energy-Based Transformers are Scalable Learners and ThinkersAlexi Gladstone, Ganesh Nanduru, Md Mofijul Islam, Peixuan Han et al.ICLR 2026 · 38 citations
Builds on30
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
Related papers
- Joint Reward and Policy Learning with Demonstrations and Human Feedback Improves AlignmentChenliang Li, Siliang Zeng, Zeyi Liao, Jiaxiang Li et al.ICLR 2025
- Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT for LLM AlignmentJiaxiang Li, Siliang Zeng, Hoi-To Wai, Chenliang Li et al.NeurIPS 2024 · 44 citations
- Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPORuizhe Shi, Minhak Song, Runlong Zhou, Zihan Zhang et al.ICML 2026
- Is On-Policy Data always the Best Choice for Direct Preference Optimization-Based LM Alignment?Zetian Sun, Dongfang Li, Xuhui Chen, Baotian Hu et al.ICLR 2026 · 1 citation
- Why DPO is a Misspecified Estimator and How to Fix ItAditya Gopalan, Sayak Ray Chowdhury, Debangshu BanerjeeICLR 2026 · 4 citations
