Direct Preference-based Policy Optimization without Reward Modeling
Gaon An, Junhyeok Lee, Xingdong Zuo, Norio Kosaka, Kyung-Min Kim, Hyun Oh Song
Abstract
Preference-based reinforcement learning (PbRL) is an approach that enables RL agents to learn from preference, which is particularly useful when formulating a reward function is challenging. Existing PbRL methods generally involve a two-step procedure: they first learn a reward model based on given preference data and then employ off-the-shelf reinforcement learning algorithms using the learned reward model. However, obtaining an accurate reward model solely from preference information, especially when the preference is from human teachers, can be difficult. Instead, we propose a PbRL algorithm that directly learns from preference without requiring any reward modeling. To achieve this, we adopt a contrastive learning framework to design a novel policy scoring metric that assigns a high score to policies that align with the given preferences. We apply our algorithm to offline RL tasks with actual human preference labels and show that our algorithm outperforms or is on par with the existing PbRL methods. Notably, on high-dimensional control tasks, our algorithm surpasses offline RL methods that learn with ground-truth reward information. Finally, we show that our algorithm can be successfully applied to fine-tune large language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8cd05347-a30a-4c0e-b9dd-ff88f71d6abfCited by top-tier papers28
- Contrastive Preference Learning: Learning from Human Feedback without Reinforcement LearningJoey Hejna, Rafael Rafailov, Harshit Sikchi, Chelsea Finn et al.ICLR 2024 · 37 citations
- PICLe: Eliciting Diverse Behaviors from Large Language Models with Persona In-Context LearningHyeong Kyu Choi, Yixuan LiICML 2024 · 31 citations
- Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human PreferencesAndi Nika, Debmalya Mandal, Parameswaran Kamalaruban, Georgios Tzannetos et al.ICML 2024 · 22 citations
- Listwise Reward Estimation for Offline Preference-based Reinforcement LearningHeewoong Choi, Sangwon Jung, Hongjoon Ahn, Taesup MoonICML 2024 · 14 citations
- Forward KL Regularized Preference Optimization for Aligning Diffusion PoliciesZhao Shan, Chenyou Fan, Shuang Qiu, Jiyuan Shi et al.AAAI 2025 · 8 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
Related papers
- From Reward-Free Representations to Preferences: Rethinking Offline Preference-Based Reinforcement LearningJun-Jie Yang, Chia-Heng Hsu, Kui-Yuan Chen, Ping-Chun HsiehICML 2026
- Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy DataFahim Tajwar, Anikait Singh, Archit Sharma, Rafael Rafailov et al.ICML 2024 · 189 citations
- Cal-DPO: Calibrated Direct Preference Optimization for Language Model AlignmentTeng Xiao, Yige Yuan, Huaisheng Zhu, Mingxiao Li et al.NeurIPS 2024 · 76 citations
- Provable Reward-Agnostic Preference-Based Reinforcement LearningWenhao Zhan, Masatoshi Uehara, Wen Sun, Jason D. LeeICLR 2024 · 16 citations
- Beyond Reward: Offline Preference-guided Policy OptimizationYachen Kang, Diyuan Shi, Jinxin Liu, Li He et al.ICML 2023 · 41 citations
