Listwise Reward Estimation for Offline Preference-based Reinforcement Learning
Heewoong Choi, Sangwon Jung, Hongjoon Ahn, Taesup Moon
摘要
In Reinforcement Learning (RL), designing precise reward functions remains to be a challenge, particularly when aligning with human intent. Preference-based RL (PbRL) was introduced to address this problem by learning reward models from human feedback. However, existing PbRL methods have limitations as they often overlook the second-order preference that indicates the relative strength of preference. In this paper, we propose Listwise Reward Estimation (LiRE), a novel approach for offline PbRL that leverages second-order preference information by constructing a Ranked List of Trajectories (RLT), which can be efficiently built by using the same ternary feedback type as traditional methods. To validate the effectiveness of LiRE, we propose a new offline PbRL dataset that objectively reflects the effect of the estimated rewards. Our extensive experiments on the dataset demonstrate the superiority of LiRE, i.e., outperforming state-of-the-art baselines even with modest feedback budgets and enjoying robustness with respect to the number of feedbacks and feedback noise. Our code is available at https://github.com/chwoong/LiRE
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- R-WoM: Retrieval-augmented World Model For Computer-use AgentsKai Mei, Jiang Guo, Shuaichen Chang, Mingwen Dong 等ICLR 2026 · 被引用 12 次
- ResponseRank: Data-Efficient Reward Modeling through Preference Strength LearningTimo Kaufmann, Yannick Metz, Daniel A. Keim, Eyke HüllermeierNeurIPS 2025 · 被引用 2 次
- Uncertainty-aware Preference Alignment for Diffusion PoliciesRunqing Miao, Sheng Xu, Runyi Zhao, Wai Kin (Victor) Chan 等NeurIPS 2025 · 被引用 2 次
- COLLIE: Guiding Skill Discovery in Semantically Coherent Latent SpaceYao Luan, Ni Mu, Hanfei Ge, Yiqin Yang 等ICML 2026 · 被引用 1 次
- Video-Based Optimal Transport for Feedback-Efficient Offline Preference-Based Reinforcement LearningMinh-Tung Luu, Hwanhee Kim, Younghwan Lee, Chang D. YooICML 2026 · 被引用 1 次
它引用的顶会 Paper15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- Preference Ranking Optimization for Human AlignmentFeifan Song, Bowen Yu, Minghao Li, Haiyang Yu 等AAAI 2024 · 被引用 357 次
- Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise ComparisonsBanghua Zhu, Michael I. Jordan, Jiantao JiaoICML 2023 · 被引用 273 次
相关 Paper
- OPRIDE: Efficient Offline Preference-based Reinforcement Learning via In-Dataset ExplorationYiqin Yang, Hao Hu, Yihuan Mao, Jin Zhang 等ICLR 2026
- From Reward-Free Representations to Preferences: Rethinking Offline Preference-Based Reinforcement LearningJun-Jie Yang, Chia-Heng Hsu, Kui-Yuan Chen, Ping-Chun HsiehICML 2026
- Provable Offline Preference-Based Reinforcement LearningWenhao Zhan, Masatoshi Uehara, Nathan Kallus, Jason D. Lee 等ICLR 2024 · 被引用 50 次
- Beyond Reward: Offline Preference-guided Policy OptimizationYachen Kang, Diyuan Shi, Jinxin Liu, Li He 等ICML 2023 · 被引用 41 次
- Direct Preference-based Policy Optimization without Reward ModelingGaon An, Junhyeok Lee, Xingdong Zuo, Norio Kosaka 等NeurIPS 2023 · 被引用 61 次
