Meta-Reward-Net: Implicitly Differentiable Reward Learning for Preference-based Reinforcement Learning
Runze Liu, Fengshuo Bai, Yali Du, Yaodong Yang
摘要
Setting up a well-designed reward function has been challenging for many reinforcement learning applications. Preference-based reinforcement learning (PbRL) provides a new framework that avoids reward engineering by leveraging human preferences (i.e., preferring apples over oranges) as the reward signal. Therefore, improving the efficacy of data usage for preference data becomes critical. In this work, we propose Meta-Reward-Net (MRN), a data-efficient PbRL framework that incorporates bi-level optimization for both reward and policy learning. The key idea of MRN is to adopt the performance of the Q-function as the learning target. Based on this, MRN learns the Q-function and the policy in the inner level while updating the reward function adaptively according to the performance of the Q-function on the preference data in the outer level. Our experiments on robotic simulated manipulation tasks and locomotion tasks demonstrate that MRN outperforms prior methods in the case of few preference labels and significantly improves data efficiency, achieving state-of-the-art in preference-based RL. Ablation studies further demonstrate that MRN learns a more accurate Q-function compared to prior work and shows obvious advantages when only a small amount of human feedback is available. The source code and videos of this project are released at https://sites.google.com/view/meta-reward-net 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative ReasoningJian Zhao, Runze Liu, Kaiyan Zhang, Zhimu Zhou 等AAAI 2026 · 被引用 68 次
- Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)Zhenjie Yang, Xiaosong Jia, Qifeng Li, Xue Yang 等NeurIPS 2025 · 被引用 65 次
- RIME: Robust Preference-based Reinforcement Learning with Noisy PreferencesJie Cheng, Gang Xiong, Xingyuan Dai, Qinghai Miao 等ICML 2024 · 被引用 42 次
- Reinforcing LLM Agents via Policy Optimization with Action DecompositionMuning Wen, Ziyu Wan, Jun Wang, Weinan Zhang 等NeurIPS 2024 · 被引用 31 次
- Sequential Preference Ranking for Efficient Reinforcement Learning from Human FeedbackMinyoung Hwang, Gunmin Lee, Hogun Kee, Chanwoo Kim 等NeurIPS 2023 · 被引用 24 次
它引用的顶会 Paper12
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- Meta Label Correction for Noisy Label LearningGuoqing Zheng, Ahmed Hassan Awadallah, Susan T. DumaisAAAI 2021 · 被引用 239 次
- SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement LearningJongjin Park, Younggyo Seo, Jinwoo Shin, Honglak Lee 等ICLR 2022 · 被引用 116 次
- Bi-Level Actor-Critic for Multi-Agent CoordinationHaifeng Zhang, Weizhe Chen, Zeren Huang, Minne Li 等AAAI 2020 · 被引用 113 次
- What Can Learned Intrinsic Rewards Capture?Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu 等ICML 2020 · 被引用 87 次
相关 Paper
- From Reward-Free Representations to Preferences: Rethinking Offline Preference-Based Reinforcement LearningJun-Jie Yang, Chia-Heng Hsu, Kui-Yuan Chen, Ping-Chun HsiehICML 2026
- Inverse Preference Learning: Preference-based RL without a Reward FunctionJoey Hejna, Dorsa SadighNeurIPS 2023 · 被引用 92 次
- Query-Policy Misalignment in Preference-Based Reinforcement LearningXiao Hu, Jianxiong Li, Xianyuan Zhan, Qing-Shan Jia 等ICLR 2024 · 被引用 15 次
- Decoding Global Preferences: Temporal and Cooperative Dependency Modeling in Multi-Agent Preference-Based Reinforcement LearningTianchen Zhu, Yue Qiu, Haoyi Zhou, Jianxin LiAAAI 2024 · 被引用 9 次
- OPRIDE: Efficient Offline Preference-based Reinforcement Learning via In-Dataset ExplorationYiqin Yang, Hao Hu, Yihuan Mao, Jin Zhang 等ICLR 2026
