Meta-Reward-Net: Implicitly Differentiable Reward Learning for Preference-based Reinforcement Learning
Runze Liu, Fengshuo Bai, Yali Du, Yaodong Yang
Abstract
Setting up a well-designed reward function has been challenging for many reinforcement learning applications. Preference-based reinforcement learning (PbRL) provides a new framework that avoids reward engineering by leveraging human preferences (i.e., preferring apples over oranges) as the reward signal. Therefore, improving the efficacy of data usage for preference data becomes critical. In this work, we propose Meta-Reward-Net (MRN), a data-efficient PbRL framework that incorporates bi-level optimization for both reward and policy learning. The key idea of MRN is to adopt the performance of the Q-function as the learning target. Based on this, MRN learns the Q-function and the policy in the inner level while updating the reward function adaptively according to the performance of the Q-function on the preference data in the outer level. Our experiments on robotic simulated manipulation tasks and locomotion tasks demonstrate that MRN outperforms prior methods in the case of few preference labels and significantly improves data efficiency, achieving state-of-the-art in preference-based RL. Ablation studies further demonstrate that MRN learns a more accurate Q-function compared to prior work and shows obvious advantages when only a small amount of human feedback is available. The source code and videos of this project are released at https://sites.google.com/view/meta-reward-net 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers29
- GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative ReasoningJian Zhao, Runze Liu, Kaiyan Zhang, Zhimu Zhou et al.AAAI 2026 · 68 citations
- Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)Zhenjie Yang, Xiaosong Jia, Qifeng Li, Xue Yang et al.NeurIPS 2025 · 65 citations
- RIME: Robust Preference-based Reinforcement Learning with Noisy PreferencesJie Cheng, Gang Xiong, Xingyuan Dai, Qinghai Miao et al.ICML 2024 · 42 citations
- Reinforcing LLM Agents via Policy Optimization with Action DecompositionMuning Wen, Ziyu Wan, Jun Wang, Weinan Zhang et al.NeurIPS 2024 · 31 citations
- Sequential Preference Ranking for Efficient Reinforcement Learning from Human FeedbackMinyoung Hwang, Gunmin Lee, Hogun Kee, Chanwoo Kim et al.NeurIPS 2023 · 24 citations
Builds on12
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
- Meta Label Correction for Noisy Label LearningGuoqing Zheng, Ahmed Hassan Awadallah, Susan T. DumaisAAAI 2021 · 239 citations
- SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement LearningJongjin Park, Younggyo Seo, Jinwoo Shin, Honglak Lee et al.ICLR 2022 · 116 citations
- Bi-Level Actor-Critic for Multi-Agent CoordinationHaifeng Zhang, Weizhe Chen, Zeren Huang, Minne Li et al.AAAI 2020 · 113 citations
- What Can Learned Intrinsic Rewards Capture?Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu et al.ICML 2020 · 87 citations
Related papers
- From Reward-Free Representations to Preferences: Rethinking Offline Preference-Based Reinforcement LearningJun-Jie Yang, Chia-Heng Hsu, Kui-Yuan Chen, Ping-Chun HsiehICML 2026
- Inverse Preference Learning: Preference-based RL without a Reward FunctionJoey Hejna, Dorsa SadighNeurIPS 2023 · 92 citations
- Query-Policy Misalignment in Preference-Based Reinforcement LearningXiao Hu, Jianxiong Li, Xianyuan Zhan, Qing-Shan Jia et al.ICLR 2024 · 15 citations
- Decoding Global Preferences: Temporal and Cooperative Dependency Modeling in Multi-Agent Preference-Based Reinforcement LearningTianchen Zhu, Yue Qiu, Haoyi Zhou, Jianxin LiAAAI 2024 · 9 citations
- OPRIDE: Efficient Offline Preference-based Reinforcement Learning via In-Dataset ExplorationYiqin Yang, Hao Hu, Yihuan Mao, Jin Zhang et al.ICLR 2026
