Query-Policy Misalignment in Preference-Based Reinforcement Learning
Xiao Hu, Jianxiong Li, Xianyuan Zhan, Qing-Shan Jia, Ya-Qin Zhang
Abstract
Preference-based reinforcement learning (PbRL) provides a natural way to align RL agents' behavior with human desired outcomes, but is often restrained by costly human feedback. To improve feedback efficiency, most existing PbRL methods focus on selecting queries to maximally improve the overall quality of the reward model, but counter-intuitively, we find that this may not necessarily lead to improved performance. To unravel this mystery, we identify a long-neglected issue in the query selection schemes of existing PbRL studies: Query-Policy Misalignment. We show that the seemingly informative queries selected to improve the overall quality of reward model actually may not align with RL agents' interests, thus offering little help on policy learning and eventually resulting in poor feedback efficiency. We show that this issue can be effectively addressed via policy-aligned query and a specially designed hybrid experience replay, which together enforce the bidirectional query-policy alignment. Simple yet elegant, our method can be easily incorporated into existing approaches by changing only a few lines of code. We showcase in comprehensive experiments that our method achieves substantial gains in both human feedback and RL sample efficiency, demonstrating the importance of addressing query-policy misalignment in PbRL tasks.Code is available at https://github.com/huxiao09/QPA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98378ffb-7763-4254-9256-2e02c952dc11Cited by top-tier papers14
- Instruction-Guided Visual MaskingJinliang Zheng, Jianxiong Li, Sijie Cheng, Yinan Zheng et al.NeurIPS 2024 · 22 citations
- Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human FeedbackYifu Yuan, Jianye Hao, Yi Ma, Zibin Dong et al.ICLR 2024 · 21 citations
- DecisionNCE: Embodied Multimodal Representations via Implicit Preference LearningJianxiong Li, Jinliang Zheng, Yinan Zheng, Liyuan Mao et al.ICML 2024 · 19 citations
- DUO: Diverse, Uncertain, On-Policy Query Generation and Selection for Reinforcement Learning from Human FeedbackXuening Feng, Zhaohui Jiang, Timo Kaufmann, Puchen Xu et al.AAAI 2025 · 7 citations
- STAIR: Addressing Stage Misalignment through Temporal-Aligned Preference Reinforcement LearningYao Luan, Ni Mu, Yiqin Yang, Bo Xu et al.NeurIPS 2025 · 3 citations
Builds on14
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto et al.NeurIPS 2020 · 833 citations
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
- The Value Equivalence Principle for Model-Based Reinforcement LearningChristopher Grimm, André Barreto, Satinder Singh, David SilverNeurIPS 2020 · 129 citations
Related papers
- Policy Likelihood-based Query Sampling and Critic-Exploited Reset for Efficient Preference-based Reinforcement LearningJongkook Heo, Jaehoon Kim, Young Jae Lee, Min Gu Kwak et al.ICLR 2026
- Improving Reward Models with Proximal Policy Exploration for Preference-Based Reinforcement LearningYiwen Zhu, Jinyi Liu, Pengjie Gu, Yifu Yuan et al.NeurIPS 2025 · 1 citation
- Reward Uncertainty for Exploration in Preference-based Reinforcement LearningXinran Liang, Katherine Shu, Kimin Lee, Pieter AbbeelICLR 2022
- STAR: Efficient Preference-based Reinforcement Learning via Dual RegularizationFengshuo Bai, Rui Zhao, Hongming Zhang, Sijia Cui et al.NeurIPS 2025 · 13 citations
- OPRIDE: Efficient Offline Preference-based Reinforcement Learning via In-Dataset ExplorationYiqin Yang, Hao Hu, Yihuan Mao, Jin Zhang et al.ICLR 2026
