RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences
Jie Cheng, Gang Xiong, Xingyuan Dai, Qinghai Miao, Yisheng Lv, Fei-Yue Wang
Abstract
Preference-based Reinforcement Learning (PbRL) circumvents the need for reward engineering by harnessing human preferences as the reward signal. However, current PbRL methods excessively depend on high-quality feedback from domain experts, which results in a lack of robustness. In this paper, we present RIME, a robust PbRL algorithm for effective reward learning from noisy preferences. Our method utilizes a sample selection-based discriminator to dynamically filter out noise and ensure robust training. To counteract the cumulative error stemming from incorrect selection, we suggest a warm start for the reward model, which additionally bridges the performance gap during the transition from pre-training to online training in PbRL. Our experiments on robotic manipulation and locomotion tasks demonstrate that RIME significantly enhances the robustness of the state-of-the-art PbRL method. Code is available at https://github.com/CJReinforce/ RIME_ICML2024 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d698dd8b-3c35-4ca0-938b-77aaeacdde2eCited by top-tier papers16
- Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for ReasoningJie Cheng, Gang Xiong, Ruixi Qiao, Lijun Li et al.NeurIPS 2025 · 56 citations
- Robust Reinforcement Learning from Corrupted Human FeedbackAlexander Bukharin, Ilgee Hong, Haoming Jiang, Zichong Li et al.NeurIPS 2024 · 30 citations
- Strategyproof Reinforcement Learning from Human FeedbackThomas Kleine Buening, Jiarui Gan, Debmalya Mandal, Marta KwiatkowskaNeurIPS 2025 · 10 citations
- VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-TrainingDingwei Zhu, Shihan Dou, Zhiheng Xi, Senjie Jin et al.ACL 2026 · 9 citations
- STAIR: Addressing Stage Misalignment through Temporal-Aligned Preference Reinforcement LearningYao Luan, Ni Mu, Yiqin Yang, Bo Xu et al.NeurIPS 2025 · 3 citations
Builds on15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Does label smoothing mitigate label noise?Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, Sanjiv KumarICML 2020 · 411 citations
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
- Robust early-learning: Hindering the memorization of noisy labelsXiaobo Xia, Tongliang Liu, Bo Han, Chen Gong et al.ICLR 2021 · 322 citations
- Behavior From the Void: Unsupervised Active Pre-TrainingHao Liu, Pieter AbbeelNeurIPS 2021 · 258 citations
Related papers
- PEARL: Zero-shot Cross-task Preference Alignment and Robust Reward Learning for Robotic ManipulationRunze Liu, Yali Du, Fengshuo Bai, Jiafei Lyu et al.ICML 2024 · 10 citations
- Policy Likelihood-based Query Sampling and Critic-Exploited Reset for Efficient Preference-based Reinforcement LearningJongkook Heo, Jaehoon Kim, Young Jae Lee, Min Gu Kwak et al.ICLR 2026
- PAWS: Preference Learning with Advantage-Weighted SegmentsAleksandar Taranovic, Onur Celik, Niklas Freymuth, Ge Li et al.ICML 2026
- From Reward-Free Representations to Preferences: Rethinking Offline Preference-Based Reinforcement LearningJun-Jie Yang, Chia-Heng Hsu, Kui-Yuan Chen, Ping-Chun HsiehICML 2026
- Limited Preference Aided Imitation Learning from Imperfect DemonstrationsXingchen Cao, Fan-Ming Luo, Junyin Ye, Tian Xu et al.ICML 2024 · 6 citations
