Rating-Based Reinforcement Learning
Devin White, Mingkang Wu, Ellen R. Novoseller, Vernon J. Lawhern, Nicholas R. Waytowich, Yongcan Cao
摘要
This paper develops a novel rating-based reinforcement learning (RbRL) approach that uses human ratings to obtain human guidance in reinforcement learning. Different from the existing preference-based and ranking-based reinforcement learning paradigms, based on human relative preferences over sample pairs, the proposed rating-based reinforcement learning approach is based on human evaluation of individual trajectories without relative comparisons between sample pairs. The rating-based reinforcement learning approach builds on a new prediction model for human ratings and a novel multiclass loss function. We finally conduct several experimental studies based on synthetic ratings and real human ratings to evaluate the performance of the new rating-based reinforcement learning approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human FeedbackYifu Yuan, Jianye Hao, Yi Ma, Zibin Dong 等ICLR 2024 · 被引用 21 次
- Listwise Reward Estimation for Offline Preference-based Reinforcement LearningHeewoong Choi, Sangwon Jung, Hongjoon Ahn, Taesup MoonICML 2024 · 被引用 14 次
- Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language ModelsTung Minh Luu, Younghwan Lee, Donghoon Lee, Sunho Kim 等ICML 2025
- CoRe: Combined Rewards with Vision-Language Model Feedback for Preference-Aligned Reinforcement LearningHexian Ni, Tao Lu, Yinghao CaiICML 2026
- Leveraging Sub-Optimal Data for Human-in-the-Loop Reinforcement LearningCalarina Muslimani, Matthew E. TaylorICLR 2025
它引用的顶会 Paper6
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement LearningJongjin Park, Younggyo Seo, Jinwoo Shin, Honglak Lee 等ICLR 2022 · 被引用 116 次
- Safe Imitation Learning via Fast Bayesian Reward Inference from PreferencesDaniel S. Brown, Russell Coleman, Ravi Srinivasan, Scott NiekumICML 2020 · 被引用 113 次
- Preference-based Reinforcement Learning with Finite-Time GuaranteesYichong Xu, Ruosong Wang, Lin F. Yang, Aarti Singh 等NeurIPS 2020 · 被引用 82 次
- BC-IRL: Learning Generalizable Reward Functions from DemonstrationsAndrew Szot, Amy Zhang, Dhruv Batra, Zsolt Kira 等ICLR 2023 · 被引用 1 次
相关 Paper
- Reward Learning through Ranking Mean Squared ErrorChaitanya Kharyal, Calarina Muslimani, Matthew TaylorICML 2026
- Beyond Reward: Offline Preference-guided Policy OptimizationYachen Kang, Diyuan Shi, Jinxin Liu, Li He 等ICML 2023 · 被引用 41 次
- Human-in-the-loop: Provably Efficient Preference-based Reinforcement Learning with General Function ApproximationXiaoyu Chen, Han Zhong, Zhuoran Yang, Zhaoran Wang 等ICML 2022 · 被引用 90 次
- Provable Reward-Agnostic Preference-Based Reinforcement LearningWenhao Zhan, Masatoshi Uehara, Wen Sun, Jason D. LeeICLR 2024 · 被引用 16 次
- Sequential Preference Ranking for Efficient Reinforcement Learning from Human FeedbackMinyoung Hwang, Gunmin Lee, Hogun Kee, Chanwoo Kim 等NeurIPS 2023 · 被引用 24 次
