DUO: Diverse, Uncertain, On-Policy Query Generation and Selection for Reinforcement Learning from Human Feedback
Xuening Feng, Zhaohui Jiang, Timo Kaufmann, Puchen Xu, Eyke Hüllermeier, Paul Weng, Yifei Zhu
摘要
Defining a reward function is usually a challenging but critical task for the system designer in reinforcement learning, especially when specifying complex behaviors. Reinforcement learning from human feedback (RLHF) emerges as a promising approach to circumvent this. In RLHF, the agent typically learns a reward function by querying a human teacher using pairwise comparisons of trajectory segments. A key question in this domain is how to reduce the number of queries necessary to learn an informative reward function, since asking a human teacher too many queries is impractical and costly. To tackle this question, most existing methods mainly focus on improving exploration, introducing data augmentation or designing sophisticated training objectives for RLHF, while the potential of query generation and selection schemes have not been fully exploited. In this paper, we propose DUO, a novel method for diverse, uncertain, on-policy query generation and selection in RLHF. Our method produces queries that are (1) more relevant for policy training (via an on-policy criterion), (2) more informative (via a principled measure of epistemic uncertainty), and (3) diverse (via a clustering-based filter). Experimental results on a variety of locomotion and robotic manipulation tasks demonstrate that our method can outperform state-of-the-art RLHF methods given the same total budget of queries while being robust to possibly irrational teachers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Policy Likelihood-based Query Sampling and Critic-Exploited Reset for Efficient Preference-based Reinforcement LearningJongkook Heo, Jaehoon Kim, Young Jae Lee, Min Gu Kwak 等ICLR 2026
- Progressive Learning with Human Feedback for Personalized Adaptive Video StreamingZhaohui Jiang, Xuening Feng, Tianchi Huang, Ruixiao Zhang 等ACM MM 2025
它引用的顶会 Paper9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Defining and Characterizing Reward GamingJoar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, David KruegerNeurIPS 2022 · 被引用 466 次
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- The Ingredients of Real World Robotic Reinforcement LearningHenry Zhu, Justin Yu, Abhishek Gupta, Dhruv Shah 等ICLR 2020 · 被引用 202 次
- SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement LearningJongjin Park, Younggyo Seo, Jinwoo Shin, Honglak Lee 等ICLR 2022 · 被引用 116 次
相关 Paper
- Comparing Comparisons: Informative and Easy Human Feedback with Distinguishability QueriesXuening Feng, Zhaohui Jiang, Timo Kaufmann, Eyke Hüllermeier 等ICML 2025
- Adaptive Preference Scaling for Reinforcement Learning with Human FeedbackIlgee Hong, Zichong Li, Alexander Bukharin, Yixiao Li 等NeurIPS 2024 · 被引用 23 次
- Mitigating Reward Overoptimization via Lightweight Uncertainty EstimationXiaoying Zhang, Jean-Francois Ton, Wei Shen, Hongning Wang 等NeurIPS 2024 · 被引用 11 次
- Efficient Preference-Based Reinforcement Learning: Randomized Exploration meets Experimental DesignAndreas Schlaginhaufen, Reda Ouhamma, Maryam KamgarpourNeurIPS 2025 · 被引用 4 次
- Sequential Preference Ranking for Efficient Reinforcement Learning from Human FeedbackMinyoung Hwang, Gunmin Lee, Hogun Kee, Chanwoo Kim 等NeurIPS 2023 · 被引用 24 次
