Preference Elicitation for Offline Reinforcement Learning
Alizée Pace, Bernhard Schölkopf, Gunnar Rätsch, Giorgia Ramponi
摘要
Applying reinforcement learning (RL) to real-world problems is often made challenging by the inability to interact with the environment and the difficulty of designing reward functions. Offline RL addresses the first challenge by considering access to an offline dataset of environment interactions labeled by the reward function. In contrast, Preference-based RL does not assume access to the reward function and learns it from preferences, but typically requires an online interaction with the environment. We bridge the gap between these frameworks by exploring efficient methods for acquiring preference feedback in a fully offline setup. We propose Sim-OPRL, an offline preference-based reinforcement learning algorithm, which leverages a learned environment model to elicit preference feedback on simulated rollouts. Drawing on insights from both the offline RL and the preference-based RL literature, our algorithm employs a pessimistic approach for out-of-distribution data, and an optimistic approach for acquiring informative preferences about the optimal policy. We provide theoretical guarantees regarding the sample complexity of our approach, dependent on how well the offline data covers the optimal policy. Finally, we demonstrate the empirical performance of Sim-OPRL in different environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Adversarial Policy Optimization for Offline Preference-based Reinforcement LearningHyungkyu Kang, Min-hwan OhICLR 2025
- Offline Preference-Based Value OptimizationHyungkyu Kang, Min-hwan OhICLR 2026
- Toward Conservative Planning from Human-AI Preferences in Reinforcement LearningHuazhong Wang, Wenzhuo ZhouICLR 2026
它引用的顶会 Paper12
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 419 次
- Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise ComparisonsBanghua Zhu, Michael I. Jordan, Jiantao JiaoICML 2023 · 被引用 273 次
- Pessimistic Model-based Offline Reinforcement Learning under Partial CoverageMasatoshi Uehara, Wen SunICLR 2022 · 被引用 176 次
相关 Paper
- From Reward-Free Representations to Preferences: Rethinking Offline Preference-Based Reinforcement LearningJun-Jie Yang, Chia-Heng Hsu, Kui-Yuan Chen, Ping-Chun HsiehICML 2026
- Beyond Reward: Offline Preference-guided Policy OptimizationYachen Kang, Diyuan Shi, Jinxin Liu, Li He 等ICML 2023 · 被引用 41 次
- OPRIDE: Efficient Offline Preference-based Reinforcement Learning via In-Dataset ExplorationYiqin Yang, Hao Hu, Yihuan Mao, Jin Zhang 等ICLR 2026
- Direct Preference-based Policy Optimization without Reward ModelingGaon An, Junhyeok Lee, Xingdong Zuo, Norio Kosaka 等NeurIPS 2023 · 被引用 61 次
- Optimistic Model Rollouts for Pessimistic Offline Policy OptimizationYuanzhao Zhai, Yiying Li, Zijian Gao, Xudong Gong 等AAAI 2024 · 被引用 4 次
