Beyond Reward: Offline Preference-guided Policy Optimization
Yachen Kang, Diyuan Shi, Jinxin Liu, Li He, Donglin Wang
Abstract
This study focuses on the topic of offline preference-based reinforcement learning (PbRL), a variant of conventional reinforcement learning that dispenses with the need for online interaction or specification of reward functions. Instead, the agent is provided with fixed offline trajectories and human preferences between pairs of trajectories to extract the dynamics and task information, respectively. Since the dynamics and task information are orthogonal, a naive approach would involve using preference-based reward learning followed by an off-the-shelf offline RL algorithm. However, this requires the separate learning of a scalar reward function, which is assumed to be an information bottleneck of the learning process. To address this issue, we propose the offline preference-guided policy optimization (OPPO) paradigm, which models offline trajectories and preferences in a one-step process, eliminating the need for separately learning a reward function. OPPO achieves this by introducing an offline hindsight information matching objective for optimizing a contextual policy and a preference modeling objective for finding the optimal context. OPPO further integrates a well-performing decision policy by optimizing the two objectives iteratively. Our empirical results demonstrate that OPPO effectively models offline preferences and outperforms prior competing baselines, including offline RL algorithms performed over either true or pseudo reward function specifications. Our code is available on the project website: https://sites.google. com/view/oppo-icml-2023 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f702ac78-1665-48b9-8edf-7d7a58c0ffd6Cited by top-tier papers29
- Inverse Preference Learning: Preference-based RL without a Reward FunctionJoey Hejna, Dorsa SadighNeurIPS 2023 · 92 citations
- ChiPFormer: Transferable Chip Placement via Offline Decision TransformerYao Lai, Jinxin Liu, Zhentao Tang, Bin Wang et al.ICML 2023 · 69 citations
- Direct Preference-based Policy Optimization without Reward ModelingGaon An, Junhyeok Lee, Xingdong Zuo, Norio Kosaka et al.NeurIPS 2023 · 61 citations
- Contrastive Preference Learning: Learning from Human Feedback without Reinforcement LearningJoey Hejna, Rafael Rafailov, Harshit Sikchi, Chelsea Finn et al.ICLR 2024 · 37 citations
- Beyond OOD State Actions: Supported Cross-Domain Offline Reinforcement LearningJinxin Liu, Ziqi Zhang, Zhenyu Wei, Zifeng Zhuang et al.AAAI 2024 · 30 citations
Builds on19
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
Related papers
- Adversarial Policy Optimization for Offline Preference-based Reinforcement LearningHyungkyu Kang, Min-hwan OhICLR 2025
- A Distributional Approach to Uncertainty-Aware Preference Alignment Using Offline DemonstrationsSheng Xu, Bo Yue, Hongyuan Zha, Guiliang LiuICLR 2025
- Toward Conservative Planning from Human-AI Preferences in Reinforcement LearningHuazhong Wang, Wenzhuo ZhouICLR 2026
- Flow to Better: Offline Preference-based Reinforcement Learning via Preferred Trajectory GenerationZhilong Zhang, Yihao Sun, Junyin Ye, Tian-Shuo Liu et al.ICLR 2024 · 23 citations
- Offline Preference-Based Value OptimizationHyungkyu Kang, Min-hwan OhICLR 2026
