Behaviour Preference Regression for Offline Reinforcement Learning
Padmanaba Srinivasan, William Knottenbelt
Abstract
Offline reinforcement learning (RL) methods aim to learn optimal policies with access only to trajectories in a fixed dataset. Policy constraint methods formulate policy learning as an optimization problem that balances maximizing reward with minimizing deviation from the behavior policy. Closed form solutions to this problem can be derived as weighted behavioral cloning objectives that, in theory, must compute an intractable partition function. Reinforcement learning has gained popularity in language modeling to align models with human preferences; some recent works consider paired completions that are ranked by a preference model following which the likelihood of the preferred completion is directly increased. We adapt this approach of paired comparison. By reformulating the paired-sample optimization problem, we fit the maximum-mode of the Q function while maximizing behavioral consistency of policy actions. This yields our algorithm, Behavior Preference Regression for offline RL (BPR). We empirically evaluate BPR on the widely used D4RL Locomotion and Antmaze datasets, as well as the more challenging V-D4RL suite, which operates in image-based state spaces. BPR demonstrates state-of-the-art performance over all domains. Our on-policy experiments suggest that BPR takes advantage of the stability of on-policy value functions with minimal perceptible performance degradation on Locomotion datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on23
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
Related papers
- Policy Regularization with Dataset Constraint for Offline Reinforcement LearningYuhang Ran, Yi-Chen Li, Fuxiang Zhang, Zongzhang Zhang et al.ICML 2023 · 49 citations
- Weighted Policy Constraints for Offline Reinforcement LearningZhiyong Peng, Changlin Han, Yadong Liu, Zongtan ZhouAAAI 2023 · 18 citations
- Offline Reinforcement Learning with Closed-Form Policy Improvement OperatorsJiachen Li, Edwin Zhang, Ming Yin, Qinxun Bai et al.ICML 2023 · 18 citations
- Beyond Uniform Sampling: Offline Reinforcement Learning with Imbalanced DatasetsZhang-Wei Hong, Aviral Kumar, Sathwik Karnik, Abhishek Bhandwaldar et al.NeurIPS 2023 · 34 citations
- Direct Preference-based Policy Optimization without Reward ModelingGaon An, Junhyeok Lee, Xingdong Zuo, Norio Kosaka et al.NeurIPS 2023 · 61 citations
