Supported Policy Optimization for Offline Reinforcement Learning
Jialong Wu, Haixu Wu, Zihan Qiu, Jianmin Wang, Mingsheng Long
Abstract
Policy constraint methods to offline reinforcement learning (RL) typically utilize parameterization or regularization that constrains the policy to perform actions within the support set of the behavior policy. The elaborative designs of parameterization methods usually intrude into the policy networks, which may bring extra inference cost and cannot take full advantage of well-established online methods. Regularization methods reduce the divergence between the learned policy and the behavior policy, which may mismatch the inherent density-based definition of support set thereby failing to avoid the out-of-distribution actions effectively. This paper presents Supported Policy OpTimization (SPOT), which is directly derived from the theoretical formalization of the density-based support constraint. SPOT adopts a VAE-based density estimator to explicitly model the support set of behavior policy and presents a simple but effective density-based regularization term, which can be plugged non-intrusively into off-the-shelf off-policy RL algorithms. SPOT achieves the state-of-the-art performance on standard benchmarks for offline RL. Benefiting from the pluggable design, offline pretrained models from SPOT can also be applied to perform online fine-tuning seamlessly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2dd460ed-9cda-43f0-bfeb-e0ade401b550Cited by top-tier papers61
- Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-TuningMitsuhiko Nakamoto, Simon Zhai, Anikait Singh, Max Sobol Mark et al.NeurIPS 2023 · 296 citations
- Revisiting the Minimalist Approach to Offline Reinforcement LearningDenis Tarasov, Vladislav Kurenkov, Alexander Nikulin, Sergey KolesnikovNeurIPS 2023 · 148 citations
- Score Regularized Policy Optimization through Diffusion BehaviorHuayu Chen, Cheng Lu, Zhengyi Wang, Hang Su et al.ICLR 2024 · 59 citations
- Counterfactual Conservative Q Learning for Offline Multi-agent Reinforcement LearningJianzhun Shao, Yun Qu, Chen Chen, Hongchang Zhang et al.NeurIPS 2023 · 56 citations
- Policy Regularization with Dataset Constraint for Offline Reinforcement LearningYuhang Ran, Yi-Chen Li, Fuxiang Zhang, Zongzhang Zhang et al.ICML 2023 · 49 citations
Builds on16
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
Related papers
- Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement LearningTenglong Liu, Yang Li, Yixing Lan, Hao Gao et al.ICML 2024 · 15 citations
- Supported Trust Region Optimization for Offline Reinforcement LearningYixiu Mao, Hongchang Zhang, Chen Chen, Yi Xu et al.ICML 2023 · 24 citations
- Supported Value Regularization for Offline Reinforcement LearningYixiu Mao, Hongchang Zhang, Chen Chen, Yi Xu et al.NeurIPS 2023 · 37 citations
- Constrained Policy Optimization with Explicit Behavior Density For Offline Reinforcement LearningJing Zhang, Chi Zhang, Wenjia Wang, Bingyi JingNeurIPS 2023 · 19 citations
- VOCE: Variational Optimization with Conservative Estimation for Offline Safe Reinforcement LearningJiayi Guan, Guang Chen, Jiaming Ji, Long Yang et al.NeurIPS 2023 · 19 citations
