Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement Learning
Tenglong Liu, Yang Li, Yixing Lan, Hao Gao, Wei Pan, Xin Xu
Abstract
In offline reinforcement learning, the challenge of out-of-distribution (OOD) is pronounced. To address this, existing methods often constrain the learned policy through policy regularization. However, these methods often suffer from the issue of unnecessary conservativeness, hampering policy improvement. This occurs due to the indiscriminate use of all actions from the behavior policy that generates the offline dataset as constraints. The problem becomes particularly noticeable when the quality of the dataset is suboptimal. Thus, we propose Adaptive Advantage-guided Policy Regularization (A2PR), obtaining high-advantage actions from an augmented behavior policy combined with VAE to guide the learned policy. A2PR can select high-advantage actions that differ from those present in the dataset, while still effectively maintaining conservatism from OOD actions. This is achieved by harnessing the VAE capacity to generate samples matching the distribution of the data points. We theoretically prove that the improvement of the behavior policy is guaranteed. Besides, it effectively mitigates value overestimation with a bounded performance gap. Empirically, we conduct a series of experiments on the D4RL benchmark, where A2PR demonstrates state-of-the-art performance. Furthermore, experimental results on additional suboptimal mixed datasets reveal that A2PR exhibits superior performance. Code is available at https://github.com/ltlhuuu/A2PR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- Is Mamba Compatible with Trajectory Optimization in Offline Reinforcement Learning?Yang Dai, Oubo Ma, Longfei Zhang, Xingxing Liang et al.NeurIPS 2024 · 10 citations
- Dual-Robust Cross-Domain Offline Reinforcement Learning Against Dynamics ShiftsZhongjian Qiao, Rui Yang, Jiafei Lyu, Xiu Li et al.ICLR 2026 · 7 citations
- ReFORM: Reflected Flows for On-support Offline RL via Noise ManipulationSongyuan Zhang, Oswin So, H. M. Sabbir Ahmad, Eric Yang Yu et al.ICLR 2026 · 5 citations
- ADG: Ambient Diffusion-Guided Dataset Recovery for Corruption-Robust Offline Reinforcement LearningZeyuan Liu, Zhihe Yang, Jiawei Xu, Rui Yang et al.NeurIPS 2025 · 3 citations
- Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement LearningQingjun Wang, Hongtu Zhou, Hang Yu, Junqiao Zhao et al.ICLR 2026 · 1 citation
Builds on13
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Offline Reinforcement Learning with Fisher Divergence Critic RegularizationIlya Kostrikov, Rob Fergus, Jonathan Tompson, Ofir NachumICML 2021 · 350 citations
Related papers
- A2PO: Towards Effective Offline Reinforcement Learning from an Advantage-aware PerspectiveYunpeng Qing, Shunyu Liu, Jingyuan Cong, Kaixuan Chen et al.NeurIPS 2024 · 16 citations
- Supported Policy Optimization for Offline Reinforcement LearningJialong Wu, Haixu Wu, Zihan Qiu, Jianmin Wang et al.NeurIPS 2022 · 113 citations
- Policy Regularization with Dataset Constraint for Offline Reinforcement LearningYuhang Ran, Yi-Chen Li, Fuxiang Zhang, Zongzhang Zhang et al.ICML 2023 · 49 citations
- Diffusion Actor-Critic: Formulating Constrained Policy Iteration as Diffusion Noise Regression for Offline Reinforcement LearningLinjiajie Fang, Ruoxue Liu, Jing Zhang, Wenjia Wang et al.ICLR 2025
- Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RLQin-Wen Luo, Ming-Kun Xie, Ye-Wen Wang, Sheng-Jun HuangICML 2025
