DASCO: Dual-Generator Adversarial Support Constrained Offline Reinforcement Learning
Quan Vuong, Aviral Kumar, Sergey Levine, Yevgen Chebotar
Abstract
In offline RL, constraining the learned policy to remain close to the data is essential 1 to prevent the policy from outputting out-of-distribution (OOD) actions with erro-2 neously overestimated values. In principle, generative adversarial networks (GAN) 3 can provide an elegant solution to do so, with the discriminator directly providing 4 a probability that quantifies distributional shift. However, in practice, GAN-based 5 offline RL methods have not outperformed alternative approaches, perhaps because 6 the generator is trained to both fool the discriminator and maximize return – two 7 objectives that are often at odds with each other. In this paper, we show that the 8 issue of conflicting objectives can be resolved by training two generators: one that 9 maximizes return, with the other capturing the “remainder” of the data distribution 10 in the offline dataset, such that the mixture of the two is close to the behavior policy. 11 We show that not only does having two generators enable an effective GAN-based 12 offline RL method, but also approximates a support constraint, where the policy 13 does not need to match the entire data distribution, but only the slice of the data 14 that leads to high long term performance. We name our method DASCO, for 15 D ual-Generator A dversarial S upport C onstrained O ffline RL. On benchmark tasks 16 that require learning from sub-optimal data, DASCO significantly outperforms 17 prior methods that enforce distribution constraint. 18
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2325947a-cccf-41c2-83d3-ec7dc5cd2615Cited by top-tier papers6
- Beyond Uniform Sampling: Offline Reinforcement Learning with Imbalanced DatasetsZhang-Wei Hong, Aviral Kumar, Sathwik Karnik, Abhishek Bhandwaldar et al.NeurIPS 2023 · 34 citations
- Offline Reinforcement Learning with Closed-Form Policy Improvement OperatorsJiachen Li, Edwin Zhang, Ming Yin, Qinxun Bai et al.ICML 2023 · 18 citations
- A2PO: Towards Effective Offline Reinforcement Learning from an Advantage-aware PerspectiveYunpeng Qing, Shunyu Liu, Jingyuan Cong, Kaixuan Chen et al.NeurIPS 2024 · 16 citations
- Rethinking Optimal Transport in Offline Reinforcement LearningArip Asadulaev, Rostislav Korst, Aleksandr Korotin, Vage Egiazarian et al.NeurIPS 2024 · 12 citations
- Enhancing Diffusion Policies with Distribution-Matching Generator in Offline Reinforcement LearningXuemin Hu, Shen Li, Yingfen Xu, Bo Tang et al.AAAI 2026 · 1 citation
Builds on13
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
Related papers
- Weighted Policy Constraints for Offline Reinforcement LearningZhiyong Peng, Changlin Han, Yadong Liu, Zongtan ZhouAAAI 2023 · 18 citations
- Reining Generalization in Offline Reinforcement Learning via Representation DistinctionYi Ma, Hongyao Tang, Dong Li, Zhaopeng MengNeurIPS 2023 · 19 citations
- Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement LearningTenglong Liu, Yang Li, Yixing Lan, Hao Gao et al.ICML 2024 · 15 citations
- Policy Regularization with Dataset Constraint for Offline Reinforcement LearningYuhang Ran, Yi-Chen Li, Fuxiang Zhang, Zongzhang Zhang et al.ICML 2023 · 49 citations
- Dual RL: Unification and New Methods for Reinforcement and Imitation LearningHarshit Sikchi, Qinqing Zheng, Amy Zhang, Scott NiekumICLR 2024 · 48 citations
