Efficient Off-Policy Learning for High-Dimensional Action Spaces
Fabian Otto, Philipp Becker, Ngo Anh Vien, Gerhard Neumann
摘要
Existing off-policy reinforcement learning algorithms often rely on an explicit stateaction-value function representation, which can be problematic in high-dimensional action spaces due to the curse of dimensionality. This reliance results in data inefficiency as maintaining a state-action-value function in such spaces is challenging. We present an efficient approach that utilizes only a state-value function as the critic for off-policy deep reinforcement learning. This approach, which we refer to as Vlearn, effectively circumvents the limitations of existing methods by eliminating the necessity for an explicit state-action-value function. To this end, we leverage a weighted importance sampling loss for learning deep value functions from off-policy data. While this is common for linear methods, it has not been combined with deep value function networks. This transfer to deep methods is not straightforward and requires novel design choices such as robust policy updates, twin value function networks to avoid an optimization bias, and importance weight clipping. We also present a novel analysis of the variance of our estimate compared to commonly used importance sampling estimators such as V-trace. Our approach improves sample complexity as well as final performance and ensures consistent and robust performance across various benchmark tasks. Eliminating the state-action-value function in Vlearn facilitates a streamlined learning process, yielding high-return agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Stop Regressing: Training Value Functions via Classification for Scalable Deep RLJesse Farebrother, Jordi Orbay, Quan Vuong, Adrien Ali Taïga 等ICML 2024 · 被引用 118 次
- CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and SimplicityAditya Bhatt, Daniel Palenicek, Boris Belousov, Max Argus 等ICLR 2024 · 被引用 106 次
- Randomized Ensembled Double Q-Learning: Learning Fast Without a ModelXinyue Chen, Che Wang, Zijian Zhou, Keith W. RossICLR 2021 · 被引用 26 次
- Differentiable Trust Region Layers for Deep Reinforcement LearningFabian Otto, Philipp Becker, Ngo Anh Vien, Hanna Carolin Maria Ziesche 等ICLR 2021 · 被引用 23 次
相关 Paper
- A Deep Reinforcement Learning Approach to Marginalized Importance Sampling with the Successor RepresentationScott Fujimoto, David Meger, Doina PrecupICML 2021 · 被引用 17 次
- VA-learning as a more efficient alternative to Q-learningYunhao Tang, Rémi Munos, Mark Rowland, Michal ValkoICML 2023 · 被引用 11 次
- Off-Policy Actor-Critic with Shared Experience ReplaySimon Schmitt, Matteo Hessel, Karen SimonyanICML 2020 · 被引用 71 次
- Doubly Robust Off-Policy Actor-Critic: Convergence and OptimalityTengyu Xu, Zhuoran Yang, Zhaoran Wang, Yingbin LiangICML 2021 · 被引用 31 次
- Variational Delayed Policy OptimizationQingyuan Wu, Simon Sinong Zhan, Yixuan Wang, Yuhui Wang 等NeurIPS 2024 · 被引用 10 次
