Damped Anderson Mixing for Deep Reinforcement Learning: Acceleration, Convergence, and Stabilization
Ke Sun, Yafei Wang, Yi Liu, Yingnan Zhao, Bo Pan, Shangling Jui, Bei Jiang, Linglong Kong
摘要
Anderson mixing has been heuristically applied to reinforcement learning (RL) algorithms for accelerating convergence and improving the sampling efficiency of deep RL. Despite its heuristic improvement of convergence, a rigorous mathematical justification for the benefits of Anderson mixing in RL has not yet been put forward. In this paper, we provide deeper insights into a class of acceleration schemes built on Anderson mixing that improve the convergence of deep RL algorithms. Our main results establish a connection between Anderson mixing and quasi-Newton methods and prove that Anderson mixing increases the convergence radius of policy iteration schemes by an extra contraction factor. The key focus of the analysis roots in the fixed-point iteration nature of RL. We further propose a stabilization strategy by introducing a stable regularization term in Anderson mixing and a differentiable, non-expansive MellowMax operator that can allow both faster convergence and more stable behavior. Extensive experiments demonstrate that our proposed method enhances the convergence, stability, and performance of RL algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Accelerating Value Iteration with AnchoringJongmin Lee, Ernest K. RyuNeurIPS 2023 · 被引用 20 次
- A Variant of Anderson Mixing with Minimal Memory SizeFuchao Wei, Chenglong Bao, Yang Liu, Guangwen YangNeurIPS 2022 · 被引用 1 次
相关 Paper
- Stochastic Anderson Mixing for Nonconvex Stochastic OptimizationFuchao Wei, Chenglong Bao, Yang LiuNeurIPS 2021 · 被引用 27 次
- A Class of Short-term Recurrence Anderson Mixing Methods and Their ApplicationsFuchao Wei, Chenglong Bao, Yang LiuICLR 2022 · 被引用 6 次
- Deep Conservative Policy IterationNino Vieillard, Olivier Pietquin, Matthieu GeistAAAI 2020 · 被引用 29 次
- Stabilizing Q Learning Via Soft Mellowmax OperatorYaozhong Gan, Zhe Zhang, Xiaoyang TanAAAI 2021 · 被引用 10 次
- Anderson Acceleration of Proximal Gradient MethodsVien V. Mai, Mikael JohanssonICML 2020 · 被引用 45 次
