PID Accelerated Value Iteration Algorithm
Amir Massoud Farahmand, Mohammad Ghavamzadeh
摘要
The convergence rate of Value Iteration (VI), a fundamental procedure in dynamic programming and reinforcement learning, for solving MDPs can be slow when the discount factor is close to one. We propose modifications to VI in order to potentially accelerate its convergence behaviour. The key insight is the realization that the evolution of the value function approximations (V k ) k≥0 in the VI procedure can be seen as a dynamical system. This opens up the possibility of using techniques from control theory to modify, and potentially accelerate, this dynamics. We present such modifications based on simple controllers, such as PD (Proportional-Derivative), PI (Proportional-Integral), and PID. We present the error dynamics of these variants of VI, and provably (for certain classes of MDPs) and empirically (for more general classes) show that the convergence rate can be significantly improved. We also propose a gain adaptation mechanism in order to automatically select the controller gains, and empirically show the effectiveness of this procedure.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Accelerating Value Iteration with AnchoringJongmin Lee, Ernest K. RyuNeurIPS 2023 · 被引用 20 次
- Sample Efficient Stochastic Policy Extragradient Algorithm for Zero-Sum Markov GameZiyi Chen, Shaocong Ma, Yi ZhouICLR 2022 · 被引用 18 次
- Operator Splitting Value IterationAmin Rakhsha, Andrew Wang, Mohammad Ghavamzadeh, Amir-massoud FarahmandNeurIPS 2022 · 被引用 11 次
- On PI Controllers for Updating Lagrange Multipliers in Constrained OptimizationMotahareh Sohrabi, Juan Ramirez, Tianyue H. Zhang, Simon Lacoste-Julien 等ICML 2024 · 被引用 7 次
相关 Paper
- PID-controlled Langevin Dynamics for Faster Sampling of Generative ModelsHongyi Chen, Jianhai Shu, Jingtao Ding, Yong Li 等NeurIPS 2025 · 被引用 1 次
- Optimal Non-Asymptotic Rates of Value Iteration for Average-Reward Markov Decision ProcessesJongmin Lee, Ernest K. RyuICLR 2025
- Faster Fixed-Point Methods for Multichain MDPsMatthew Zurek, Yudong ChenNeurIPS 2025 · 被引用 3 次
- Value Iteration in Continuous Actions, States and TimeMichael Lutter, Shie Mannor, Jan Peters, Dieter Fox 等ICML 2021 · 被引用 40 次
- Quantum algorithms for reinforcement learning with a generative modelDaochen Wang, Aarthi Sundaram, Robin Kothari, Ashish Kapoor 等ICML 2021 · 被引用 38 次
