PID Accelerated Value Iteration Algorithm
Amir Massoud Farahmand, Mohammad Ghavamzadeh
Abstract
The convergence rate of Value Iteration (VI), a fundamental procedure in dynamic programming and reinforcement learning, for solving MDPs can be slow when the discount factor is close to one. We propose modifications to VI in order to potentially accelerate its convergence behaviour. The key insight is the realization that the evolution of the value function approximations (V k ) k≥0 in the VI procedure can be seen as a dynamical system. This opens up the possibility of using techniques from control theory to modify, and potentially accelerate, this dynamics. We present such modifications based on simple controllers, such as PD (Proportional-Derivative), PI (Proportional-Integral), and PID. We present the error dynamics of these variants of VI, and provably (for certain classes of MDPs) and empirically (for more general classes) show that the convergence rate can be significantly improved. We also propose a gain adaptation mechanism in order to automatically select the controller gains, and empirically show the effectiveness of this procedure.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c08a6801-0042-4d8d-92c6-548a2a53e70eCited by top-tier papers4
- Accelerating Value Iteration with AnchoringJongmin Lee, Ernest K. RyuNeurIPS 2023 · 20 citations
- Sample Efficient Stochastic Policy Extragradient Algorithm for Zero-Sum Markov GameZiyi Chen, Shaocong Ma, Yi ZhouICLR 2022 · 18 citations
- Operator Splitting Value IterationAmin Rakhsha, Andrew Wang, Mohammad Ghavamzadeh, Amir-massoud FarahmandNeurIPS 2022 · 11 citations
- On PI Controllers for Updating Lagrange Multipliers in Constrained OptimizationMotahareh Sohrabi, Juan Ramirez, Tianyue H. Zhang, Simon Lacoste-Julien et al.ICML 2024 · 7 citations
Related papers
- PID-controlled Langevin Dynamics for Faster Sampling of Generative ModelsHongyi Chen, Jianhai Shu, Jingtao Ding, Yong Li et al.NeurIPS 2025 · 1 citation
- Optimal Non-Asymptotic Rates of Value Iteration for Average-Reward Markov Decision ProcessesJongmin Lee, Ernest K. RyuICLR 2025
- Faster Fixed-Point Methods for Multichain MDPsMatthew Zurek, Yudong ChenNeurIPS 2025 · 3 citations
- Value Iteration in Continuous Actions, States and TimeMichael Lutter, Shie Mannor, Jan Peters, Dieter Fox et al.ICML 2021 · 40 citations
- Quantum algorithms for reinforcement learning with a generative modelDaochen Wang, Aarthi Sundaram, Robin Kothari, Ashish Kapoor et al.ICML 2021 · 38 citations
