Operator Splitting Value Iteration
Amin Rakhsha, Andrew Wang, Mohammad Ghavamzadeh, Amir-massoud Farahmand
摘要
We introduce new planning and reinforcement learning algorithms for discounted MDPs that utilize an approximate model of the environment to accelerate the convergence of the value function. Inspired by the splitting approach in numerical linear algebra, we introduce Operator Splitting Value Iteration (OS-VI) for both Policy Evaluation and Control problems. OS-VI achieves a much faster convergence rate when the model is accurate enough. We also introduce a sample-based version of the algorithm called OS-Dyna. Unlike the traditional Dyna architecture, OS-Dyna still converges to the correct value function in presence of model approximation error.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Maximum Entropy Model Correction in Reinforcement LearningAmin Rakhsha, Mete Kemertas, Mohammad Ghavamzadeh, Amir-massoud FarahmandICLR 2024 · 被引用 3 次
- Rank-One Modified Value IterationArman Sharifi Kolarijani, Tolga Ok, Peyman Mohajerin Esfahani, Mohamad Amin Sharifi KolarijaniICML 2025
它引用的顶会 Paper6
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang 等ICML 2020 · 被引用 324 次
- The Value Equivalence Principle for Model-Based Reinforcement LearningChristopher Grimm, André Barreto, Satinder Singh, David SilverNeurIPS 2020 · 被引用 129 次
- Control-Oriented Model-Based Reinforcement Learning with Implicit DifferentiationEvgenii Nikishin, Romina Abachi, Rishabh Agarwal, Pierre-Luc BaconAAAI 2022 · 被引用 47 次
- Gradient-Aware Model-Based Policy SearchPierluca D'Oro, Alberto Maria Metelli, Andrea Tirinzoni, Matteo Papini 等AAAI 2020 · 被引用 40 次
- Value Gradient weighted Model-Based Reinforcement LearningClaas Voelcker, Victor Liao, Animesh Garg, Amir-massoud FarahmandICLR 2022 · 被引用 37 次
相关 Paper
- PID Accelerated Value Iteration AlgorithmAmir Massoud Farahmand, Mohammad GhavamzadehICML 2021 · 被引用 16 次
- Optimistic Planning by Regularized Dynamic ProgrammingAntoine Moulin, Gergely NeuICML 2023 · 被引用 8 次
- Self-Consistent Models and ValuesGregory Farquhar, Kate Baumli, Zita Marinho, Angelos Filos 等NeurIPS 2021 · 被引用 10 次
- Average-Reward Learning and Planning with OptionsYi Wan, Abhishek Naik, Richard S. SuttonNeurIPS 2021 · 被引用 12 次
- Learning and Planning in Average-Reward Markov Decision ProcessesYi Wan, Abhishek Naik, Richard S. SuttonICML 2021 · 被引用 82 次
