Multi-step Greedy Reinforcement Learning Algorithms
Manan Tomar, Yonathan Efroni, Mohammad Ghavamzadeh
摘要
Multi-step greedy policies have been extensively used in model-based reinforcement learning (RL), both when a model of the environment is available (e.g., in the game of Go) and when it is learned. In this paper, we explore their benefits in model-free RL, when employed using multi-step dynamic programming algorithms: -Policy Iteration (-PI) and -Value Iteration (-VI). These methods iteratively compute the next policy (-PI) and value function (-VI) by solving a surrogate decision problem with a shaped reward and a smaller discount factor. We derive model-free RL algorithms based on -PI and -VI in which the surrogate problem can be solved by any discrete or continuous action RL method, such as DQN and TRPO. We identify the importance of a hyper-parameter that controls the extent to which the surrogate problem is solved and suggest a way to set this parameter. When evaluated on a range of Atari and MuJoCo benchmark tasks, our results indicate that for the right range of , our algorithms outperform DQN and TRPO. This shows that our multi-step greedy algorithms are general enough to be applied over any existing RL algorithm and can significantly improve its performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Generalised Policy Improvement with Geometric Policy CompositionShantanu Thakoor, Mark Rowland, Diana Borsa, Will Dabney 等ICML 2022 · 被引用 11 次
- Policy Mirror Descent with LookaheadKimon Protopapas, Anas BarakatNeurIPS 2024 · 被引用 7 次
- DoMo-AC: Doubly Multi-step Off-policy Actor-Critic AlgorithmYunhao Tang, Tadashi Kozuno, Mark Rowland, Anna Harutyunyan 等ICML 2023
相关 Paper
- Structured Policy Iteration for Linear Quadratic RegulatorYoungsuk Park, Ryan A. Rossi, Zheng Wen, Gang Wu 等ICML 2020 · 被引用 23 次
- Muesli: Combining Improvements in Policy OptimizationMatteo Hessel, Ivo Danihelka, Fabio Viola, Arthur Guez 等ICML 2021 · 被引用 69 次
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 被引用 164 次
- Successor Feature Sets: Generalizing Successor Representations Across PoliciesKianté Brantley, Soroush Mehri, Geoffrey J. GordonAAAI 2021 · 被引用 11 次
- Generalized Hidden Parameter MDPs: Transferable Model-Based RL in a Handful of TrialsChristian F. Perez, Felipe Petroski Such, Theofanis KaraletsosAAAI 2020 · 被引用 39 次
