Learning and Planning in Complex Action Spaces
Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Mohammadamin Barekatain, Simon Schmitt, David Silver
摘要
Many important real-world problems have action spaces that are high-dimensional, continuous or both, making full enumeration of all possible actions infeasible. Instead, only small subsets of actions can be sampled for the purpose of policy evaluation and improvement. In this paper, we propose a general framework to reason in a principled way about policy evaluation and improvement over such sampled action subsets. This sample-based policy iteration framework can in principle be applied to any reinforcement learning algorithm based upon policy iteration. Concretely, we propose Sampled MuZero, an extension of the MuZero algorithm that is able to learn in domains with arbitrarily complex action spaces by planning over sampled actions. We demonstrate this approach on the classical board game of Go and on two continuous control benchmark domains: DeepMind Control Suite and Real-World RL Suite.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper49
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 被引用 388 次
- Mastering Atari Games with Limited DataWeirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel 等NeurIPS 2021 · 被引用 345 次
- AlphaZero-Like Tree-Search can Guide Large Language Model Decoding and TrainingZiyu Wan, Xidong Feng, Muning Wen, Stephen Marcus McAleer 等ICML 2024 · 被引用 325 次
- Online and Offline Reinforcement Learning by Planning with a Learned ModelJulian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain 等NeurIPS 2021 · 被引用 149 次
- Policy improvement by planning with GumbelIvo Danihelka, Arthur Guez, Julian Schrittwieser, David SilverICLR 2022 · 被引用 84 次
它引用的顶会 Paper5
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Discretizing Continuous Action Space for On-Policy OptimizationYunhao Tang, Shipra AgrawalAAAI 2020 · 被引用 150 次
- A Self-Tuning Actor-Critic AlgorithmTom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel 等NeurIPS 2020 · 被引用 106 次
- Monte-Carlo Tree Search as Regularized Policy OptimizationJean-Bastien Grill, Florent Altché, Yunhao Tang, Thomas Hubert 等ICML 2020 · 被引用 84 次
- An operator view of policy gradient methodsDibya Ghosh, Marlos C. Machado, Nicolas Le RouxNeurIPS 2020 · 被引用 30 次
相关 Paper
- On the role of planning in model-based deep reinforcement learningJessica B. Hamrick, Abram L. Friesen, Feryal M. P. Behbahani, Arthur Guez 等ICLR 2021 · 被引用 77 次
- Multiagent Gumbel MuZero: Efficient Planning in Combinatorial Action SpacesXiaotian Hao, Jianye Hao, Chenjun Xiao, Kai Li 等AAAI 2024 · 被引用 5 次
- Planning in Stochastic Environments with a Learned ModelIoannis Antonoglou, Julian Schrittwieser, Sherjil Ozair, Thomas K. Hubert 等ICLR 2022 · 被引用 79 次
- EfficientZero V2: Mastering Discrete and Continuous Control with Limited DataShengjie Wang, Shaohuai Liu, Weirui Ye, Jiacheng You 等ICML 2024 · 被引用 36 次
- Efficient Offline Policy Optimization with a Learned ModelZichen Liu, Siyi Li, Wee Sun Lee, Shuicheng Yan 等ICLR 2023
