Learning and Planning in Complex Action Spaces
Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Mohammadamin Barekatain, Simon Schmitt, David Silver
Abstract
Many important real-world problems have action spaces that are high-dimensional, continuous or both, making full enumeration of all possible actions infeasible. Instead, only small subsets of actions can be sampled for the purpose of policy evaluation and improvement. In this paper, we propose a general framework to reason in a principled way about policy evaluation and improvement over such sampled action subsets. This sample-based policy iteration framework can in principle be applied to any reinforcement learning algorithm based upon policy iteration. Concretely, we propose Sampled MuZero, an extension of the MuZero algorithm that is able to learn in domains with arbitrarily complex action spaces by planning over sampled actions. We demonstrate this approach on the classical board game of Go and on two continuous control benchmark domains: DeepMind Control Suite and Real-World RL Suite.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 110140be-fb1e-4803-ade7-dc95b6549177Cited by top-tier papers49
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 388 citations
- Mastering Atari Games with Limited DataWeirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel et al.NeurIPS 2021 · 345 citations
- AlphaZero-Like Tree-Search can Guide Large Language Model Decoding and TrainingZiyu Wan, Xidong Feng, Muning Wen, Stephen Marcus McAleer et al.ICML 2024 · 325 citations
- Online and Offline Reinforcement Learning by Planning with a Learned ModelJulian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain et al.NeurIPS 2021 · 149 citations
- Policy improvement by planning with GumbelIvo Danihelka, Arthur Guez, Julian Schrittwieser, David SilverICLR 2022 · 84 citations
Builds on5
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Discretizing Continuous Action Space for On-Policy OptimizationYunhao Tang, Shipra AgrawalAAAI 2020 · 150 citations
- A Self-Tuning Actor-Critic AlgorithmTom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel et al.NeurIPS 2020 · 106 citations
- Monte-Carlo Tree Search as Regularized Policy OptimizationJean-Bastien Grill, Florent Altché, Yunhao Tang, Thomas Hubert et al.ICML 2020 · 84 citations
- An operator view of policy gradient methodsDibya Ghosh, Marlos C. Machado, Nicolas Le RouxNeurIPS 2020 · 30 citations
Related papers
- On the role of planning in model-based deep reinforcement learningJessica B. Hamrick, Abram L. Friesen, Feryal M. P. Behbahani, Arthur Guez et al.ICLR 2021 · 77 citations
- Multiagent Gumbel MuZero: Efficient Planning in Combinatorial Action SpacesXiaotian Hao, Jianye Hao, Chenjun Xiao, Kai Li et al.AAAI 2024 · 5 citations
- Planning in Stochastic Environments with a Learned ModelIoannis Antonoglou, Julian Schrittwieser, Sherjil Ozair, Thomas K. Hubert et al.ICLR 2022 · 79 citations
- EfficientZero V2: Mastering Discrete and Continuous Control with Limited DataShengjie Wang, Shaohuai Liu, Weirui Ye, Jiacheng You et al.ICML 2024 · 36 citations
- Efficient Offline Policy Optimization with a Learned ModelZichen Liu, Siyi Li, Wee Sun Lee, Shuicheng Yan et al.ICLR 2023
