Growing Action Spaces
Gregory Farquhar, Laura Gustafson, Zeming Lin, Shimon Whiteson, Nicolas Usunier, Gabriel Synnaeve
Abstract
In complex tasks, such as those with large combinatorial action spaces, random exploration may be too inefficient to achieve meaningful learning progress. In this work, we use a curriculum of progressively growing action spaces to accelerate learning. We assume the environment is out of our control, but that the agent may set an internal curriculum by initially restricting its action space. Our approach uses off-policy reinforcement learning to estimate optimal value functions for multiple action spaces simultaneously and efficiently transfers data, value estimates, and state representations from restricted action spaces to the full task. We show the efficacy of our approach in proof-of-concept control tasks and on challenging large-scale StarCraft micromanagement tasks with large, multi-agent action spaces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80e7fd7a-dc1f-4ab6-8a86-9c643fd328f7Cited by top-tier papers9
- RODE: Learning Roles to Decompose Multi-Agent TasksTonghan Wang, Tarun Gupta, Anuj Mahajan, Bei Peng et al.ICLR 2021 · 60 citations
- Is Bang-Bang Control All You Need? Solving Continuous Control with Bernoulli PoliciesTim Seyde, Igor Gilitschenski, Wilko Schwarting, Bartolomeo Stellato et al.NeurIPS 2021 · 59 citations
- Learning to Represent Action Values as a Hypergraph on the Action VerticesArash Tavakoli, Mehdi Fatemi, Petar KormushevICLR 2021 · 25 citations
- Stochastic Q-learning for Large Discrete Action SpacesFares Fourati, Vaneet Aggarwal, Mohamed-Slim AlouiniICML 2024 · 9 citations
- DEP-RL: Embodied Exploration for Reinforcement Learning in Overactuated and Musculoskeletal SystemsPierre Schumacher, Daniel F. B. Haeufle, Dieter Büchler, Syn Schmitt et al.ICLR 2023 · 6 citations
Related papers
- Self-Paced Deep Reinforcement LearningPascal Klink, Carlo D'Eramo, Jan Peters, Joni PajarinenNeurIPS 2020 · 83 citations
- Understanding the Complexity Gains of Single-Task RL with a CurriculumQiyang Li, Yuexiang Zhai, Yi Ma, Sergey LevineICML 2023 · 21 citations
- From Few to More: Large-Scale Dynamic Multiagent Curriculum LearningWeixun Wang, Tianpei Yang, Yong Liu, Jianye Hao et al.AAAI 2020 · 138 citations
- Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMsGeorgios Tzannetos, Parameswaran Kamalaruban, Adish SinglaNeurIPS 2025 · 3 citations
- PORTAL: Automatic Curricula Generation for Multiagent Reinforcement LearningJizhou Wu, Jianye Hao, Tianpei Yang, Xiaotian Hao et al.AAAI 2024 · 12 citations
