Emergent Tool Use From Multi-Agent Autocurricula
Bowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu, Glenn Powell, Bob McGrew, Igor Mordatch
Abstract
Through multi-agent competition, the simple objective of hide-and-seek, and standard reinforcement learning algorithms at scale, we find that agents create a selfsupervised autocurriculum inducing multiple distinct rounds of emergent strategy, many of which require sophisticated tool use and coordination. We find clear evidence of six emergent phases in agent strategy in our environment, each of which creates a new pressure for the opposing team to adapt; for instance, agents learn to build multi-object shelters using moveable boxes which in turn leads to agents discovering that they can overcome obstacles using ramps. We further provide evidence that multi-agent competition may scale better with increasing environment complexity and leads to behavior that centers around far more human-relevant skills than other self-supervised reinforcement learning methods such as intrinsic motivation. Finally, we propose transfer and fine-tuning as a way to quantitatively evaluate targeted capabilities, and we compare hide-and-seek agents to both intrinsic motivation and random initialization baselines in a suite of domain-specific intelligence tests. * This was a large project and many people made significant contributions. Bowen, Bob, and Igor conceived the project and provided guidance through all stages of the work. Bowen created the initial environment, infrastructure and models, and obtained the first results of sequential skill progression. Ingmar obtained the first results of tool use, contributed to environment variants, created domain-specific statistics, and with Bowen created the final environment. Todor created the manipulation tasks in the transfer suite, helped Yi with the RND baseline, and prepared code for open-sourcing. Yi created the navigation tasks in the transfer suite, intrinsic motivation comparisons, and contributed to environment variants. Glenn contributed to designing the final environment and created final renderings and project video. Igor provided research supervision and team leadership. † Work performed while at OpenAI
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers109
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga et al.NeurIPS 2022 · 458 citations
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen et al.NeurIPS 2020 · 362 citations
- Building Cooperative Embodied Agents Modularly with Large Language ModelsHongxin Zhang, Weihua Du, Jiaming Shan, Qinhong Zhou et al.ICLR 2024 · 303 citations
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned ModelsAlexander Pan, Kush Bhatia, Jacob SteinhardtICLR 2022 · 293 citations
- ROMA: Multi-Agent Reinforcement Learning with Emergent RolesTonghan Wang, Heng Dong, Victor R. Lesser, Chongjie ZhangICML 2020 · 286 citations
Related papers
- Variational Automatic Curriculum Learning for Sparse-Reward Cooperative Multi-Agent ProblemsJiayu Chen, Yuanxin Zhang, Yuanfan Xu, Huimin Ma et al.NeurIPS 2021 · 48 citations
- Learning with AMIGo: Adversarially Motivated Intrinsic GoalsAndres Campero, Roberta Raileanu, Heinrich Küttler, Joshua B. Tenenbaum et al.ICLR 2021 · 48 citations
- UltraHorizon: Benchmarking LLM-Agent Capabilities in Ultra Long-Horizon ScenariosHaotian Luo, Huaisong Zhang, Xuelin Zhang, Haoyu Wang et al.ICML 2026 · 21 citations
- Unsupervised Reinforcement Learning with Contrastive Intrinsic ControlMichael Laskin, Hao Liu, Xue Bin Peng, Denis Yarats et al.NeurIPS 2022 · 62 citations
- Emergent Social Learning via Multi-agent Reinforcement LearningKamal Ndousse, Douglas Eck, Sergey Levine, Natasha JaquesICML 2021 · 61 citations
