Emergent Tool Use From Multi-Agent Autocurricula
Bowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu, Glenn Powell, Bob McGrew, Igor Mordatch
摘要
Through multi-agent competition, the simple objective of hide-and-seek, and standard reinforcement learning algorithms at scale, we find that agents create a selfsupervised autocurriculum inducing multiple distinct rounds of emergent strategy, many of which require sophisticated tool use and coordination. We find clear evidence of six emergent phases in agent strategy in our environment, each of which creates a new pressure for the opposing team to adapt; for instance, agents learn to build multi-object shelters using moveable boxes which in turn leads to agents discovering that they can overcome obstacles using ramps. We further provide evidence that multi-agent competition may scale better with increasing environment complexity and leads to behavior that centers around far more human-relevant skills than other self-supervised reinforcement learning methods such as intrinsic motivation. Finally, we propose transfer and fine-tuning as a way to quantitatively evaluate targeted capabilities, and we compare hide-and-seek agents to both intrinsic motivation and random initialization baselines in a suite of domain-specific intelligence tests. * This was a large project and many people made significant contributions. Bowen, Bob, and Igor conceived the project and provided guidance through all stages of the work. Bowen created the initial environment, infrastructure and models, and obtained the first results of sequential skill progression. Ingmar obtained the first results of tool use, contributed to environment variants, created domain-specific statistics, and with Bowen created the final environment. Todor created the manipulation tasks in the transfer suite, helped Yi with the RND baseline, and prepared code for open-sourcing. Yi created the navigation tasks in the transfer suite, intrinsic motivation comparisons, and contributed to environment variants. Glenn contributed to designing the final environment and created final renderings and project video. Igor provided research supervision and team leadership. † Work performed while at OpenAI
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper109
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga 等NeurIPS 2022 · 被引用 458 次
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen 等NeurIPS 2020 · 被引用 362 次
- Building Cooperative Embodied Agents Modularly with Large Language ModelsHongxin Zhang, Weihua Du, Jiaming Shan, Qinhong Zhou 等ICLR 2024 · 被引用 303 次
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned ModelsAlexander Pan, Kush Bhatia, Jacob SteinhardtICLR 2022 · 被引用 293 次
- ROMA: Multi-Agent Reinforcement Learning with Emergent RolesTonghan Wang, Heng Dong, Victor R. Lesser, Chongjie ZhangICML 2020 · 被引用 286 次
相关 Paper
- Variational Automatic Curriculum Learning for Sparse-Reward Cooperative Multi-Agent ProblemsJiayu Chen, Yuanxin Zhang, Yuanfan Xu, Huimin Ma 等NeurIPS 2021 · 被引用 48 次
- Learning with AMIGo: Adversarially Motivated Intrinsic GoalsAndres Campero, Roberta Raileanu, Heinrich Küttler, Joshua B. Tenenbaum 等ICLR 2021 · 被引用 48 次
- UltraHorizon: Benchmarking LLM-Agent Capabilities in Ultra Long-Horizon ScenariosHaotian Luo, Huaisong Zhang, Xuelin Zhang, Haoyu Wang 等ICML 2026 · 被引用 21 次
- Unsupervised Reinforcement Learning with Contrastive Intrinsic ControlMichael Laskin, Hao Liu, Xue Bin Peng, Denis Yarats 等NeurIPS 2022 · 被引用 62 次
- Emergent Social Learning via Multi-agent Reinforcement LearningKamal Ndousse, Douglas Eck, Sergey Levine, Natasha JaquesICML 2021 · 被引用 61 次
