Cross-environment Cooperation Enables Zero-shot Multi-agent Coordination
Kunal Jha, Wilka Carvalho, Yancheng Liang, Simon Shaolei Du, Max Kleiman-Weiner, Natasha Jaques
Abstract
Zero-shot coordination (ZSC), the ability to adapt to a new partner in a cooperative task, is a critical component of human-compatible AI. While prior work has focused on training agents to cooperate on a single task, these specialized models do not generalize to new tasks, even if they are highly similar. Here, we study how reinforcement learning on a distribution of environments with a single partner enables learning general cooperative skills that support ZSC with many new partners on many new problems. We introduce two Jax-based, procedural generators that create billions of solvable coordination challenges. We develop a new paradigm called Cross-Environment Cooperation (CEC), and show that it outperforms competitive baselines quantitatively and qualitatively when collaborating with real people. Our findings suggest that learning to collaborate across many unique scenarios encourages agents to develop general norms, which prove effective for collaboration with different partners. Together, our results suggest a new route toward designing generalist cooperative agents capable of interacting with humans without requiring human data. Code for environment, training, and testing scripts and more can be found at https://kjha02.github.io/ publication/cross-env-coop .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa9852f0-004c-4fc6-8d2b-b94a399d1f6fCited by top-tier papers5
- Evaluating LLMs in Open-Source GamesSwadesh Sistla, Max Kleiman-WeinerNeurIPS 2025 · 5 citations
- SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent SystemZhiyu Xu, Weilong Yan, Yufei Shi, Xin Meng et al.CVPR 2026 · 4 citations
- Unsupervised Partner Design Enables Robust Ad-hoc TeamworkConstantin Ruhdorfer, Matteo Bortoletto, Victor Oei, Anna Penzkofer et al.ICML 2026 · 3 citations
- Role-Level Inductive Bias for Cross-Task Generalization in Multi-Agent Reinforcement LearningChang Yao, Youfang Lin, Shoucheng Song, Hao Wu et al.ICML 2026
- CooT: Learning to Coordinate In-Context with Coordination TransformersHuai-Chih Wang, Hsiang-Chun Chuang, Hsi-Chun Cheng, Dai-Jie Wu et al.ICML 2026
Builds on15
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 685 citations
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen et al.NeurIPS 2020 · 362 citations
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 271 citations
- Collaborating with Humans without Human DataDJ Strouse, Kevin R. McKee, Matt M. Botvinick, Edward Hughes et al.NeurIPS 2021 · 239 citations
- Personalizing Reinforcement Learning from Human Feedback with Variational Preference LearningSriyash Poddar, Yanming Wan, Hamish Ivison, Abhishek Gupta et al.NeurIPS 2024 · 188 citations
Related papers
- Adaptive Coordination in Social Embodied RearrangementAndrew Szot, Unnat Jain, Dhruv Batra, Zsolt Kira et al.ICML 2023 · 20 citations
- Learning to Cooperate with Humans using Generative AgentsYancheng Liang, Daphne Chen, Abhishek Gupta, Simon S. Du et al.NeurIPS 2024 · 32 citations
- An Efficient End-to-End Training Approach for Zero-Shot Human-AI CoordinationXue Yan, Jiaxian Guo, Xingzhou Lou, Jun Wang et al.NeurIPS 2023 · 37 citations
- Efficient Reinforcement Learning for Zero-Shot Coordination in Evolving GamesBingyu Hui, Lebin Yu, Quanming Yao, Yunpeng Qu et al.AAAI 2026 · 1 citation
- OvercookedV2: Rethinking Overcooked for Zero-Shot CoordinationTobias Gessler, Tin Dizdarevic, Ani Calinescu, Benjamin Ellis et al.ICLR 2025
