An Efficient End-to-End Training Approach for Zero-Shot Human-AI Coordination
Xue Yan, Jiaxian Guo, Xingzhou Lou, Jun Wang, Haifeng Zhang, Yali Du
摘要
The goal of zero-shot human-AI coordination is to develop an agent capable of collaborating with humans without relying on human data. Prevailing two-stage population-based methods require a diverse population of mutually distinct policies to simulate diverse human behaviors. The necessity of such populations severely limits their computational efficiency. To address this issue, we pro-pose E3T, an E fficient E nd-to-E nd T raining approach for zero-shot human-AI coordination. E3T employs a mixture of ego policy and random policy to construct the partner policy, making it both skilled in coordination and diverse. This way, the ego agent is trained end-to-end with this mixture policy, eliminating the need for a pre-trained population, and thus significantly improving training efficiency. In addition, we introduce a partner modeling module designed to predict the partner’s actions based on historical contexts. With the predicted partner’s action, the ego policy can adapt its strategy and take actions accordingly when collaborating with humans exhibiting different behavior patterns. Empirical results on the Overcooked environment demonstrate that our method substantially improves the training efficiency while preserving comparable or superior performance than the population-based baselines. Demo videos are available at https://sites.google.com/view/e3t-overcooked .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- MEAL: A Benchmark for Continual Multi-Agent Reinforcement LearningTristan Tomilin, Luka van den Boogaard, Samuel Garcin, Constantin Ruhdorfer 等ICML 2026 · 被引用 9 次
- Fast Peer Adaptation with Context-aware ExplorationLong Ma, Yuanfei Wang, Fangwei Zhong, Song-Chun Zhu 等ICML 2024 · 被引用 9 次
- Adaptively Coordinating with Novel Partners via Learned Latent StrategiesBenjamin Li, Shuyang Shi, Lucia Romero, Huao Li 等NeurIPS 2025 · 被引用 4 次
- Safe Exploitative Play with Untrusted Type BeliefsTongxin Li, Tinashe Handina, Shaolei Ren, Adam WiermanNeurIPS 2024 · 被引用 3 次
- Unsupervised Partner Design Enables Robust Ad-hoc TeamworkConstantin Ruhdorfer, Matteo Bortoletto, Victor Oei, Anna Penzkofer 等ICML 2026 · 被引用 3 次
它引用的顶会 Paper9
- A Simple Language Model for Task-Oriented DialogueEhsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz 等NeurIPS 2020 · 被引用 590 次
- Google Research Football: A Novel Reinforcement Learning EnvironmentKarol Kurach, Anton Raichuk, Piotr Stanczyk, Michal Zajac 等AAAI 2020 · 被引用 496 次
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 被引用 271 次
- Collaborating with Humans without Human DataDJ Strouse, Kevin R. McKee, Matt M. Botvinick, Edward Hughes 等NeurIPS 2021 · 被引用 239 次
- Trajectory Diversity for Zero-Shot CoordinationAndrei Lupu, Brandon Cui, Hengyuan Hu, Jakob N. FoersterICML 2021 · 被引用 157 次
相关 Paper
- Learning to Cooperate with Humans using Generative AgentsYancheng Liang, Daphne Chen, Abhishek Gupta, Simon S. Du 等NeurIPS 2024 · 被引用 32 次
- OvercookedV2: Rethinking Overcooked for Zero-Shot CoordinationTobias Gessler, Tin Dizdarevic, Ani Calinescu, Benjamin Ellis 等ICLR 2025
- Cross-environment Cooperation Enables Zero-shot Multi-agent CoordinationKunal Jha, Wilka Carvalho, Yancheng Liang, Simon Shaolei Du 等ICML 2025
- Efficient Reinforcement Learning for Zero-Shot Coordination in Evolving GamesBingyu Hui, Lebin Yu, Quanming Yao, Yunpeng Qu 等AAAI 2026 · 被引用 1 次
- Cooperative Open-ended Learning Framework for Zero-Shot CoordinationYang Li, Shao Zhang, Jichen Sun, Yali Du 等ICML 2023 · 被引用 35 次
