Diverse Conventions for Human-AI Collaboration
Bidipta Sarkar, Andy Shih, Dorsa Sadigh
Abstract
Conventions are crucial for strong performance in cooperative multi-agent games, because they allow players to coordinate on a shared strategy without explicit communication. Unfortunately, standard multi-agent reinforcement learning techniques, such as self-play, converge to conventions that are arbitrary and non-diverse, leading to poor generalization when interacting with new partners. In this work, we present a technique for generating diverse conventions by (1) maximizing their rewards during self-play, while (2) minimizing their rewards when playing with previously discovered conventions (cross-play), stimulating conventions to be semantically different. To ensure that learned policies act in good faith despite the adversarial optimization of cross-play, we introduce mixed-play, where an initial state is randomly generated by sampling self-play and cross-play transitions and the player learns to maximize the self-play reward from this initial state. We analyze the benefits of our technique on various multi-agent collaborative games, including Overcooked, and find that our technique can adapt to the conventions of humans, surpassing human-level performance when paired with real users. 1 Unfortunately, self-play results in incredibly brittle policies that cannot work well with humans who may have different conventions [29] , including conventions that are more intuitive for people to use. In the cooking task, self-play could converge to a convention that expects player 2 to bring onions to AI player 1 from the right side. If a human player instead decides to bring onions from the left side to the AI player 1, the AI's policy may not react appropriately since this is not an interaction 1 Supplemental videos can be found on our website along with source code and anonymized user study data. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce8031a7-b2df-4b7e-9ebc-6b248330517dCited by top-tier papers10
- Learning to Cooperate with Humans using Generative AgentsYancheng Liang, Daphne Chen, Abhishek Gupta, Simon S. Du et al.NeurIPS 2024 · 32 citations
- Adaptively Coordinating with Novel Partners via Learned Latent StrategiesBenjamin Li, Shuyang Shi, Lucia Romero, Huao Li et al.NeurIPS 2025 · 4 citations
- Robust and Diverse Multi-Agent Learning via Rational Policy GradientNiklas Lauffer, Ameesh Shah, Micah Carroll, Sanjit A. Seshia et al.NeurIPS 2025 · 4 citations
- Diversity Is Not All You Need: Training A Robust Cooperative Agent Needs Specialist PartnersRujikorn Charakorn, Poramate Manoonpong, Nat DilokthanakulNeurIPS 2024 · 3 citations
- Who Is Helping Whom? Analyzing Inter-Dependencies to Evaluate Cooperation in Human-AI TeamingUpasana Biswas, Vardhan Palod, Siddhant Bhambri, Subbarao KambhampatiAAAI 2026 · 3 citations
Builds on12
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 271 citations
- Collaborating with Humans without Human DataDJ Strouse, Kevin R. McKee, Matt M. Botvinick, Edward Hughes et al.NeurIPS 2021 · 239 citations
- Trajectory Diversity for Zero-Shot CoordinationAndrei Lupu, Brandon Cui, Hengyuan Hu, Jakob N. FoersterICML 2021 · 157 citations
- Reincarnating Reinforcement Learning: Reusing Prior Computation to Accelerate ProgressRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2022 · 95 citations
- Evaluation of Human-AI Teams for Learned and Rule-Based Agents in HanabiHo Chit Siu, Jaime Daniel Peña, Edenna Chen, Yutai Zhou et al.NeurIPS 2021 · 78 citations
Related papers
- Unsupervised Partner Design Enables Robust Ad-hoc TeamworkConstantin Ruhdorfer, Matteo Bortoletto, Victor Oei, Anna Penzkofer et al.ICML 2026 · 3 citations
- Learning Zero-Shot Cooperation with Humans, Assuming Humans Are BiasedChao Yu, Jiaxuan Gao, Weilin Liu, Botian Xu et al.ICLR 2023 · 4 citations
- Maximum Entropy Population-Based Training for Zero-Shot Human-AI CoordinationRui Zhao, Jinming Song, Yufeng Yuan, Haifeng Hu et al.AAAI 2023 · 94 citations
- Generalized Beliefs for Cooperative AIDarius Muglich, Luisa M. Zintgraf, Christian A. Schröder de Witt, Shimon Whiteson et al.ICML 2022 · 11 citations
- Beyond Single Stationary Policies: Meta-Task Players as Naturally Superior CollaboratorsHaoming Wang, Zhaoming Tian, Yunpeng Song, Xiangliang Zhang et al.NeurIPS 2024 · 3 citations
