Maximum Entropy Population-Based Training for Zero-Shot Human-AI Coordination
Rui Zhao, Jinming Song, Yufeng Yuan, Haifeng Hu, Yang Gao, Yi Wu, Zhongqian Sun, Wei Yang
摘要
We study the problem of training a Reinforcement Learning (RL) agent that is collaborative with humans without using human data. Although such agents can be obtained through self-play training, they can suffer significantly from the distributional shift when paired with unencountered partners, such as humans. In this paper, we propose Maximum Entropy Population-based training (MEP) to mitigate such distributional shift. In MEP, agents in the population are trained with our derived Population Entropy bonus to promote the pairwise diversity between agents and the individual diversity of agents themselves. After obtaining this diversified population, a common best agent is trained by paring with agents in this population via prioritized sampling, where the prioritization is dynamically adjusted based on the training progress. We demonstrate the effectiveness of our method MEP, with comparison to Self-Play PPO (SP), Population-Based Training (PBT), Trajectory Diversity (TrajeDi), and Fictitious Co-Play (FCP) in both matrix game and Overcooked game environments, with partners being human proxy models and real humans. A supplementary video showing experimental results is available at https://youtu.be/Xh-FKD0AAKE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- ProAgent: Building Proactive Cooperative Agents with Large Language ModelsCeyao Zhang, Kaijie Yang, Siyi Hu, Zihao Wang 等AAAI 2024 · 被引用 141 次
- An Efficient End-to-End Training Approach for Zero-Shot Human-AI CoordinationXue Yan, Jiaxian Guo, Xingzhou Lou, Jun Wang 等NeurIPS 2023 · 被引用 37 次
- Cooperative Open-ended Learning Framework for Zero-Shot CoordinationYang Li, Shao Zhang, Jichen Sun, Yali Du 等ICML 2023 · 被引用 35 次
- Learning to Cooperate with Humans using Generative AgentsYancheng Liang, Daphne Chen, Abhishek Gupta, Simon S. Du 等NeurIPS 2024 · 被引用 32 次
- Adaptive Coordination in Social Embodied RearrangementAndrew Szot, Unnat Jain, Dhruv Batra, Zsolt Kira 等ICML 2023 · 被引用 20 次
它引用的顶会 Paper7
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 被引用 271 次
- Collaborating with Humans without Human DataDJ Strouse, Kevin R. McKee, Matt M. Botvinick, Edward Hughes 等NeurIPS 2021 · 被引用 239 次
- Effective Diversity in Population Based Reinforcement LearningJack Parker-Holder, Aldo Pacchiano, Krzysztof Marcin Choromanski, Stephen J. RobertsNeurIPS 2020 · 被引用 195 次
- Trajectory Diversity for Zero-Shot CoordinationAndrei Lupu, Brandon Cui, Hengyuan Hu, Jakob N. FoersterICML 2021 · 被引用 157 次
- Modelling Behavioural Diversity for Learning in Open-Ended GamesNicolas Perez Nieves, Yaodong Yang, Oliver Slumbers, David Henry Mguni 等ICML 2021 · 被引用 80 次
相关 Paper
- Unsupervised Partner Design Enables Robust Ad-hoc TeamworkConstantin Ruhdorfer, Matteo Bortoletto, Victor Oei, Anna Penzkofer 等ICML 2026 · 被引用 3 次
- Learning Zero-Shot Cooperation with Humans, Assuming Humans Are BiasedChao Yu, Jiaxuan Gao, Weilin Liu, Botian Xu 等ICLR 2023 · 被引用 4 次
- Diverse Conventions for Human-AI CollaborationBidipta Sarkar, Andy Shih, Dorsa SadighNeurIPS 2023 · 被引用 23 次
- Adaptively Coordinating with Novel Partners via Learned Latent StrategiesBenjamin Li, Shuyang Shi, Lucia Romero, Huao Li 等NeurIPS 2025 · 被引用 4 次
- Diversity Is Not All You Need: Training A Robust Cooperative Agent Needs Specialist PartnersRujikorn Charakorn, Poramate Manoonpong, Nat DilokthanakulNeurIPS 2024 · 被引用 3 次
