Beyond Single Stationary Policies: Meta-Task Players as Naturally Superior Collaborators
Haoming Wang, Zhaoming Tian, Yunpeng Song, Xiangliang Zhang, Zhongmin Cai
摘要
In human-AI collaborative tasks, the distribution of human behavior, influenced by mental models, is non-stationary, manifesting in various levels of initiative and different collaborative strategies. A significant challenge in human-AI collaboration is determining how to collaborate effectively with humans exhibiting non-stationary dynamics. Current collaborative agents involve initially running self-play (SP) multiple times to build a policy pool, followed by training the final adaptive policy against this pool. These agents themselves are a single policy network, which is insufficient for handling non-stationary human dynamics . We discern that despite the inherent diversity in human behaviors, the underlying meta-tasks within specific collaborative contexts tend to be strikingly similar . Accordingly, we propose C ollaborative B ayesian P olicy R euse ( CBPR 1 ), a novel Bayesian-based framework that adaptively selects optimal collaborative policies matching the current meta-task from multiple policy networks instead of just selecting actions relying on a single policy network. We provide theoretical guarantees for CBPR’s rapid convergence to the optimal policy once human partners alter their policies. This framework shifts from directly modeling human behavior to identifying various meta-tasks that support human decision-making and training meta-task playing (MTP) agents tailored to enhance collaboration. Our method undergoes rigorous testing in a well-recognized collaborative cooking simulator, Overcooked . Both empirical results and user studies demonstrate CBPR’s superior competitiveness compared to existing baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 被引用 271 次
- Collaborating with Humans without Human DataDJ Strouse, Kevin R. McKee, Matt M. Botvinick, Edward Hughes 等NeurIPS 2021 · 被引用 239 次
- Unpacking Reward Shaping: Understanding the Benefits of Reward Engineering on Sample ComplexityAbhishek Gupta, Aldo Pacchiano, Yuexiang Zhai, Sham M. Kakade 等NeurIPS 2022 · 被引用 115 次
- Maximum Entropy Population-Based Training for Zero-Shot Human-AI CoordinationRui Zhao, Jinming Song, Yufeng Yuan, Haifeng Hu 等AAAI 2023 · 被引用 94 次
- Towards Safe Policy Improvement for Non-Stationary MDPsYash Chandak, Scott M. Jordan, Georgios Theocharous, Martha White 等NeurIPS 2020 · 被引用 47 次
相关 Paper
- Adaptively Coordinating with Novel Partners via Learned Latent StrategiesBenjamin Li, Shuyang Shi, Lucia Romero, Huao Li 等NeurIPS 2025 · 被引用 4 次
- Partner Modelling Emerges in Recurrent Agents (But Only When It Matters)Ruaridh Mon-Williams, Max Taylor-Davies, Elizabeth Mieczkowski, Natalia Vélez 等NeurIPS 2025 · 被引用 6 次
- Learning to Cooperate with Humans using Generative AgentsYancheng Liang, Daphne Chen, Abhishek Gupta, Simon S. Du 等NeurIPS 2024 · 被引用 32 次
- Learning Zero-Shot Cooperation with Humans, Assuming Humans Are BiasedChao Yu, Jiaxuan Gao, Weilin Liu, Botian Xu 等ICLR 2023 · 被引用 4 次
- Diverse Conventions for Human-AI CollaborationBidipta Sarkar, Andy Shih, Dorsa SadighNeurIPS 2023 · 被引用 23 次
