Minimum Coverage Sets for Training Robust Ad Hoc Teamwork Agents
Muhammad Rahman, Jiaxun Cui, Peter Stone
摘要
Robustly cooperating with unseen agents and human partners presents significant challenges due to the diverse cooperative conventions these partners may adopt. Existing Ad Hoc Teamwork (AHT) methods address this challenge by training an agent with a population of diverse teammate policies obtained through maximizing specific diversity metrics. However, prior heuristic-based diversity metrics do not always maximize the agent's robustness in all cooperative problems. In this work, we first propose that maximizing an AHT agent's robustness requires it to emulate policies in the minimum coverage set (MCS), the set of best-response policies to any partner policies in the environment. We then introduce the L-BRDiv algorithm that generates a set of teammate policies that, when used for AHT training, encourage agents to emulate policies from the MCS. L-BRDiv works by solving a constrained optimization problem to jointly train teammate policies for AHT training and approximating AHT agent policies that are members of the MCS. We empirically demonstrate that L-BRDiv produces more robust AHT agents than state-of-the-art methods in a broader range of two-player cooperative problems without the need for extensive hyperparameter tuning for its objectives. Our study shows that L-BRDiv outperforms the baseline methods by prioritizing discovering distinct members of the MCS instead of repeatedly finding redundant policies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Diversity Is Not All You Need: Training A Robust Cooperative Agent Needs Specialist PartnersRujikorn Charakorn, Poramate Manoonpong, Nat DilokthanakulNeurIPS 2024 · 被引用 3 次
- Improving Human-AI Coordination through Online Adversarial Training and Generative ModelsParesh R. Chaudhary, Yancheng Liang, Daphne Chen, Simon Shaolei Du 等ICLR 2026 · 被引用 2 次
- Improving Cooperation in Language Games with Bayesian Inference and the Cognitive HierarchyJoseph Bills, Christopher Archibald, Diego BlaylockAAAI 2025 · 被引用 1 次
- LLM-Assisted Semantically Diverse Teammate Generation for Efficient Multi-agent CoordinationLihe Li, Lei Yuan, Pengsen Liu, Tao Jiang 等ICML 2025
- Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc TeamworkYuheng Jing, Kai Li, Jiajun Zhang, Zeyao Ma 等ICML 2026
它引用的顶会 Paper10
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 被引用 271 次
- Collaborating with Humans without Human DataDJ Strouse, Kevin R. McKee, Matt M. Botvinick, Edward Hughes 等NeurIPS 2021 · 被引用 239 次
- Shared Experience Actor-Critic for Multi-Agent Reinforcement LearningFilippos Christianos, Lukas Schäfer, Stefano V. AlbrechtNeurIPS 2020 · 被引用 238 次
- Trajectory Diversity for Zero-Shot CoordinationAndrei Lupu, Brandon Cui, Hengyuan Hu, Jakob N. FoersterICML 2021 · 被引用 157 次
- Agent Modelling under Partial Observability for Deep Reinforcement LearningGeorgios Papoudakis, Filippos Christianos, Stefano V. AlbrechtNeurIPS 2021 · 被引用 110 次
相关 Paper
- Generating Diverse Cooperative Agents by Learning Incompatible PoliciesRujikorn Charakorn, Poramate Manoonpong, Nat DilokthanakulICLR 2023
- PADiff: Predictive and Adaptive Diffusion Policies for Ad Hoc TeamworkHohei Chan, Xinzhi Zhang, Antao Xiang, Weinan Zhang 等AAAI 2026
- Unsupervised Partner Design Enables Robust Ad-hoc TeamworkConstantin Ruhdorfer, Matteo Bortoletto, Victor Oei, Anna Penzkofer 等ICML 2026 · 被引用 3 次
- N-agent Ad Hoc TeamworkCaroline Wang, Arrasy Rahman, Ishan Durugkar, Elad Liebman 等NeurIPS 2024 · 被引用 21 次
- Robust Multi-Agent Coordination via Evolutionary Generation of Auxiliary Adversarial AttackersLei Yuan, Ziqian Zhang, Ke Xue, Hao Yin 等AAAI 2023 · 被引用 31 次
