Open Ad Hoc Teamwork with Cooperative Game Theory
Jianhong Wang, Yang Li, Yuan Zhang, Wei Pan, Samuel Kaski
Abstract
Ad hoc teamwork poses a challenging problem, requiring the design of an agent to collaborate with teammates without prior coordination or joint training. Open ad hoc teamwork (OAHT) further complicates this challenge by considering environments with a changing number of teammates, referred to as open teams. One promising solution in practice to this problem is leveraging the generalizability of graph neural networks to handle an unrestricted number of agents with various agenttypes, named graph-based policy learning (GPL). However, its joint Q-value representation over a coordination graph lacks convincing explanations. In this paper, we establish a new theory to understand the representation of the joint Q-value for OAHT and its learning paradigm, through the lens of cooperative game theory. Building on our theory, we propose a novel algorithm named CIAO, based on GPL's framework, with additional provable implementation tricks that can facilitate learning. The demos of experimental results are available on https://sites.google. com/view/ciao2024 , and the code of experiments is published on https://github. com/hsvgbkhgbv/CIAO .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b0a725a9-b5cf-4c49-9026-ba6e92935b8fCited by top-tier papers3
- A Principle of Targeted Intervention for Multi-Agent Reinforcement LearningAnjie Liu, Jianhong Wang, Samuel Kaski, Jun Wang et al.NeurIPS 2025 · 3 citations
- Efficient Multi-agent Offline Coordination via Diffusion-based Trajectory StitchingLei Yuan, Yuqi Bian, Lihe Li, Ziqian Zhang et al.ICLR 2025
- Ad Hoc Teamwork via Offline Goal-Based Decision TransformersXinzhi Zhang, Hohei Chan, Deheng Ye, Yi Cai et al.ICML 2025
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 1,960 citations
- Multi-Agent Reinforcement Learning for Active Voltage Control on Power Distribution NetworksJianhong Wang, Wangkun Xu, Yunjie Gu, Wenbin Song et al.NeurIPS 2021 · 216 citations
- Shapley Q-Value: A Local Reward Approach to Solve Global Reward GamesJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuAAAI 2020 · 159 citations
- Towards Open Ad Hoc Teamwork Using Graph-based Policy LearningArrasy Rahman, Niklas Höpner, Filippos Christianos, Stefano V. AlbrechtICML 2021 · 75 citations
Related papers
- N-agent Ad Hoc TeamworkCaroline Wang, Arrasy Rahman, Ishan Durugkar, Elad Liebman et al.NeurIPS 2024 · 21 citations
- PADiff: Predictive and Adaptive Diffusion Policies for Ad Hoc TeamworkHohei Chan, Xinzhi Zhang, Antao Xiang, Weinan Zhang et al.AAAI 2026
- Deep Coordination GraphsWendelin Boehmer, Vitaly Kurin, Shimon WhitesonICML 2020 · 209 citations
- Cooperative Open-ended Learning Framework for Zero-Shot CoordinationYang Li, Shao Zhang, Jichen Sun, Yali Du et al.ICML 2023 · 35 citations
- Online Ad Hoc Teamwork under Partial ObservabilityPengjie Gu, Mengchen Zhao, Jianye Hao, Bo AnICLR 2022 · 35 citations
