Open Ad Hoc Teamwork with Cooperative Game Theory
Jianhong Wang, Yang Li, Yuan Zhang, Wei Pan, Samuel Kaski
摘要
Ad hoc teamwork poses a challenging problem, requiring the design of an agent to collaborate with teammates without prior coordination or joint training. Open ad hoc teamwork (OAHT) further complicates this challenge by considering environments with a changing number of teammates, referred to as open teams. One promising solution in practice to this problem is leveraging the generalizability of graph neural networks to handle an unrestricted number of agents with various agenttypes, named graph-based policy learning (GPL). However, its joint Q-value representation over a coordination graph lacks convincing explanations. In this paper, we establish a new theory to understand the representation of the joint Q-value for OAHT and its learning paradigm, through the lens of cooperative game theory. Building on our theory, we propose a novel algorithm named CIAO, based on GPL's framework, with additional provable implementation tricks that can facilitate learning. The demos of experimental results are available on https://sites.google. com/view/ciao2024 , and the code of experiments is published on https://github. com/hsvgbkhgbv/CIAO .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- A Principle of Targeted Intervention for Multi-Agent Reinforcement LearningAnjie Liu, Jianhong Wang, Samuel Kaski, Jun Wang 等NeurIPS 2025 · 被引用 3 次
- Efficient Multi-agent Offline Coordination via Diffusion-based Trajectory StitchingLei Yuan, Yuqi Bian, Lihe Li, Ziqian Zhang 等ICLR 2025
- Ad Hoc Teamwork via Offline Goal-Based Decision TransformersXinzhi Zhang, Hohei Chan, Deheng Ye, Yi Cai 等ICML 2025
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 被引用 1,960 次
- Multi-Agent Reinforcement Learning for Active Voltage Control on Power Distribution NetworksJianhong Wang, Wangkun Xu, Yunjie Gu, Wenbin Song 等NeurIPS 2021 · 被引用 216 次
- Shapley Q-Value: A Local Reward Approach to Solve Global Reward GamesJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuAAAI 2020 · 被引用 159 次
- Towards Open Ad Hoc Teamwork Using Graph-based Policy LearningArrasy Rahman, Niklas Höpner, Filippos Christianos, Stefano V. AlbrechtICML 2021 · 被引用 75 次
相关 Paper
- N-agent Ad Hoc TeamworkCaroline Wang, Arrasy Rahman, Ishan Durugkar, Elad Liebman 等NeurIPS 2024 · 被引用 21 次
- PADiff: Predictive and Adaptive Diffusion Policies for Ad Hoc TeamworkHohei Chan, Xinzhi Zhang, Antao Xiang, Weinan Zhang 等AAAI 2026
- Deep Coordination GraphsWendelin Boehmer, Vitaly Kurin, Shimon WhitesonICML 2020 · 被引用 209 次
- Cooperative Open-ended Learning Framework for Zero-Shot CoordinationYang Li, Shao Zhang, Jichen Sun, Yali Du 等ICML 2023 · 被引用 35 次
- Online Ad Hoc Teamwork under Partial ObservabilityPengjie Gu, Mengchen Zhao, Jianye Hao, Bo AnICLR 2022 · 被引用 35 次
