Self-Organized Polynomial-Time Coordination Graphs
Qianlan Yang, Weijun Dong, Zhizhou Ren, Jianhao Wang, Tonghan Wang, Chongjie Zhang
Abstract
Coordination graph is a promising approach to model agent collaboration in multi-agent reinforcement learning. It conducts a graph-based value factorization and induces explicit coordination among agents to complete complicated tasks. However, one critical challenge in this paradigm is the complexity of greedy action selection with respect to the factorized values. It refers to the decentralized constraint optimization problem (DCOP), which and whose constant-ratio approximation are NP-hard problems. To bypass this systematic hardness, this paper proposes a novel method, named Self-Organized Polynomial-time Coordination Graphs (SOP-CG), which uses structured graph classes to guarantee the accuracy and the computational efficiency of collaborated action selection. SOP-CG employs dynamic graph topology to ensure sufficient value function expressiveness. The graph selection is unified into an end-to-end learning paradigm. In experiments, we show that our approach learns succinct and well-adapted graph topologies, induces effective coordination, and improves performance across a variety of cooperative multi-agent tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47e3fbbb-282c-4a7c-b64e-6212422650efCited by top-tier papers9
- Bayesian Ego-graph Inference for Networked Multi-Agent Reinforcement LearningWei Duan, Jie Lu, Junyu XuanNeurIPS 2025 · 15 citations
- Non-Linear Coordination GraphsYipeng Kang, Tonghan Wang, Qianlan Yang, Xiaoran Wu et al.NeurIPS 2022 · 14 citations
- Multi-Sender Persuasion: A Computational PerspectiveSafwan Hossain, Tonghan Wang, Tao Lin, Yiling Chen et al.ICML 2024 · 14 citations
- The Bandit Whisperer: Communication Learning for Restless BanditsYunfan Zhao, Tonghan Wang, Dheeraj Mysore Nagaraj, Aparna Taneja et al.AAAI 2025 · 6 citations
- More Centralized Training, Still Decentralized Execution: Multi-Agent Conditional Policy FactorizationJiangxing Wang, Deheng Ye, Zongqing LuICLR 2023 · 5 citations
Builds on6
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 1,960 citations
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu et al.ICLR 2021 · 595 citations
- Deep Coordination GraphsWendelin Boehmer, Vitaly Kurin, Shimon WhitesonICML 2020 · 209 citations
- Learning Nearly Decomposable Value Functions Via Communication MinimizationTonghan Wang, Jianhao Wang, Chongyi Zheng, Chongjie ZhangICLR 2020 · 170 citations
- Context-Aware Sparse Deep Coordination GraphsTonghan Wang, Liang Zeng, Weijun Dong, Qianlan Yang et al.ICLR 2022 · 40 citations
Related papers
- Multi Agent Reinforcement Learning for Sequential Satellite Assignment ProblemsJoshua Holder, Natasha Jaques, Mehran MesbahiAAAI 2025 · 5 citations
- Unified and Generalizable Reinforcement Learning for Facility Location Problems on GraphsWenxuan Guo, Runzhong Wang, Yanyan Xu, Yaohui JinWWW 2025 · 2 citations
- Learning to Simulate Self-driven Particles System with Coordinated Policy OptimizationZhenghao Peng, Quanyi Li, Ka-Ming Hui, Chunxiao Liu et al.NeurIPS 2021 · 88 citations
- CSO: Constraint-Guided Space Optimization for Active Scene MappingXuefeng Yin, Chenyang Zhu, Shanglai Qu, Yuqi Li et al.ACM MM 2024
- Stateful Active Facilitator: Coordination and Environmental Heterogeneity in Cooperative Multi-Agent Reinforcement LearningDianbo Liu, Vedant Shah, Oussama Boussif, Cristian Meo et al.ICLR 2023 · 1 citation
