Context-Aware Bayesian Network Actor-Critic Methods for Cooperative Multi-Agent Reinforcement Learning
Dingyang Chen, Qi Zhang
Abstract
Executing actions in a correlated manner is a common strategy for human coordination that often leads to better cooperation, which is also potentially beneficial for cooperative multi-agent reinforcement learning (MARL). However, the recent success of MARL relies heavily on the convenient paradigm of purely decentralized execution, where there is no action correlation among agents for scalability considerations. In this work, we introduce a Bayesian network to inaugurate correlations between agents' action selections in their joint policy. Theoretically, we establish a theoretical justification for why action dependencies are beneficial by deriving the multi-agent policy gradient formula under such a Bayesian network joint policy and proving its global convergence to Nash equilibria under tabular softmax policy parameterization in cooperative Markov games. Further, by equipping existing MARL algorithms with a recent method of differentiable directed acyclic graphs (DAGs), we develop practical algorithms to learn the context-aware Bayesian network policies in scenarios with partial observability and various difficulty. We also dynamically decrease the sparsity of the learned DAG throughout the training process, which leads to weakly or even purely independent policies for decentralized execution. Empirical results on a range of MARL benchmarks show the benefits of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f1b4cf7-cee2-4d93-bd47-fd79e25d63bfCited by top-tier papers1
Ask how each one uses itBuilds on7
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu et al.ICLR 2020 · 751 citations
- Deep Coordination GraphsWendelin Boehmer, Vitaly Kurin, Shimon WhitesonICML 2020 · 209 citations
- Global Convergence of Multi-Agent Policy Gradient in Markov Potential GamesStefanos Leonardos, Will Overman, Ioannis Panageas, Georgios PiliourasICLR 2022 · 158 citations
- Discovering Diverse Multi-Agent Strategic Behavior via Reward RandomizationZhenggang Tang, Chao Yu, Boyuan Chen, Huazhe Xu et al.ICLR 2021 · 63 citations
- Differentiable DAG SamplingBertrand Charpentier, Simon Kibler, Stephan GünnemannICLR 2022 · 51 citations
Related papers
- Bayesian Ego-graph Inference for Networked Multi-Agent Reinforcement LearningWei Duan, Jie Lu, Junyu XuanNeurIPS 2025 · 15 citations
- Bridging Training and Execution via Dynamic Directed Graph-Based Communication in Cooperative Multi-Agent SystemsZhuohui Zhang, Bin He, Bin Cheng, Gang LiAAAI 2025 · 10 citations
- MARLIN: Multi-Agent Reinforcement Learning for Incremental DAG DiscoveryDong Li, Zhengzhang Chen, Xujiang Zhao, Linlin Yu et al.AAAI 2026
- Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task SimilarityWeichao Mao, Haoran Qiu, Chen Wang, Hubertus Franke et al.NeurIPS 2023 · 17 citations
- Multi-Agent Guided Policy OptimizationYueheng Li, Guangming Xie, Zongqing LuICLR 2026 · 4 citations
