Deconfounded Value Decomposition for Multi-Agent Reinforcement Learning
Jiahui Li, Kun Kuang, Baoxiang Wang, Furui Liu, Long Chen, Changjie Fan, Fei Wu, Jun Xiao
摘要
Value decomposition (VD) methods have been widely used in cooperative multi-agent reinforcement learning (MARL), where credit assignment plays an important role in guiding the agents' decentralized execution. In this paper, we investigate VD from a novel perspective of causal inference. We first show that the environment in existing VD methods is an unobserved confounder as the common cause factor of the global state and the joint value function, which leads to the confounding bias on learning credit assignment. We then present our approach, deconfounded value decomposition (DVD), which cuts off the backdoor confounding path from the global state to the joint value function. The cut is implemented by introducing the trajectory graph, which depends only on the local trajectories, as a proxy confounder. DVD is general enough to be applied to various VD methods, and extensive experiments show that DVD can consistently achieve significant performance gains over different stateof-the-art VD methods on StarCraft II and MACO benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Heterogeneous Agent Q-weighted Policy OptimizationBor-Jiun Lin, Chun-Yi LeeICLR 2026 · 被引用 102 次
- NA2Q: Neural Attention Additive Model for Interpretable Multi-Agent Q-LearningZichuan Liu, Yuanyang Zhu, Chunlin ChenICML 2023 · 被引用 25 次
- ConfounderGAN: Protecting Image Data Privacy with Causal ConfounderQi Tian, Kun Kuang, Kelu Jiang, Furui Liu 等NeurIPS 2022 · 被引用 11 次
- Two Heads are Better Than One: A Simple Exploration Framework for Efficient Multi-Agent Reinforcement LearningJiahui Li, Kun Kuang, Baoxiang Wang, Xingchen Li 等NeurIPS 2023 · 被引用 7 次
- High-order Interactions Modeling for Interpretable Multi-Agent Q-LearningQinyu Xu, Yuanyang Zhu, Xuefei Wu, Chunlin ChenNeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper8
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu 等ICLR 2021 · 被引用 595 次
- Deep Coordination GraphsWendelin Boehmer, Vitaly Kurin, Shimon WhitesonICML 2020 · 被引用 209 次
- Learning Nearly Decomposable Value Functions Via Communication MinimizationTonghan Wang, Jianhao Wang, Chongyi Zheng, Chongjie ZhangICLR 2020 · 被引用 170 次
- Counterfactual Critic Multi-Agent Training for Scene Graph GenerationLong Chen, Hanwang Zhang, Jun Xiao, Xiangnan He 等ICCV 2019 · 被引用 165 次
- Q-value Path Decomposition for Deep Multiagent Reinforcement LearningYaodong Yang, Jianye Hao, Guangyong Chen, Hongyao Tang 等ICML 2020 · 被引用 64 次
相关 Paper
- Contrastive Identity-Aware Learning for Multi-Agent Value DecompositionShunyu Liu, Yihe Zhou, Jie Song, Tongya Zheng 等AAAI 2023 · 被引用 43 次
- Shapley Counterfactual Credits for Multi-Agent Reinforcement LearningJiahui Li, Kun Kuang, Baoxiang Wang, Furui Liu 等KDD 2021 · 被引用 49 次
- Dual Self-Awareness Value Decomposition Framework without Individual Global Max for Cooperative MARLZhiwei Xu, Bin Zhang, Dapeng Li, Guangchong Zhou 等NeurIPS 2023 · 被引用 12 次
- Towards Understanding Cooperative Multi-Agent Q-Learning with Value FactorizationJianhao Wang, Zhizhou Ren, Beining Han, Jianing Ye 等NeurIPS 2021 · 被引用 50 次
- Solving Homogeneous and Heterogeneous Cooperative Tasks with Greedy Sequential ExecutionShanqi Liu, Dong Xing, Pengjie Gu, Xinrun Wang 等ICLR 2024 · 被引用 2 次
