Potentially Optimal Joint Actions Recognition for Cooperative Multi-Agent Reinforcement Learning
Chang Huang, Shatong Zhu, Junqiao Zhao, Hongtu Zhou, Hai Zhang, Di Zhang, Chen Ye, Ziqiao Wang, Guang Chen
摘要
Value function factorization is widely used in cooperative multi-agent reinforcement learning (MARL). Existing approaches often impose monotonicity constraints between the joint action value and individual action values to enable decentralized execution. However, such constraints limit the expressiveness of value factorization, restricting the range of joint action values that can be represented and hindering the learning of optimal policies. To address this, we propose Potentially Optimal Joint Actions Weighting (POW), a method that ensures optimal policy recovery where existing approximate weighting strategies may fail. POW iteratively identifies potentially optimal joint actions and assigns them higher training weights through a theoretically grounded iterative weighted training process. We prove that this mechanism guarantees recovery of the true optimal policy, overcoming the limitations of prior heuristic weighting strategies. POW is architecture-agnostic and can be seamlessly integrated into existing value factorization algorithms. Extensive experiments on matrix games, difficulty-enhanced predator-prey tasks, SMAC, SMACv2, and a highway-env intersection scenario show that POW substantially improves stability and consistently surpasses state-of-the-art value-based MARL methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 被引用 1,960 次
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu 等ICLR 2021 · 被引用 595 次
- PettingZoo: Gym for Multi-Agent Reinforcement LearningJ. K. Terry, Benjamin Black, Nathaniel Grammel, Mario Jayakumar 等NeurIPS 2021 · 被引用 478 次
- FOP: Factorizing Optimal Joint Policy of Maximum-Entropy Multi-Agent Reinforcement LearningTianhao Zhang, Yueheng Li, Chen Wang, Guangming Xie 等ICML 2021 · 被引用 88 次
- Contrastive Identity-Aware Learning for Multi-Agent Value DecompositionShunyu Liu, Yihe Zhou, Jie Song, Tongya Zheng 等AAAI 2023 · 被引用 43 次
相关 Paper
- More Centralized Training, Still Decentralized Execution: Multi-Agent Conditional Policy FactorizationJiangxing Wang, Deheng Ye, Zongqing LuICLR 2023 · 被引用 5 次
- ResQ: A Residual Q Function-based Approach for Multi-Agent Reinforcement Learning Value FactorizationSiqi Shen, Mengwei Qiu, Jun Liu, Weiquan Liu 等NeurIPS 2022 · 被引用 35 次
- DFAC Framework: Factorizing the Value Function via Quantile Mixture for Multi-Agent Distributional Q-LearningWei-Fang Sun, Cheng-Kuang Lee, Chun-Yi LeeICML 2021 · 被引用 56 次
- In-Context Fully Decentralized Cooperative Multi-Agent Reinforcement LearningChao Li, Bingkun Bao, Yang GaoNeurIPS 2025 · 被引用 2 次
- PAC: Assisted Value Factorization with Counterfactual Predictions in Multi-Agent Reinforcement LearningHanhan Zhou, Tian Lan, Vaneet AggarwalNeurIPS 2022 · 被引用 47 次
