Regularized Softmax Deep Multi-Agent Q-Learning
Ling Pan, Tabish Rashid, Bei Peng, Longbo Huang, Shimon Whiteson
Abstract
Tackling overestimation in -learning is an important problem that has been extensively studied in single-agent reinforcement learning, but has received comparatively little attention in the multi-agent setting. In this work, we empirically demonstrate that QMIX, a popular -learning algorithm for cooperative multi-agent reinforcement learning (MARL), suffers from a more severe overestimation in practice than previously acknowledged, and is not mitigated by existing approaches. We rectify this with a novel regularization-based update scheme that penalizes large joint action-values that deviate from a baseline and demonstrate its effectiveness in stabilizing learning. Furthermore, we propose to employ a softmax operator, which we efficiently approximate in a novel way in the multi-agent setting, to further reduce the potential overestimation bias. Our approach, Regularized Softmax (RES) Deep Multi-Agent -Learning, is general and can be applied to any -learning based MARL algorithm. We demonstrate that, when applied to QMIX, RES avoids severe overestimation and significantly improves performance, yielding state-of-the-art results in a variety of cooperative multi-agent tasks, including the challenging StarCraft II micromanagement benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b62948eb-943f-4299-929f-28886f9ea2ebCited by top-tier papers4
- Towards a Standardised Performance Evaluation Protocol for Cooperative MARLRihab Gorsane, Omayma Mahjoub, Ruan de Kock, Roland Dubb et al.NeurIPS 2022 · 79 citations
- Counterfactual Conservative Q Learning for Offline Multi-agent Reinforcement LearningJianzhun Shao, Yun Qu, Chen Chen, Hongchang Zhang et al.NeurIPS 2023 · 56 citations
- Divergence-Regularized Multi-Agent Actor-CriticKefan Su, Zongqing LuICML 2022 · 31 citations
- Value-Based Deep Multi-Agent Reinforcement Learning with Dynamic Sparse TrainingPihe Hu, Shaolong Li, Zhuoran Li, Ling Pan et al.NeurIPS 2024 · 2 citations
Builds on8
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 1,960 citations
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu et al.ICLR 2021 · 595 citations
- Maxmin Q-learning: Controlling the Estimation Bias of Q-learningQingfeng Lan, Yangchen Pan, Alona Fyshe, Martha WhiteICLR 2020 · 213 citations
- Softmax Deep Double Deterministic Policy GradientsLing Pan, Qingpeng Cai, Longbo HuangNeurIPS 2020 · 138 citations
- UneVEn: Universal Value Exploration for Multi-Agent Reinforcement LearningTarun Gupta, Anuj Mahajan, Bei Peng, Wendelin Boehmer et al.ICML 2021 · 59 citations
Related papers
- FACMAC: Factored Multi-Agent Centralised Policy GradientsBei Peng, Tabish Rashid, Christian Schröder de Witt, Pierre-Alexandre Kamienny et al.NeurIPS 2021 · 399 citations
- Towards Understanding Cooperative Multi-Agent Q-Learning with Value FactorizationJianhao Wang, Zhizhou Ren, Beining Han, Jianing Ye et al.NeurIPS 2021 · 50 citations
- Inverse Factorized Soft Q-Learning for Cooperative Multi-agent Imitation LearningThe Viet Bui, Tien Mai, Thanh Hong NguyenNeurIPS 2024 · 10 citations
- Retaining Suboptimal Actions to Follow Shifting Optima in Multi-Agent Reinforcement LearningYonghyeon Jo, Sunwoo Lee, Seungyul HanICLR 2026 · 5 citations
- Value-Decomposition Multi-Agent Actor-CriticsJianyu Su, Stephen C. Adams, Peter A. BelingAAAI 2021 · 140 citations
