Retaining Suboptimal Actions to Follow Shifting Optima in Multi-Agent Reinforcement Learning
Yonghyeon Jo, Sunwoo Lee, Seungyul Han
Abstract
Value decomposition is a core approach for cooperative multi-agent reinforcement learning (MARL). However, existing methods still rely on a single optimal action and struggle to adapt when the underlying value function shifts during training, often converging to suboptimal policies. To address this limitation, we propose Successive Sub-value Q-learning (S2Q), which learns multiple sub-value functions to retain alternative high-value actions. Incorporating these sub-value functions into a Softmax-based behavior policy, S2Q encourages persistent exploration and enables to adjust quickly to the changing optima. Experiments on challenging MARL benchmarks confirm that S2Q consistently outperforms various MARL algorithms, demonstrating improved adaptability and overall performance. Our code is available at https://github.com/hyeon1996/S2Q.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4471444-de6f-426b-91a9-472ecc30908bCited by top-tier papers4
- Self-Improving Skill Learning for Robust Skill-based Meta-Reinforcement LearningSeungyul Han, Sanghyeon Lee, Sangjun Bae, Yisak ParkICLR 2026 · 5 citations
- Strict Subgoal Execution: Reliable Long-Horizon Planning in Hierarchical Reinforcement LearningSeungyul Han, Jaebak Hwang, Sanghyeon Lee, Jeongmo KimICLR 2026 · 3 citations
- LLM-Guided Communication for Cooperative Multi-Agent Reinforcement LearningSangjun Bae, Yisak Park, Sanghyeon Lee, Seungyul HanICML 2026 · 2 citations
- Interaction-Breaking Adversarial Learning Framework for Robust Multi-Agent Reinforcement LearningSunwoo Lee, Mingu Kang, Yonghyeon Jo, Seungyul HanICML 2026 · 1 citation
Builds on39
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 1,960 citations
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu et al.ICLR 2021 · 595 citations
- Google Research Football: A Novel Reinforcement Learning EnvironmentKarol Kurach, Anton Raichuk, Piotr Stanczyk, Michal Zajac et al.AAAI 2020 · 496 citations
- Multi-Agent Reinforcement Learning is a Sequence Modeling ProblemMuning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang et al.NeurIPS 2022 · 408 citations
- FACMAC: Factored Multi-Agent Centralised Policy GradientsBei Peng, Tabish Rashid, Christian Schröder de Witt, Pierre-Alexandre Kamienny et al.NeurIPS 2021 · 399 citations
Related papers
- O-MAPL: Offline Multi-agent Preference LearningThe Viet Bui, Tien Mai, Thanh Hong NguyenICML 2025
- Solving Homogeneous and Heterogeneous Cooperative Tasks with Greedy Sequential ExecutionShanqi Liu, Dong Xing, Pengjie Gu, Xinrun Wang et al.ICLR 2024 · 2 citations
- UneVEn: Universal Value Exploration for Multi-Agent Reinforcement LearningTarun Gupta, Anuj Mahajan, Bei Peng, Wendelin Boehmer et al.ICML 2021 · 59 citations
- Resilient Multi-Agent Reinforcement Learning with Adversarial Value DecompositionThomy Phan, Lenz Belzner, Thomas Gabor, Andreas Sedlmeier et al.AAAI 2021 · 34 citations
- NA2Q: Neural Attention Additive Model for Interpretable Multi-Agent Q-LearningZichuan Liu, Yuanyang Zhu, Chunlin ChenICML 2023 · 25 citations
