Retaining Suboptimal Actions to Follow Shifting Optima in Multi-Agent Reinforcement Learning
Yonghyeon Jo, Sunwoo Lee, Seungyul Han
摘要
Value decomposition is a core approach for cooperative multi-agent reinforcement learning (MARL). However, existing methods still rely on a single optimal action and struggle to adapt when the underlying value function shifts during training, often converging to suboptimal policies. To address this limitation, we propose Successive Sub-value Q-learning (S2Q), which learns multiple sub-value functions to retain alternative high-value actions. Incorporating these sub-value functions into a Softmax-based behavior policy, S2Q encourages persistent exploration and enables to adjust quickly to the changing optima. Experiments on challenging MARL benchmarks confirm that S2Q consistently outperforms various MARL algorithms, demonstrating improved adaptability and overall performance. Our code is available at https://github.com/hyeon1996/S2Q.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Self-Improving Skill Learning for Robust Skill-based Meta-Reinforcement LearningSeungyul Han, Sanghyeon Lee, Sangjun Bae, Yisak ParkICLR 2026 · 被引用 5 次
- Strict Subgoal Execution: Reliable Long-Horizon Planning in Hierarchical Reinforcement LearningSeungyul Han, Jaebak Hwang, Sanghyeon Lee, Jeongmo KimICLR 2026 · 被引用 3 次
- LLM-Guided Communication for Cooperative Multi-Agent Reinforcement LearningSangjun Bae, Yisak Park, Sanghyeon Lee, Seungyul HanICML 2026 · 被引用 2 次
- Interaction-Breaking Adversarial Learning Framework for Robust Multi-Agent Reinforcement LearningSunwoo Lee, Mingu Kang, Yonghyeon Jo, Seungyul HanICML 2026 · 被引用 1 次
它引用的顶会 Paper39
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 被引用 1,960 次
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu 等ICLR 2021 · 被引用 595 次
- Google Research Football: A Novel Reinforcement Learning EnvironmentKarol Kurach, Anton Raichuk, Piotr Stanczyk, Michal Zajac 等AAAI 2020 · 被引用 496 次
- Multi-Agent Reinforcement Learning is a Sequence Modeling ProblemMuning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang 等NeurIPS 2022 · 被引用 408 次
- FACMAC: Factored Multi-Agent Centralised Policy GradientsBei Peng, Tabish Rashid, Christian Schröder de Witt, Pierre-Alexandre Kamienny 等NeurIPS 2021 · 被引用 399 次
相关 Paper
- O-MAPL: Offline Multi-agent Preference LearningThe Viet Bui, Tien Mai, Thanh Hong NguyenICML 2025
- Solving Homogeneous and Heterogeneous Cooperative Tasks with Greedy Sequential ExecutionShanqi Liu, Dong Xing, Pengjie Gu, Xinrun Wang 等ICLR 2024 · 被引用 2 次
- UneVEn: Universal Value Exploration for Multi-Agent Reinforcement LearningTarun Gupta, Anuj Mahajan, Bei Peng, Wendelin Boehmer 等ICML 2021 · 被引用 59 次
- Resilient Multi-Agent Reinforcement Learning with Adversarial Value DecompositionThomy Phan, Lenz Belzner, Thomas Gabor, Andreas Sedlmeier 等AAAI 2021 · 被引用 34 次
- NA2Q: Neural Attention Additive Model for Interpretable Multi-Agent Q-LearningZichuan Liu, Yuanyang Zhu, Chunlin ChenICML 2023 · 被引用 25 次
