Stabilizing Q Learning Via Soft Mellowmax Operator
Yaozhong Gan, Zhe Zhang, Xiaoyang Tan
Abstract
Learning complicated value functions in high dimensional state space by function approximation is a challenging task, partially due to that the max-operator used in temporal difference updates can theoretically cause instability for most linear or non-linear approximation schemes. Mellowmax is a recently proposed differentiable and non-expansion softmax operator that allows a convergent behavior in learning and planning. Unfortunately, the performance bound for the fixed point it converges to remains unclear, and in practice, its parameter is sensitive to various domains and has to be tuned case by case. Finally, the Mellowmax operator may suffer from oversmoothing as it ignores the probability being taken for each action when aggregating them. In this paper, we address all the above issues with an enhanced Mellowmax operator, named SM2 (Soft Mellowmax). Particularly, the proposed operator is reliable, easy to implement, and has provable performance guarantee, while preserving all the advantages of Mellowmax. Furthermore, we show that our SM2 operator can be applied to the challenging multi-agent reinforcement learning scenarios, leading to stable value function approximation and state of the art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Regularized Softmax Deep Multi-Agent Q-LearningLing Pan, Tabish Rashid, Bei Peng, Longbo Huang et al.NeurIPS 2021 · 50 citations
- Bridging the performance-gap between target-free and target-based reinforcement learningThéo Vincent, Yogesh Tripathi, Tim Lukas Faust, Abdullah Akgül et al.ICLR 2026 · 6 citations
- Use the Online Network If You Can: Towards Fast and Stable Reinforcement LearningAhmed Hendawy, Henrik Metternich, Théo Vincent, Mahdi Kallel et al.ICLR 2026 · 4 citations
- Smoothing Advantage LearningYaozhong Gan, Zhe Zhang, Xiaoyang TanAAAI 2022 · 3 citations
- Value-Based Deep Multi-Agent Reinforcement Learning with Dynamic Sparse TrainingPihe Hu, Shaolong Li, Zhuoran Li, Ling Pan et al.NeurIPS 2024 · 2 citations
Related papers
- Damped Anderson Mixing for Deep Reinforcement Learning: Acceleration, Convergence, and StabilizationKe Sun, Yafei Wang, Yi Liu, Yingnan Zhao et al.NeurIPS 2021 · 17 citations
- Softmax Deep Double Deterministic Policy GradientsLing Pan, Qingpeng Cai, Longbo HuangNeurIPS 2020 · 138 citations
- Retaining Suboptimal Actions to Follow Shifting Optima in Multi-Agent Reinforcement LearningYonghyeon Jo, Sunwoo Lee, Seungyul HanICLR 2026 · 5 citations
- The In-Sample Softmax for Offline Reinforcement LearningChenjun Xiao, Han Wang, Yangchen Pan, Adam White et al.ICLR 2023 · 3 citations
- Controlling Underestimation Bias in Reinforcement Learning via Quasi-median OperationWei Wei, Yujia Zhang, Jiye Liang, Lin Li et al.AAAI 2022 · 20 citations
