RMIX: Learning Risk-Sensitive Policies for Cooperative Reinforcement Learning Agents
Wei Qiu, Xinrun Wang, Runsheng Yu, Rundong Wang, Xu He, Bo An, Svetlana Obraztsova, Zinovi Rabinovich
摘要
Current value-based multi-agent reinforcement learning methods optimize individual Q values to guide individuals' behaviours via centralized training with decentralized execution (CTDE). However, such expected, i.e., risk-neutral, Q value is not sufficient even with CTDE due to the randomness of rewards and the uncertainty in environments, which causes the failure of these methods to train coordinating agents in complex environments. To address these issues, we propose RMIX, a novel cooperative MARL method with the Conditional Value at Risk (CVaR) measure over the learned distributions of individuals' Q values. Specifically, we first learn the return distributions of individuals to analytically calculate CVaR for decentralized execution. Then, to handle the temporal nature of the stochastic outcomes during executions, we propose a dynamic risk level predictor for risk level tuning. Finally, we optimize the CVaR policies with CVaR values used to estimate the target in TD error during centralized training and the CVaR values are used as auxiliary local rewards to update the local distribution via Quantile Regression loss. Empirically, we show that our method significantly outperforms state-of-the-art methods on challenging StarCraft II tasks, demonstrating enhanced coordination and improved sample efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Efficient Risk-Averse Reinforcement LearningIdo Greenberg, Yinlam Chow, Mohammad Ghavamzadeh, Shie MannorNeurIPS 2022 · 被引用 61 次
- ResQ: A Residual Q Function-based Approach for Multi-Agent Reinforcement Learning Value FactorizationSiqi Shen, Mengwei Qiu, Jun Liu, Weiquan Liu 等NeurIPS 2022 · 被引用 35 次
- Sample and Communication-Efficient Decentralized Actor-Critic Algorithms with Finite-Time AnalysisZiyi Chen, Yi Zhou, Rong-Rong Chen, Shaofeng ZouICML 2022 · 被引用 35 次
- RiskQ: Risk-sensitive Multi-Agent Reinforcement Learning Value FactorizationSiqi Shen, Chennan Ma, Chao Li, Weiquan Liu 等NeurIPS 2023 · 被引用 34 次
- Exploring both Individuality and Cooperation for Air-Ground Spatial Crowdsourcing by Multi-Agent Deep Reinforcement LearningYuxiao Ye, Chi Harold Liu, Zipeng Dai, Jianxin Zhao 等ICDE 2023 · 被引用 26 次
它引用的顶会 Paper3
- Maxmin Q-learning: Controlling the Estimation Bias of Q-learningQingfeng Lan, Yangchen Pan, Alona Fyshe, Martha WhiteICLR 2020 · 被引用 213 次
- Being Optimistic to Be Conservative: Quickly Learning a CVaR PolicyRamtin Keramati, Christoph Dann, Alex Tamkin, Emma BrunskillAAAI 2020 · 被引用 86 次
- Randomized Ensembled Double Q-Learning: Learning Fast Without a ModelXinyue Chen, Che Wang, Zijian Zhou, Keith W. RossICLR 2021 · 被引用 26 次
相关 Paper
- SMIX(λ): Enhancing Centralized Value Functions for Cooperative Multi-Agent Reinforcement LearningChao Wen, Xinghu Yao, Yuhui Wang, Xiaoyang TanAAAI 2020 · 被引用 57 次
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 被引用 1,960 次
- Shapley Counterfactual Credits for Multi-Agent Reinforcement LearningJiahui Li, Kun Kuang, Baoxiang Wang, Furui Liu 等KDD 2021 · 被引用 49 次
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu 等ICLR 2021 · 被引用 595 次
- Distributionally Robust Cooperative Multi-agent Reinforcement Learning with Value FactorizationChengrui Qu, Christopher Yeh, Kishan Panaganti, Eric Mazumdar 等ICLR 2026
