Distributional Reward Estimation for Effective Multi-agent Deep Reinforcement Learning
Jifeng Hu, Yanchao Sun, Hechang Chen, Sili Huang, Haiyin Piao, Yi Chang, Lichao Sun
Abstract
Multi-agent reinforcement learning has drawn increasing attention in practice, e.g., robotics and automatic driving, as it can explore optimal policies using samples generated by interacting with the environment. However, high reward uncertainty still remains a problem when we want to train a satisfactory model, because obtaining high-quality reward feedback is usually expensive and even infeasible. To handle this issue, previous methods mainly focus on passive reward correction. At the same time, recent active reward estimation methods have proven to be a recipe for reducing the effect of reward uncertainty. In this paper, we propose a novel Distributional Reward Estimation framework for effective Multi-Agent Reinforcement Learning (DRE-MARL). Our main idea is to design the multi-action-branch reward estimation and policy-weighted reward aggregation for stabilized training. Specifically, we design the multi-action-branch reward estimation to model reward distributions on all action branches. Then we utilize reward aggregation to obtain stable updating signals during training. Our intuition is that consideration of all possible consequences of actions could be useful for learning policies. The superiority of the DRE-MARL is demonstrated using benchmark multi-agent scenarios, compared with the SOTA baselines in terms of both effectiveness and robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f5f49d7-bb26-43ac-adc3-c98a4abb83d0Cited by top-tier papers7
- Learning Generalizable Agents via Saliency-guided Features DecorrelationSili Huang, Yanchao Sun, Jifeng Hu, Siyuan Guo et al.NeurIPS 2023 · 13 citations
- Analytic Energy-Guided Policy Optimization for Offline Reinforcement LearningJifeng Hu, Sili Huang, Zhejian Yang, Shengchao Hu et al.NeurIPS 2025 · 4 citations
- A Reward-Free Viewpoint on Multi-Objective Reinforcement LearningYing-Tu Chen, Wei Hung, Bing-Shu Wu, Zhang-Wei Hong et al.ICLR 2026 · 2 citations
- Tackling Continual Offline RL through Selective Weights Activation on Aligned SpacesJifeng Hu, Sili Huang, Li Shen, Zhejian Yang et al.NeurIPS 2025 · 2 citations
- Safe Multi-Agent Reinforcement Learning via Distributional Safety Critic and Maximum Entropy OptimizationQiwei Liu, Ye Yuan, Lingyue Zhang, Kaitian Chen et al.AAAI 2026
Builds on15
- Multi-Agent Game Abstraction via Graph Attention Neural NetworkYong Liu, Weixun Wang, Yujing Hu, Jianye Hao et al.AAAI 2020 · 316 citations
- Provably End-to-end Label-noise Learning without Anchor PointsXuefeng Li, Tongliang Liu, Bo Han, Gang Niu et al.ICML 2021 · 161 citations
- Reinforcement Learning with Perturbed RewardsJingkang Wang, Yang Liu, Bo LiAAAI 2020 · 161 citations
- Robust Multi-Agent Reinforcement Learning with Model UncertaintyKaiqing Zhang, Tao Sun, Yunzhe Tao, Sahika Genc et al.NeurIPS 2020 · 118 citations
- Clusterability as an Alternative to Anchor Points When Learning with Noisy LabelsZhaowei Zhu, Yiwen Song, Yang LiuICML 2021 · 112 citations
Related papers
- The Distributional Reward Critic Framework for Reinforcement Learning Under Perturbed RewardsXi Chen, Zhihui Zhu, Andrew PerraultAAAI 2025
- Proactive Multi-Camera Collaboration for 3D Human Pose EstimationHai Ci, Mickel Liu, Xuehai Pan, Fangwei Zhong et al.ICLR 2023 · 6 citations
- ADDQ: Adaptive distributional double Q-learningLeif Döring, Benedikt Wille, Maximilian Birr, Mihail Bîrsan et al.ICML 2025
- Non-Crossing Quantile Regression for Distributional Reinforcement LearningFan Zhou, Jianing Wang, Xingdong FengNeurIPS 2020 · 63 citations
- Variance Control for Distributional Reinforcement LearningQi Kuang, Zhoufan Zhu, Liwen Zhang, Fan ZhouICML 2023 · 4 citations
