Individual Reward Assisted Multi-Agent Reinforcement Learning
Li Wang, Yupeng Zhang, Yujing Hu, Weixun Wang, Chongjie Zhang, Yang Gao, Jianye Hao, Tangjie Lv, Changjie Fan
Abstract
In many real-world multi-agent systems, the sparsity of team rewards often makes it difficult for an algorithm to successfully learn a cooperative team policy. At present, the common way for solving this problem is to design some dense individual rewards for the agents to guide the cooperation. However, most existing works utilize individual rewards in ways that do not always promote teamwork and sometimes are even counterproductive. In this paper, we propose Individual Reward Assisted Team Policy Learning (IRAT), which learns two policies for each agent from the dense individual reward and the sparse team reward with discrepancy constraints for updating the two policies mutually. Experimental results in different scenarios, such as the Multi-Agent Particle Environment and the Google Research Football Environment, show that IRAT significantly outperforms the baseline methods and can greatly promote team policy learning without deviating from the original team objective, even when the individual rewards are misleading or conflict with the team rewards.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68145f9c-d507-4901-abed-35d577a37befCited by top-tier papers3
- FoX: Formation-Aware Exploration in Multi-Agent Reinforcement LearningYonghyeon Jo, Sunwoo Lee, Junghyuk Yeom, Seungyul HanAAAI 2024 · 22 citations
- Gradient-Guided Credit Assignment and Joint Optimization for Dependency-Aware Spatial CrowdsourcingYafei Li, Wei Chen, Jinxing Yan, Huiling Li et al.AAAI 2025 · 3 citations
- M³HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed QualityZiyan Wang, Zhicheng Zhang, Fei Fang, Yali DuICML 2025
Builds on6
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu et al.ICLR 2021 · 595 citations
- Google Research Football: A Novel Reinforcement Learning EnvironmentKarol Kurach, Anton Raichuk, Piotr Stanczyk, Michal Zajac et al.AAAI 2020 · 496 citations
- Towards Playing Full MOBA Games with Deep Reinforcement LearningDeheng Ye, Guibin Chen, Wen Zhang, Sheng Chen et al.NeurIPS 2020 · 225 citations
- DOP: Off-Policy Multi-Agent Decomposed Policy GradientsYihan Wang, Beining Han, Tonghan Wang, Heng Dong et al.ICLR 2021 · 208 citations
Related papers
- Promoting Coordination through Policy Regularization in Multi-Agent Deep Reinforcement LearningJulien Roy, Paul Barde, Félix G. Harvey, Derek Nowrouzezahrai et al.NeurIPS 2020 · 25 citations
- ELIGN: Expectation Alignment as a Multi-Agent Intrinsic RewardZixian Ma, Rose E. Wang, Fei-Fei Li, Michael S. Bernstein et al.NeurIPS 2022 · 22 citations
- Lazy Agents: A New Perspective on Solving Sparse Reward Problem in Multi-agent Reinforcement LearningBoyin Liu, Zhiqiang Pu, Yi Pan, Jianqiang Yi et al.ICML 2023 · 34 citations
- Backpropagation Through AgentsZhiyuan Li, Wenshuai Zhao, Lijun Wu, Joni PajarinenAAAI 2024 · 3 citations
- Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement LearningWei Fu, Chao Yu, Zelai Xu, Jiaqi Yang et al.ICML 2022 · 49 citations
