Backpropagation Through Agents
Zhiyuan Li, Wenshuai Zhao, Lijun Wu, Joni Pajarinen
Abstract
A fundamental challenge in multi-agent reinforcement learning (MARL) is to learn the joint policy in an extremely large search space, which grows exponentially with the number of agents. Moreover, fully decentralized policy factorization significantly restricts the search space, which may lead to sub-optimal policies. In contrast, the auto-regressive joint policy can represent a much richer class of joint policies by factorizing the joint policy into the product of a series of conditional individual policies. While such factorization introduces the action dependency among agents explicitly in sequential execution, it does not take full advantage of the dependency during learning. In particular, the subsequent agents do not give the preceding agents feedback about their decisions. In this paper, we propose a new framework Back-Propagation Through Agents (BPTA) that directly accounts for both agents' own policy updates and the learning of their dependent counterparts. This is achieved by propagating the feedback through action chains. With the proposed framework, our Bidirectional Proximal Policy Optimisation (BPPO) outperforms the state-of-the-art methods. Extensive experiments on matrix games, StarCraftII v2, Multi-agent MuJoCo, and Google Research Football demonstrate the effectiveness of the proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 778afb8d-6f16-470c-a0be-404fdcf397b2Cited by top-tier papers2
- AgentMixer: Multi-Agent Correlated Policy FactorizationZhiyuan Li, Wenshuai Zhao, Lijun Wu, Joni PajarinenAAAI 2025 · 7 citations
- Learning Progress Driven Multi-Agent CurriculumWenshuai Zhao, Zhiyuan Li, Joni PajarinenICML 2025
Builds on10
- Multi-Agent Reinforcement Learning is a Sequence Modeling ProblemMuning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang et al.NeurIPS 2022 · 408 citations
- FACMAC: Factored Multi-Agent Centralised Policy GradientsBei Peng, Tabish Rashid, Christian Schröder de Witt, Pierre-Alexandre Kamienny et al.NeurIPS 2021 · 399 citations
- DOP: Off-Policy Multi-Agent Decomposed Policy GradientsYihan Wang, Beining Han, Tonghan Wang, Heng Dong et al.ICLR 2021 · 208 citations
- FOP: Factorizing Optimal Joint Policy of Maximum-Entropy Multi-Agent Reinforcement LearningTianhao Zhang, Yueheng Li, Chen Wang, Guangming Xie et al.ICML 2021 · 88 citations
- Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement LearningWei Fu, Chao Yu, Zelai Xu, Jiaqi Yang et al.ICML 2022 · 49 citations
Related papers
- ACE: Cooperative Multi-Agent Q-learning with Bidirectional Action-DependencyChuming Li, Jie Liu, Yinmin Zhang, Yuhong Wei et al.AAAI 2023 · 38 citations
- Order Matters: Agent-by-agent Policy OptimizationXihuai Wang, Zheng Tian, Ziyu Wan, Ying Wen et al.ICLR 2023 · 3 citations
- Trust Region Policy Optimisation in Multi-Agent Reinforcement LearningJakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen et al.ICLR 2022 · 367 citations
- Coordinated Proximal Policy OptimizationZifan Wu, Chao Yu, Deheng Ye, Junge Zhang et al.NeurIPS 2021 · 73 citations
- Individual Reward Assisted Multi-Agent Reinforcement LearningLi Wang, Yupeng Zhang, Yujing Hu, Weixun Wang et al.ICML 2022 · 41 citations
