Dealing with Non-Stationarity in MARL via Trust-Region Decomposition
Wenhao Li, Xiangfeng Wang, Bo Jin, Junjie Sheng, Hongyuan Zha
摘要
Non-stationarity is one thorny issue in cooperative multi-agent reinforcement learning (MARL). One of the reasons is the policy changes of agents during the learning process. Some existing works have discussed various consequences caused by non-stationarity with several kinds of measurement indicators. This makes the objectives or goals of existing algorithms are inevitably inconsistent and disparate. In this paper, we introduce a novel notion, the -measurement, to explicitly measure the non-stationarity of a policy sequence, which can be further proved to be bounded by the KL-divergence of consecutive joint policies. A straightforward but highly non-trivial way is to control the joint policies' divergence, which is difficult to estimate accurately by imposing the trust-region constraint on the joint policy. Although it has lower computational complexity to decompose the joint policy and impose trust-region constraints on the factorized policies, simple policy factorization like mean-field approximation will lead to more considerable policy divergence, which can be considered as the trust-region decomposition dilemma. We model the joint policy as a pairwise Markov random field and propose a trust-region decomposition network (TRD-Net) based on message passing to estimate the joint policy divergence more accurately. The Multi-Agent Mirror descent policy algorithm with Trust region decomposition, called MAMT, is established by adjusting the trust-region of the local policies adaptively in an end-to-end manner. MAMT can approximately constrain the consecutive joint policies' divergence to satisfy -stationarity and alleviate the non-stationarity problem. Our method can bring noticeable and stable performance improvement compared with baselines in cooperative tasks of different complexity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Order Matters: Agent-by-agent Policy OptimizationXihuai Wang, Zheng Tian, Ziyu Wan, Ying Wen 等ICLR 2023 · 被引用 3 次
- Distributed Influence-Augmented Local Simulators for Parallel MARL in Large Networked SystemsMiguel Suau, Jinke He, Mustafa Mert Çelikok, Matthijs T. J. Spaan 等NeurIPS 2022 · 被引用 2 次
它引用的顶会 Paper13
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu 等ICLR 2020 · 被引用 751 次
- Trust Region Policy Optimisation in Multi-Agent Reinforcement LearningJakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen 等ICLR 2022 · 被引用 367 次
- Towards Playing Full MOBA Games with Deep Reinforcement LearningDeheng Ye, Guibin Chen, Wen Zhang, Sheng Chen 等NeurIPS 2020 · 被引用 225 次
相关 Paper
- On the Hardness of Constrained Cooperative Multi-Agent Reinforcement LearningZiyi Chen, Yi Zhou, Heng HuangICLR 2024 · 被引用 6 次
- A Policy Gradient Algorithm for Learning to Learn in Multiagent Reinforcement LearningDong-Ki Kim, Miao Liu, Matthew Riemer, Chuangchuang Sun 等ICML 2021 · 被引用 66 次
- Scalable Constrained Policy Optimization for Safe Multi-agent Reinforcement LearningLijun Zhang, Lin Li, Wei Wei, Huizhong Song 等NeurIPS 2024 · 被引用 22 次
- In-Context Fully Decentralized Cooperative Multi-Agent Reinforcement LearningChao Li, Bingkun Bao, Yang GaoNeurIPS 2025 · 被引用 2 次
- Multi-agent Reinforcement Learning for Networked System ControlTianshu Chu, Sandeep Chinchali, Sachin KattiICLR 2020 · 被引用 134 次
