Cooperative Policy Agreement: Learning Diverse Policy for Offline MARL
Yihe Zhou, Yuxuan Zheng, Yue Hu, Kaixuan Chen, Tongya Zheng, Jie Song, Mingli Song, Shunyu Liu
Abstract
Offline Multi-Agent Reinforcement Learning (MARL) aims to learn optimal joint policies from pre-collected datasets without further interaction with the environment. Despite the encouraging results achieved so far, we identify the policy mismatch problem that arises from employing diverse offline MARL datasets, a highly important ingredient for cooperative generalization yet largely overlooked by existing literature. Specifically, in the case that offline datasets exhibit various optimal joint policies, policy mismatch often occurs when individual actions from different optimal joint actions are combined in a way that results in a suboptimal joint action. In this paper, we introduce a novel Cooperative Policy Agreement (CPA) method, that not only mitigates the policy mismatch problem but also learns to generate diverse joint policies. CPA firstly introduces an autoregressive decision-making mechanism among agents during offline training. This mechanism enables agents to access the actions previously taken by other agents, thereby facilitating effective joint policy matching. Moreover, diverse joint policies can be directly obtained through sequential action sampling from the autoregressive model. Then we further incorporate a policy agreement mechanism to convert these autoregressive joint policies into decentralized policies with a non-autoregressive form, while still ensuring the diversity of the generated policies. This mechanism guarantees that the proposed CPA adheres to the Centralized Training with Decentralized Execution (CTDE) constraint. Experiments conducted on various benchmarks demonstrate that CPA yields superior performance to state-of-the-art competitors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on18
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Multi-Agent Reinforcement Learning is a Sequence Modeling ProblemMuning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang et al.NeurIPS 2022 · 408 citations
- Trust Region Policy Optimisation in Multi-Agent Reinforcement LearningJakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen et al.ICLR 2022 · 367 citations
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 271 citations
- Multi-Agent Reinforcement Learning for Active Voltage Control on Power Distribution NetworksJianhong Wang, Wangkun Xu, Yunjie Gu, Wenbin Song et al.NeurIPS 2021 · 216 citations
Related papers
- Offline Multi-Agent Reinforcement Learning via In-Sample Sequential Policy OptimizationZongkai Liu, Qian Lin, Chao Yu, Xiawei Wu et al.AAAI 2025 · 3 citations
- Enhancing Diffusion Policies with Distribution-Matching Generator in Offline Reinforcement LearningXuemin Hu, Shen Li, Yingfen Xu, Bo Tang et al.AAAI 2026 · 1 citation
- ComaDICE: Offline Cooperative Multi-Agent Reinforcement Learning with Stationary Distribution Shift RegularizationThe Viet Bui, Thanh Hong Nguyen, Tien Anh MaiICLR 2025
- Multi-Agent Guided Policy OptimizationYueheng Li, Guangming Xie, Zongqing LuICLR 2026 · 4 citations
- Offline Multi-Agent Reinforcement Learning with Knowledge DistillationWei-Cheng Tseng, Tsun-Hsuan Johnson Wang, Yen-Chen Lin, Phillip IsolaNeurIPS 2022 · 62 citations
