Who Matters Matters: Agent-Specific Conservative Offline MARL
Haosheng Chen, Yun Hua, Wenhao Li, Shiqin Wang, Xiangfeng Wang
Abstract
Offline Multi-Agent Reinforcement Learning (MARL) enables policy learning from static datasets in multi-agent systems, eliminating the need for risky or costly environment interactions during training. A central challenge in offline MARL lies in achieving effective collaboration among heterogeneous agents under the constraints of fixed datasets, where conservatism is introduced to restrict behaviors to data-supported distributions. Agents with distinct roles and capabilities require individualized conservatism - yet must maintain cohesive team performance. However, existing approaches often apply uniform conservatism across all agents, leading to over-constraining critical agents and under-constraining others, which hampers effective collaboration. To address this issue, a novel framework, OMCDA, is proposed, where the degree of conservatism is dynamically adjusted for individual agents based on their impact on overall system performance. The framework is characterized by two key innovations: (1) A decomposed Q-function architecture is introduced to disentangle return computation from policy deviation assessment, allowing precise evaluations of each agent's contribution; and (2) An adaptive conservatism mechanism is developed to scale constraint strength according to both behavior policy divergence and the estimated importance of agents to the system. Experiments on MuJoCo and SMAC show OMCDA outperforms existing offline MARL methods, effectively balancing the flexibility and conservatism across agents while ensuring fair credit assignment and better collaboration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- Trust Region Policy Optimisation in Multi-Agent Reinforcement LearningJakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen et al.ICLR 2022 · 367 citations
Related papers
- Partial Action Replacement: Tackling Distribution Shift in Offline MARLYue Jin, Giovanni MontanaAAAI 2026 · 1 citation
- ComaDICE: Offline Cooperative Multi-Agent Reinforcement Learning with Stationary Distribution Shift RegularizationThe Viet Bui, Thanh Hong Nguyen, Tien Anh MaiICLR 2025
- Learning from Good Trajectories in Offline Multi-Agent Reinforcement LearningQi Tian, Kun Kuang, Furui Liu, Baoxiang WangAAAI 2023 · 14 citations
- Believe What You See: Implicit Constraint Approach for Offline Multi-Agent Reinforcement LearningYiqin Yang, Xiaoteng Ma, Chenghao Li, Zewu Zheng et al.NeurIPS 2021 · 133 citations
- AlberDICE: Addressing Out-Of-Distribution Joint Actions in Offline Multi-Agent RL via Alternating Stationary Distribution Correction EstimationDaiki E. Matsunaga, Jongmin Lee, Jaeseok Yoon, Stefanos Leonardos et al.NeurIPS 2023 · 11 citations
