AgentMixer: Multi-Agent Correlated Policy Factorization
Zhiyuan Li, Wenshuai Zhao, Lijun Wu, Joni Pajarinen
摘要
In multi-agent reinforcement learning, centralized training with decentralized execution (CTDE) methods typically assumes that agents make decisions based on their local observations independently, which may not lead to a correlated joint policy with coordination. Coordination can be explicitly encouraged during training and individual policies can be trained to imitate the correlated joint policy. However, this may lead to an asymmetric learning failure due to the observation mismatch between the joint and individual policies. Inspired by the concept of correlated equilibrium, we introduce a strategy modification called AgentMixer that allows agents to correlate their policies. AgentMixer combines individual partially observable policies into a joint fully observable policy non-linearly. To enable decentralized execution, we introduce Individual-Global-Consistency to guarantee mode consistency during joint training of the centralized and decentralized policies and prove that AgentMixer converges to an ϵ-approximate Correlated Equilibrium. In the Multi-Agent MuJoCo, SMAC-v2, Matrix Game, and Predator-Prey benchmarks, AgentMixer outperforms or matches state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper15
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Multi-Agent Reinforcement Learning is a Sequence Modeling ProblemMuning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang 等NeurIPS 2022 · 被引用 408 次
- FACMAC: Factored Multi-Agent Centralised Policy GradientsBei Peng, Tabish Rashid, Christian Schröder de Witt, Pierre-Alexandre Kamienny 等NeurIPS 2021 · 被引用 399 次
- Trust Region Policy Optimisation in Multi-Agent Reinforcement LearningJakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen 等ICLR 2022 · 被引用 367 次
- DOP: Off-Policy Multi-Agent Decomposed Policy GradientsYihan Wang, Beining Han, Tonghan Wang, Heng Dong 等ICLR 2021 · 被引用 208 次
相关 Paper
- Intrinsic Action Tendency Consistency for Cooperative Multi-Agent Reinforcement LearningJunkai Zhang, Yifan Zhang, Xi Sheryl Zhang, Yifan Zang 等AAAI 2024 · 被引用 9 次
- More Centralized Training, Still Decentralized Execution: Multi-Agent Conditional Policy FactorizationJiangxing Wang, Deheng Ye, Zongqing LuICLR 2023 · 被引用 5 次
- Multi-Agent Guided Policy OptimizationYueheng Li, Guangming Xie, Zongqing LuICLR 2026 · 被引用 4 次
- Rethinking Individual Global Max in Cooperative Multi-Agent Reinforcement LearningYitian Hong, Yaochu Jin, Yang TangNeurIPS 2022 · 被引用 40 次
- Agent-Centric Actor-Critic for Asynchronous Multi-Agent Reinforcement LearningWhiyoung Jung, Sunghoon Hong, Deunsol Yoon, Kanghoon Lee 等ICML 2025
