Multi-Agent Imitation by Learning and Sampling from Factorized Soft Q-Function
Yi-Chen Li, Zhongxiang Ling, Tao Jiang, Fuxiang Zhang, Pengyuan Wang, Lei Yuan, Zongzhang Zhang, Yang Yu
摘要
Learning from multi-agent expert demonstrations, known as Multi-Agent Imitation Learning (MAIL), provides a promising approach to sequential decision-making. However, existing MAIL methods including Behavior Cloning (BC) and Adversarial Imitation Learning (AIL) face significant challenges: BC suffers from the compounding error issue, while the very nature of adversarial optimization makes AIL prone to instability. In this work, we propose M ulti-A gent imitation by learning and sampling from F actor I zed S oft Q-function (MAFIS), a novel method that addresses these limitations for both online and offline MAIL settings. Built upon the single-agent IQ-Learn framework, MAFIS introduces the value decomposition network to factorize the imitation objective at agent level, thus enabling scalable training for multi-agent systems. Moreover, we observe that the soft Q-function implicitly defines the optimal policy as an energy-based model, from which we can sample actions via stochastic gradient Langevin dynamics. This allows us to estimate the gradient of the factorized optimization objective for continuous control tasks, avoiding the adversarial optimization between the soft Q-function and the policy required by prior work. By doing so, we obtain a tractable and non-adversarial objective for both discrete and continuous multi-agent control. Experiments on common benchmarks including the discrete control tasks StarCraft Multi-Agent Challenge v2 (SMACv2), Gold Miner, and Multi Particle Environments (MPE), as well as the continuous control task Multi-Agent MuJoCo (MaMuJoCo), demonstrate that MAFIS achieves superior performance compared with baselines. Our code is available at https://github.com/LAMDA-RL/MAFIS .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Multi-agent imitation learning with function approximation: linear Markov games and beyondLuca Viano, Till Freihaut, Emanuele Nevali, Volkan Cevher 等ICML 2026 · 被引用 1 次
- Non-Adversarial Imitation Learning Provably Free of Compounding Errors: The Value Flow MechanismTian Xu, Chenyang Wang, Xiaochen Zhai, Ziniu Li 等ICML 2026
它引用的顶会 Paper10
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu 等ICLR 2021 · 被引用 595 次
- PettingZoo: Gym for Multi-Agent Reinforcement LearningJ. K. Terry, Benjamin Black, Nathaniel Grammel, Mario Jayakumar 等NeurIPS 2021 · 被引用 478 次
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song 等NeurIPS 2021 · 被引用 271 次
- Error Bounds of Imitating Policies and EnvironmentsTian Xu, Ziniu Li, Yang YuNeurIPS 2020 · 被引用 141 次
- On the Loss Landscape of Adversarial Training: Identifying Challenges and How to Overcome ThemChen Liu, Mathieu Salzmann, Tao Lin, Ryota Tomioka 等NeurIPS 2020 · 被引用 103 次
相关 Paper
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 被引用 299 次
- Inverse Factorized Soft Q-Learning for Cooperative Multi-agent Imitation LearningThe Viet Bui, Tien Mai, Thanh Hong NguyenNeurIPS 2024 · 被引用 10 次
- Bayesian Multi-type Mean Field Multi-agent Imitation LearningFan Yang, Alina Vereshchaka, Changyou Chen, Wen DongNeurIPS 2020 · 被引用 21 次
- Multi-Agent Interactions Modeling with Correlated PoliciesMinghuan Liu, Ming Zhou, Weinan Zhang, Yuzheng Zhuang 等ICLR 2020 · 被引用 22 次
- FACMAC: Factored Multi-Agent Centralised Policy GradientsBei Peng, Tabish Rashid, Christian Schröder de Witt, Pierre-Alexandre Kamienny 等NeurIPS 2021 · 被引用 399 次
