Inverse Factorized Soft Q-Learning for Cooperative Multi-agent Imitation Learning
The Viet Bui, Tien Mai, Thanh Hong Nguyen
Abstract
This paper concerns imitation learning (IL) in cooperative multi-agent systems. The learning problem under consideration poses several challenges, characterized by high-dimensional state and action spaces and intricate inter-agent dependencies. In a single-agent setting, IL was shown to be done efficiently via an inverse soft-Q learning process. However, extending this framework to a multi-agent context introduces the need to simultaneously learn both local value functions to capture local observations and individual actions, and a joint value function for exploiting centralized learning. In this work, we introduce a new multi-agent IL algorithm designed to address these challenges. Our approach enables the centralized learning by leveraging mixing networks to aggregate decentralized Q functions. We further establish conditions for the mixing networks under which the multi-agent IL objective function exhibits convexity within the Q function space. We present extensive experiments conducted on some challenging multi-agent game environments, including an advanced version of the Star-Craft multi-agent challenge ( SMACv2 ), which demonstrates the effectiveness of our algorithm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation LearningTill Freihaut, Luca Viano, Volkan Cevher, Matthieu Geist et al.NeurIPS 2025 · 4 citations
- Multi-Agent Imitation by Learning and Sampling from Factorized Soft Q-FunctionYi-Chen Li, Zhongxiang Ling, Tao Jiang, Fuxiang Zhang et al.NeurIPS 2025 · 3 citations
- Multi-agent imitation learning with function approximation: linear Markov games and beyondLuca Viano, Till Freihaut, Emanuele Nevali, Volkan Cevher et al.ICML 2026 · 1 citation
- MisoDICE: Multi-Agent Imitation from Mixed-Quality DemonstrationsThe Viet Bui, Tien Anh Mai, Thanh Hong NguyenNeurIPS 2025
- O-MAPL: Offline Multi-agent Preference LearningThe Viet Bui, Tien Mai, Thanh Hong NguyenICML 2025
Builds on3
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song et al.NeurIPS 2021 · 271 citations
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 239 citations
- Adversarial Cooperative Imitation Learning for Dynamic Treatment Regimes✱Lu Wang, Wenchao Yu, Xiaofeng He, Wei Cheng et al.WWW 2020 · 33 citations
Related papers
- Mimicking To Dominate: Imitation Learning Strategies for Success in Multiagent GamesThe Viet Bui, Tien Mai, Thanh Hong NguyenNeurIPS 2024 · 5 citations
- SMIX(λ): Enhancing Centralized Value Functions for Cooperative Multi-Agent Reinforcement LearningChao Wen, Xinghu Yao, Yuhui Wang, Xiaoyang TanAAAI 2020 · 57 citations
- MANSA: Learning Fast and Slow in Multi-Agent SystemsDavid Henry Mguni, Haojun Chen, Taher Jafferjee, Jianhong Wang et al.ICML 2023 · 4 citations
- FACMAC: Factored Multi-Agent Centralised Policy GradientsBei Peng, Tabish Rashid, Christian Schröder de Witt, Pierre-Alexandre Kamienny et al.NeurIPS 2021 · 399 citations
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 1,960 citations
