Multi-Agent Imitation Learning: Value is Easy, Regret is Hard
Jingwu Tang, Gokul Swamy, Fei Fang, Zhiwei Steven Wu
Abstract
We study a multi-agent imitation learning (MAIL) problem where we take the perspective of a learner attempting to coordinate a group of agents based on demonstrations of an expert doing so. Most prior work in MAIL essentially reduces the problem to matching the behavior of the expert within the support of the demonstrations. While doing so is sufficient to drive the value gap between the learner and the expert to zero under the assumption that agents are non-strategic, it does not guarantee robustness to deviations by strategic agents. Intuitively, this is because strategic deviations can depend on a counterfactual quantity: the coordinator's recommendations outside of the state distribution their recommendations induce. In response, we initiate the study of an alternative objective for MAIL in Markov Games we term the regret gap that explicitly accounts for potential deviations by agents in the group. We first perform an in-depth exploration of the relationship between the value and regret gaps. First, we show that while the value gap can be efficiently minimized via a direct extension of single-agent IL algorithms, even value equivalence can lead to an arbitrarily large regret gap. This implies that achieving regret equivalence is harder than achieving value equivalence in MAIL. We then provide a pair of efficient reductions to no-regret online convex optimization that are capable of minimizing the regret gap (a) under a coverage assumption on the expert (MALICE) or (b) with access to a queryable expert (BLADES).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f384097-8f48-4073-9284-87cdf5dd2389Cited by top-tier papers5
- MAPF-GPT: Imitation Learning for Multi-Agent Pathfinding at ScaleAnton Andreychuk, Konstantin S. Yakovlev, Aleksandr Panov, Alexey SkrynnikAAAI 2025 · 19 citations
- On Feasible Rewards in Multi-Agent Inverse Reinforcement LearningTill Freihaut, Giorgia RamponiNeurIPS 2025 · 5 citations
- Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation LearningTill Freihaut, Luca Viano, Volkan Cevher, Matthieu Geist et al.NeurIPS 2025 · 4 citations
- Multi-agent imitation learning with function approximation: linear Markov games and beyondLuca Viano, Till Freihaut, Emanuele Nevali, Volkan Cevher et al.ICML 2026 · 1 citation
- Matching Multiple Experts: On the Exploitability of Multi-Agent Imitation LearningAntoine Bergerault, Volkan Cevher, Negar MehrICLR 2026
Builds on9
- The Value Equivalence Principle for Model-Based Reinforcement LearningChristopher Grimm, André Barreto, Satinder Singh, David SilverNeurIPS 2020 · 129 citations
- Of Moments and Matching: A Game-Theoretic Framework for Closing the Imitation GapGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Steven WuICML 2021 · 90 citations
- Inverse Reinforcement Learning without Reinforcement LearningGokul Swamy, David Wu, Sanjiban Choudhury, Drew Bagnell et al.ICML 2023 · 49 citations
- Sequence Model Imitation Learning with Unobserved ContextsGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Zhiwei Steven WuNeurIPS 2022 · 39 citations
- Causal Imitation Learning under Temporally Correlated NoiseGokul Swamy, Sanjiban Choudhury, Drew Bagnell, Steven WuICML 2022 · 36 citations
Related papers
- Regret Minimization and Convergence to Equilibria in General-sum Markov GamesLiad Erez, Tal Lancewicki, Uri Sherman, Tomer Koren et al.ICML 2023 · 35 citations
- Online Learning in Unknown Markov GamesYi Tian, Yuanhao Wang, Tiancheng Yu, Suvrit SraICML 2021 · 48 citations
- On Efficient Online Imitation Learning via ClassificationYichen Li, Chicheng ZhangNeurIPS 2022 · 7 citations
- Offline Learning in Markov Games with General Function ApproximationYuheng Zhang, Yu Bai, Nan JiangICML 2023 · 17 citations
- Multi-Objective Online LearningJiyan Jiang, Wenpeng Zhang, Shiji Zhou, Lihong Gu et al.ICLR 2023 · 35 citations
