Mixed-Initiative Multiagent Apprenticeship Learning for Human Training of Robot Teams
Esmaeil Seraj, Jerry Xiong, Mariah Schrum, Matthew C. Gombolay
Abstract
Extending recent advances in Learning from Demonstration (LfD) frameworks to multi-robot settings poses critical challenges such as environment non-stationarity due to partial observability which is detrimental to the applicability of existing methods. Although prior work has shown that enabling communication among agents of a robot team can alleviate such issues, creating inter-agent communication under existing Multi-Agent LfD (MA-LfD) frameworks requires the human expert to provide demonstrations for both environment actions and communication actions, which necessitates an efficient communication strategy on a known message space. To address this problem, we propose Mixed-Initiative Multi-Agent Apprenticeship Learning (MixTURE). MixTURE enables robot teams to learn from a human expert-generated data a preferred policy to accomplish a collaborative task, while simultaneously learning emergent inter-agent communication to enhance team coordination. The key ingredient to MixTURE’s success is automatically learning a communication policy, enhanced by a mutual-information maximizing reverse model that rationalizes the underlying expert demonstrations without the need for human generated data or an auxiliary reward function. MixTURE outperforms a variety of relevant baselines on diverse data generated by human experts in complex heterogeneous domains. MixTURE is the first MA-LfD framework to enable learning multi-robot collaborative policies directly from real human data, resulting in 44% less human workload, and 46% higher usability score.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 995d40fb-563b-4a51-8730-13161e79e25fCited by top-tier papers3
- EMOS: Embodiment-aware Heterogeneous Multi-robot Operating System with LLM AgentsJunting Chen, Checheng Yu, Xunzhe Zhou, Tianqi Xu et al.ICLR 2025
- DLM: Unified Decision Language Models for Offline Multi-Agent Sequential Decision MakingZhuohui Zhang, Bin Cheng, Bin HeICML 2026
- Reinforcement Learning with Fuzzy Human Attention-Guided Graph for Heterogeneous Multiagent SystemsDingbang Liu, Fenghui Ren, Jun Yan, Guoxin Su et al.AAAI 2026
Builds on4
- On the Expressivity of Markov RewardDavid Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho et al.NeurIPS 2021 · 107 citations
- Interpretable and Personalized Apprenticeship Scheduling: Learning Interpretable Scheduling Policies from Heterogeneous User DemonstrationsRohan R. Paleja, Andrew Silva, Letian Chen, Matthew C. GombolayNeurIPS 2020 · 41 citations
- Iterated Reasoning with Mutual Information in Cooperative and Byzantine Decentralized TeamingSachin G. Konan, Esmaeil Seraj, Matthew C. GombolayICLR 2022 · 27 citations
- Bayesian Multi-type Mean Field Multi-agent Imitation LearningFan Yang, Alina Vereshchaka, Changyou Chen, Wen DongNeurIPS 2020 · 21 citations
Related papers
- Model Predictive Adversarial Imitation Learning for Planning from ObservationTyler Han, Yanda Bao, Bhaumik Mehta, Gabriel Guo et al.ICLR 2026 · 4 citations
- Meta-Imitation Learning by Watching Video DemonstrationsJiayi Li, Tao Lu, Xiaoge Cao, Yinghao Cai et al.ICLR 2022 · 25 citations
- Language Grounded Multi-agent Reinforcement Learning with Human-interpretable CommunicationHuao Li, Hossein Nourkhiz Mahjoub, Behdad Chalaki, Vaishnav Tadiparthi et al.NeurIPS 2024 · 31 citations
- Cheap Talk Discovery and Utilization in Multi-Agent Reinforcement LearningYat Long Lo, Christian Schröder de Witt, Samuel Sokota, Jakob Nicolaus Foerster et al.ICLR 2023
- ELEMENTAL: Interactive Learning from Demonstrations and Vision-Language Models for Reward Design in RoboticsLetian Chen, Nina Marie Moorman, Matthew Craig GombolayICML 2025
