Learning Multi-agent Behaviors from Distributed and Streaming Demonstrations
Shicheng Liu, Minghui Zhu
Abstract
This paper considers the problem of inferring the behaviors of multiple interacting experts by estimating their reward functions and constraints where the distributed demonstrated trajectories are sequentially revealed to a group of learners. We formulate the problem as a distributed online bi-level optimization problem where the outer-level problem is to estimate the reward functions and the inner-level problem is to learn the constraints and corresponding policies. We propose a novel “multi-agent behavior inference from distributed and streaming demonstrations" (MA-BIRDS) algorithm that allows the learners to solve the outer-level and inner-level problems in a single loop through intermittent communications. We formally guarantee that the distributed learners achieve consensus on reward functions, constraints, and policies, the average local regret (over N online iterations) decreases at the rate of O (1 /N 1 − η 1 +1 /N 1 − η 2 +1 /N ) , and the cumulative constraint violation increases sub-linearly at the rate of O ( N η 2 + 1) where η 1 , η 2 ∈ (1 / 2 , 1) .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e4146dd-03e0-45ee-88bd-bea661f47fa4Cited by top-tier papers18
- Meta Inverse Constrained Reinforcement Learning: Convergence Guarantee and Generalization AnalysisShicheng Liu, Minghui ZhuICLR 2024 · 26 citations
- MAPF-GPT: Imitation Learning for Multi-Agent Pathfinding at ScaleAnton Andreychuk, Konstantin S. Yakovlev, Aleksandr Panov, Alexey SkrynnikAAAI 2025 · 19 citations
- In-Trajectory Inverse Reinforcement Learning: Learn Incrementally Before an Ongoing Trajectory TerminatesShicheng Liu, Minghui ZhuNeurIPS 2024 · 11 citations
- Meta-Reinforcement Learning with Universal Policy Adaptation: Provable Near-Optimality under All-task Optimum ComparatorSiyuan Xu, Minghui ZhuNeurIPS 2024 · 8 citations
- Robust Inverse Constrained Reinforcement Learning under Model MisspecificationSheng Xu, Guiliang LiuICML 2024 · 7 citations
Builds on9
- Bilevel Optimization: Convergence Analysis and Enhanced DesignKaiyi Ji, Junjie Yang, Yingbin LiangICML 2021 · 343 citations
- Neural Policy Gradient Methods: Global Optimality and Rates of ConvergenceLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2020 · 270 citations
- Maximum Likelihood Constraint Inference for Inverse Reinforcement LearningDexter R. R. Scobee, S. Shankar SastryICLR 2020 · 74 citations
- Maximum-Likelihood Inverse Reinforcement Learning with Finite-Time GuaranteesSiliang Zeng, Chenliang Li, Alfredo García, Mingyi HongNeurIPS 2022 · 60 citations
- Distributed Inverse Constrained Reinforcement Learning for Multi-agent SystemsShicheng Liu, Minghui ZhuNeurIPS 2022 · 41 citations
Related papers
- Multi-Agent Learning from LearnersMine Melodi Caliskan, Francesco Chini, Setareh MaghsudiICML 2023
- Multi-Modal Inverse Constrained Reinforcement Learning from a Mixture of DemonstrationsGuanren Qiao, Guiliang Liu, Pascal Poupart, Zhiqiang XuNeurIPS 2023 · 28 citations
- Sub-optimal Experts mitigate Ambiguity in Inverse Reinforcement LearningRiccardo Poiani, Gabriele Curti, Alberto Maria Metelli, Marcello RestelliNeurIPS 2024 · 2 citations
- Stronger Benchmarks for Prediction as a Service with ConstraintsYahav Bechavod, Jiuyao Lu, Aaron RothICML 2026
- Queue Up Your Regrets: Achieving the Dynamic Capacity Region of Multiplayer BanditsIlai Bistritz, Nicholas BambosNeurIPS 2022 · 5 citations
