Meta Inverse Constrained Reinforcement Learning: Convergence Guarantee and Generalization Analysis
Shicheng Liu, Minghui Zhu
Abstract
This paper considers the problem of learning the reward function and constraints of an expert from few demonstrations. This problem can be considered as a metalearning problem where we first learn meta-priors over reward functions and constraints from other distinct but related tasks and then adapt the learned meta-priors to new tasks from only few expert demonstrations. We formulate a bi-level optimization problem where the upper level aims to learn a meta-prior over reward functions and the lower level is to learn a meta-prior over constraints. We propose a novel algorithm to solve this problem and formally guarantee that the algorithm reaches the set of ϵ-stationary points at the iteration complexity O( 1 ϵ 2 ). We also quantify the generalization error to an arbitrary new task. Experiments are used to validate that the learned meta-priors can adapt to new tasks with good performance from only few demonstrations. ∞ t=0 γ t log π ω;θ (a t |s t )] where π ω;θ is the constrained soft Bellman policy (see the expression in Appendix A.2) (Liu & Zhu, 2022; 2024) under the reward function r θ and cost function c ω . The constrained soft Bellman policy is an extension of soft Bellman policy (Ziebart et al., 2010; Zhou et al., 2017) to CMDPs. The soft Bellman policy is widely used in soft Q-learning (Haarnoja et al., 2017) and soft actor-critic (Haarnoja et al., 2018) .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c912571-198b-4147-a512-e1ef894165d1Cited by top-tier papers16
- In-Trajectory Inverse Reinforcement Learning: Learn Incrementally Before an Ongoing Trajectory TerminatesShicheng Liu, Minghui ZhuNeurIPS 2024 · 11 citations
- Meta-Reinforcement Learning with Universal Policy Adaptation: Provable Near-Optimality under All-task Optimum ComparatorSiyuan Xu, Minghui ZhuNeurIPS 2024 · 8 citations
- Efficient Safe Meta-Reinforcement Learning: Provable Near-Optimality and Anytime SafetySiyuan Xu, Minghui ZhuNeurIPS 2025 · 8 citations
- Robust Inverse Constrained Reinforcement Learning under Model MisspecificationSheng Xu, Guiliang LiuICML 2024 · 7 citations
- Confidence Aware Inverse Constrained Reinforcement LearningSriram Ganapathi Subramanian, Guiliang Liu, Mohammed Elmahgiubi, Kasra Rezaee et al.ICML 2024 · 5 citations
Builds on15
- Bilevel Optimization: Convergence Analysis and Enhanced DesignKaiyi Ji, Junjie Yang, Yingbin LiangICML 2021 · 343 citations
- Neural Policy Gradient Methods: Global Optimality and Rates of ConvergenceLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2020 · 270 citations
- CRPO: A New Approach for Safe Reinforcement Learning with Convergence GuaranteeTengyu Xu, Yingbin Liang, Guanghui LanICML 2021 · 171 citations
- Maximum Likelihood Constraint Inference for Inverse Reinforcement LearningDexter R. R. Scobee, S. Shankar SastryICLR 2020 · 74 citations
- Generalization of Model-Agnostic Meta-Learning Algorithms: Recurring and Unseen TasksAlireza Fallah, Aryan Mokhtari, Asuman E. OzdaglarNeurIPS 2021 · 63 citations
Related papers
- Distributed Inverse Constrained Reinforcement Learning for Multi-agent SystemsShicheng Liu, Minghui ZhuNeurIPS 2022 · 41 citations
- Learning Soft Constraints From Constrained Expert DemonstrationsAshish Gaurav, Kasra Rezaee, Guiliang Liu, Pascal PoupartICLR 2023 · 4 citations
- Learning Multi-agent Behaviors from Distributed and Streaming DemonstrationsShicheng Liu, Minghui ZhuNeurIPS 2023 · 34 citations
- Accelerating Safe Reinforcement Learning with Constraint-mismatched Baseline PoliciesTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICML 2021 · 20 citations
- Online Constrained Meta-Learning: Provable Guarantees for GeneralizationSiyuan Xu, Minghui ZhuNeurIPS 2023 · 10 citations
