Meta Inverse Constrained Reinforcement Learning: Convergence Guarantee and Generalization Analysis
Shicheng Liu, Minghui Zhu
摘要
This paper considers the problem of learning the reward function and constraints of an expert from few demonstrations. This problem can be considered as a metalearning problem where we first learn meta-priors over reward functions and constraints from other distinct but related tasks and then adapt the learned meta-priors to new tasks from only few expert demonstrations. We formulate a bi-level optimization problem where the upper level aims to learn a meta-prior over reward functions and the lower level is to learn a meta-prior over constraints. We propose a novel algorithm to solve this problem and formally guarantee that the algorithm reaches the set of ϵ-stationary points at the iteration complexity O( 1 ϵ 2 ). We also quantify the generalization error to an arbitrary new task. Experiments are used to validate that the learned meta-priors can adapt to new tasks with good performance from only few demonstrations. ∞ t=0 γ t log π ω;θ (a t |s t )] where π ω;θ is the constrained soft Bellman policy (see the expression in Appendix A.2) (Liu & Zhu, 2022; 2024) under the reward function r θ and cost function c ω . The constrained soft Bellman policy is an extension of soft Bellman policy (Ziebart et al., 2010; Zhou et al., 2017) to CMDPs. The soft Bellman policy is widely used in soft Q-learning (Haarnoja et al., 2017) and soft actor-critic (Haarnoja et al., 2018) .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- In-Trajectory Inverse Reinforcement Learning: Learn Incrementally Before an Ongoing Trajectory TerminatesShicheng Liu, Minghui ZhuNeurIPS 2024 · 被引用 11 次
- Meta-Reinforcement Learning with Universal Policy Adaptation: Provable Near-Optimality under All-task Optimum ComparatorSiyuan Xu, Minghui ZhuNeurIPS 2024 · 被引用 8 次
- Efficient Safe Meta-Reinforcement Learning: Provable Near-Optimality and Anytime SafetySiyuan Xu, Minghui ZhuNeurIPS 2025 · 被引用 8 次
- Robust Inverse Constrained Reinforcement Learning under Model MisspecificationSheng Xu, Guiliang LiuICML 2024 · 被引用 7 次
- Confidence Aware Inverse Constrained Reinforcement LearningSriram Ganapathi Subramanian, Guiliang Liu, Mohammed Elmahgiubi, Kasra Rezaee 等ICML 2024 · 被引用 5 次
它引用的顶会 Paper15
- Bilevel Optimization: Convergence Analysis and Enhanced DesignKaiyi Ji, Junjie Yang, Yingbin LiangICML 2021 · 被引用 343 次
- Neural Policy Gradient Methods: Global Optimality and Rates of ConvergenceLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2020 · 被引用 270 次
- CRPO: A New Approach for Safe Reinforcement Learning with Convergence GuaranteeTengyu Xu, Yingbin Liang, Guanghui LanICML 2021 · 被引用 171 次
- Maximum Likelihood Constraint Inference for Inverse Reinforcement LearningDexter R. R. Scobee, S. Shankar SastryICLR 2020 · 被引用 74 次
- Generalization of Model-Agnostic Meta-Learning Algorithms: Recurring and Unseen TasksAlireza Fallah, Aryan Mokhtari, Asuman E. OzdaglarNeurIPS 2021 · 被引用 63 次
相关 Paper
- Distributed Inverse Constrained Reinforcement Learning for Multi-agent SystemsShicheng Liu, Minghui ZhuNeurIPS 2022 · 被引用 41 次
- Learning Soft Constraints From Constrained Expert DemonstrationsAshish Gaurav, Kasra Rezaee, Guiliang Liu, Pascal PoupartICLR 2023 · 被引用 4 次
- Learning Multi-agent Behaviors from Distributed and Streaming DemonstrationsShicheng Liu, Minghui ZhuNeurIPS 2023 · 被引用 34 次
- Accelerating Safe Reinforcement Learning with Constraint-mismatched Baseline PoliciesTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICML 2021 · 被引用 20 次
- Online Constrained Meta-Learning: Provable Guarantees for GeneralizationSiyuan Xu, Minghui ZhuNeurIPS 2023 · 被引用 10 次
