Learning Constraints from Offline Demonstrations via Superior Distribution Correction Estimation
Guorui Quan, Zhiqiang Xu, Guiliang Liu
摘要
An effective approach for learning both safety constraints and control policies is Inverse Constrained Reinforcement Learning (ICRL). Previous ICRL algorithms commonly employ an online learning framework that permits unlimited sampling from an interactive environment. This setting, however, is infeasible in many realistic applications where data collection is dangerous and expensive. To address this challenge, we propose Inverse Constrained Superior Distribution Correction Estimation (ICSDICE) as an offline ICRL solver. ICSDICE extracts feasible constraints from superior distributions, thereby highlighting policies with expert-exceeding rewards maximization ability. To estimate these distributions, ICSDICE solves a regularized dual optimization problem for safe control by exploiting the observed reward signals and expert preferences. Striving for transferable constraints and unbiased estimations, ICSDICE actively encourages sparsity and incorporates a discounting effect within the learned and observed distributions. Empirical studies show that ICSDICE outperforms other baselines by accurately recovering the constraints and adapting to high-dimensional environments. The code is available at https: //github.com/quangr/ICSDICE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Understanding Constraint Inference in Safety-Critical Inverse Reinforcement LearningBo Yue, Shufan Wang, Ashish Gaurav, Jian Li 等ICLR 2025
- Toward Exploratory Inverse Constraint Inference with Generative Diffusion VerifiersRunyi Zhao, Sheng Xu, Bo Yue, Guiliang LiuICLR 2025
它引用的顶会 Paper21
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran 等NeurIPS 2021 · 被引用 549 次
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song 等NeurIPS 2021 · 被引用 271 次
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 被引用 239 次
相关 Paper
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction EstimationJongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess 等ICLR 2022 · 被引用 84 次
- Robust Inverse Constrained Reinforcement Learning under Model MisspecificationSheng Xu, Guiliang LiuICML 2024 · 被引用 7 次
- Confidence Aware Inverse Constrained Reinforcement LearningSriram Ganapathi Subramanian, Guiliang Liu, Mohammed Elmahgiubi, Kasra Rezaee 等ICML 2024 · 被引用 5 次
- Distributed Inverse Constrained Reinforcement Learning for Multi-agent SystemsShicheng Liu, Minghui ZhuNeurIPS 2022 · 被引用 41 次
- Simplifying Constraint Inference with Inverse Reinforcement LearningAdriana Hugessen, Harley Wiltzer, Glen BersethNeurIPS 2024 · 被引用 6 次
