SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred Trajectories
Returaj Burnwal, Nirav Pravinbhai Bhatt, Balaraman Ravindran
摘要
In this work, we study the problem of offline safe imitation learning (IL). In many real-world settings, online interactions can be risky, and accurately specifying the reward and the safety cost information at each timestep can be difficult. However, it is often feasible to collect trajectories reflecting undesirable or risky behavior, implicitly conveying the behavior the agent should avoid. We refer to these trajectories as non-preferred trajectories. Unlike standard IL, which aims to mimic demonstrations, our agent must also learn to avoid risky behavior using non-preferred trajectories. In this paper, we propose a novel approach, SafeMIL, to learn a parameterized cost that predicts if the state-action pair is risky via Multiple Instance Learning. The learned cost is then used to avoid non-preferred behaviors, resulting in a policy that prioritizes safety. We empirically demonstrate that our approach can learn a safer policy that satisfies cost constraints without degrading the reward performance, thereby outperforming several baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu 等ICLR 2020 · 被引用 751 次
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- DemoDICE: Offline Imitation Learning with Supplementary Imperfect DemonstrationsGeon-Hyeong Kim, Seokin Seo, Jongmin Lee, Wonseok Jeon 等ICLR 2022 · 被引用 111 次
- Discriminator-Weighted Offline Imitation Learning from Suboptimal DemonstrationsHaoran Xu, Xianyuan Zhan, Honglei Yin, Huiling QinICML 2022 · 被引用 105 次
相关 Paper
- SafeDICE: Offline Safe Imitation Learning with Non-Preferred DemonstrationsYoungsoo Jang, Geon-Hyeong Kim, Jongmin Lee, Sungryull Sohn 等NeurIPS 2023 · 被引用 9 次
- Offline Safe Reinforcement Learning Using Trajectory ClassificationZe Gong, Akshat Kumar, Pradeep VarakanthamAAAI 2025 · 被引用 6 次
- DualCOIL: Offline Imitation Learning from Contrasting DemonstrationsHuy Hoang, Tien Mai, Pradeep Varakantham, Tanvi VermaICML 2026
- No Experts, No Problem: Avoidance Learning from Bad DemonstrationsHuy Hoang, Tien Mai, Pradeep VarakanthamNeurIPS 2025 · 被引用 2 次
- Constraints Penalized Q-learning for Safe Offline Reinforcement LearningHaoran Xu, Xianyuan Zhan, Xiangyu ZhuAAAI 2022 · 被引用 127 次
