Safety through feedback in Constrained RL
Shashank Reddy Chirra, Pradeep Varakantham, Praveen Paruchuri
摘要
In safety-critical RL settings, the inclusion of an additional cost function is often favoured over the arduous task of modifying the reward function to ensure the agent's safe behaviour. However, designing or evaluating such a cost function can be prohibitively expensive. For instance, in the domain of self-driving, designing a cost function that encompasses all unsafe behaviours (e.g. aggressive lane changes) is inherently complex. In such scenarios, the cost function can be learned from feedback collected offline in between training rounds. This feedback can be system generated or elicited from a human observing the training process. Previous approaches have not been able to scale to complex environments and are constrained to receiving feedback at the state level which can be expensive to collect. To this end, we introduce an approach that scales to more complex domains and extends to beyond state-level feedback, thus, reducing the burden on the evaluator. Inferring the cost function in such settings poses challenges, particularly in assigning credit to individual states based on trajectory-level feedback. To address this, we propose a surrogate objective that transforms the problem into a state-level supervised classification task with noisy labels, which can be solved efficiently. Additionally, it is often infeasible to collect feedback on every trajectory generated by the agent, hence, two fundamental questions arise: (1) Which trajectories should be presented to the human? and (2) How many trajectories are necessary for effective learning? To address these questions, we introduce novelty-based sampling that selectively involves the evaluator only when the the agent encounters a novel trajectory. We showcase the efficiency of our method through experimentation on several benchmark Safety Gymnasium environments and realistic self-driving scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- TraCeS: Learning Per-Timestep Constraint-Violation Credit from Sparse Trajectory-Level LabelsSiow Meng Low, Ze Gong, Akshat KumarICML 2026
- Safe Reinforcement Learning with Preference-based Constraint InferenceChenglin Li, Grant Ruan, Hua GengICML 2026
它引用的顶会 Paper12
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song 等NeurIPS 2021 · 被引用 271 次
- First Order Constrained Optimization in Policy SpaceYiming Zhang, Quan Vuong, Keith W. RossNeurIPS 2020 · 被引用 238 次
- Safe Reinforcement Learning in Constrained Markov Decision ProcessesAkifumi Wachi, Yanan SuiICML 2020 · 被引用 190 次
- Constrained Update Projection Approach to Safe Policy OptimizationLong Yang, Jiaming Ji, Juntao Dai, Linrui Zhang 等NeurIPS 2022 · 被引用 95 次
相关 Paper
- Offline Safe Reinforcement Learning Using Trajectory ClassificationZe Gong, Akshat Kumar, Pradeep VarakanthamAAAI 2025 · 被引用 6 次
- Learning Human Objectives by Evaluating Hypothetical BehaviorSiddharth Reddy, Anca D. Dragan, Sergey Levine, Shane Legg 等ICML 2020 · 被引用 81 次
- Imitate the Good and Avoid the Bad: An Incremental Approach to Safe Reinforcement LearningHuy Hoang, Tien Mai, Pradeep VarakanthamAAAI 2024 · 被引用 8 次
- Exploring Safer Behaviors for Deep Reinforcement LearningEnrico Marchesini, Davide Corsi, Alessandro FarinelliAAAI 2022 · 被引用 36 次
- From Text to Trajectory: Exploring Complex Constraint Representation and Decomposition in Safe Reinforcement LearningPusen Dong, Tianchen Zhu, Yue Qiu, Haoyi Zhou 等NeurIPS 2024 · 被引用 2 次
