Safety through feedback in Constrained RL
Shashank Reddy Chirra, Pradeep Varakantham, Praveen Paruchuri
Abstract
In safety-critical RL settings, the inclusion of an additional cost function is often favoured over the arduous task of modifying the reward function to ensure the agent's safe behaviour. However, designing or evaluating such a cost function can be prohibitively expensive. For instance, in the domain of self-driving, designing a cost function that encompasses all unsafe behaviours (e.g. aggressive lane changes) is inherently complex. In such scenarios, the cost function can be learned from feedback collected offline in between training rounds. This feedback can be system generated or elicited from a human observing the training process. Previous approaches have not been able to scale to complex environments and are constrained to receiving feedback at the state level which can be expensive to collect. To this end, we introduce an approach that scales to more complex domains and extends to beyond state-level feedback, thus, reducing the burden on the evaluator. Inferring the cost function in such settings poses challenges, particularly in assigning credit to individual states based on trajectory-level feedback. To address this, we propose a surrogate objective that transforms the problem into a state-level supervised classification task with noisy labels, which can be solved efficiently. Additionally, it is often infeasible to collect feedback on every trajectory generated by the agent, hence, two fundamental questions arise: (1) Which trajectories should be presented to the human? and (2) How many trajectories are necessary for effective learning? To address these questions, we introduce novelty-based sampling that selectively involves the evaluator only when the the agent encounters a novel trajectory. We showcase the efficiency of our method through experimentation on several benchmark Safety Gymnasium environments and realistic self-driving scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 542aaabe-ebf7-42fb-b93c-20523087363fCited by top-tier papers2
- TraCeS: Learning Per-Timestep Constraint-Violation Credit from Sparse Trajectory-Level LabelsSiow Meng Low, Ze Gong, Akshat KumarICML 2026
- Safe Reinforcement Learning with Preference-based Constraint InferenceChenglin Li, Grant Ruan, Hua GengICML 2026
Builds on12
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song et al.NeurIPS 2021 · 271 citations
- First Order Constrained Optimization in Policy SpaceYiming Zhang, Quan Vuong, Keith W. RossNeurIPS 2020 · 238 citations
- Safe Reinforcement Learning in Constrained Markov Decision ProcessesAkifumi Wachi, Yanan SuiICML 2020 · 190 citations
- Constrained Update Projection Approach to Safe Policy OptimizationLong Yang, Jiaming Ji, Juntao Dai, Linrui Zhang et al.NeurIPS 2022 · 95 citations
Related papers
- Offline Safe Reinforcement Learning Using Trajectory ClassificationZe Gong, Akshat Kumar, Pradeep VarakanthamAAAI 2025 · 6 citations
- Learning Human Objectives by Evaluating Hypothetical BehaviorSiddharth Reddy, Anca D. Dragan, Sergey Levine, Shane Legg et al.ICML 2020 · 81 citations
- Imitate the Good and Avoid the Bad: An Incremental Approach to Safe Reinforcement LearningHuy Hoang, Tien Mai, Pradeep VarakanthamAAAI 2024 · 8 citations
- Exploring Safer Behaviors for Deep Reinforcement LearningEnrico Marchesini, Davide Corsi, Alessandro FarinelliAAAI 2022 · 36 citations
- From Text to Trajectory: Exploring Complex Constraint Representation and Decomposition in Safe Reinforcement LearningPusen Dong, Tianchen Zhu, Yue Qiu, Haoyi Zhou et al.NeurIPS 2024 · 2 citations
