Offline Safe Reinforcement Learning Using Trajectory Classification
Ze Gong, Akshat Kumar, Pradeep Varakantham
Abstract
Offline safe reinforcement learning (RL) has emerged as a promising approach for learning safe behaviors without engaging in risky online interactions with the environment. Most existing methods in offline safe RL rely on cost constraints at each time step (derived from global cost constraints) and this can result in either overly conservative policies or violation of safety constraints. In this paper, we propose to learn a policy that generates desirable trajectories and avoids undesirable trajectories. To be specific, we first partition the pre-collected dataset of state-action trajectories into desirable and undesirable subsets. Intuitively, the desirable set contains high reward and safe trajectories, and undesirable set contains unsafe trajectories and low-reward safe trajectories. Second, we learn a policy that generates desirable trajectories and avoids undesirable trajectories, where (un)desirability scores are provided by a classifier learnt from the dataset of desirable and undesirable trajectories. This approach bypasses the computational complexity and stability issues of a min-max objective that is employed in existing methods. Theoretically, we also show our approach’s strong connections to existing learning paradigms involving human feedback. Finally, we extensively evaluate our method using the DSRL benchmark for offline safe RL. Empirically, our method outperforms competitive baselines, achieving higher rewards and better constraint satisfaction across a wide variety of benchmark tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d5138c31-1382-4458-943d-6894bb17816eCited by top-tier papers3
- Boundary-to-Region Supervision for Offline Safe Reinforcement LearningHuikang Su, Dengyun Peng, Zifeng Zhuang, Yuhan Liu et al.NeurIPS 2025 · 2 citations
- Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RLJunyu Guo, Zhi Zheng, Donghao Ying, Ming Jin et al.NeurIPS 2025 · 2 citations
- DualCOIL: Offline Imitation Learning from Contrasting DemonstrationsHuy Hoang, Tien Mai, Pradeep Varakantham, Tanvi VermaICML 2026
Builds on10
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 403 citations
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 252 citations
- OptiDICE: Offline Policy Optimization via Stationary Distribution Correction EstimationJongmin Lee, Wonseok Jeon, Byung-Jun Lee, Joelle Pineau et al.ICML 2021 · 137 citations
- Constraints Penalized Q-learning for Safe Offline Reinforcement LearningHaoran Xu, Xianyuan Zhan, Xiangyu ZhuAAAI 2022 · 127 citations
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction EstimationJongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess et al.ICLR 2022 · 84 citations
Related papers
- Online Optimization for Offline Safe Reinforcement LearningYassine Chemingui, Aryan Deshwal, Alan Fern, Thanh Nguyen-Tang et al.NeurIPS 2025 · 3 citations
- SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred TrajectoriesReturaj Burnwal, Nirav Pravinbhai Bhatt, Balaraman RavindranAAAI 2026
- Q-Supervised Contrastive Representation: A State Decoupling Framework for Safe Offline Reinforcement LearningZhihe Yang, Yunjian Xu, Yang ZhangICML 2025
- C2IQL: Constraint-Conditioned Implicit Q-learning for Safe Offline Reinforcement LearningZifan Liu, Xinran Li, Jun ZhangICML 2025
- Constraint-Adaptive Policy Switching for Offline Safe Reinforcement LearningYassine Chemingui, Aryan Deshwal, Honghao Wei, Alan Fern et al.AAAI 2025 · 12 citations
