Constrained Decision Transformer for Offline Safe Reinforcement Learning
Zuxin Liu, Zijian Guo, Yihang Yao, Zhepeng Cen, Wenhao Yu, Tingnan Zhang, Ding Zhao
Abstract
Safe reinforcement learning (RL) trains a constraint satisfaction policy by interacting with the environment. We aim to tackle a more challenging problem: learning a safe policy from an offline dataset. We study the offline safe RL problem from a novel multi-objective optimization perspective and propose the -reducible concept to characterize problem difficulties. The inherent trade-offs between safety and task performance inspire us to propose the constrained decision transformer (CDT) approach, which can dynamically adjust the trade-offs during deployment. Extensive experiments show the advantages of the proposed method in learning an adaptive, safe, robust, and high-reward policy. CDT outperforms its variants and strong offline safe RL baselines by a large margin with the same hyperparameters across all tasks, while keeping the zero-shot adaptation capability to different constraint thresholds, making our approach more suitable for real-world RL under constraints. The code is available at https://github.com/liuzuxin/OSRL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55f9b9f7-cb61-42ee-8e89-39ac501f7117Cited by top-tier papers47
- Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion ModelYinan Zheng, Jianxiong Li, Dongjie Yu, Yujie Yang et al.ICLR 2024 · 72 citations
- TAIL: Task-specific Adapters for Imitation Learning with Large Pretrained ModelsZuxin Liu, Jesse Zhang, Kavosh Asadi, Yao Liu et al.ICLR 2024 · 46 citations
- Safe Offline Reinforcement Learning with Real-Time Budget ConstraintsQian Lin, Bo Tang, Zifan Wu, Chao Yu et al.ICML 2023 · 31 citations
- Survival Instinct in Offline Reinforcement LearningAnqi Li, Dipendra Misra, Andrey Kolobov, Ching-An ChengNeurIPS 2023 · 26 citations
- Decision Mamba: A Multi-Grained State Space Model with Self-Evolution Regularization for Offline RLQi Lv, Xiang Deng, Gongwei Chen, Michael Yu Wang et al.NeurIPS 2024 · 25 citations
Builds on20
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 403 citations
- Offline Reinforcement Learning with Fisher Divergence Critic RegularizationIlya Kostrikov, Rob Fergus, Jonathan Tompson, Ofir NachumICML 2021 · 350 citations
Related papers
- Adaptable Safe Policy Learning from Multi-task Data with Constraint Prioritized Decision TransformerRuiqi Xue, Ziqian Zhang, Lihe Li, Cong Guan et al.NeurIPS 2025 · 2 citations
- Constraint-Adaptive Policy Switching for Offline Safe Reinforcement LearningYassine Chemingui, Aryan Deshwal, Honghao Wei, Alan Fern et al.AAAI 2025 · 12 citations
- Temporal Logic Specification-Conditioned Decision Transformer for Offline Safe Reinforcement LearningZijian Guo, Weichao Zhou, Wenchao LiICML 2024 · 7 citations
- Constraints Penalized Q-learning for Safe Offline Reinforcement LearningHaoran Xu, Xianyuan Zhan, Xiangyu ZhuAAAI 2022 · 127 citations
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction EstimationJongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess et al.ICLR 2022 · 84 citations
