Q-Supervised Contrastive Representation: A State Decoupling Framework for Safe Offline Reinforcement Learning
Zhihe Yang, Yunjian Xu, Yang Zhang
Abstract
Safe offline reinforcement learning (RL), which aims to learn the safety-guaranteed policy without risky online interaction with environments, has attracted growing recent attention for safetycritical scenarios. However, existing approaches encounter out-of-distribution problems during the testing phase, which can result in potentially unsafe outcomes. This issue arises due to the infinite possible combinations of reward-related and costrelated states. In this work, we propose State Decoupling with Q-supervised Contrastive representation (SDQC), a novel framework that decouples the global observations into reward-and cost-related representations for decision-making, thereby improving the generalization capability for unfamiliar global observations. Compared with the classical representation learning methods, which typically require model-based estimation (e.g., bisimulation), we theoretically prove that our Q-supervised method generates a coarser representation while preserving the optimal policy, resulting in improved generalization performance. Experiments on DSRL benchmark provide compelling evidence that SDQC surpasses other baseline algorithms, especially for its exceptional ability to achieve almost zero violations in more than half of the tasks. Further, we demonstrate that SDQC possesses superior generalization ability when confronted with unseen environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69b8d2b3-5687-424a-8d45-1613d299babaCited by top-tier papers1
Ask how each one uses itBuilds on28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen et al.NeurIPS 2022 · 2,653 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
Related papers
- C2IQL: Constraint-Conditioned Implicit Q-learning for Safe Offline Reinforcement LearningZifan Liu, Xinran Li, Jun ZhangICML 2025
- Offline Safe Reinforcement Learning Using Trajectory ClassificationZe Gong, Akshat Kumar, Pradeep VarakanthamAAAI 2025 · 6 citations
- Constraint-Adaptive Policy Switching for Offline Safe Reinforcement LearningYassine Chemingui, Aryan Deshwal, Honghao Wei, Alan Fern et al.AAAI 2025 · 12 citations
- Reining Generalization in Offline Reinforcement Learning via Representation DistinctionYi Ma, Hongyao Tang, Dong Li, Zhaopeng MengNeurIPS 2023 · 19 citations
- VOCE: Variational Optimization with Conservative Estimation for Offline Safe Reinforcement LearningJiayi Guan, Guang Chen, Jiaming Ji, Long Yang et al.NeurIPS 2023 · 19 citations
