Look Beneath the Surface: Exploiting Fundamental Symmetry for Sample-Efficient Offline RL
Peng Cheng, Xianyuan Zhan, Zhi-Hao Wu, Wenjia Zhang, Youfang Lin, Shoucheng Song, Han Wang, Li Jiang
Abstract
Offline reinforcement learning (RL) offers an appealing approach to real-world tasks by learning policies from pre-collected datasets without interacting with the environment. However, the performance of existing offline RL algorithms heavily depends on the scale and state-action space coverage of datasets. Real-world data collection is often expensive and uncontrollable, leading to small and narrowly covered datasets and posing significant challenges for practical deployments of offline RL. In this paper, we provide a new insight that leveraging the fundamental symmetry of system dynamics can substantially enhance offline RL performance under small datasets. Specifically, we propose a Time-reversal symmetry (T-symmetry) enforced Dynamics Model (TDM), which establishes consistency between a pair of forward and reverse latent dynamics. TDM provides both well-behaved representations for small datasets and a new reliability measure for OOD samples based on compliance with the T-symmetry. These can be readily used to construct a new offline RL algorithm (TSRL) with less conservative policy constraints and a reliable latent space data augmentation procedure. Based on extensive experiments, we find TSRL achieves great performance on small benchmark datasets with as few as 1% of the original samples, which significantly outperforms the recent offline RL algorithms in terms of data efficiency and generalizability.Code is available at: https://github.com/pcheng2/TSRL
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7227ee14-d286-4891-bad6-a728e8db0b8dCited by top-tier papers16
- Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion ModelYinan Zheng, Jianxiong Li, Dongjie Yu, Yujie Yang et al.ICLR 2024 · 72 citations
- Offline Multi-Agent Reinforcement Learning with Implicit Global-to-Local Value RegularizationXiangsen Wang, Haoran Xu, Yinan Zheng, Xianyuan ZhanNeurIPS 2023 · 65 citations
- ODICE: Revealing the Mystery of Distribution Correction Estimation via Orthogonal-gradient UpdateLiyuan Mao, Haoran Xu, Weinan Zhang, Xianyuan ZhanICLR 2024 · 23 citations
- Physics-informed Value Learner for Offline Goal-Conditioned Reinforcement LearningVittorio Giammarino, Ruiqi Ni, Ahmed H. QureshiNeurIPS 2025 · 15 citations
- Are Expressive Models Truly Necessary for Offline RL?Guan Wang, Haoyi Niu, Jianxiong Li, Li Jiang et al.AAAI 2025 · 9 citations
Builds on29
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
Related papers
- Sample Efficient Offline RL via T-Symmetry Enforced Latent State-StitchingPeng Cheng, Zhihao Wu, Jianxiong Li, Ziteng He et al.ICLR 2026
- Koopman Q-learning: Offline Reinforcement Learning via Symmetries of DynamicsMatthias Weissenbacher, Samarth Sinha, Animesh Garg, Yoshinobu KawaharaICML 2022 · 33 citations
- DARA: Dynamics-Aware Reward Augmentation in Offline Reinforcement LearningJinxin Liu, Hongyin Zhang, Donglin WangICLR 2022 · 47 citations
- Time Reversal Symmetry for Efficient Robotic Manipulations in Deep Reinforcement LearningYunpeng Jiang, Jianshu Hu, Paul Weng, Yutong BanNeurIPS 2025 · 1 citation
- Double Check Your State Before Trusting It: Confidence-Aware Bidirectional Offline Model-Based ImaginationJiafei Lyu, Xiu Li, Zongqing LuNeurIPS 2022 · 35 citations
