Sample Efficient Offline RL via T-Symmetry Enforced Latent State-Stitching
Peng Cheng, Zhihao Wu, Jianxiong Li, Ziteng He, Haoran Xu, Wei Sun, Youfang Lin, Yunxin Liu, Xianyuan Zhan
摘要
Offline reinforcement learning (RL) has achieved notable progress in recent years. However, most existing offline RL methods require a large amount of training data to achieve reasonable performance and offer limited out-of-distribution (OOD) generalization capability due to conservative data-related regularizations. This seriously hinders the usability of offline RL in solving many real-world applications, where the available data are often limited. In this study, we introduce TELS, a highly sample-efficient offline RL algorithm that enables state-stitching in a compact latent space regulated by the fundamental time-reversal symmetry (T-symmetry) of dynamical systems. Specifically, we introduce a T-symmetry enforced inverse dynamics model (TS-IDM) to derive well-regulated latent state representations that greatly facilitate OOD generalization. A guide-policy can then be learned entirely in the latent space to optimize for the reward-maximizing next state, bypassing the conservative action-level behavioral regularization adopted in most offline RL methods. Finally, the optimized action can be extracted using the learned TS-IDM, together with the optimized latent next state from the guide-policy. We conducted comprehensive experiments on both the D4RL benchmark tasks and a real-world industrial control test environment, TELS achieves superior sample efficiency and OOD generalization performance, significantly outperforming existing offline RL methods in a wide range of challenging small-sample tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper44
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
相关 Paper
- Look Beneath the Surface: Exploiting Fundamental Symmetry for Sample-Efficient Offline RLPeng Cheng, Xianyuan Zhan, Zhi-Hao Wu, Wenjia Zhang 等NeurIPS 2023 · 被引用 23 次
- Recovering from Out-of-sample States via Inverse Dynamics in Offline Reinforcement LearningKe Jiang, Jia-Yu Yao, Xiaoyang TanNeurIPS 2023 · 被引用 12 次
- Time Reversal Symmetry for Efficient Robotic Manipulations in Deep Reinforcement LearningYunpeng Jiang, Jianshu Hu, Paul Weng, Yutong BanNeurIPS 2025 · 被引用 1 次
- Koopman Q-learning: Offline Reinforcement Learning via Symmetries of DynamicsMatthias Weissenbacher, Samarth Sinha, Animesh Garg, Yoshinobu KawaharaICML 2022 · 被引用 33 次
- Iteratively Refined Behavior Regularization for Offline Reinforcement LearningYi Ma, Jianye Hao, Xiaohan Hu, Yan Zheng 等NeurIPS 2024 · 被引用 11 次
