Learning State Representations via Retracing in Reinforcement Learning
Changmin Yu, Dong Li, Jianye Hao, Jun Wang, Neil Burgess
摘要
We propose learning via retracing, a novel self-supervised approach for learning the state representation (and the associated dynamics model) for reinforcement learning tasks. In addition to the predictive (reconstruction) supervision in the forward direction, we propose to include "retraced" transitions for representation/model learning, by enforcing the cycle-consistency constraint between the original and retraced states, hence improve upon the sample efficiency of learning. Moreover, learning via retracing explicitly propagates information about future transitions backward for inferring previous states, thus facilitates stronger representation learning for the downstream reinforcement learning tasks. We introduce Cycle-Consistency World Model (CCWM), a concrete model-based instantiation of learning via retracing. Additionally we propose a novel adaptive "truncation" mechanism for counteracting the negative impacts brought by "irreversible" transitions such that learning via retracing can be maximally effective. Through extensive empirical studies on visual-based continuous control benchmarks, we demonstrate that CCWM achieves state-of-the-art performance in terms of sample efficiency and asymptotic performance, whilst exhibiting behaviours that are indicative of stronger representation learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Look Beneath the Surface: Exploiting Fundamental Symmetry for Sample-Efficient Offline RLPeng Cheng, Xianyuan Zhan, Zhi-Hao Wu, Wenjia Zhang 等NeurIPS 2023 · 被引用 23 次
- Successor-Predecessor Intrinsic ExplorationChangmin Yu, Neil Burgess, Maneesh Sahani, Samuel J. GershmanNeurIPS 2023 · 被引用 12 次
它引用的顶会 Paper10
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 被引用 911 次
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto 等NeurIPS 2020 · 被引用 833 次
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 被引用 457 次
相关 Paper
- Dynamics-Aware EmbeddingsWilliam F. Whitney, Rajat Agarwal, Kyunghyun Cho, Abhinav GuptaICLR 2020
- Learning Temporal Dynamics from Cycles in Narrated VideoDave Epstein, Jiajun Wu, Cordelia Schmid, Chen SunICCV 2021 · 被引用 15 次
- PlayVirtual: Augmenting Cycle-Consistent Virtual Trajectories for Reinforcement LearningTao Yu, Cuiling Lan, Wenjun Zeng, Mingxiao Feng 等NeurIPS 2021 · 被引用 64 次
- Simplified Temporal Consistency Reinforcement LearningYi Zhao, Wenshuai Zhao, Rinu Boney, Juho Kannala 等ICML 2023 · 被引用 19 次
- Become a Proficient Player with Limited Data through Watching Pure VideosWeirui Ye, Yunsheng Zhang, Pieter Abbeel, Yang GaoICLR 2023
