PlayVirtual: Augmenting Cycle-Consistent Virtual Trajectories for Reinforcement Learning
Tao Yu, Cuiling Lan, Wenjun Zeng, Mingxiao Feng, Zhizheng Zhang, Zhibo Chen
摘要
Learning good feature representations is important for deep reinforcement learning (RL). However, with limited experience, RL often suffers from data inefficiency for training. For un-experienced or less-experienced trajectories (i.e., state-action sequences), the lack of data limits the use of them for better feature learning. In this work, we propose a novel method, dubbed PlayVirtual, which augments cycle-consistent virtual trajectories to enhance the data efficiency for RL feature representation learning. Specifically, PlayVirtual predicts future states in the latent space based on the current state and action by a dynamics model and then predicts the previous states by a backward dynamics model, which forms a trajectory cycle. Based on this, we augment the actions to generate a large amount of virtual state-action trajectories. Being free of groudtruth state supervision, we enforce a trajectory to meet the cycle consistency constraint, which can significantly enhance the data efficiency. We validate the effectiveness of our designs on the Atari and DeepMind Control Suite benchmarks. Our method achieves the state-of-the-art performance on both benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Reinforcement Learning with Action-Free Pre-Training from VideosYounggyo Seo, Kimin Lee, Stephen James, Pieter AbbeelICML 2022 · 被引用 150 次
- Pre-Trained Image Encoder for Generalizable Visual Reinforcement LearningZhecheng Yuan, Zhengrong Xue, Bo Yuan, Xueqian Wang 等NeurIPS 2022 · 被引用 112 次
- Mask-based Latent Reconstruction for Reinforcement LearningTao Yu, Zhizheng Zhang, Cuiling Lan, Yan Lu 等NeurIPS 2022 · 被引用 80 次
- PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement LearningHojoon Lee, Hanseul Cho, Hyunseung Kim, Daehoon Gwak 等NeurIPS 2023 · 被引用 50 次
- Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels?Xiang Li, Jinghuan Shang, Srijan Das, Michael S. RyooNeurIPS 2022 · 被引用 43 次
它引用的顶会 Paper21
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
相关 Paper
- Become a Proficient Player with Limited Data through Watching Pure VideosWeirui Ye, Yunsheng Zhang, Pieter Abbeel, Yang GaoICLR 2023
- Value-Consistent Representation Learning for Data-Efficient Reinforcement LearningYang Yue, Bingyi Kang, Zhongwen Xu, Gao Huang 等AAAI 2023 · 被引用 19 次
- Pretraining Representations for Data-Efficient Reinforcement LearningMax Schwarzer, Nitarshan Rajkumar, Michael Noukhovitch, Ankesh Anand 等NeurIPS 2021 · 被引用 151 次
- On the Importance of Feature Decorrelation for Unsupervised Representation Learning in Reinforcement LearningHojoon Lee, Koanho Lee, Dongyoon Hwang, Hyunho Lee 等ICML 2023 · 被引用 11 次
- Pretraining a Shared Q-Network for Data-Efficient Offline Reinforcement LearningJongchan Park, Mingyu Park, Donghwan LeeNeurIPS 2025 · 被引用 2 次
